A large model migration attack method, device, equipment, medium and program product
Patent Information
- Application Number
- CN202611186277.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-06
- Publication Date
- 2026-09-04
AI Technical Summary
[0004]然而,现有迁移攻击方法主要存在以下技术缺陷:一是输入变换策略单一,变换参数固定或缺乏有效的自适应优化机制,导致生成的对抗样本模式趋同,多样性匮乏;二是优化过程过度依赖固定结构的替代模型和静态优化算法,在复杂损失曲面中易陷入局部最优,迁移成功率不高
[0019] Compared with existing technologies, the large model transfer attack method, apparatus, device, medium, and program products disclosed in this invention integrate multi-dimensional visual joint transformation to achieve simultaneous attacks on multiple perceptual dimensions of images. Compared with the single transformation or simple superposition strategies in existing technologies, the multi-dimensional joint transformation of this invention can more comprehensively cover the vulnerable areas of the target model's decision boundary, significantly improving the success rate of cross-model transfer of adversarial examples. Furthermore, this invention constructs a parameter search model to adaptively generate image transformation weights, realizing adaptive dynamic adjustment of attack parameters, and designs a learnable search mechanism that can continuously optimize transformation weights based on attack feedback during iteration. Compared with the parameter configuration methods in existing technologies that rely on manual setting or static optimization, this invention effectively avoids the defect of traditional gradient-based attack methods that are prone to getting trapped in local minima, enhancing the robustness and versatility of the attack.
Smart Images

Figure CN122698379A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method, apparatus, device, medium, and program product for large model migration attacks. Background Technology
[0002] Adversarial examples are input data that introduces subtle, imperceptible perturbations into the original input data, inducing machine learning models to make incorrect predictions or classifications. The existence of adversarial examples reveals the inherent vulnerability of deep learning models in terms of robustness. The process of generating adversarial examples is called an adversarial attack, while a transferable attack refers to an attacker using a known-structured surrogate model to generate adversarial examples that can successfully attack a target model with an unknown structure. Because transferable attacks do not require knowledge of the target model's internal parameters, they pose a significant real-world threat in black-box attack scenarios and can lead to severe economic losses in critical applications such as facial recognition and autonomous driving.
[0003] Existing transfer attack techniques can be broadly categorized into two types: model optimization-based methods and input transformation-based methods. In model optimization-based transfer attacks, existing techniques generate transferable adversarial examples by designing specific loss functions or intervening in the model's internal feature layers. In input transformation-based transfer attacks, existing techniques enrich the diversity of adversarial examples by performing data augmentation on the input image.
[0004] However, existing transfer attack methods mainly suffer from the following technical defects: First, the input transformation strategy is singular, the transformation parameters are fixed or lack an effective adaptive optimization mechanism, resulting in the convergence of generated adversarial sample patterns and a lack of diversity; Second, the optimization process relies excessively on fixed-structure alternative models and static optimization algorithms, which are prone to getting trapped in local optima in complex loss surfaces, resulting in a low transfer success rate. Summary of the Invention
[0005] The purpose of this invention is to provide a method, apparatus, device, medium, and program product for large model migration attacks, which can generate attack samples through a collaborative multi-dimensional visual transformation and search network adaptive feedback mechanism, thereby enhancing the effectiveness of migration attacks.
[0006] To achieve the above objectives, embodiments of the present invention provide a large model migration attack method, comprising: The original sample image is input into a preset parameter search model to obtain the image transformation weights currently generated by the parameter search model; Based on the image transformation weights, the original sample image is subjected to multi-dimensional visual transformation to obtain the transformed sample image; The transformed sample image is input into a preset alternative model for iterative attack to generate adversarial examples; Based on the attack effect of the adversarial sample, the network parameters of the parameter search model are updated accordingly. Repeat the above steps until the preset iteration termination condition is met to obtain the final target adversarial sample.
[0007] As an improvement to the above scheme, the parameter search model includes a first parameter search model and a second parameter search model; The step of inputting the original sample image into a preset parameter search model and obtaining the image transformation weights currently generated by the parameter search model includes: In the coarse-grained search stage, global features of the original sample image are extracted and input into the first parameter search model to generate initial transformation weights. Sampling and attack loss evaluation are performed within the parameter range of the initial transformation weights to filter and obtain image transformation weights. In the fine-grained adjustment stage, local features of the original sample image are extracted, and the local features and the image transformation weights generated in the coarse-grained search stage are input into the second parameter search model for fine-tuning to obtain the image transformation weights. The phase switching probability is calculated in real time based on the attack loss change rate and weight entropy, and the coarse-grained search phase and the fine-grained adjustment phase are dynamically switched according to the magnitude of the phase switching probability.
[0008] As an improvement to the above scheme, the first parameter search model is a global search multilayer perceptron (MLP) network. The process of extracting global features from the original sample image and inputting them into the first parameter search model to generate initial transformation weights, and then sampling and evaluating the attack loss within the parameter range of the initial transformation weights to obtain image transformation weights, includes: Extract the mean feature of the HSV space from the original sample image; The mean feature is input into the global search MLP network to generate initial transformation weights; Within the parameter range where the initial transformation weights are located, a number of candidate transformation weights are generated using Halton sequence sampling; Based on each candidate transform weight, a corresponding candidate transform image is generated and input into a preset alternative model for attack. Calculate the attack success rate of each candidate transformed image, and select the candidate transformed weight corresponding to the highest attack success rate as the transformed weight of the currently generated image.
[0009] As an improvement to the above scheme, the second parameter search model is a two-layer MLP network; The step of extracting local features from the original sample image and inputting these local features and the image transformation weights generated in the coarse-grained search stage into the second parameter search model for fine-tuning to obtain the image transformation weights includes: Extract local texture features from the original sample image; The local texture features and the image transformation weights obtained in the coarse-grained adjustment stage are concatenated; The concatenated features are input into the two-layer MLP network, fine-tuned using a gradient-based optimization algorithm, and the transformed weights of the currently generated image are output.
[0010] As an improvement to the above scheme, the image transformation weight is a weight combination consisting of brightness transformation weight, contrast transformation weight, and saturation transformation weight; The step of performing a multi-dimensional visual transformation on the original sample image according to the image transformation weights to obtain a transformed sample image includes: Based on the saturation transformation weights, the saturation of the original sample image is adjusted in the color space; Based on the contrast transformation weights, pixel-level adaptive perturbation is used to adjust the contrast of the original sample image. Based on the brightness transformation weights, a differentiable nonlinear transformation function is used to adjust the brightness of the original sample image to obtain a transformed sample image.
[0011] As an improvement to the above scheme, the step of updating the network parameters of the parameter search model based on the attack effect of the adversarial example includes: A multi-objective joint optimization strategy is adopted to calculate a joint loss function; wherein, the joint loss function is used to simultaneously optimize the attack success rate and the concealment of the adversarial sample; Based on the joint loss function, the network parameters of the parameter search model are updated using feedback.
[0012] As an improvement to the above scheme, the joint loss includes a first loss term and a second loss term; the first loss term is used to maximize the distance between the adversarial sample and the true label of the original sample image, and the second loss term is used to constrain the visual consistency between the adversarial sample and the original sample image.
[0013] As an improvement to the above scheme, after obtaining the final target adversarial sample, the method further includes: The target adversarial sample is input into the target model to perform a cross-model transfer attack; The success of the migration attack is verified based on the output of the target model to the target adversarial sample.
[0014] As an improvement to the above scheme, the target model includes one of a convolutional neural network model, a visual encoder model, or a large visual language model; When the target model is a large visual-language model, the step of inputting the target adversarial example into the target model to perform a cross-model transfer attack includes: The target adversarial sample is combined with a preset text prompt to form a multimodal input pair; The multimodal input pairs are input into the visual language large model for processing; Based on the output of the target model to the target adversarial sample, verify whether the migration attack was successful, including: Obtain the response result output by the visual large language model; Determine whether the response result contains the true label corresponding to the original sample image; If the response does not contain the real label, the migration attack is considered successful.
[0015] This invention also provides a large model migration attack device, comprising: The image transformation weight acquisition module is used to input the original sample image into a preset parameter search model and obtain the image transformation weights currently generated by the parameter search model; The original sample image transformation module is used to perform multi-dimensional visual transformation on the original sample image according to the image transformation weights to obtain a transformed sample image; The adversarial example generation module is used to input the transformed sample image into a preset alternative model for iterative attack in order to generate adversarial examples; The network parameter update module is used to update the network parameters of the parameter search model based on the attack effect of the adversarial sample. The target adversarial sample output module is used to re-execute the above steps until the preset iteration termination condition is met, so as to obtain the final target adversarial sample.
[0016] This invention also provides a large model migration attack device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the large model migration attack method as described in any of the above embodiments.
[0017] This invention also provides a computer-readable storage medium comprising a stored computer program, wherein, when the computer program is executed, it controls the device where the computer-readable storage medium is located to execute the large model migration attack method as described in any of the preceding embodiments.
[0018] This invention also provides a computer program product, which includes a computer program or computer instructions. When the computer program or computer instructions are executed by a processor, they implement the large model migration attack method as described above.
[0019] Compared with existing technologies, the large model transfer attack method, apparatus, device, medium, and program products disclosed in this invention integrate multi-dimensional visual joint transformation to achieve simultaneous attacks on multiple perceptual dimensions of images. Compared with the single transformation or simple superposition strategies in existing technologies, the multi-dimensional joint transformation of this invention can more comprehensively cover the vulnerable areas of the target model's decision boundary, significantly improving the success rate of cross-model transfer of adversarial examples. Furthermore, this invention constructs a parameter search model to adaptively generate image transformation weights, realizing adaptive dynamic adjustment of attack parameters, and designs a learnable search mechanism that can continuously optimize transformation weights based on attack feedback during iteration. Compared with the parameter configuration methods in existing technologies that rely on manual setting or static optimization, this invention effectively avoids the defect of traditional gradient-based attack methods that are prone to getting trapped in local minima, enhancing the robustness and versatility of the attack. Attached Figure Description
[0020] Figure 1 This is a flowchart illustrating a large model migration attack method provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the process for generating adversarial examples in an embodiment of the present invention; Figure 3 This is a schematic diagram of a large model migration attack device provided in an embodiment of the present invention. Detailed Implementation
[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0022] In the description of this application, it should be understood that the terms "center", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this application.
[0023] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, unless otherwise stated, "a plurality of" means two or more.
[0024] In the description of this application, it should be noted that, unless otherwise expressly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection between two components. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances.
[0025] See Figure 1 This is a flowchart illustrating a large model migration attack method provided by an embodiment of the present invention. The embodiment of the present invention provides a large model migration attack method, the method comprising steps S11 to S15: S11. Input the original sample image into the preset parameter search model and obtain the image transformation weights currently generated by the parameter search model; S12. Based on the image transformation weights, perform multi-dimensional visual transformation on the original sample image to obtain a transformed sample image; S13. Input the transformed sample image into a preset alternative model for iterative attack to generate adversarial examples; S14. Based on the attack effect of the adversarial sample, update the network parameters of the parameter search model accordingly. S15. Repeat the above steps until the preset iteration termination condition is met to obtain the final target adversarial sample.
[0026] This invention automatically searches for the optimal image transformation weights, performs multi-dimensional visual transformations on sample images based on these weights, and iteratively generates adversarial examples with high transferability by combining alternative models.
[0027] See Figure 2 This is a flowchart illustrating the process of generating adversarial examples in an embodiment of the present invention. In the process of generating adversarial examples, a parameter search model is pre-built and trained to dynamically generate image transformation weights for the current input image. The parameter search model can be a neural network built based on a multilayer perceptron (MLP).
[0028] A clean, original sample image x with dimensions H×W×C (height, width, number of channels) is acquired. The original image sample x is input into the parameter search model. The parameter search model outputs the image transformation weights at the current iteration step based on the global or local features of the image, which are used to control the intensity of subsequent visual transformations.
[0029] Preferably, the multi-dimensional visual transformation includes adjustments to three dimensions: brightness, contrast, and saturation. The image transformation weights are determined by the brightness transformation weights. Contrast transformation weight Saturation transformation weights The weighted combination.
[0030] The image transformation weights output by the parameter search model These are used to control the intensity of brightness transformation, contrast transformation, and saturation transformation, respectively, to perform a joint transformation on the original sample image x. Specifically, brightness adjustment, contrast transformation, and saturation adjustment are represented as Brightness(…). ), Contrast Saturation The transformed sample image is denoted as . Its expression is:
[0031] Transformed sample image As the initial point, input to the preset alternative model In neural networks like ResNet-50 convolutional neural networks, a gradient-based iterative attack method is used to generate adversarial examples. In each iteration, the image pixel values are updated within a perturbation budget constraint ε based on the gradient direction of the loss function relative to the input image. After M iterations, the final adversarial example is obtained. .
[0032] The generated adversarial examples Input to alternative models By calculating the original image samples and adversarial examples The loss between the two is used to evaluate the attack effect, and the network parameters of the parameter search model are updated through the backpropagation algorithm so that it can generate image transformation weights that are more conducive to the attack in subsequent iterations.
[0033] Repeat the steps until a preset iteration termination condition is met, such as reaching a preset total number of iterations T, or the attack success rate no longer significantly improving. The final output is the adversarial example. This is the target adversarial sample with high migration attack capability generated in this embodiment. Subsequently, the generated adversarial sample can be input into target models with different structures to achieve migration attacks.
[0034] By employing the technical means of this invention, a multi-dimensional visual joint transformation is integrated to achieve simultaneous attacks on multiple perceptual dimensions of an image. Compared to existing strategies that rely on single transformations or simple superposition, this invention's multi-dimensional joint transformation can more comprehensively cover vulnerable areas of the target model's decision boundary, significantly improving the cross-model transfer success rate of adversarial examples. Furthermore, this invention constructs a parameter search model to adaptively generate image transformation weights, achieving adaptive dynamic adjustment of attack parameters. A learnable search mechanism is designed to continuously optimize the transformation weights based on attack feedback during iteration. Compared to existing parameter configuration methods that rely on manual setting or static optimization, this invention effectively avoids the vulnerability of traditional gradient-based attack methods to local minima, enhancing the robustness and versatility of the attack.
[0035] As a preferred embodiment, this invention further refines the structure and dynamic switching mechanism of the parameter search model based on the above embodiments. The parameter search model includes a first parameter search model and a second parameter search model, corresponding to the coarse-grained search stage and the fine-grained adjustment stage, respectively.
[0036] Then step S11, which is to input the original sample image into the preset parameter search model and obtain the image transformation weights currently generated by the parameter search model, includes steps S111 to S113: S111. In the coarse-grained search stage, global features of the original sample image are extracted and input into the first parameter search model to generate initial transformation weights. Sampling and attack loss evaluation are performed within the parameter range of the initial transformation weights to filter and obtain image transformation weights. S112. In the fine-grained adjustment stage, local features of the original sample image are extracted, and the local features and the image transformation weights generated in the coarse-grained search stage are input into the second parameter search model for fine-tuning to obtain the image transformation weights. S113. Calculate the stage switching probability in real time, and dynamically switch the coarse-grained search stage and the fine-grained adjustment stage according to the magnitude of the stage switching probability.
[0037] In this embodiment of the invention, the process first enters a coarse-grained search stage, extracting global features from the original sample image x, such as the mean of the three HSV channels. The global feature vector is input into the first parameter search model (such as a global search MLP) to generate initial transformation weights. Subsequently, sampling is performed within a preset parameter range around the initial weights to generate multiple candidate weight combinations. Each candidate weight combination is then subjected to a rapid attack evaluation, and the weight combination with the highest attack success rate is selected as the output of this stage.
[0038] When the preset stage switching conditions are met, the process switches to the fine-grained adjustment stage, extracting local texture features from the original sample image, such as the Sobel gradient magnitude G(x). The weights obtained from the coarse-grained stage are then... After concatenating with this local feature, the weights are input into a second parameter search model (such as a two-layer MLP) for fine-tuning, and the optimized weights are output. The fine-tuning process uses a gradient-based optimizer (such as Adam) to adjust the weights in a small range (±0.1) to approximate a local optimum.
[0039] In a preferred embodiment, the real-time calculation of the stage switching probability includes: The phase switching probability is calculated in real time based on the attack loss change rate and weight entropy.
[0040] In this embodiment of the invention, a gating network is used for dynamic switching. The gating network G( Implement adaptive switching of stages:
[0041] In each iteration, the rate of change of attack loss is monitored in real time. Weight entropy of the current weight combination Input the above information into the gating network. After normalization using the Sigmoid function, the stage switching probability p is obtained.
[0042] When the phase switching probability p is greater than a preset probability threshold, for example, p > 0.5, a fine-grained adjustment phase is triggered; otherwise, the coarse-grained search phase continues. The gating network is trained using a reinforcement learning algorithm (REINFORCE), with the historical average loss as a baseline, to ensure that the switching decision improves the overall attack performance. The reward function R is expressed as follows:
[0043] in, Attack losses at the end of the current search phase The value, Attack losses at the start of the current search phase The value, This represents the number of iterations performed in the current search phase. The baseline term is the average of the attack losses over the past several iterations.
[0044] By employing the technical means of this invention, a two-stage collaborative optimization approach combining coarse-grained and fine-grained optimization is adopted, balancing broad coverage of the search space with local fine-tuning, effectively improving the efficiency and accuracy of weighted search. The dynamic switching of the gating mechanism enables the algorithm to adaptively adjust the search strategy according to the current attack state, avoiding resource waste or insufficient search caused by fixed-stage division, further improving the quality of adversarial sample generation and the success rate of migration attacks.
[0045] As a preferred embodiment, based on the above embodiments, the specific implementation of the coarse-grained search stage is further defined in detail, and the first parameter search model is a global search MLP network.
[0046] The process of extracting global features from the original sample image and inputting them into the first parameter search model to generate initial transformation weights, and then sampling and evaluating the attack loss within the parameter range of the initial transformation weights to obtain image transformation weights, includes: Extract the mean feature of the HSV space from the original sample image; The mean feature is input into the global search MLP network to generate initial transformation weights; Within the parameter range where the initial transformation weights are located, a number of candidate transformation weights are generated using Halton sequence sampling; Based on each candidate transform weight, a corresponding candidate transform image is generated and input into a preset alternative model for attack. Calculate the attack success rate of each candidate transformed image, and select the candidate transformed weight corresponding to the highest attack success rate as the transformed weight of the currently generated image.
[0047] In this embodiment of the invention, a global search MLP network is constructed, containing a hidden layer with 128 neurons. The original sample image x is converted from the RGB color space to the HSV color space, and the global mean values of the hue (H), saturation (S), and brightness (V) channels are calculated to obtain the feature vector. .
[0048] Will The input is fed into a global search MLP network, and the network outputs the initial transformation weights. The value ranges for each weight are as follows: Brightness transformation weight ∈[0.5,1.5], contrast transformation weights ∈[0.7,1.3], saturation transformation weights ∈[0.6,1.4]. That is:
[0049] Initial network weights Within the surrounding parameter range, deterministic sampling is performed using Halton sequences from low-dispersion sequences to generate N (e.g., N=50) candidate weight combinations. Halton sequences can guarantee a uniform distribution of sampling points in the parameter space, thus providing a more comprehensive coverage of the potential optimal region.
[0050] For each candidate weight combination, the original sample image x is transformed in terms of brightness, contrast, and saturation to obtain the corresponding candidate transformed image. Input all candidate transformed images into the substitution model respectively. A fast attack is performed using a uniform number of iterations (e.g., 5 times) to obtain adversarial examples for each candidate.
[0051] The attack success rate (i.e., the proportion of predictions as incorrect categories) of each candidate adversarial example on the alternative model is calculated. The candidate weight combination with the highest attack success rate is selected as the output of the coarse-grained search stage, denoted as . .
[0052] By employing the technical means of this invention, the overall color distribution characteristics of an image can be quickly captured by using the HSV mean feature as a global representation; Halton sequence sampling has lower bias and better spatial coverage than random sampling, which helps to efficiently discover high-quality candidate weights within the parameter range; the screening mechanism based on attack success rate is directly guided by the attack effect, ensuring the effectiveness of the output weights in the coarse-grained stage and providing a good starting point for subsequent fine-grained fine-tuning.
[0053] As a preferred embodiment, based on the above embodiments, the specific implementation of the fine-grained adjustment stage is defined in detail, and the second parameter search model is a two-layer MLP network.
[0054] The step of extracting local features from the original sample image and inputting these local features and the image transformation weights generated in the coarse-grained search stage into the second parameter search model for fine-tuning to obtain the image transformation weights includes: Extract local texture features from the original sample image; The local texture features and the image transformation weights obtained in the coarse-grained adjustment stage are concatenated; The concatenated features are input into the two-layer MLP network, fine-tuned using a gradient-based optimization algorithm, and the transformed weights of the currently generated image are output.
[0055] In this embodiment of the invention, a two-layer MLP network is constructed, with 64 neurons in the first layer and 64 neurons in the second layer. After grayscale processing of the original sample image x, the Sobel operator is used to calculate the gradient magnitudes in the horizontal and vertical directions to obtain a gradient magnitude map. The global average value of this gradient magnitude map is further calculated as a local texture feature G(x) to characterize the spatial detail richness of the image.
[0056] Weights obtained from the coarse-grained stage The vector is concatenated with the local texture feature G(x) to form a joint feature vector. This joint feature vector is then input into a two-layer MLP network, which outputs fine-tuned network transformation weights. .
[0057]
[0058] The momentum Adam optimizer is used during fine-tuning, with a learning rate of [missing information]. First-order moment estimation of attenuation coefficient With joint loss function To optimize the objective, gradient updates are performed within a coarse-grained weight range of ±0.1. The joint loss function includes attack loss and weight difference constraint terms, and its expression is:
[0059] By employing the technical means of this invention, local texture features are introduced at the fine-grained stage, enabling weight fine-tuning to perceive the spatial structure information of the image and avoiding the neglect of local details due to relying solely on global features. The dual-layer MLP network has stronger nonlinear fitting capabilities and can capture the complex interaction relationships between weights. The gradient-based Adam optimizer realizes fine-tuning of weights, while preventing performance degradation caused by excessive fine-tuning through weight difference constraints, thereby approximating the optimal weight configuration in local regions and effectively alleviating the local optimum problem.
[0060] As a preferred embodiment, the present invention further implements the above embodiments, providing detailed definitions for the specific implementation of multi-dimensional visual transformation. The step of performing multi-dimensional visual transformation on the original sample image according to the image transformation weights to obtain a transformed sample image includes: Based on the saturation transformation weights, the saturation of the original sample image is adjusted in the color space; Based on the contrast transformation weights, pixel-level adaptive perturbation is used to adjust the contrast of the original sample image. Based on the brightness transformation weights, a differentiable nonlinear transformation function is used to adjust the brightness of the original sample image to obtain a transformed sample image.
[0061] In this embodiment of the invention, a collaborative attack strategy integrating brightness, contrast, and saturation is proposed. This strategy achieves efficient generation of adversarial examples by establishing a joint optimization framework. The multi-dimensional visual transformation process of the sample image is as follows: first, saturation is adjusted in the HSV color space; then, contrast is adaptively perturbed; and finally, brightness is adjusted using a nonlinear function.
[0062] Appropriate brightness perturbations can lead to incorrect predictions from the model. The brightness transformation employs a differentiable piecewise nonlinear function:
[0063] in, It is the brightness scaling factor, determined by the brightness transformation weight. control, is the brightness shift, k is the non-linear curvature factor, and C is the output compensation term, which ensures that the transformed pixel values are distributed within a reasonable range.
[0064] In adversarial attacks, altering contrast can significantly impact the ability of deep learning models to recognize image features. The contrast transformation employs pixel-level adaptive Gamma perturbation:
[0065]
[0066] in, A spatial weight map is generated using a lightweight CNN, and the weights are transformed by contrast. control, The region sensitivity coefficient is based on image block segmentation. =0.1, StdDev is the local standard deviation calculation. The standard deviation of the 8×8 local region centered on each pixel is used to enhance the contrast variation in the edge region.
[0067] Because deep learning models are sensitive to color purity, attackers can generate adversarial examples by increasing or decreasing saturation. Saturation transformations are performed using HSV space adjustment.
[0068] in Indicates hue, Indicates brightness, The saturation adjustment coefficient is determined by the saturation transformation weight. control.
[0069] Final transformed sample image The expression is obtained by combining the above three transformations in sequence:
[0070] By employing the technical means of this invention, this embodiment achieves joint perturbation of the multi-dimensional perceptual channels of an image through the coordinated transformation of three core visual attributes: brightness, contrast, and saturation. By using refined transformation techniques such as nonlinear Sigmoid brightness transformation, pixel-level adaptive Gamma contrast correction, and HSV spatial saturation adjustment, the visual naturalness and concealment of adversarial examples are effectively maintained while ensuring attack effectiveness, thereby enhancing the diversity and cross-model transferability of adversarial examples.
[0071] As a preferred embodiment, the present invention further implements the above embodiments, wherein the step of updating the network parameters of the parameter search model based on the attack effect of the adversarial sample includes: A multi-objective joint optimization strategy is adopted to calculate a joint loss function; wherein, the joint loss function is used to simultaneously optimize the attack success rate and the concealment of the adversarial sample; Based on the joint loss function, the network parameters of the parameter search model are updated using feedback.
[0072] In this embodiment of the invention, a joint loss function is constructed. This function takes into account both attack effectiveness and visual concealment. The specific expression is:
[0073] in, The attack loss is used to ensure that adversarial examples can deceive alternative models; The fidelity loss is used to constrain the visual similarity between adversarial examples and the original examples.
[0074] Joint loss function As the optimization objective, the gradients of all trainable parameters in the parameter search model are calculated using the backpropagation algorithm. The Adam optimizer is then used to update the network parameters, ensuring that the transformed weights generated in subsequent iterations reduce the perceptibility of adversarial examples while maintaining the attack effectiveness.
[0075] The technical means employed in this invention utilize a multi-objective joint optimization strategy, avoiding the problem of traditional methods focusing solely on attack success rate while neglecting sample concealment. Through the synergistic constraint of attack loss and fidelity loss in the joint loss function, the parameter search model can automatically balance attack effectiveness and visual quality during iterative updates. The generated adversarial examples possess both high transfer attack capability and maintain a high degree of visual consistency with the original examples, enhancing the practicality and concealment of the attack.
[0076] In a preferred embodiment, the joint loss includes a first loss term and a second loss term; the first loss term is used to maximize the distance between the adversarial sample and the true label of the original sample image, and the second loss term is used to constrain the visual consistency between the adversarial sample and the original sample image.
[0077] First loss item The goal of using cross-entropy loss or margin loss is to maximize the alternative model. Adversarial examples The difference between the predicted result and the true label y. The specific expression is:
[0078] Among them, losses Maximize the distance between the adversarial sample and the clean label to ensure that the generated adversarial sample can successfully attack. For the model to include adversarial examples The probability of predicting the true class y. By maximizing the negative value of this probability, the model is forced to make incorrect predictions.
[0079] Second loss item This is used to ensure visual consistency between adversarial examples and clean examples. L2 norm distance metric is used to constrain adversarial examples. The pixel-level difference between the original sample x and the actual sample x. Furthermore, to further enhance visual naturalness, perceptual similarity can be calculated in the perceptual space. The specific expression is:
[0080] Among them, coefficient Used to balance attack effectiveness and stealth. A constraint threshold ε is set to ensure... A value of 0.1 makes adversarial perturbations visually imperceptible to humans. Simultaneously, the x and y values of the images before and after the transformation should be guaranteed. Consistency at both the visual and semantic levels.
[0081] In addition, a transformation consistency constraint is introduced, that is, the transformation sample image... Semantic consistency should also be maintained with the original sample image x to enhance the naturalness of the generated sample.
[0082] By employing the technical means of this invention, the attack process is directly driven by a first loss term, ensuring that the adversarial sample can effectively deceive the target model; the visual perturbation amplitude is strictly constrained by a second loss term, ensuring the naturalness of the adversarial sample. The synergistic effect of the two losses enables the parameter search model to balance attack effectiveness and stealth during the optimization process, preventing the adversarial sample from being easily identified by the defense system due to excessive perturbation, thus improving the practicality and stealth of the attack.
[0083] The overall algorithm flow is as follows:
[0084] As a preferred embodiment, the present invention is further implemented based on any of the above embodiments. After obtaining the final target adversarial sample, the method further includes steps S16 and S17: S16. Input the target adversarial sample into the target model to perform a cross-model transfer attack; S17. Verify whether the migration attack is successful based on the output of the target model to the target adversarial sample.
[0085] In this embodiment of the invention, the final generated adversarial sample The inputs are fed into multiple different target models, and the structures and alternative models of these target models are described. Different. Based on the target model, the adversarial examples... The output results are used to verify whether the migration attack was successful.
[0086] It should be noted that existing technologies lack effective migration attack schemes for novel multimodal architectures such as large visual language models, thus limiting the applicability of the technology.
[0087] In a preferred embodiment, the target model includes one of a convolutional neural network model, a visual encoder model, or a large visual language model.
[0088] The target model may include, but is not limited to: different types of convolutional neural networks (such as VGG, DenseNet), visual encoder models (such as ViT, Swin Transformer), and large visual language models (such as CLIP, Qwen-VL, DeepSeek-VL).
[0089] In a preferred embodiment, when the target model is a large visual-language model, the step of inputting the target adversarial sample into the target model to perform a cross-model transfer attack includes: The target adversarial sample is combined with a preset text prompt to form a multimodal input pair; The multimodal input pairs are input into the visual language large model for processing; Based on the output of the target model to the target adversarial sample, verify whether the migration attack was successful, including: Obtain the response result output by the visual large language model; Determine whether the response result contains the true label corresponding to the original sample image; If the response does not contain the real label, the migration attack is considered successful.
[0090] Specifically, for large visual language models, specific text prompts are also needed, such as combining adversarial examples with the text "What's in the picture?" to form an image-text pair input model.
[0091] For unimodal classification models such as CNN or ViT, if the predicted class output by the model is inconsistent with the true label y of the original sample, the attack is considered successful.
[0092] For large visual language models, the target VLLM model is represented as: The generated adversarial examples The image and prompt form an image-text pair. Enter The response r obtained is represented as:
[0093] The generated response text *r* is retrieved. If the response does not contain the label *y* corresponding to the original sample, or contains an incorrect description, the attack is considered successful. The attack success rate on each target model is calculated as the final metric for evaluating the performance of adversarial example transfer.
[0094] Using the technical means of this invention, this embodiment verifies the cross-model transferability of adversarial samples by applying them to target models with various architectures. In particular, the verification of attacks on large visual-language models fills the technical gap in the current field of transfer attacks regarding the security assessment of large models, providing strong experimental support for the practical application of this invention and further demonstrating the advantages of this invention's method in terms of generalization and applicability compared to existing technologies.
[0095] This invention utilizes an end-to-end learnable optimization framework to automate the search and dynamic adjustment of attack parameters, adapting to different target models and attack scenarios without manual intervention. Compared to existing technologies, this invention significantly lowers the implementation threshold for migration attacks while possessing excellent model generalization capabilities, demonstrating broad application prospects and economic benefits.
[0096] See Figure 3 This is a schematic diagram of a large model migration attack device provided in an embodiment of the present invention. The embodiment of the present invention provides a large model migration attack device 10, comprising: The image transformation weight acquisition module 11 is used to input the original sample image into a preset parameter search model and obtain the image transformation weights currently generated by the parameter search model. The original sample image transformation module 12 is used to perform multi-dimensional visual transformation on the original sample image according to the image transformation weight to obtain a transformed sample image; Adversarial example generation module 13 is used to input the transformed sample image into a preset alternative model for iterative attack in order to generate adversarial examples; The network parameter update module 14 is used to update the network parameters of the parameter search model based on the attack effect of the adversarial sample. The target adversarial sample output module 15 is used to re-execute the above steps until the preset iteration termination condition is met, so as to obtain the final target adversarial sample.
[0097] In a preferred embodiment, the large model migration attack device 10 further includes: The migration attack module is used to input the target adversarial sample into the target model to perform a cross-model migration attack; and to verify whether the migration attack is successful based on the output result of the target model on the target adversarial sample.
[0098] It should be noted that the large model migration attack device provided in this embodiment of the invention is used to execute all the process steps of the large model migration attack method in the above embodiment. The working principle and beneficial effect of the two are one-to-one, so they will not be described again.
[0099] This invention also provides a large model migration attack device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the large model migration attack method as described in any of the above embodiments.
[0100] This invention also provides a computer-readable storage medium comprising a stored computer program, wherein, when the computer program is executed, it controls the device where the computer-readable storage medium is located to perform the large model migration attack method as described in any of the above embodiments.
[0101] This invention also provides a computer program product, which includes a computer program or computer instructions. When the computer program or computer instructions are executed by a processor, they implement the large model migration attack method as described in any of the above embodiments.
[0102] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.
[0103] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications are also considered to be within the scope of protection of the present invention.
Claims
1. A method for large model migration attack, characterized in that, include: The original sample image is input into a preset parameter search model to obtain the image transformation weights currently generated by the parameter search model; Based on the image transformation weights, the original sample image is subjected to multi-dimensional visual transformation to obtain the transformed sample image; The transformed sample image is input into a preset alternative model for iterative attack to generate adversarial examples; Based on the attack effect of the adversarial sample, the network parameters of the parameter search model are updated accordingly. Repeat the above steps until the preset iteration termination condition is met to obtain the final target adversarial sample.
2. The large model migration attack method as described in claim 1, characterized in that, The parameter search model includes a first parameter search model and a second parameter search model; The step of inputting the original sample image into a preset parameter search model and obtaining the image transformation weights currently generated by the parameter search model includes: In the coarse-grained search stage, global features of the original sample image are extracted and input into the first parameter search model to generate initial transformation weights. Sampling and attack loss evaluation are performed within the parameter range of the initial transformation weights to filter and obtain image transformation weights. In the fine-grained adjustment stage, local features of the original sample image are extracted, and the local features and the image transformation weights generated in the coarse-grained search stage are input into the second parameter search model for fine-tuning to obtain the image transformation weights. The phase switching probability is calculated in real time based on the attack loss change rate and weight entropy, and the coarse-grained search phase and the fine-grained adjustment phase are dynamically switched according to the magnitude of the phase switching probability.
3. The large model migration attack method as described in claim 2, characterized in that, The first parameter search model is a global search multilayer perceptron (MLP) network; The process of extracting global features from the original sample image and inputting them into the first parameter search model to generate initial transformation weights, and then sampling and evaluating the attack loss within the parameter range of the initial transformation weights to obtain image transformation weights, includes: Extract the mean feature of the HSV space from the original sample image; The mean feature is input into the global search MLP network to generate initial transformation weights; Within the parameter range where the initial transformation weights are located, a number of candidate transformation weights are generated using Halton sequence sampling; Based on each candidate transform weight, a corresponding candidate transform image is generated and input into a preset alternative model for attack. Calculate the attack success rate of each candidate transformed image, and select the candidate transformed weight corresponding to the highest attack success rate as the transformed weight of the currently generated image.
4. The large model migration attack method as described in claim 2, characterized in that, The second parameter search model is a two-layer MLP network; The step of extracting local features from the original sample image and inputting these local features and the image transformation weights generated in the coarse-grained search stage into the second parameter search model for fine-tuning to obtain the image transformation weights includes: Extract local texture features from the original sample image; The local texture features and the image transformation weights obtained in the coarse-grained adjustment stage are concatenated; The concatenated features are input into the two-layer MLP network, fine-tuned using a gradient-based optimization algorithm, and the transformed weights of the currently generated image are output.
5. The large model migration attack method as described in claim 1, characterized in that, The image transformation weight is a weight combination consisting of brightness transformation weight, contrast transformation weight, and saturation transformation weight; The step of performing a multi-dimensional visual transformation on the original sample image according to the image transformation weights to obtain a transformed sample image includes: Based on the saturation transformation weights, the saturation of the original sample image is adjusted in the color space; Based on the contrast transformation weights, pixel-level adaptive perturbation is used to adjust the contrast of the original sample image. Based on the brightness transformation weights, a differentiable nonlinear transformation function is used to adjust the brightness of the original sample image to obtain a transformed sample image.
6. The large model migration attack method as described in claim 1, characterized in that, The step of updating the network parameters of the parameter search model based on the attack effect of the adversarial example includes: A multi-objective joint optimization strategy is adopted to calculate a joint loss function; wherein, the joint loss function is used to simultaneously optimize the attack success rate and the concealment of the adversarial sample; Based on the joint loss function, the network parameters of the parameter search model are updated using feedback.
7. The large model migration attack method as described in claim 6, characterized in that, The joint loss includes a first loss term and a second loss term; the first loss term is used to maximize the distance between the adversarial sample and the true label of the original sample image, and the second loss term is used to constrain the visual consistency between the adversarial sample and the original sample image.
8. The large model migration attack method as described in claim 1, characterized in that, After obtaining the final target adversarial sample, the method further includes: The target adversarial sample is input into the target model to perform a cross-model transfer attack; The success of the migration attack is verified based on the output of the target model to the target adversarial sample.
9. The large model migration attack method as described in claim 8, characterized in that, The target model includes one of the following: a convolutional neural network model, a visual encoder model, or a large visual language model; When the target model is a large visual-language model, the step of inputting the target adversarial example into the target model to perform a cross-model transfer attack includes: The target adversarial sample is combined with a preset text prompt to form a multimodal input pair; The multimodal input pairs are input into the visual language large model for processing; Based on the output of the target model to the target adversarial sample, verify whether the migration attack was successful, including: Obtain the response result output by the visual large language model; Determine whether the response result contains the true label corresponding to the original sample image; If the response does not contain the real label, the migration attack is considered successful.
10. A large-scale model migration attack device, characterized in that, include: The image transformation weight acquisition module is used to input the original sample image into a preset parameter search model and obtain the image transformation weights currently generated by the parameter search model; The original sample image transformation module is used to perform multi-dimensional visual transformation on the original sample image according to the image transformation weights to obtain a transformed sample image; The adversarial example generation module is used to input the transformed sample image into a preset alternative model for iterative attack in order to generate adversarial examples; The network parameter update module is used to update the network parameters of the parameter search model based on the attack effect of the adversarial sample. The target adversarial sample output module is used to re-execute the above steps until the preset iteration termination condition is met, so as to obtain the final target adversarial sample.
11. A large-scale model migration attack device, characterized in that, It includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor, when executing the computer program, implements the large model migration attack method as described in any one of claims 1 to 9.
12. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored computer program, wherein, when the computer program is executed, it controls the device where the computer-readable storage medium is located to perform the large model migration attack method as described in any one of claims 1 to 9.
13. A computer program product, characterized in that, The computer program product includes a computer program or computer instructions, which, when executed by a processor, implement the large model migration attack method as described in any one of claims 1 to 9.