Black box adaptive safety assessment method and system for artificial intelligence model

By integrating multiple black-box adversarial evaluation and robustness enhancement algorithms, and combining the differences in multi-model architectures, parameters are dynamically adjusted to generate adversarial examples and perform adaptive evaluation. This solves the problems of low attack efficiency and weak robustness in black-box scenarios, and achieves efficient vulnerability mining and robustness enhancement of visual models.

CN121033484APending Publication Date: 2025-11-28SOUTH CHINA AGRICULTURAL UNIVERSITY +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510954815.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-07-10
Filing Date
2025-07-11
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

In black-box scenarios, existing adversarial evaluation algorithms are inefficient and fragile, unable to effectively deal with attacks on visual models with different internal structures, resulting in low attack efficiency and insufficient adaptability.

Method used

It integrates multiple black-box adversarial evaluation algorithms (such as NES, Sign-opt, BA, TA) and robustness enhancement algorithms (such as LGS, PRN, RS, FFDnet), and combines the differences of multiple model architectures (CNN, ViT) to dynamically adjust the evaluation strategy and robustness enhancement parameters. Through metaheuristic algorithms and transferability enhancement techniques, it generates adversarial examples and performs adaptive evaluation.

Benefits of technology

It significantly improves the efficiency of adversarial example generation and the robustness of visual models, realizes the systematic mining of vulnerabilities in visual models and enhances their robustness, and provides user-friendly API interfaces and operation interfaces, reducing the technical threshold for use.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121033484A_ABST
    Figure CN121033484A_ABST
Patent Text Reader

Abstract

The invention discloses an artificial intelligence model-oriented black box adaptive security assessment method and system, and the method comprises the steps: enabling the system to dynamically generate an adversarial sample through integrating a plurality of black box adversarial assessment algorithms, carrying out the robustness enhancement of the adversarial sample through a robustness enhancement algorithm, and obtaining an enhanced adversarial sample; and adaptively evaluating the vulnerability of the visual model by using the adversarial sample and the enhanced adversarial sample, and automatically storing all operation and detection records. The invention aims at integrating an existing excellent black box confrontation evaluation algorithm and a robustness enhancement algorithm in a black box scene, and solving the problems that the confrontation evaluation algorithm is low in attack efficiency and the robustness of the confrontation evaluation algorithm is fragile for visual models of different internal structures, so that a user can better explore vulnerabilities of current visual models, and the robustness of the visual models is enhanced. And a method for improving the robustness is searched.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of computer vision, and particularly relates to a black-box adaptive security evaluation method and system for an artificial intelligence model. BACKGROUND

[0002] At present, visual AI systems have been widely applied to key fields such as face recognition, image classification, object detection, etc. However, with the continuous emergence of adversarial evaluation algorithms, the security of visual AI systems is facing unprecedented challenges. Attackers can introduce tiny and difficult-to-detect perturbations in the input data of the model through carefully designed adversarial samples, thereby inducing the model to produce incorrect output results, and thus causing serious security risks. In the real world, since the internal parameters and architecture of most artificial intelligence models are strictly protected, the black-box adversarial evaluation scenario is more realistic. Adversarial evaluation algorithm refers to adding a segment of artificially designed noise to the image, i.e. making subtle pixel point changes to the image, and such changes cannot be detected by the naked eye, and finally can deceive the visual model and make it classify the changed image incorrectly. In this process, the image that makes the visual model classify incorrectly after adding noise is called an adversarial sample. The internal architecture of the model cannot be known in reality, and this scenario is called a black-box scenario.

[0003] Unfortunately, the black-box adversarial evaluation assumes that the attacker can only obtain the input and output information of the model, and cannot directly access the internal structure and parameters of the model. Therefore, in order to access the internal model and still maintain the correct classification ability when facing perturbations, a robustness enhancement method is generally used to improve the model's resistance to adversarial samples through different mechanisms, to ensure the security and reliability of the model in actual application. Robustness enhancement algorithm refers to a technology that identifies and eliminates malicious perturbations in adversarial samples, or enhances the robustness of the model itself, to ensure that the visual model can maintain correct classification ability when facing perturbations. The core goal is to resist the deception of the model by human-designed tiny pixel changes.

[0004] Although the prior art has made some progress in adversarial evaluation and robustness enhancement, there are still many deficiencies in the existing methods under the black-box scenario, such as the following problems: (1) low attack efficiency, the internal structure and calculation method of different models are different, and their response speed and degree to input perturbation are also different. If these differences are ignored, a lot of time may be wasted on meaningless perturbation attempts on some models, and effective adversarial samples that make the model misjudge cannot be found, thereby reducing the attack efficiency. (2) Poor adaptability to unknown attacks, LGS, PRN and other algorithms rely on the characteristics of specific adversarial samples. When the characteristics of unknown attacks are inconsistent with the patterns in the training data, they may not be able to effectively identify and resist. The robustness enhancement effect of LGS is limited for global attacks. The robustness enhancement effect of PRN depends on the diversity and representativeness of the adversarial samples used during training.

[0005] Therefore, how to solve the problems of low attack efficiency of adversarial evaluation algorithm and vulnerability of robustness of different internal structure visual models, so that users can better explore the vulnerabilities of current visual models, is a problem to be solved. SUMMARY

[0006] The main purpose of the present application is to overcome the shortcomings and deficiencies of the prior art, provide a black-box adaptive security evaluation method and system for artificial intelligence models, which aims to integrate existing excellent black-box adversarial evaluation algorithms and robustness enhancement algorithms under the black-box scenario, solve the problems of low attack efficiency of adversarial evaluation algorithm and vulnerability of robustness of different internal structure visual models, so that users can better explore the vulnerabilities of current visual models and find ways to improve their robustness.

[0007] In order to achieve the above purpose, the technical scheme adopted by the present application is as follows: In a first aspect, the present application provides a black-box adaptive security evaluation method for artificial intelligence models, comprising the following steps: The system selects a suitable black-box adversarial evaluation algorithm according to the characteristics of the target visual model, dynamically generates adversarial samples, and the characteristics of the target visual model include the target model architecture and the data set type; The system selects a robustness enhancement algorithm according to the characteristics of the adversarial samples, obtains the enhanced adversarial samples, and the characteristics of the adversarial samples include the adversarial noise type and the frequency distribution; Adaptive evaluation of the vulnerability of the visual model is performed using the adversarial samples and the enhanced adversarial samples, and all operations and detection records are automatically saved.

[0008] As a preferred technical scheme, the system selects a suitable black-box adversarial evaluation algorithm according to the characteristics of the target visual model, dynamically generates adversarial samples, which comprises: For the CNN-based target visual model, the Sign-opt algorithm based on gradient estimation or the BA algorithm based on decision boundary is selected; For the ViT-based target visual model, the TA algorithm based on frequency domain transformation is selected; The visual model includes a convolutional neural network and a ViT, and the two models are used in cooperation; The convolutional neural network includes at least one of the following models: Alexnet and EfficientNet; The ViT includes at least one of the following models: T2T-ViT and FastViT.

[0009] As a preferred technical solution, it also includes adjusting the algorithm parameters according to the structure parameters of the target visual model, the structure parameters including the number of convolutional layers and the type of activation function.

[0010] As a preferred technical solution, the system selects a robust enhancement algorithm according to the features of the adversarial samples to obtain the enhanced adversarial samples, including: If the adversarial sample is a high-frequency noise adversarial sample, a local gradient smoothing algorithm is used; If the adversarial sample is a global perturbation adversarial sample, a random smoothing algorithm is used; The system dynamically adjusts the parameters of the robustness enhancement algorithm according to the influence of the enhanced adversarial samples on the classification accuracy of the model, specifically: comparing the accuracy of the enhanced adversarial samples and the adversarial samples, if the accuracy decreases, automatically increasing the processing intensity of the robustness enhancement algorithm, if the accuracy recovers, reducing the enhancement intensity, the system uses a sliding window to record the enhancement effect, and combines a threshold judgment mechanism to dynamically adjust the parameters of the enhancement algorithm.

[0011] As a preferred technical solution, the use of adversarial samples and enhanced adversarial samples to adaptively evaluate the vulnerability of the visual model includes a vulnerability detection strategy based on the differences in the multi-model architecture, specifically: Quantize the sensitive features of the input sample, extract the structure parameters of the visual model through the API interface, the sensitive features including edge sensitivity and color sensitivity, the edge sensitivity is used to locate the high-frequency detail concentrated area of the input sample, and the color sensitivity is used to identify the hue or saturation vulnerability channel of the input sample, the structure parameters including the number of convolutional layers, the existence of attention modules, the proportion of ReLU activation function, and the pooling strategy; Map the structure parameters to the vulnerability mode, select the attack strategy according to the current vulnerability mode, and calculate the success rate of each attack in real time; save the operation and detection records, including the structure parameters of the visual model, the vulnerability mode and the attack strategy; The structure parameters are mapped to the vulnerability mode, an attack strategy is selected according to the current vulnerability mode, and specifically: When the convolution layer is detected and the convolution layer is higher than a set threshold, it is marked as high-frequency vulnerability, and the TA algorithm is selected for frequency domain attack; When the attention module is detected, it is marked as semantic boundary vulnerability, and the BA algorithm is used for boundary attack; When the proportion of ReLU activation function is higher than a set threshold, it is marked as gradient saturation vulnerability, and the NES algorithm is used for confidence attack; When the proportion of the pooling layer is higher than a set threshold, it is marked as local strong feature dependence, and the Sign-opt algorithm is used for local perturbation attack.

[0012] As a preferred technical solution, the vulnerability of the visual model is adaptively evaluated by using the adversarial sample and the enhanced adversarial sample, which includes a dynamic parameter adjustment method for different adversarial evaluation algorithms, and specifically: When the frequency domain attack rate exceeds the set threshold for two consecutive times, the low frequency disturbance ratio is reduced when the frequency domain attack is selected; When the boundary attack is selected, if the boundary approximation iteration is more than 800 steps and still unsuccessful, the neighbor constraint range is relaxed according to the exponential decay rule; When the confidence attack is selected, if the gradient symbol matching rate is less than the set threshold for three consecutive times, the number of randomly sampled directions is increased, and the gradient estimation accuracy is improved; When the perturbation attack is selected, when the expected loss variance is less than a set threshold, the Gaussian noise scale is increased to improve the sampling diversity.

[0013] As a preferred technical solution, the vulnerability of the visual model is adaptively evaluated by using the adversarial sample and the enhanced adversarial sample, which includes a real-time feedback mechanism for the success rate of adversarial evaluation and the robustness enhancement effect, and specifically: Attack side: for the same attack strategy, the success rate of the adversarial sample attacking the visual model is calculated in real time, when the success rate of a single attack exceeds a set value, the corresponding vulnerability of the visual model is marked, when the success rate of the attack is less than a set value for multiple times, the defense detection is triggered, and the attack task is continued in other cases; Defense side: when the defense detection is triggered, the enhanced adversarial sample is input to the visual model for defense, and the success rate of the attack side attack is calculated again, if the success rate of a single attack side attack exceeds a set value, it is determined to be invalid, and the vulnerable area is located; According to the comparison of the success rates of the attacks before and after triggering the defense detection, the attack strategy of the attack side and the defense strategy of the defense side are optimized.

[0014] As a preferred technical solution, the optimization of the attack strategy of the attack side comprises: when the success rate of the attack after triggering the defense detection once is less than a set value, and the corresponding enhanced adversarial sample meets the requirement of the defense side, the perturbation intensity of the attack side is increased; when the success rate of the attack after triggering the defense detection for 2 or 3 times continuously exceeds the set value, the perturbation intensity of the attack side is compressed.

[0015] As a preferred technical solution, the optimization of the defense strategy of the defense side comprises: The rate of increase of the success rate of the attack before and after triggering the defense detection is calculated, if the rate of increase is not greater than a set value, the number of types of input enhanced adversarial samples is increased, after increasing the number of types of input enhanced adversarial samples, if the rate of increase is less than the set value and is reduced to a number of times of the set value, corresponding robustness enhancement processing is performed on the corresponding increased types, and when the rate of increase is greater than the set value, the attack strategy of the attack side is switched.

[0016] In a second aspect, the present application further provides a black-box adaptive security evaluation system for an artificial intelligence model, which is applied to the black-box adaptive security evaluation method for the artificial intelligence model and comprises a first processing module, a second processing module and a task execution module. The first processing module is used for selecting a suitable black-box adversarial evaluation algorithm according to the characteristics of a target visual model, and dynamically generating adversarial samples, wherein the characteristics of the target visual model include a target model architecture and a data set type. The second processing module is used for selecting a robustness enhancement algorithm according to the characteristics of the adversarial samples, and obtaining enhanced adversarial samples, wherein the characteristics of the adversarial samples include an adversarial noise type and a frequency distribution. The task execution module is used for adaptively evaluating the vulnerable vulnerabilities of the visual model by using the adversarial samples and the enhanced adversarial samples, and automatically saving all operations and detection records.

[0017] Compared with the prior art, the present application has the following advantages and beneficial effects: (1) The present application integrates various black-box adversarial evaluation algorithms (such as NES, Sign-opt, BA and TA) and robustness enhancement algorithms (such as LGS, PRN, RS and FFDnet), combines the differences of multiple model architectures (CNN, ViT, etc.), dynamically adjusts the evaluation strategy and the robustness enhancement parameters, realizes the systematic mining and robustness enhancement of the vulnerable points of the visual model, simultaneously introduces meta-heuristic algorithms (such as natural evolution strategy NES) and migration enhancement technology, combines gradient sign estimation, frequency domain perturbation optimization and other means, and significantly improves the generation efficiency and visual concealment of the adversarial samples.

[0018] (2) The application integrates multi-level and multi-dimensional robustness enhancement algorithms (such as gradient suppression, disturbance correction, noise smoothing, and feature denoising), optimizes the combination scheme by evaluating the synergistic effect among the robustness enhancement algorithms, and forms a comprehensive robustness enhancement algorithm for diversified adversarial evaluation.

[0019] (3) The application automates the whole process of adversarial evaluation generation, robustness enhancement implementation, and effect evaluation, provides user-friendly API interfaces and operation interfaces, supports the storage and analysis of key data such as evaluation parameters, success rate, and robustness enhancement effect, and reduces the technical use threshold. BRIEF DESCRIPTION OF DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0021] Figure 1 The flowchart of the black-box adaptive security evaluation method for artificial intelligence model of the embodiment of the present application; Figure 2 The structural schematic diagram of the black-box adaptive security evaluation system for artificial intelligence model of the embodiment of the present application. DETAILED DESCRIPTION

[0022] In order to make the personnel in the technical field better understand the present application scheme, the technical solutions in the embodiments of the present application will be described clearly and completely in the following with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0023] In the present application, "embodiment" means that the specific features, structures or characteristics described in conjunction with the embodiment can be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily mean the same embodiment, nor is it an independent or alternative embodiment to other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described in the present application can be combined with other embodiments.

[0024] Please refer to Figure 1 The embodiment provides a black-box adaptive security evaluation method for artificial intelligence model, comprising the following steps: S1, the system generates adversarial samples by integrating multiple black-box adversarial evaluation algorithms.

[0025] In this embodiment, considering that different black-box adversarial evaluation algorithms focus on different points when attacked, they are divided into: sampling-based, gradient estimation-based, transferability-based and decision-based black-box adversarial attacks. Since the transferability-based needs a substitute model and has poor interpretability, no algorithm is selected from it. This embodiment integrates four algorithms: NES, Sign-opt, BA and TA, which have great differences in evaluation focus. Their main formulas are as follows: (1) NES: The following are its Gaussian search distribution (centered on the current parameter), expected loss and gradient estimation:

[0026]

[0027]

[0028] where is the perturbed parameter, is the current parameter value, is the step size, F is the loss function, and E is the expected loss.

[0029] It uses the estimated gradient to update the parameters, usually combined with PGD (Projected Gradient Descent) to ensure that the solution is within a reasonable range.

[0030] (2) Sign-opt: Optimization goal:

[0031] Gradient update:

[0032] where, in each iteration, the algorithm first randomly samples multiple directions then updates the search direction according to the gradient sign of these directions, is the objective function, f is the hard label black box function, is a very small smoothing parameter, and η is the step size.

[0033] (3) BA:

[0034] where represents the adversarial sample after t iterations, x is the original sample, and η is the step size parameter, represents the perturbed sample limited within the field of x, and to classify a sample different from x.

[0035] (4) TA:

[0036] wherein is an initial input sample, is an adversarial generated sample, is a perturbation, is a loss function, is a classification model.

[0037] By frequency domain transformation and optimization of the perturbation, the input sample is misled in the model classification, while the sample is kept similar to the original input, and the perturbation parameters are updated gradually.

[0038] For the uploaded original image, the above algorithm is used to introduce a small noise that is difficult for the human eye to detect, so that the model is fooled and the robustness of the model is destroyed, thereby generating the adversarial sample. In the black box scenario, the internal parameters of the visual model cannot be obtained, so this embodiment will dynamically attack the visual model from multiple strategies and aspects, so as to find the most vulnerable and easiest to fool vulnerability of the visual model, and adaptively evaluate the weakest angle of the model, and test the robustness of the model.

[0039] Specifically, when generating an adversarial sample, first consider what type of target visual model it is, such as a convolutional neural network and ViT. In this embodiment, the convolutional neural network and ViT are combined and applied to adapt to different attack scenarios, wherein the convolutional neural network includes at least one of the following models: Alexnet and EfficientNet, and the ViT includes at least one of the following models: T2T-ViT and FastViT. When generating an adversarial sample, for a target visual model based on CNN, the Sign-opt algorithm based on gradient estimation or the BA algorithm based on decision boundary is preferred, and for a target visual model based on ViT, the TA algorithm based on frequency domain transformation or the NES algorithm based on sampling is preferred.

[0040] In order to save computing expenses and improve attack efficiency, this embodiment will also consider adjusting the algorithm parameters according to the structure parameters of the target visual model, for example, adjusting the attack times or attack strength according to the number of convolutional layers and the type of activation function.

[0041] In addition, if the computing power of the computer or other conditions are considered, the embodiments can also consider the following black box adversarial evaluation algorithm: Adversarial sample generation based on genetic algorithm: considering adversarial sample generation as an optimization problem, using the population evolution mechanism of genetic algorithm, iteratively optimizing the disturbance parameters through selection, crossover and mutation operations, so that the disturbance maximizes the model misclassification probability while maintaining visual concealment. Reinforcement learning driven evaluation strategy: construct Markov decision process (MDP), use the model classification result as the reward signal, train the agent to generate adversarial samples that minimize the disturbance and maximize the evaluation success rate. Bayesian optimization-based method: use Gaussian process to model the response function of the target model, efficiently search for the optimal disturbance parameter through Bayesian optimization, and reduce the number of queries.

[0042] S2, the system uses a robustness enhancement algorithm to enhance the robustness of the adversarial sample, and obtains an enhanced adversarial sample.

[0043] The system integrates four typical robustness enhancement algorithms: LGS, PRN, RS and FFDnet. These methods have significant differences in the design of robustness enhancement, and the core robustness enhancement principle can be summarized as: (1) Local Gradients Smoothing (LGS):

[0044] Where x is the original image, g(x) is the normalized gradient size map used to locate the high gradient area of the adversarial noise, λ is the smoothing factor that controls the degree of smoothing, and ⊙ represents the pixel-wise multiplication operation. This formula reduces the influence of adversarial noise by suppressing high gradient areas while preserving important features of the image.

[0045] (3) Perturbation Rectifying Network (PRN)

[0046] Where θ p and b p are the weight and bias parameters of PRN, N is the number of training samples, and L(ℓi∗,ℓi) is the loss between the predicted label of the joint network QUOTE and the predicted label ℓ i of the target network. This formula trains PRN to learn how to correct the image so that the target network can correctly classify the corrected image by minimizing the difference between the predicted label of the joint network and the predicted label of the target network.

[0047] (3) RS (Randomized Smoothing):

[0048] Randomized smoothing is a probability-based robustness enhancement algorithm that constructs a new "smoothed" classifier g from an arbitrary base classifier f. g(x) returns a class c such that the preimage {x∈Rd:f(x)=c} has the largest probability measure under the distribution N(x,σ2I). The basic principle is to generate multiple smoothed image samples by adding random noise of a certain distribution on the basis of the original image, then make predictions on these samples, and make decisions based on the prediction results. For example, in an image classification task, for an input image, randomized smoothing will generate multiple copies of the image with different noise added, input these copies into the classification model for prediction, and finally determine the final classification result by counting the prediction results. (4) FFDnet: FFDnet is built based on a deep learning framework and uses an end-to-end training method, which can learn the denoised image directly from the noisy image. FFDnet combines the feature extraction capability of convolutional neural network (CNN) and the idea of residual learning. In the network, the noisy image is feature-extracted through multiple convolutional layers to capture the noise and image feature information, and then the original noisy image and the noise features learned by the network are combined through residual connection to obtain the denoised image.

[0049] Optionally, if the adversarial sample is a high-frequency noise adversarial sample, a local gradient smoothing algorithm is used, and if the adversarial sample is a global perturbation adversarial sample, a randomized smoothing algorithm is used.

[0050] In order to improve the robustness of the enhanced adversarial samples, the system dynamically adjusts the parameters of the robustness enhancement algorithm according to the influence of the enhanced adversarial samples on the classification accuracy of the model. Specifically, by comparing the accuracy of the enhanced adversarial samples and the original samples, if the accuracy decreases significantly, it means that the enhancement effect is insufficient, and the processing strength of the robustness enhancement algorithm is automatically increased (such as increasing the smoothing degree or the number of noise sampling); if the accuracy is similar or the same, the enhancement strength is appropriately reduced to avoid over-processing. The system uses a sliding window to record the enhancement effect and combines a threshold judgment mechanism to dynamically adjust the parameters of the enhancement algorithm to achieve adaptive optimization of different adversarial samples. In addition, if the computer's computing power or other conditions are considered, the following robustness enhancement schemes can also be considered: Adversarial training: inject adversarial samples during model training to optimize the loss function of normal samples and adversarial samples to improve model robustness. Randomly initialized FGSM or PGD can be combined to generate adversarial samples for training. Noise suppression based on attention mechanism: introduce an attention module (such as the multi-head attention in Vision Transformer) in the robustness enhancement network to dynamically identify and suppress high-frequency disturbance areas in the input image and preserve low-frequency semantic information. Sample repair of generative adversarial network: train a generative adversarial network (generator + discriminator) to restore adversarial samples to clean samples, and the discriminator distinguishes between repaired samples and original samples to optimize the repair effect through adversarial training.

[0051] S3, using the adversarial samples and the enhanced adversarial samples to adaptively evaluate the vulnerability of the visual model, automatically saving all operations and detection records.

[0052] After the above processing of the adversarial samples, the system will automatically attack the visual model from multiple angles and strategies, calculate the success probability of each attack in real time, then adaptively evaluate the vulnerability of the model according to the probability, and adaptively adjust the parameters, strategies and angles of the attack according to the difference in attack parameters and success rate of each algorithm, so as to help users discover the vulnerabilities of the attacked model and amplify the defects, and improve the probability of model being deceived.

[0053] Specifically, please refer to the following description: First, the visual model involved in this embodiment is CNN and ViT, which are used together; among them, the convolutional neural network includes but is not limited to Alexnet, EfficientNet, etc., and the ViT-based architecture includes but is not limited to T2T-ViT, FastViT, etc.

[0054] Then, the adaptive evaluation method of this embodiment involves multi-angle and multi-strategy attacks and detection, for example: Multi-model angle: As described in the above visual model, this embodiment adopts different model architectures as target models, such as convolutional neural networks (CNN), ViT (Vision Transformer), etc., which exhibit different performance characteristics and vulnerabilities when processing various types of data. The adaptive evaluation method fully considers these architectural differences and designs corresponding evaluation strategies for each architecture, thereby more accurately discovering potential security vulnerabilities. Moreover, different models of the same architecture may have different underlying implementations, which can also affect their sensitivity to adversarial evaluation, so this method can also identify sensitive parts and provide robustness enhancement references.

[0055] Multiple black-box adversarial evaluation algorithm angles: Each black-box adversarial evaluation algorithm has its unique evaluation focus and approach, such as based on transferability, gradient estimation, sampling, and decision-making. The adaptive evaluation method classifies and studies these different evaluation techniques, analyzes their effectiveness in different scenarios, and provides references for selecting appropriate robustness enhancement strategies.

[0056] Multiple robustness enhancement algorithm angles: Different robustness enhancement algorithms have different focuses in the evaluation process, such as adversarial training to improve model robustness and anomaly detection to filter adversarial samples. The adaptive evaluation method integrates these multi-level and multi-dimensional robustness enhancement methods to build a comprehensive robustness enhancement system. At the same time, it evaluates the synergistic effect between robustness enhancement algorithms, optimizes the combination scheme, effectively improves the overall robustness enhancement effect, and avoids performance degradation caused by mutual interference.

[0057] Next, based on the above angles and corresponding attack strategies, this embodiment will use adversarial samples and enhanced adversarial samples to perform adaptive vulnerability evaluation on the target model (i.e., the visual model). The specific implementation of the adaptive evaluation method includes vulnerability detection strategies based on multi-model architecture differences, dynamic parameter adjustment methods for different adversarial evaluation algorithms, and real-time feedback mechanisms for adversarial evaluation success rates and robustness enhancement effects.

[0058] (1) Vulnerability detection strategy based on multi-model architecture differences.

[0059] For the vulnerability detection strategy, the inherent characteristics of the visual model are deeply analyzed to provide attack basis and defense baseline for sensitive feature quantification and vulnerability mapping. To this end, this embodiment establishes a decision chain of "sensitive features → attack angles → attack strategies" to detect vulnerabilities.

[0060] Specifically, the vulnerability detection strategy includes the following steps: S3001, quantize the original input image sensitive features, that is, measure the characteristics of the input image itself, in this embodiment, Sobel gradient variance is used to mark the high frequency detail area in the image, and HSV saturation standard deviation is used to identify the hue or saturation fragile channel of the image, so as to locate the object or attack target of the image. Then, the data of the visual model is received through the API interface, and the structure parameters of the visual model are extracted therefrom, including the number of convolution layers, the existence of attention module, the proportion of ReLU activation function and the pooling strategy.

[0061] S3002, map the structure parameters to the vulnerability mode, select the attack strategy according to the current vulnerability mode, attack in the located or constrained area according to the selected attack strategy, calculate the success rate of each attack in real time, and finally save the operation and detection record, including the structure parameters of the visual model, the vulnerability mode and the attack strategy.

[0062] The structure parameters are mapped to the vulnerability mode, and the attack strategy is selected according to the current vulnerability mode, specifically: When the convolution layer is detected and the convolution layer is greater than 10 layers, it is marked as high frequency vulnerability, and TA algorithm is selected for frequency domain attack; When the attention module is detected, it is marked as semantic boundary vulnerability, and BA algorithm is used for boundary attack; When the proportion of ReLU activation function is higher than 80%, it is marked as gradient saturation vulnerability, and NES algorithm is used for confidence attack; When the proportion of pooling layer is higher than 70%, it is marked as feature disturbance vulnerability, and Sign-opt algorithm is used for global or local disturbance attack, wherein the pooling layer includes maximum pooling layer and average pooling layer.

[0063] Further explanation is that the strategy of pooling mapping can also be divided into the proportion of maximum pooling layer and average pooling layer respectively, for example: when the proportion of maximum pooling is higher than 70%, it is marked as local strong feature dependence, that is, excessive dependence on local extreme value, and the attack side uses Sign-opt algorithm to modify the key pixel points for once attack. When the average pooling proportion is higher than 70%, it is marked as weak global disturbance resistance, because the features are uniformly distributed, global disturbance is needed to destroy the statistical characteristics.

[0064] Taking convolutional neural network CNN as an example, when the convolution layer of CNN is detected to be more than 10 layers, the feature is marked as high frequency vulnerability, TA algorithm is started for frequency domain attack, and the disturbance intensity is adjusted according to 0.05*convolution layer number, and other features are the same.

[0065] (2) Dynamic parameter adjustment method for different adversarial evaluation algorithms.

[0066] Based on the perspective of multiple black-box adversarial evaluation algorithms, after completing the vulnerability detection strategy, the success rate of each attack is analyzed, and the attack parameters are dynamically adjusted according to the analysis structure as follows: When selecting frequency domain attack, if the frequency domain attack rate exceeds 75% for two consecutive times, then reduce the low frequency disturbance ratio; When selecting boundary attack, if the boundary approximation iteration exceeds 800 steps and is still unsuccessful, then relax the neighborhood constraint range according to the exponential decay rule; When selecting confidence attack, if the gradient symbol matching rate is less than 70% for three consecutive times, increase the number of randomly sampled directions and improve the gradient estimation accuracy; When selecting perturbation attack, when the expected loss variance is less than 0.02, increase the Gaussian noise scale to improve sampling diversity.

[0067] (3) Real-time feedback mechanism of adversarial evaluation success rate and robustness enhancement effect.

[0068] In this real-time feedback mechanism, dynamic optimization is achieved through a three-stage closed loop: Attack side: For the same attack strategy, the success rate of the adversarial sample against the visual model attack is calculated in real time. When the success rate of a single attack exceeds the set value, the corresponding vulnerability of the visual model is marked. When the success rate of the attack is less than the set value for multiple consecutive times, defense detection is triggered. Otherwise, the attack task is continued; Defense side: When defense detection is triggered, the enhanced adversarial sample is input to the visual model for defense, and the success rate of the attack side attack is calculated again. If the success rate of a single attack side attack exceeds the set value, it is determined to be invalid, and the vulnerable area is located; According to the comparison of the success rates of attacks before and after triggering defense detection, the attack strategy of the attack side and the defense strategy of the defense side are optimized.

[0069] It is worth explaining that the key of this mechanism lies in the real-time data based on attack detection and defense feedback, and the precise execution of adversarial strategy iteration, especially the comparison of the success rates of attacks before and after defense detection. The attack improvement rate can be calculated through the success rates of attacks before and after defense detection, and the effectiveness of attacks before and after defense detection can be clearly known through the improvement rate. If the improvement rate increases, it means that the enhanced adversarial sample will excavate the vulnerable place. When it reaches a certain threshold of the embodiment, the best enhancement range can be found, which meets the system parameter optimization expectation. Conversely, if the improvement rate does not increase but decreases, it means that under the attack and defense of the attack strategy, the vulnerable place cannot be effectively identified. In order to further excavate the vulnerable vulnerabilities of the visual model or excavate the vulnerable vulnerabilities of different visual models, the selection of attack strategy or the selection of defense strategy needs to be further optimized.

[0070] Further, the optimization of the attack strategy of the attack side includes: when the success rate of the attack after triggering the defense detection once is less than a set value (such as 80%), and the corresponding enhanced adversarial sample meets the requirements of the defense side (such as the sample after denoising, correction, smoothing and denoising meets the corresponding requirements), the perturbation intensity of the attack side is increased; when the success rate of the attack after triggering the defense detection for two or three times is greater than the set value, the perturbation intensity of the attack side is compressed.

[0071] Further, the optimization of the defense strategy of the defense side includes: The promotion rate of the success rate of the attack before and after triggering the defense detection is calculated, if the promotion rate is not greater than a set value, the number of types of input enhanced adversarial samples is increased; after increasing the number of types of input enhanced adversarial samples, if the promotion rate is less than the set value and is reduced to a number of times of the set value, the corresponding robustness enhancement processing is performed on the corresponding increased types; when the promotion rate is greater than the set value, the attack strategy of the attack side is switched.

[0072] Finally, the system automatically saves the attack parameters, attack methods, labels, original images, adversarial samples and other key information that help to improve the success rate and facilitate the user to analyze the model vulnerability after the attack is completed. And the system will automatically classify and evaluate which method is the most threatening to the model. Through this kind of persistent storage strategy, the robustness of the model can be better detected, and new angles are provided for robustness enhancement algorithms.

[0073] It should be noted that, for the foregoing method embodiments, in order to facilitate description, they are all expressed as a series of action combinations, but those skilled in the art should know that the present application is not limited by the action sequence described, because according to the present application, certain steps can be performed in other sequences or simultaneously.

[0074] Based on the same idea as the black-box adaptive security evaluation method for artificial intelligence models in the above embodiments, the present application also provides a black-box adaptive security evaluation system for artificial intelligence models, which can be used to execute the black-box adaptive security evaluation method for artificial intelligence models. For the convenience of description, in the structural schematic diagram of the embodiment of the black-box adaptive security evaluation system for artificial intelligence models, only the part related to the embodiment of the present application is shown, and those skilled in the art can understand that the illustrated structure does not constitute a limitation on the device, and can include more or fewer components than the illustrated, or combine certain components, or different component arrangements.

[0075] Please refer to Figure 2 In another embodiment of the present application, a black-box adaptive security evaluation system 10 for artificial intelligence models is provided, which includes a first processing module 11, a second processing module 12 and a task execution module 13. The first processing module 11 is configured to select a suitable black-box adversarial evaluation algorithm according to characteristics of the target visual model, and dynamically generate an adversarial sample, wherein the characteristics of the target visual model include a target model architecture and a data set type. The second processing module 12 is configured to select a robust enhancement algorithm according to characteristics of the adversarial sample, and obtain an enhanced adversarial sample, wherein the characteristics of the adversarial sample include an adversarial noise type and a frequency distribution. The task execution module 13 is configured to adaptively evaluate a vulnerability of the visual model by using the adversarial sample and the enhanced adversarial sample, and automatically save all operations and detection records.

[0076] It should be noted that the black-box adaptive security evaluation system for the artificial intelligence model of the present application corresponds to the black-box adaptive security evaluation method for the artificial intelligence model of the present application. The technical features and advantages described in the above embodiments of the black-box adaptive security evaluation method for the artificial intelligence model are applicable to the embodiments of the black-box adaptive security evaluation method for the artificial intelligence model. For specific content, please refer to the description in the method embodiments of the present application. Here, no further description is given, and it is hereby declared.

[0077] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combinations of the technical features do not contradict each other, they should be considered as within the scope of the present application.

[0078] The above embodiments are the preferred embodiments of the present application, but the embodiments of the present application are not limited to the above embodiments. Any changes, modifications, substitutions, combinations, simplifications made without departing from the spirit and principles of the present application should be considered as equivalent replacement modes and should be included in the protection scope of the present application.

Claims

1. A black-box adaptive security assessment method for artificial intelligence models, characterized in that, include: The system selects an appropriate black-box adversarial evaluation algorithm based on the characteristics of the target visual model and dynamically generates adversarial examples. The characteristics of the target visual model include the target model architecture and the dataset type. The system selects a robust enhancement algorithm based on the characteristics of the adversarial examples to obtain enhanced adversarial examples. The characteristics of the adversarial examples include the type and frequency distribution of adversarial noise. Adaptively evaluate the vulnerabilities of the visual model using adversarial examples and enhanced adversarial examples, and automatically save all operation and detection records.

2. The black-box adaptive security assessment method for artificial intelligence models according to claim 1, characterized in that, The system selects a suitable black-box adversarial evaluation algorithm based on the characteristics of the target visual model and dynamically generates adversarial examples, including: For target vision models based on convolutional neural networks, choose either the Sign-opt algorithm based on gradient estimation or the BA algorithm based on decision boundaries. For ViT-based target vision models, choose either the frequency domain transform-based TA algorithm or the sampling-based NES algorithm. The visual model includes convolutional neural networks and ViT, and the two models are used together. The convolutional neural network includes at least one of the following models: AlexNet and EfficientNet; The ViT includes at least one of the following models: T2T-ViT and FastViT.

3. The black-box adaptive security assessment method for artificial intelligence models according to claim 2, characterized in that, include: The algorithm parameters are adjusted based on the structural parameters of the target visual model, including the number of convolutional layers and the type of activation function.

4. The black-box adaptive security assessment method for artificial intelligence models according to claim 1, characterized in that, The system selects a robust enhancement algorithm based on the characteristics of the adversarial examples to obtain enhanced adversarial examples, including: If the adversarial sample is a high-frequency noise adversarial sample, then the local gradient smoothing algorithm is used; If the adversarial sample is a globally perturbative adversarial sample, then a random smoothing algorithm is used; The system dynamically adjusts the parameters of the robustness enhancement algorithm based on the impact of the enhanced adversarial examples on the model's classification accuracy. Specifically, it compares the accuracy of the enhanced adversarial examples with that of the adversarial examples. If the accuracy decreases, the processing intensity of the robustness enhancement algorithm is automatically increased. If the accuracy is similar or the same, the enhancement intensity is reduced. The system uses a sliding window to record the enhancement effect and combines it with a threshold judgment mechanism to dynamically adjust the parameters of the enhancement algorithm.

5. The black-box adaptive security assessment method for artificial intelligence models according to claim 1, characterized in that, The adaptive evaluation of visual model vulnerabilities using adversarial examples and enhanced adversarial examples includes a vulnerability detection strategy based on multi-model architecture differences, specifically: The sensitive features of the input sample are quantified, and the structural parameters of the visual model are extracted through the API interface. The sensitive features include edge sensitivity and color sensitivity. Edge sensitivity is used to locate the high-frequency detail concentration area of ​​the input sample, and color sensitivity is used to identify the hue or saturation weak channels of the input sample. The structural parameters include the number of convolutional layers, the presence of attention modules, the proportion of ReLU activation functions, and the pooling strategy. The structural parameters are mapped to vulnerability patterns, and an attack strategy is selected based on the current vulnerability pattern. The success rate of each attack is calculated in real time. Save operation and detection records, including the structural parameters of the visual model, vulnerability patterns, and attack strategies; The process of mapping structural parameters to vulnerability patterns and selecting an attack strategy based on the current vulnerability pattern specifically involves: When a convolutional layer is detected and its value exceeds a set threshold, it is marked as a high-frequency vulnerability, and the TA algorithm is selected for frequency domain attack. When an attention module is detected, it is marked as a semantic boundary vulnerability, and the BA algorithm is used for boundary attack. When the proportion of ReLU activation function exceeds a set threshold, it is marked as a gradient saturation vulnerability, and a confidence attack is performed using the NES algorithm. When the pooling layer ratio is detected to be higher than a set threshold, it is marked as a local strong feature dependency, and the Sign-opt algorithm is used to perform a local perturbation attack.

6. The black-box adaptive security assessment method for artificial intelligence models according to claim 5, characterized in that, The method for adaptively evaluating the vulnerabilities of visual models using adversarial examples and enhanced adversarial examples includes dynamic parameter adjustment methods for different adversarial evaluation algorithms, specifically: When selecting frequency domain attacks, if the frequency domain attack rate exceeds the set threshold twice consecutively, the low-frequency perturbation ratio will be reduced. When choosing a boundary attack, if the boundary approach iterations fail after more than 800 steps, the neighborhood constraint range is relaxed according to the exponential decay rule. When choosing a confidence attack, if the gradient sign matching rate is lower than the set threshold three times in a row, the number of random sampling directions is increased and the gradient estimation accuracy is improved. When selecting a perturbation attack, if the expected loss variance is less than a set threshold, the Gaussian noise scale is increased to improve sampling diversity.

7. The black-box adaptive security assessment method for artificial intelligence models according to claim 1, characterized in that, The adaptive evaluation of visual model vulnerabilities using adversarial examples and enhanced adversarial examples includes a real-time feedback mechanism for adversarial evaluation success rate and robustness enhancement effect, specifically: On the attack side: For the same attack strategy, the success rate of the adversarial sample attacking the visual model is calculated in real time. When the success rate of a single attack exceeds the set value, the corresponding vulnerability of the visual model is marked. When the success rate of the attack is less than the set value multiple times in a row, the defense detection is triggered. Otherwise, the attack task continues to be executed. Defense side: When defense detection is triggered, the enhanced adversarial sample is input to defend the visual model, and the success rate of the attack side is recalculated. If the success rate of a single attack side exceeds the set value, it is judged as a failure, and the vulnerable area is located. Based on a comparison of the success rates of attacks before and after triggering defense detection, the attack strategy on the attacking side and the defense strategy on the defending side are optimized.

8. The black-box adaptive security assessment method for artificial intelligence models according to claim 7, characterized in that, The optimization of the attack strategy on the attack side includes: if the success rate of the attack after a single trigger of the defense detection is less than a set value, and the corresponding enhanced adversarial sample meets the requirements of the defense side, then the perturbation intensity of the attack side is increased; if the success rate of the attack after two or three consecutive triggers of the defense detection exceeds a set value, then the perturbation intensity of the attack side is reduced.

9. The black-box adaptive security assessment method for artificial intelligence models according to claim 7, characterized in that, The optimization of the defense strategy on the defense side includes: Calculate the improvement rate of the attack success rate before and after triggering the defense detection. If the improvement rate is not greater than the set value, increase the number of types of enhanced adversarial samples. If the improvement rate is less than the set value and shrinks to several times the set value after increasing the number of enhanced adversarial samples, perform corresponding robustness enhancement processing on the corresponding added types. If the improvement rate is greater than the set value, switch the attack strategy on the attack side.

10. A black-box adaptive security assessment system for artificial intelligence models, characterized in that, The black-box adaptive security assessment method for artificial intelligence models applied to any one of claims 1-9 includes a first processing module, a second processing module, and a task execution module; The first processing module is used by the system to select a suitable black-box adversarial evaluation algorithm based on the characteristics of the target visual model and dynamically generate adversarial examples. The characteristics of the target visual model include the target model architecture and the dataset type. The second processing module is used by the system to select a robust enhancement algorithm based on the characteristics of the adversarial sample and obtain the enhanced adversarial sample. The characteristics of the adversarial sample include the type and frequency distribution of adversarial noise. The task execution module is used to adaptively evaluate the vulnerabilities of the visual model using adversarial examples and enhanced adversarial examples, and automatically saves all operation and detection records.

Citation Information

Cited By

  • Large model security evaluation method based on multi-dimensional adversarial attack

    CN121547305A