A method, apparatus, device and medium for enhancing robustness of an image classification model

CN122676261APending Publication Date: 2026-09-01HEFEI UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610997533.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-06
Publication Date
2026-09-01

AI Technical Summary

Technical Problem

由于无法获知模型在面对特定扰动时内部注意力区域的偏移情况与特征分布的变化程度,随机组合的扰动往往无法命中模型的最薄弱环节,使得增强训练消耗大量算力却难以针对脆弱特征进行定向修正,最终导致模型鲁棒性提升效率低下且效果不稳定,无法实现针对模型真实弱点的量化评估与精准增强

Benefits of technology

本发明从模型注意力机制漂移、预测决策一致性、特征分布偏差三个核心内部维度联合量化扰动对模型的破坏能力,三类指标分别对应模型中间层关注区域偏移、输出层决策稳定性、整体特征空间割裂程度,可多角度地量化多维扰动对模型内在推理机制与外在分类性能的综合破坏力;基于该三个维度构建综合鲁棒性适应度指标,并以该指标最大化作为优化目标、以多维扰动参数为变量搭建优化模型求解最优扰动参数;该优化逻辑可锁定能够最大化激发模型内在缺陷、引发模型注意力质心极值偏移、样本特征分布显著割裂的扰动组合,实现了扰动参数在多维连续空间的迭代搜索,相较于随机扰动、固定参数扰动方式,这种方式生成的扰动数据针对性、破坏性与靶向性更强,能够精准聚焦模型脆弱的特征响应区间,为后续定向鲁棒性训练提供高质量、高价值的扰动样本支撑,实现了模型的定向鲁棒性修复与强化。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122676261A_ABST
    Figure CN122676261A_ABST
Patent Text Reader

Abstract

This invention provides a method, apparatus, device, and medium for robustness enhancement of an image classification model. It relates to the field of image processing technology. The method includes: applying a perturbation to the original image to generate a perturbation map; inputting the perturbation map into the model to be enhanced to obtain attention heatmaps and classification results; constructing consistent and inconsistent sample sets based on the consistency of the classification results; calculating attention shifts based on heatmap differences, calculating prediction consistency based on classification results, and calculating feature distribution differences based on the distance between the two sample sets; weighted summing the three to construct a comprehensive robustness fitness index; optimizing the optimal perturbation parameter combination with the goal of maximizing this index; performing directional perturbation on the original image to generate an enhanced image, which is then used to train the model to be enhanced, ultimately obtaining a robustly enhanced image classification model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and in particular to a method, apparatus, device, and medium for enhancing the robustness of an image classification model. Background Technology

[0002] Robustness assessment and enhancement of image classification models are core components of deep learning engineering deployment, directly impacting the model's reliability and security in complex environments. Random data augmentation is currently the mainstream baseline technique for improving model robustness. Its core principle is to randomly sample and combine various image transformation operations during the training phase to generate augmented samples with different perturbation features, forcing the model to learn feature representations insensitive to these transformations. However, because the shift in the model's internal attention region and the degree of change in feature distribution when facing specific perturbations are unknown, randomly combined perturbations often fail to target the model's weakest points. This results in computationally intensive augmentation training that struggles to target vulnerable features, ultimately leading to inefficient and unstable robustness improvements, and failing to achieve quantitative assessment and precise enhancement of the model's true weaknesses. Summary of the Invention

[0003] Therefore, it is necessary to provide a method, apparatus, device, and medium for enhancing the robustness of image classification models to address the aforementioned technical problems.

[0004] The following technical solution is adopted in this specification: This specification provides a method for enhancing the robustness of an image classification model, including: Obtain the original image; Initial perturbation parameters are generated within a preset range of multidimensional perturbation parameters. Based on the initial perturbation parameters, corresponding image transformation operations are performed on the original image data to generate perturbed image data. The original image and the perturbation image are respectively input into the image classification model to be robustly enhanced, and the attention heatmap and classification result of the original image and the perturbation image are obtained. Based on the consistency between the classification results of the original image and the perturbation image, the mixed original image and the perturbation image are divided into a consistent sample set and a non-consistent sample set. The attention shift of the image classification model is obtained based on the difference between the attention heatmap of the original image and the attention heatmap of the perturbed image. The prediction consistency of the image classification model is obtained based on the classification results of the original image and the classification results of the perturbed image. The feature distribution difference of the image classification model is obtained based on the empirical cumulative distribution function distance between the consistent sample set and the non-consistent sample set. The attention shift, prediction consistency and feature distribution difference are fused to obtain a comprehensive robustness fitness index. An optimization model is constructed with the goal of maximizing the comprehensive robustness fitness index and the perturbation parameters as the optimization variables. The optimal perturbation parameters are then obtained by solving the optimization model. The original image is augmented by directional perturbation based on the optimal perturbation parameters to obtain an augmented image. The augmented image is then used to train the image classification model to be robustly augmented, resulting in a robustly augmented image classification model.

[0005] Furthermore, obtaining the attention shift of the image classification model based on the difference between the original image attention heatmap and the perturbed image attention heatmap includes: The attention heatmaps of the original image and the attention heatmaps of the perturbed image are normalized and the weighted centroid positions are calculated. Perform an inverse geometric transformation on the weighted centroid position corresponding to the attention heatmap of the perturbation image to map it back to the original image coordinate system; The distance between the weighted centroid position mapped back to the original image coordinate system and the weighted centroid position of the original image attention heatmap is used as the attention offset.

[0006] Further, the step of dividing the mixed original image and perturbed image into a consistent sample set and a inconsistent sample set based on the consistency between the classification results of the original image and the perturbed image includes: Samples with correct prediction results before and after the perturbation are assigned to a consistent sample set, while samples with changed prediction results after the perturbation are assigned to a non-consistent sample set. The method of obtaining the prediction consistency of the image classification model based on the original image classification result and the perturbed image classification result includes: The number of consistent samples is obtained by counting the number of images whose classification model predictions are correct before and after the perturbation and whose categories remain consistent. The ratio of the number of consistent samples to the total number of samples is used as the prediction consistency.

[0007] Furthermore, obtaining the feature distribution difference of the image classification model based on the empirical cumulative distribution function distance between the consistent sample set and the inconsistent sample set includes: The empirical cumulative distribution functions of the prediction confidence of the consistent sample set and the inconsistent sample set are obtained respectively; The statistical distance between the empirical cumulative distribution function of a consistent sample set and the empirical cumulative distribution function of a non-consistent sample set is used as the feature distribution difference.

[0008] Furthermore, the process of constructing an optimization model with the goal of maximizing the comprehensive robustness fitness index and the perturbation parameters as optimization variables, and solving the optimization model to obtain the optimal perturbation parameters, includes: A genetic algorithm is used to initialize and generate a population of perturbation parameters within a pre-defined multidimensional perturbation parameter space; Using the comprehensive robustness fitness index as the fitness function, the fitness value of each individual in the current perturbation parameter population is calculated; Based on the fitness value, perform selection, crossover, and mutation operations on the current perturbation parameter population to generate a new generation of perturbation parameter population; The fitness value calculation and population update operations are performed iteratively until the preset iteration termination condition is met, and the perturbation parameter combination with the largest fitness value in the last generation population is taken as the optimal perturbation parameter combination.

[0009] Furthermore, the initial perturbation parameters include rotation angle, translation offset, blur level, noise intensity, contrast adjustment value, and scaling ratio.

[0010] Furthermore, the image classification model to be robustly enhanced is a convolutional neural network model containing at least one convolutional layer.

[0011] This specification provides a robustness enhancement device for an image classification model, including: The data acquisition module is used to obtain the raw image; The image perturbation module is used to generate initial perturbation parameters within a preset multidimensional perturbation parameter range, and to perform corresponding image transformation operations on the original image data based on the initial perturbation parameters to generate perturbed image data. The sample set construction module is used to input the original image and the perturbation image into the image classification model to be robustly enhanced, respectively, to obtain the original image attention heatmap and original image classification result corresponding to the original image, and the perturbation image attention heatmap and perturbation image classification result corresponding to the perturbation image; based on the consistency between the original image classification result and the perturbation image classification result, the mixed original image and perturbation image are divided into a consistent sample set and a non-consistent sample set; The fitness index module is used to obtain the attention shift of the image classification model based on the difference between the attention heatmap of the original image and the attention heatmap of the perturbed image, to obtain the prediction consistency of the image classification model based on the classification results of the original image and the classification results of the perturbed image, and to obtain the feature distribution difference of the image classification model based on the empirical cumulative distribution function distance between the consistent sample set and the inconsistent sample set. The attention shift, prediction consistency and feature distribution difference are fused to obtain a comprehensive robustness fitness index. The optimal perturbation parameter combination acquisition module is used to construct an optimization model with the goal of maximizing the comprehensive robustness fitness index and the perturbation parameters as optimization variables, solve the optimization model, and obtain the optimal perturbation parameters. The classification model enhancement module is used to perform directional perturbation enhancement on the original image based on the optimal perturbation parameters to obtain an enhanced image. The enhanced image is then used to train the image classification model to be robustly enhanced, resulting in a robustly enhanced image classification model.

[0012] This specification provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the robustness enhancement method for the image classification model described above.

[0013] This specification provides a computer device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the robustness enhancement method of the image classification model described above.

[0014] The above-mentioned technical solutions adopted in this specification can achieve the following beneficial effects: This invention quantifies the destructive power of perturbations on the model from three core internal dimensions: model attention mechanism drift, prediction decision consistency, and feature distribution bias. These three indices correspond to the shift in the model's intermediate layer attention region, the stability of the output layer decision, and the degree of fragmentation in the overall feature space, respectively. This allows for multi-faceted quantification of the comprehensive destructive power of multidimensional perturbations on the model's internal inference mechanism and external classification performance. Based on these three dimensions, a comprehensive robustness fitness index is constructed. Maximizing this index is used as the optimization objective, and an optimization model is built using multidimensional perturbation parameters as variables to solve for the optimal perturbation parameters. This optimization logic can identify perturbation combinations that maximize the activation of the model's inherent defects, induce extreme shifts in the model's attention centroid, and significantly fragment the sample feature distribution. This enables iterative search of perturbation parameters in a multidimensional continuous space. Compared to random or fixed-parameter perturbations, this method generates perturbation data that is more targeted, destructive, and focused, accurately targeting the model's vulnerable feature response intervals. This provides high-quality, high-value perturbation samples to support subsequent targeted robustness training, achieving targeted robustness repair and enhancement of the model. Attached Figure Description

[0015] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0016] Figure 1 This is one of the flowcharts illustrating a robustness enhancement method for an image classification model provided in this specification. Figure 2 The second flowchart illustrates a robustness enhancement method for an image classification model provided in this specification. Figure 3A schematic diagram of a robustness enhancement device for an image classification model provided in this specification; Figure 4 This is a schematic diagram of a computer device provided for this specification. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of this specification clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments in this specification without creative effort are within the scope of protection of this application.

[0018] The technical solution provided by this invention can be applied to urban low-altitude logistics drone navigation scenarios. Existing image classification models generally suffer from sensitivity to natural disturbances in real-world applications. These natural disturbances include, but are not limited to, image rotation, translation, blurring, noise, brightness variations, and compression distortion. While these disturbances do not alter the semantic information of the image, they significantly reduce the model's classification accuracy, severely impacting the model's reliability and security in engineering deployments. To address these issues, this application provides a robustness enhancement method for image classification models.

[0019] The robustness enhancement method of the image classification model of the present invention is described below with reference to the accompanying drawings.

[0020] Figure 1 This document provides a flowchart illustrating a robustness enhancement method for an image classification model. Figure 1 As shown, the method includes the following: S100, Obtain the original image.

[0021] For example, the original image includes the original image data and its corresponding category label. Specifically, the data acquisition module reads the image data from local storage devices, databases, or dataset interfaces, and performs basic preprocessing operations such as uniform size adjustment and normalization on the image to meet the input requirements of subsequent model inference. This step outputs the preprocessed original image data and corresponding label information. Its function and beneficial effect are: to provide data input in a uniform format for subsequent perturbation generation and model inference, avoiding the impact of inconsistent data formats on robustness evaluation results.

[0022] S200. Generate initial perturbation parameters within a preset multidimensional perturbation parameter range, and perform corresponding image transformation operations on the original image data based on the initial perturbation parameters to generate perturbed image data.

[0023] For example, initial perturbation parameters are generated within a preset range of multidimensional perturbation parameters, and corresponding image transformation operations are performed on the original image data based on the initial perturbation parameters to generate perturbed image data. Here, the perturbation parameter vector is defined as... ,in: Indicates rotation angle, and Indicates horizontal and vertical offset, Indicates the degree of ambiguity, Indicates noise intensity, Indicates contrast, This indicates the scaling ratio, and the constraint is... , It is the first The minimum value of each disturbance parameter, It is the first The perturbation generation module calculates the maximum value of each perturbation parameter. Based on the input perturbation parameters, it performs corresponding image transformation operations on the original image, including but not limited to: image rotation, image translation, image blurring, image contrast adjustment, and image noise injection. Each set of perturbation parameters generates a corresponding perturbed image. This step outputs the perturbed image data. Its function and beneficial effects are: to uniformly describe natural perturbations through parameterization, to achieve controllable generation of multi-dimensional perturbation combinations, and to provide a technical foundation for systematic perturbation search.

[0024] S300. Input the original image and the perturbation image into the image classification model to be robustly enhanced, respectively, to obtain the original image attention heatmap and original image classification result corresponding to the original image, and the perturbation image attention heatmap and perturbation image classification result corresponding to the perturbation image; based on the consistency between the original image classification result and the perturbation image classification result, divide the mixed original image and perturbation image into a consistent sample set and a non-consistent sample set.

[0025] For example, the original image and the perturbed image are input into the image classification model to be robustly enhanced, respectively, to obtain the original image attention heatmap and classification result, and the perturbed image attention heatmap and classification result. Then, based on the consistency between the original image classification result and the perturbed image classification result, the mixed original image and perturbed image are divided into a consistent sample set and a non-consistent sample set. Specifically, forward inference is performed on the original image and the perturbed image to obtain prediction results. Simultaneously, feature maps are extracted at the model's preset convolutional layers, and combined with the gradient information from backpropagation, corresponding attention heatmaps are generated (e.g., using the Grad-CAM method). Subsequently, the attention heatmaps are normalized, and the weighted centroid positions of the heatmaps are calculated. Finally, the prediction results of the original image, the prediction results of the perturbed image, the attention heatmap of the original image and its centroid coordinates, and the attention heatmap of the perturbed image and its centroid coordinates are output. This process not only obtains the model output results but also obtains information about the model's internal regions of interest, providing quantifiable internal feature evidence for robustness evaluation.

[0026] S400: Obtain the attention shift of the image classification model based on the difference between the attention heatmap of the original image and the attention heatmap of the perturbed image; obtain the prediction consistency of the image classification model based on the classification results of the original image and the classification results of the perturbed image; obtain the feature distribution difference of the image classification model based on the empirical cumulative distribution function distance between the consistent sample set and the non-consistent sample set; and integrate the attention shift, prediction consistency and feature distribution difference to obtain a comprehensive robustness fitness index.

[0027] S410, Attention Shift.

[0028] The attention heatmaps of the original and perturbed images are normalized and their weighted centroid positions are calculated. An inverse geometric transformation is performed on the weighted centroid positions corresponding to the perturbed image attention heatmaps to map them back to the original image coordinate system. The distance between the weighted centroid positions mapped back to the original image coordinate system and the weighted centroid positions of the original image attention heatmaps is used as the attention offset. Specifically:

[0029] For the original image With perturbation image Extract attention heatmaps and calculate centroids respectively: Define Original image The weighted centroid coordinates of the attention heatmap, where , These are the x and y coordinates of the centroid in the original image coordinate system, respectively; (Definition) For perturbed images The weighted centroid coordinates of the attention heatmap, where , Let x and y be the x and y coordinates of the centroid in the perturbed image coordinate system, respectively. These coordinates have been mapped back to the original image coordinate system through inverse geometric transformation. Based on the above definition, the attention offset is defined as: ; In the formula, This is the attention shift. The value represents the Euclidean distance, which quantifies the positional deviation of the model's attention region before and after the perturbation. The larger the offset, the more significantly the model's attention is affected by the perturbation, and the weaker its robustness.

[0030] S420, Prediction Consistency.

[0031] Samples with correct predictions before and after the perturbation are assigned to a consistent sample set, while samples whose predictions change after the perturbation are assigned to a inconsistent sample set. The number of samples where the image classification model's predictions are correct both before and after the perturbation and maintain the same category is counted to obtain the consistent sample count. The ratio of the consistent sample count to the total sample count is used as the prediction consistency. Specifically:

[0032] Let the image classification model be Then, define the prediction consistency criterion: ; In the formula, The result is the prediction consistency judgment of the sample; For the model, the original image The prediction category; For the model to perturb the image The predicted category; when When =1, it indicates that the model's prediction of the category is consistent between the original image and the perturbed image. When = 0, it indicates that the model's predictions for the two categories are inconsistent. Definition To ensure that the proportion of samples with accurate and consistent predictions before and after the perturbation is consistent, there are a total of If there are 100 samples participating in the evaluation, then:

[0033] ; In the formula, To predict consistency indicators, The total number of samples whose predicted class remains consistent across all evaluation samples. The total number of samples participating in the evaluation; the range of values ​​for this indicator. The larger the value, the stronger the model's predictive stability and the better its robustness under perturbation conditions.

[0034] S430, differences in characteristic distribution.

[0035] The empirical cumulative distribution functions of prediction confidence scores for the consistent sample set and the inconsistent sample set are obtained separately; the statistical distance between the empirical cumulative distribution functions of the consistent sample set and the inconsistent sample set is used as the feature distribution difference. Specifically:

[0036] Specifically, the samples are divided into a consistent sample set. Non-consistent sample sets Let the empirical distribution functions corresponding to the two be defined as follows: and The difference in characteristic distribution is defined as follows: ; In the formula, It is a characteristic distribution difference index based on the empirical cumulative distribution function; The set of samples whose prediction results are consistent before and after the perturbation; This refers to the set of samples whose prediction results are inconsistent before and after the perturbation. For a consistent sample set The empirical cumulative distribution function corresponding to the prediction confidence of all samples in the dataset; Non-consistent sample set The empirical cumulative distribution function corresponding to the prediction confidence of all samples in the dataset; This is used to predict the confidence level variable. This index quantifies the difference in confidence level distribution between the two types of samples. The larger the value, the more significant the impact of the perturbation on the model's confidence level distribution, and the worse the model's robustness.

[0037] S440, Comprehensive Robustness Fitness Index.

[0038] Finally, a comprehensive robustness evaluation index is constructed by combining the above three types of indicators; the output includes attention shift distance, feature distribution difference index, and comprehensive robustness evaluation result. This achieves a multi-dimensional quantitative evaluation of model robustness, avoiding reliance solely on accuracy as a single evaluation method. It comprehensively reflects the model's resistance to perturbations from multiple levels, including prediction results, attention mechanisms, and feature distribution, providing a reliable quantitative basis for subsequent perturbation parameter optimization and model robustness enhancement. Specifically:

[0039] ; in, , and The preset weighting coefficients are used to adjust the contribution ratio of attention shift, feature distribution difference and prediction consistency in the comprehensive index, respectively. The value range is (0,1], and can be dynamically adjusted according to the evaluation scenario and the importance of the index.

[0040] S500: With the goal of maximizing the comprehensive robustness fitness index and the perturbation parameters as the optimization variables, an optimization model is constructed, and the optimal perturbation parameters are obtained by solving the optimization model.

[0041] For example, taking the maximization of the comprehensive robustness fitness index as the overall optimization objective, various perturbation-related parameters such as perturbation amplitude, perturbation location, and perturbation intensity are set as optimization variables, and a mathematical optimization model for finding the optimal perturbation parameters is constructed accordingly. Then, a genetic algorithm is used to iteratively solve the constructed optimization model, ultimately converging to obtain the optimal combination of perturbation parameters suitable for the current image classification model. The inputs to this step are the current set of perturbation parameters and the comprehensive evaluation result output by the robustness assessment module. The specific processing method is as follows:

[0042] The perturbation search module employs a genetic algorithm to perform global search and iterative optimization of perturbation parameters. It sequentially executes standard evolutionary operations such as perturbation parameter population initialization, individual fitness calculation, population selection, gene crossover, and gene mutation. The comprehensive robustness evaluation result output by the robustness assessment module is directly used as the fitness function of the genetic algorithm. Based on the fitness value, the perturbation parameter set is continuously updated iteratively, continuously screening and retaining the perturbation parameter combinations that have the strongest destructive effect on model classification performance and best expose model vulnerabilities. This step ultimately outputs the optimal perturbation parameter combination. Its function and beneficial effects are: it achieves intelligent and automated global search for perturbation parameters targeting model weaknesses, eliminating the need for manual experience in setting perturbation parameters and avoiding the blindness and inefficiency of random perturbations. It can accurately identify the perturbation patterns and parameter configurations that pose the greatest threat to the model, providing reliable parameter basis for subsequent targeted perturbation enhancement and model robustness improvement.

[0043] S600. Perform directional perturbation enhancement on the original image based on the optimal perturbation parameters to obtain an enhanced image. Use the enhanced image to train the image classification model to be robustly enhanced to obtain a robustly enhanced image classification model.

[0044] For example, based on the optimal perturbation parameter combination, the original training data, and the parameters of the benchmark model to be optimized, the original image is subjected to targeted perturbation enhancement processing to generate enhanced images with targeted perturbation features in batches. Based on the optimal perturbation parameters, the vulnerable areas and susceptible feature dimensions of the benchmark model are accurately located, and targeted and controllable perturbation enhancement is performed on some original training samples to generate enhanced images. Then, the perturbation-enhanced samples and the original clean samples are mixed to construct a new training set, which is then fed into the image classification model for supervised training.

[0045] The generated enhanced images are integrated into the original training dataset, and a hybrid training method combining original samples and perturbation-enhanced samples is adopted. These are then input into the image classification model to be robustly enhanced for iterative training, ultimately completing model parameter updates and convergence to obtain the robustly enhanced image classification model. During model enhancement, the previously defined prediction consistency constraint and attention consistency constraint are introduced during the iterative training process. This dual constraint on the model learning process, both at the level of the model's classification output and the level of its internal attention regions, encourages the model to reduce its dependence on perturbation-sensitive features and strengthen its ability to extract robust core features, effectively improving the model's feature representation stability under complex perturbation scenarios.

[0046] In one specific embodiment, the present invention provides a robustness evaluation and enhancement method based on model attention consistency and perturbation parameter search. This method achieves systematic evaluation and targeted improvement of the robustness of image classification models by identifying weaknesses and designing closed-loop enhancement strategies. Figure 2 This is the second flowchart illustrating a robustness enhancement method for an image classification model provided in this specification. Figure 2 As shown, the overall process is as follows:

[0047] Step 1: Baseline Model Training. The image classification model is fully trained using the original training data to obtain converged baseline model parameters. Since subsequent perturbation analysis heavily relies on the stability of the model's internal feature representations, if the model has not yet converged, its attention distribution itself will exhibit random fluctuations, leading to distorted evaluation results. Therefore, this step aims to provide a stable and reliable reference baseline for perturbation analysis, ensuring that the calculation of subsequent attention shifts and feature distribution differences is based on the premise that the model has reached a steady state.

[0048] Step 2: Construction of the perturbation parameter space. A parameter space is constructed containing multi-dimensional perturbation parameters such as rotation, translation, blur, contrast, and noise. Each set of parameters constitutes a perturbation combination, forming the initial perturbation parameter population. Given that image degradation in real-world applications typically manifests as a combination of multiple natural perturbations rather than the independent effect of a single perturbation, evaluating only a specific perturbation type would fail to reflect the model's true performance in complex environments. Therefore, by constructing a multidimensional composite perturbation space, we can achieve comprehensive coverage of perturbation forms in real-world scenarios, providing a sufficient range of parameters for subsequent searches of the most destructive perturbation combinations.

[0049] Step 3: Attention Heatmap Extraction. Model inference is performed on both the original and perturbed images, and a Grad-CAM attention heatmap is generated at the last convolutional layer of the model. Since the model's decisions rely on spatial attention to key regions of the image, simply outputting the prediction results cannot reveal the internal mechanism of "where the model focuses." Therefore, by extracting the attention heatmaps before and after perturbing, the explicit spatial distribution of the model's internal attention regions is obtained, providing fundamental data for subsequent quantitative analysis of the degree of attention shift.

[0050] Step 4: Attention Consistency Calculation. The attention heatmaps generated before and after the perturbation are normalized, and their weighted centroid positions are calculated. After performing an inverse geometric transformation on the attention heatmap of the perturbation image, the distance between the centroids before and after the perturbation is calculated as the attention offset. Traditional visualization methods can only subjectively judge whether attention has drifted, lacking quantitative measurement tools. Therefore, by transforming the attention distribution into centroid coordinates and calculating the spatial offset distance, the degree of change in model attention is accurately quantified, elevating attention consistency assessment from subjective observation to objective observation.

[0051] Step 5: Feature Distribution Difference and Robustness Index Construction. Based on the model's prediction results before and after the perturbation, the samples are divided into consistent prediction samples (predicted categories are the same before and after the perturbation) and inconsistent prediction samples (predicted categories change before and after the perturbation). The distribution difference of the two types of samples in the feature space is calculated separately, and a comprehensive robustness index is constructed. Since using the change in prediction accuracy alone can only reflect the stability of the model output and cannot characterize the degree of drift in the internal feature representation, robustness is comprehensively evaluated from two dimensions—the model behavior layer and the representation layer—by simultaneously measuring prediction stability (prediction consistency) and internal feature changes (feature distribution difference), thus improving the comprehensiveness and depth of the evaluation results.

[0052] Step 6: Perturbation Search Based on Genetic Algorithm. Using the comprehensive robustness index constructed in Step 5 as the fitness function, a genetic algorithm is used to perform multi-generational evolutionary search on the perturbation parameters, outputting the perturbation parameter combination that minimizes the comprehensive robustness index (i.e., the most destructive). Manually setting perturbation parameters is often limited by subjective experience, making it difficult to systematically explore extremely vulnerable regions in the parameter space. Therefore, by leveraging the global search capability of the genetic algorithm, the most destructive perturbation combination for the model is automatically evolved in the multi-dimensional parameter space, improving the systematicness, objectivity, and exploration depth of the perturbation search, and avoiding evaluation blind spots caused by human bias.

[0053] Step 7: Targeted Augmentation Training Based on the Strongest Perturbation. The combination of parameters for the strongest perturbation found in Step 6 is applied to the training data to perform targeted augmentation training on a subset of samples, thereby improving the model's robustness in vulnerable regions. Since traditional data augmentation typically employs random or uniform sampling strategies, it lacks specificity for the model's truly vulnerable areas, resulting in low augmentation efficiency. Therefore, by directly using the strongest perturbation discovered in the aforementioned search process for sample augmentation during the training phase, a closed-loop augmentation cycle of "evaluating and identifying weaknesses—training to target weaknesses" is achieved. This concentrates limited training resources on the model's weakest link, thereby improving the model's robustness in key vulnerable regions with higher efficiency.

[0054] Finally, the output is the model parameters after targeted augmentation training and their corresponding robustness evaluation metrics, completing the entire process from robustness evaluation to robustness augmentation.

[0055] The robustness enhancement device for the image classification model provided by the present invention will be described below. The robustness enhancement device for the image classification model described below can be referred to in correspondence with the robustness enhancement method for the image classification model described above.

[0056] Figure 3 This is a schematic diagram of the robustness enhancement device for an image classification model provided by the present invention. For example, please refer to [link / reference]. Figure 3 As shown, the robustness enhancement device for this image classification model may include: The data acquisition module is used to obtain the raw images.

[0057] The image perturbation module is used to generate initial perturbation parameters within a preset range of multidimensional perturbation parameters, and to perform corresponding image transformation operations on the original image data based on the initial perturbation parameters to generate perturbed image data.

[0058] The sample set construction module is used to input the original image and the perturbation image into the image classification model to be robustly enhanced, respectively, to obtain the original image attention heatmap and original image classification result corresponding to the original image, as well as the perturbation image attention heatmap and perturbation image classification result corresponding to the perturbation image; based on the consistency between the original image classification result and the perturbation image classification result, the mixed original image and perturbation image are divided into a consistent sample set and a non-consistent sample set.

[0059] The fitness index module is used to obtain the attention shift of the image classification model based on the difference between the attention heatmap of the original image and the attention heatmap of the perturbed image, to obtain the prediction consistency of the image classification model based on the classification results of the original image and the classification results of the perturbed image, and to obtain the feature distribution difference of the image classification model based on the empirical cumulative distribution function distance between the consistent sample set and the inconsistent sample set. The attention shift, prediction consistency and feature distribution difference are fused to obtain a comprehensive robustness fitness index.

[0060] The optimal perturbation parameter combination acquisition module is used to construct an optimization model with the goal of maximizing the comprehensive robustness fitness index and the perturbation parameters as optimization variables, solve the optimization model, and obtain the optimal perturbation parameters.

[0061] The classification model enhancement module is used to perform directional perturbation enhancement on the original image based on the optimal perturbation parameters to obtain an enhanced image. The enhanced image is then used to train the image classification model to be robustly enhanced, resulting in a robustly enhanced image classification model.

[0062] Specific limitations regarding the robustness enhancement device for image classification models can be found in the above section on robustness enhancement of image classification models, and will not be repeated here. Each module in the aforementioned robustness enhancement device for image classification models can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the corresponding operations of each module.

[0063] This specification also provides a computer-readable storage medium storing a computer program that can be used to execute the above-described... Figure 1 A robustness enhancement method for the provided image classification model.

[0064] This instruction manual also provides Figure 4 The schematic diagram of the computer device shown is as follows: Figure 4 As shown, at the hardware level, this computer device includes a processor, internal bus, network interface, memory, and non-volatile memory, and may also include other hardware required for business operations. The processor reads the corresponding computer program from the non-volatile memory into memory and then executes it to achieve the above. Figure 1 A robustness enhancement method for the provided image classification model.

[0065] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical storage, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0066] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

Claims

1. A method for enhancing the robustness of an image classification model, characterized in that, include: Obtain the original image; Initial perturbation parameters are generated within a preset range of multidimensional perturbation parameters. Based on the initial perturbation parameters, corresponding image transformation operations are performed on the original image data to generate perturbed image data. The original image and the perturbed image are respectively input into the image classification model to be robustly enhanced, and the original image attention heatmap and original image classification result are obtained, as well as the perturbed image attention heatmap and perturbed image classification result are obtained. Based on the consistency between the classification results of the original image and the perturbation image, the mixed original image and perturbation image are divided into a consistent sample set and a non-consistent sample set; The attention shift of the image classification model is obtained based on the difference between the attention heatmap of the original image and the attention heatmap of the perturbed image. The prediction consistency of the image classification model is obtained based on the classification results of the original image and the classification results of the perturbed image. The feature distribution difference of the image classification model is obtained based on the empirical cumulative distribution function distance between the consistent sample set and the non-consistent sample set. The attention shift, prediction consistency and feature distribution difference are fused to obtain a comprehensive robustness fitness index. An optimization model is constructed with the goal of maximizing the comprehensive robustness fitness index and the perturbation parameters as the optimization variables. The optimal perturbation parameters are then obtained by solving the optimization model. The original image is augmented by directional perturbation based on the optimal perturbation parameters to obtain an augmented image. The augmented image is then used to train the image classification model to be robustly augmented, resulting in a robustly augmented image classification model.

2. The robustness enhancement method for the image classification model as described in claim 1, characterized in that, The method of obtaining the attention shift of the image classification model based on the difference between the attention heatmap of the original image and the attention heatmap of the perturbed image includes: The attention heatmaps of the original image and the attention heatmaps of the perturbed image are normalized and the weighted centroid positions are calculated. Perform an inverse geometric transformation on the weighted centroid position corresponding to the attention heatmap of the perturbation image to map it back to the original image coordinate system; The distance between the weighted centroid position mapped back to the original image coordinate system and the weighted centroid position of the original image attention heatmap is used as the attention offset.

3. The robustness enhancement method for the image classification model as described in claim 1, characterized in that, The step of dividing the mixed original image and perturbed image into a consistent sample set and a inconsistent sample set based on the consistency between the classification results of the original image and the perturbed image includes: Samples with correct prediction results before and after the perturbation are assigned to a consistent sample set, while samples with changed prediction results after the perturbation are assigned to a non-consistent sample set. The method of obtaining the prediction consistency of the image classification model based on the original image classification result and the perturbed image classification result includes: The number of consistent samples is obtained by counting the number of images whose classification model predictions are correct before and after the perturbation and whose categories remain consistent. The ratio of the number of consistent samples to the total number of samples is used as the prediction consistency.

4. The robustness enhancement method for the image classification model as described in claim 3, characterized in that, The method of obtaining the feature distribution difference of the image classification model based on the empirical cumulative distribution function distance between consistent and inconsistent sample sets includes: The empirical cumulative distribution functions of the prediction confidence of the consistent sample set and the inconsistent sample set are obtained respectively; The statistical distance between the empirical cumulative distribution function of a consistent sample set and the empirical cumulative distribution function of a non-consistent sample set is used as the feature distribution difference.

5. The robustness enhancement method for the image classification model as described in claim 1, characterized in that, The optimization model is constructed with the goal of maximizing the comprehensive robustness fitness index and the perturbation parameters as the optimization variables. The optimal perturbation parameters are obtained by solving the optimization model. This includes: A genetic algorithm is used to initialize and generate a population of perturbation parameters within a pre-defined multidimensional perturbation parameter space; Using the comprehensive robustness fitness index as the fitness function, the fitness value of each individual in the current perturbation parameter population is calculated; Based on the fitness value, perform selection, crossover, and mutation operations on the current perturbation parameter population to generate a new generation of perturbation parameter population; The fitness value calculation and population update operations are performed iteratively until the preset iteration termination condition is met, and the perturbation parameter combination with the largest fitness value in the last generation population is taken as the optimal perturbation parameter combination.

6. The robustness enhancement method for the image classification model as described in claim 1, characterized in that, The initial perturbation parameters include rotation angle, translation offset, blur level, noise intensity, contrast adjustment value, and scaling ratio.

7. The robustness enhancement method for the image classification model as described in claim 1, characterized in that, The image classification model to be robustly enhanced is a convolutional neural network model containing at least one convolutional layer.

8. A robustness enhancement device for an image classification model, characterized in that, include: The data acquisition module is used to obtain the raw image; The image perturbation module is used to generate initial perturbation parameters within a preset multidimensional perturbation parameter range, and to perform corresponding image transformation operations on the original image data based on the initial perturbation parameters to generate perturbed image data. The sample set construction module is used to input the original image and the perturbed image into the image classification model to be robustly enhanced, respectively, to obtain the original image attention heatmap and original image classification result corresponding to the original image, as well as the perturbed image attention heatmap and perturbed image classification result corresponding to the perturbed image; Based on the consistency between the classification results of the original image and the perturbation image, the mixed original image and perturbation image are divided into a consistent sample set and a non-consistent sample set; The fitness index module is used to obtain the attention shift of the image classification model based on the difference between the attention heatmap of the original image and the attention heatmap of the perturbed image, to obtain the prediction consistency of the image classification model based on the classification results of the original image and the classification results of the perturbed image, and to obtain the feature distribution difference of the image classification model based on the empirical cumulative distribution function distance between the consistent sample set and the inconsistent sample set. The attention shift, prediction consistency and feature distribution difference are fused to obtain a comprehensive robustness fitness index. The optimal perturbation parameter combination acquisition module is used to construct an optimization model with the goal of maximizing the comprehensive robustness fitness index and the perturbation parameters as optimization variables, solve the optimization model, and obtain the optimal perturbation parameters. The classification model enhancement module is used to perform directional perturbation enhancement on the original image based on the optimal perturbation parameters to obtain an enhanced image. The enhanced image is then used to train the image classification model to be robustly enhanced, resulting in a robustly enhanced image classification model.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the robustness enhancement method for the image classification model as described in any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the robustness enhancement method for the image classification model as described in any one of claims 1 to 7.