Method, system and equipment for testing robustness of data classification model for disturbing hidden layer neuron activation value and medium

By obtaining the activation vectors of the key hidden layers of a CNN model and applying multidimensional differential perturbations, and combining this with gradient optimization algorithms to generate test cases, the problem of insufficient coverage and defect detection in existing CNN robustness testing methods is solved, and a more comprehensive model robustness evaluation is achieved.

CN120803943APending Publication Date: 2025-10-17GUANGXI POWER GRID CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510990681.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-18
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing CNN robustness testing methods suffer from insufficient test coverage and inadequate quantity and quality of defect detection, especially the lack of diversity in random perturbation-based methods and the failure of gradient optimization-based methods to fully assess the combined impact of hidden layer neuron activation values.

Method used

By obtaining the original activation vectors of the key hidden layers of the data classification model, dividing the activation intervals and applying differential perturbations, and combining them with gradient optimization algorithms to generate test cases, we ensure that the test cases can effectively cover the multidimensional perturbations of multiple hidden layers. We then use gradient optimization algorithms to adjust the samples to generate test cases that can trigger model decision errors.

Benefits of technology

It improves the coverage of tests and the quantity and quality of robustness defects discovered, enabling a more comprehensive evaluation of the model's performance under complex perturbations, and significantly enhances the depth and comprehensiveness of robustness testing for CNN models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120803943A_ABST
    Figure CN120803943A_ABST
Patent Text Reader

Abstract

The invention discloses a robustness test method, system and device for a data classification model disturbing hidden layer neuron activation values and a medium, and belongs to the technical field of data privacy security, and the method comprises the steps: obtaining an original activation vector of a to-be-tested sample in at least one target hidden layer in the data classification model; according to the original activation vector, dividing each numerical component of the original activation vector into a plurality of activation intervals, and applying preset differential perturbation to the numerical components in different activation intervals to construct a target activation vector; based on a preset distance between the target activation vector and the original activation vector, adopting a gradient optimization algorithm to adjust the to-be-tested sample to generate a test case; and inputting the test case into the data classification model for reasoning, and when a prediction result corresponding to the test case is inconsistent with an original prediction result corresponding to the to-be-tested sample, determining that the test case is an abnormal decision test case.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data privacy security, in particular to a data classification model robustness test method, system, device and medium for hiding neuron activation values of a perturbation hidden layer. BACKGROUND

[0002] Convolutional Neural Networks (CNN) are widely used in data classification tasks in fields such as finance, medicine, and industrial power. However, in the context of data transactions, the robustness of CNN is particularly prominent. With the increasing importance of data as a new type of production factor in national strategies, the data transaction market has developed rapidly, and there are higher requirements for the stability of models under noise, adversarial samples, and other perturbations. The core of CNN robustness testing is to evaluate the decision reliability of the model under abnormal conditions through perturbed samples, which is of great significance to ensuring the security of sensitive data processing in data transactions. Currently, domestic data transaction platforms still have problems such as poor data quality and unclear ownership in terms of security protection. If CNN models lack robustness testing, it may lead to data misjudgment and trust crisis, hindering the healthy development of the data factor market. Therefore, researching an effective CNN robustness testing method can not only improve the security and stability of models in data transaction scenarios, but also has practical value for promoting the release of data value and the digital transformation of industries.

[0003] According to the different ways of generating test cases, the robustness testing methods of classification models in recent years can be roughly divided into random perturbation-based methods and gradient optimization-based methods.

[0004] 1. Random perturbation-based method. The random perturbation-based method generates test cases by adding random perturbation vectors to input data and evaluates the robustness of CNN classification models. This method relies on randomly selecting perturbation ranges and amplitudes to affect the activation values of internal neurons in CNN models through different perturbation patterns, thereby detecting the model's response to perturbations. Although this method is simple and easy to implement, its main problem is the lack of diversity and effectiveness of test cases. Due to the randomness of the perturbation method, it cannot guarantee that the adversarial perturbation can effectively cover all possible ranges of neuron activation values, and the perturbation patterns of test cases lack pertinence. In addition, the random perturbation method usually perturbs part of the neuron activation values, and the perturbation direction is fixed (such as only increasing or decreasing the activation value), which makes the change direction of neuron activation values too single, resulting in the inability to construct multiple combinations of neuron activation values. Therefore, this method has limited coverage in robustness testing, and cannot fully test the performance of the model under complex perturbations, ultimately affecting the comprehensiveness of defect discovery and test coverage.

[0005] 2. Gradient-based optimization method. Gradient-based optimization method generates perturbations to effectively expose the robustness defects of the model by calculating the gradient information of the input data or neuron activation value of the model. Through gradient backpropagation, the input data or activation value is optimized so that the perturbed sample can maximize the abnormal decision or misclassification of the model. This method can generate high-quality adversarial samples, thereby improving the accuracy of model robustness testing. However, although the gradient-based optimization method can effectively improve the quality of adversarial samples, the existing method usually focuses on reducing the classification confidence of the model to the test seed as an optimization strategy to find abnormal decisions when generating test cases. This strategy does not fully consider the diversity of each hidden layer neuron activation value inside the CNN model and their comprehensive impact on the decision, especially under the perturbation of different activation value combinations. The robustness performance of the model is often inconsistent. As a result, although the gradient-based optimization method can generate test cases that trigger decision errors, it fails to fully evaluate the perturbation effect of all neuron activation values, so the types and number of robustness defects discovered are relatively small, resulting in insufficient comprehensiveness and depth of model robustness testing. SUMMARY

[0006] In view of the above problems, the present application is proposed.

[0007] Therefore, the technical problem solved by the present application is how to solve the problems of insufficient test coverage of the random perturbation-based method and insufficient number and quality of test defects of the gradient optimization-based method.

[0008] To solve the above technical problems, the present application provides the following technical solutions: a data classification model robustness testing method for perturbing hidden layer neuron activation values, comprising: obtaining an original activation vector of a to-be-tested sample in at least one target hidden layer of a data classification model; according to the original activation vector, constructing a target activation vector by dividing each numerical component of the original activation vector into a plurality of activation intervals and applying a preset differential perturbation to the numerical components in different activation intervals; based on a preset distance between the target activation vector and the original activation vector, adjusting the to-be-tested sample using a gradient optimization algorithm to generate a test case; inputting the test case into the data classification model for inference, and when the prediction result corresponding to the test case is inconsistent with the original prediction result corresponding to the to-be-tested sample, determining that the test case is an abnormal decision test case.

[0009] As a preferred scheme of the data classification model robustness test method for disturbing the neuron activation value of the hidden layer, wherein: the step of obtaining the original activation vector of the target hidden layer in the data classification model comprises: calculating the gradient significance and neuron sparsity of each hidden layer in the data classification model to determine the layer importance score of each hidden layer; and selecting one or more hidden layers as the target hidden layer according to the layer importance score. The beneficial effect of the preferred technical scheme is that the importance score of each hidden layer is quantified by calculating the gradient significance and neuron sparsity, so that the test method is no longer blindly selecting the disturbance object. This can accurately concentrate the limited test resources on the hidden layer that has the greatest impact on the final decision of the model and is the most critical feature expression, greatly improving the pertinence and efficiency of the test, so that the deep robustness defects of the model can be found at a faster speed and with a higher success rate.

[0010] As a preferred scheme of the data classification model robustness test method for disturbing the neuron activation value of the hidden layer, wherein: the step of obtaining the original activation vector further comprises: if the target hidden layer is a fully connected layer, directly splicing the neuron activation value thereof as a first activation component; if the target hidden layer is a three-dimensional hidden layer, calculating the activation representative value of each channel feature map thereof, and splicing the activation representative value as a second activation component; and fusing the first activation component and the second activation component to generate the original activation vector.

[0011] As a preferred scheme of the data classification model robustness test method for disturbing the neuron activation value of the hidden layer, wherein: the step of applying the preset differentiated disturbance comprises: dividing the activation interval into a high activation interval, a second high activation interval, a middle interval and a low activation interval; applying zeroing disturbance to the neuron activation value in the high activation interval, applying high-amplitude disturbance to the neuron activation value in the low activation interval, and applying scaling disturbance to the neuron activation value in the second high activation interval and the middle interval. The beneficial effect of the preferred technical scheme is that by dividing the neuron activation value into different activation intervals such as high, medium and low, and applying different types of disturbances such as zeroing, scaling and high amplitude, a structured and highly diverse disturbance strategy is constructed. This completely overcomes the defects of the traditional random disturbance method, such as single direction and incomplete coverage, and can simulate more complex and realistic noise or adversarial attack scenarios, so as to more effectively expose the potential weaknesses of the model under different activation state combinations, significantly increasing the types and quality of defects found.

[0012] As a preferred scheme of the data classification model robustness test method for disturbing the activation value of the hidden layer neuron, wherein: the step of dividing the activation interval further comprises: determining the division ratio of the plurality of activation intervals according to the depth information of the data classification model; and distributing the neurons sorted by the absolute size of the activation value to the corresponding activation interval according to the division ratio.

[0013] As a preferred scheme of the data classification model robustness test method for disturbing the activation value of the hidden layer neuron, wherein: when the target hidden layer is multiple, the step of adjusting the test sample based on the preset distance between the target activation vector and the original activation vector by using the gradient optimization algorithm comprises: determining the corresponding optimization target for each target hidden layer; determining the corresponding weight coefficient according to the gradient sensitivity of each target hidden layer; weighting and aggregating the plurality of optimization targets based on the weight coefficient to obtain a total optimization target; and adjusting the test sample according to the total optimization target by back propagation.

[0014] As a preferred scheme of the data classification model robustness test method for disturbing the activation value of the hidden layer neuron, wherein: the step of adjusting the test sample by using the gradient optimization algorithm to generate a test case comprises: calculating the gradient of the preset distance on the test sample; and updating the test sample according to the gradient and a preset learning rate to generate the test case.

[0015] The present application provides a data classification model robustness test system for disturbing the activation value of the hidden layer neuron.

[0016] To solve the above technical problems, the application further provides the following technical solutions: a data classification model robustness test system for disturbing hidden layer neuron activation values, comprising: an acquisition module configured to acquire an original activation vector of a to-be-tested sample in at least one target hidden layer of a data classification model; a construction module configured to construct a target activation vector by dividing each numerical component of the original activation vector into a plurality of activation intervals and applying a preset differential disturbance to the numerical components in different activation intervals according to the original activation vector; a generation module configured to generate a test case by adjusting the to-be-tested sample based on a preset distance between the target activation vector and the original activation vector and using a gradient optimization algorithm; and a test module configured to input the test case into the data classification model for reasoning, and determine that the test case is an abnormal decision test case when a prediction result corresponding to the test case is inconsistent with an original prediction result corresponding to the to-be-tested sample.

[0017] The application provides a computer device, comprising a memory and a processor, and the memory stores a computer program, characterized in that the processor implements the steps of the data classification model robustness test method for disturbing hidden layer neuron activation values when executing the computer program.

[0018] The application provides a computer readable storage medium, which stores a computer program, characterized in that the computer program is executed by a processor to implement the steps of the data classification model robustness test method for disturbing hidden layer neuron activation values.

[0019] The application has the following beneficial effects: compared with a method based on random disturbance, the application relieves the problem of insufficient diversity of the method based on random disturbance by performing multi-dimensional disturbance on a plurality of hidden layer neuron activation values of a CNN classification model. The application can not only fully cover each stage of model decision by precisely selecting a plurality of hidden layers and performing diversified disturbance on neuron activation values of each layer, but also ensure that the generated test case has high effectiveness through a gradient optimization strategy. This method does not depend on a predefined simple disturbance mode, but dynamically generates test cases according to the structure of the model and actual requirements, thereby greatly improving the coverage of the test and ensuring that the adversarial disturbance can effectively reveal potential robustness defects of the model.

[0020] Compared with the gradient optimization-based method, the present application improves the limitations of the gradient optimization method for discovering robustness defects through the strategy of multi-dimensional perturbation and comprehensive perturbation of neuron activation values. The existing gradient optimization-based method usually focuses on single neuron or local perturbation, which fails to fully consider the comprehensive influence of the entire hidden layer neuron activation values on model decision, resulting in limited number and types of discovered robustness defects. The present application can more effectively capture the complex interaction effects between neuron activation values by applying perturbations of different amplitudes and directions to comprehensively evaluate the model at multiple levels and dimensions. This strategy not only improves the depth of model robustness testing, but also increases the types and quality of robustness defects, so that the test results are more comprehensive and can more accurately reveal the potential weaknesses of the model when facing complex perturbations. BRIEF DESCRIPTION OF DRAWINGS

[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0022] Figure 1 The principle diagram of the data classification model robustness test method for perturbing hidden layer neuron activation values provided by an embodiment of the present application;

[0023] Figure 2 The principle diagram of the data classification model robustness test method for perturbing hidden layer neuron activation values provided by an embodiment of the present application. DETAILED DESCRIPTION

[0024] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the specific embodiments of the present application will be described in detail below with reference to the drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should be within the scope of protection of the present application.

[0025] Embodiment 1, refer to Figure 1 As the first embodiment of the present application, the embodiment provides a data classification model robustness test method for perturbing hidden layer neuron activation values, comprising:

[0026] S100: obtaining the original activation vector of the at least one target hidden layer of the sample to be tested in the data classification model;

[0027] S200: According to the original activation vector, by dividing the numerical components of the original activation vector into a plurality of activation intervals, and applying a preset differential disturbance to the numerical components in different activation intervals, a target activation vector is constructed;

[0028] S300: Based on the preset distance between the target activation vector and the original activation vector, a gradient optimization algorithm is used to adjust the to-be-tested sample to generate a test case;

[0029] S400: The test case is input into the data classification model for inference, and when the prediction result corresponding to the test case is inconsistent with the original prediction result corresponding to the to-be-tested sample, the test case is determined as an abnormal decision test case.

[0030] It should be noted that the wide application of data classification models such as convolutional neural networks (CNN) in the fields of finance, medicine, etc. makes their robustness a key to ensuring system safety and reliability. However, existing robustness testing methods have obvious shortcomings. On the one hand,

[0031] The method based on random disturbance cannot guarantee effective coverage of all possible neuron activation states due to the randomness and singularity of the disturbance, resulting in insufficient diversity of test cases and limited test coverage. On the other hand,

[0032] The method based on gradient optimization usually only focuses on reducing the classification confidence of the model to the outside, and fails to fully explore the comprehensive influence of the diversified combination of hidden layer neuron activation values inside the model on the decision, resulting in relatively fewer types and quantities of robustness defects discovered.

[0033] Therefore, in view of the above-mentioned problems of insufficient test coverage and insufficient number and quality of discovered defects, through the steps of S100 to S400, the present application proposes a completely new testing paradigm. First, by

[0034] S100 obtains the original activation vector of the key hidden layer inside the model as a reference for manipulating the internal state of the model. Then, in S200, by dividing the activation values into intervals and applying bidirectional, multi-dimensional differential disturbance, a target activation vector with significant differences from the original state but structured is actively constructed, which directly solves the problem of insufficient diversity of random disturbance methods. Subsequently, in S300, the gradient optimization algorithm is used to find input test cases that can make the internal state of the model approach the target activation vector. This strategy of taking the internal activation state as the optimization target can more deeply and comprehensively explore the decision boundary of the model compared to the traditional gradient optimization method that only focuses on the final output, thereby discovering more and more diverse robustness defects. Finally, through S400, the generated test cases are verified and screened, and the adversarial samples that can stably trigger decision errors of the model are efficiently located.

[0035] Embodiment 2, Reference Figure 1 and Figure 2 For the second embodiment of the present application, a data classification model robustness test method for perturbing the activation values of hidden layer neurons is provided.

[0036] S100: Obtain the original activation vector of the at least one target hidden layer of the sample to be tested in the data classification model.

[0037] In the embodiments of the present application, step S100 includes step S101 and step S102:

[0038] Step S101: Calculate the gradient significance and neuron sparsity of each hidden layer in the data classification model to determine the layer importance score of each hidden layer. Step S102: According to the layer importance score, select one or more hidden layers as the target hidden layer.

[0039] Specifically, in step S101, the layer importance score is calculated as follows: First, calculate the gradient significance of a layer by taking the average of the L2 norm of the gradient of the input sample with respect to the activation output of all neurons in the layer. The larger the value, the higher the sensitivity of the layer to the model decision, and the calculation formula is shown in (1).

[0040]

[0041] where Sensitivity l is the sensitivity of the lth layer in the CNN model, N is the total number of neurons in the lth layer of the model, a i is the activation output of the ith neuron in the layer, and x is the input sample.

[0042] Second, calculate the neuron sparsity of the layer, that is, count the number of non-zero activation neurons in the layer and divide by the total number of neurons in the layer, and the calculation formula is as follows (2).

[0043]

[0044] where Sparsity l is the neuron sparsity of the lth layer in the CNN model, N is the total number of neurons in the lth layer of the model, and Num zero represents the number of zero-activation neurons in the layer.

[0045] Finally, multiply the gradient significance by the exponential function value with the neuron sparsity as the parameter to obtain the final comprehensive score, and the calculation formula is shown in (3), and the hidden layer with the highest score is selected as the target.

[0046]

[0047] wherein Comprehensive l represents the comprehensive score of the l-th layer in the CNN model.

[0048] In an alternative embodiment, the method of determining the importance of layers in step S101 can also be based on activation value analysis. For example, the average activation value or the variance of activation values of each hidden layer on a representative dataset can be calculated, and the layer with higher average activation value or larger variance is considered to contain richer information and is more important for model decision.

[0049] It should be noted that obtaining the original activation vector further comprises steps A1 to A3:

[0050] Step A1: if the target hidden layer is a fully connected layer, the neuron activation values thereof are directly spliced as the first activation component;

[0051] Step A2: if the target hidden layer is a three-dimensional hidden layer (such as a convolutional layer or a pooling layer), the activation representative values of the feature maps of each channel thereof are calculated, and the activation representative values are spliced as the second activation component;

[0052] Step A3: the first activation component and the second activation component are fused to generate the original activation vector.

[0053] Specifically, in step A2, the activation representative value of a channel feature map can be obtained by calculating the global average of all activation values on the two-dimensional feature map. In step A3, the activation components from different layers can be aligned in dimension by principal component analysis (PCA) or the like, and then spliced and fused to obtain the final original activation vector.

[0054] S200: according to the original activation vector, a target activation vector is constructed by dividing the numerical components of the original activation vector into a plurality of activation intervals, and applying a preset differential perturbation to the numerical components in different activation intervals.

[0055] In this embodiment, the step of applying a preset differential perturbation in step S200 comprises steps B1 and B2:

[0056] Step B1: dividing the activation interval into a high activation interval, a sub-high activation interval, an intermediate interval and a low activation interval;

[0057] Step B2: applying a zeroing perturbation to the neuron activation values in the high activation interval, a high-amplitude perturbation to the neuron activation values in the low activation interval, and a scaling perturbation to the neuron activation values in the sub-high activation interval and the intermediate interval.

[0058] It should be noted that the activation interval is further divided, which further comprises steps C1 and C2:

[0059] Step C1: determining the division ratio of the plurality of activation intervals according to the depth information of the data classification model;

[0060] Step C2: distributing the neurons sorted by the absolute value of the activation value to the corresponding activation interval according to the division ratio.

[0061] Specifically, in step C1, the division ratio parameter is determined according to a linear function of the total number of layers of the model and the current layer depth, ensuring that the deeper layers in the model have slightly larger high and low activation interval division ratios. The calculation formula is shown in (4).

[0062]

[0063] wherein P top (l), P sec (l), P mid (l), P bottom (l) respectively represent the division ratio parameters of the high activation zone, the second high activation zone, the intermediate zone and the low activation zone, α is the basic division ratio, γ is the depth compensation coefficient, and L is the total number of layers of the model.

[0064] In steps B1 and C2, the neurons in a hidden layer are arranged in descending order according to the absolute value of their activation values, and then the high activation zone, the second high activation zone, the low activation zone and the intermediate zone are divided according to the above calculated division ratio. In step B2, the specific perturbation includes: setting the activation value of the high activation zone to zero; multiplying the activation value of the second high activation zone by a perturbation parameter for scaling; and setting the activation value of the low activation zone to a higher value determined by the maximum activation absolute value of the layer and the perturbation parameter.

[0065] In an optional embodiment, the method of constructing the target activation vector in step S200 can also be based on a genetic algorithm. The original activation vector is used as the initial population, new activation vectors are generated through operations such as crossover and mutation, and the optimal target activation vector is evolved through iteration based on the fitness function of triggering model error classification or reducing confidence.

[0066] S300: adjusting the to-be-tested sample based on the preset distance between the target activation vector and the original activation vector using a gradient optimization algorithm to generate a test case.

[0067] In this embodiment, when there are multiple target hidden layers, adjusting the to-be-tested sample based on the preset distance between the target activation vector and the original activation vector using a gradient optimization algorithm includes steps D1 to D4:

[0068] Step D1: determining the corresponding optimization target for each target hidden layer;

[0069] Step D2: determining the weight coefficient of each target hidden layer according to the gradient sensitivity of the target hidden layer;

[0070] Step D3: weighting and aggregating the multiple optimization targets based on the weight coefficient to obtain a total optimization target;

[0071] Step D4: adjusting the to-be-tested sample according to the total optimization target through back propagation.

[0072] Specifically, in step D1, the optimization target of each hidden layer is the L2 norm distance between the original activation vector and the target activation vector of the hidden layer. In step D2, the weight coefficient is determined by calculating the Frobenius norm of the Jacobian matrix of the input sample with respect to the optimization target of the layer, and comparing it with the sum of the norms of all selected layers, which effectively quantifies the gradient sensitivity of the layer. In step D3, the total optimization target is the weighted sum of the optimization targets of the layers.

[0073] The step of adjusting the to-be-tested sample in step S300 to generate the test case includes steps S301 and S302:

[0074] Step S301: calculating the gradient of the preset distance (i.e., the total optimization target) with respect to the to-be-tested sample. First, forward propagation is performed to obtain the activation output of all hidden layers with respect to the input sample, and then back propagation is performed to obtain the joint gradient, as shown in formula (5);

[0075]

[0076] wherein L l represents the optimization target of the model l layer, a l represents the original activation vector of the lth hidden layer, represents the target activation vector, L total represents the total optimization target, is the back propagation joint gradient, and x represents the input sample.

[0077] Step S302: updating the to-be-tested sample according to the gradient and the preset learning rate, as shown in formula (6), to generate the test case.

[0078]

[0079] wherein x new is the updated sample, x is the original sample, and η is the preset learning rate.

[0080] Specifically, in step S302, the new test case is generated by subtracting a perturbation from the original input sample, and the perturbation is equal to the product of the preset learning rate and the gradient of the total optimization target with respect to the input sample.

[0081] In an alternative embodiment, the gradient optimization algorithm employed in step S300 can be more advanced optimizers such as Adam or RMSprop, which are able to adaptively adjust the learning rate, thus generating high-quality test cases faster and more stably.

[0082] S400: input the test case into the data classification model for inference, and when the prediction result corresponding to the test case is inconsistent with the original prediction result corresponding to the to-be-tested sample, determine that the test case is an abnormal decision test case.

[0083] For example, the specific application of the method of the present application is as follows:

[0084] The present method is applied to three different sizes of data classification models: LeNet-5, ResNet20 and VGG19, which are trained on MNIST, CIFAR-10 and ImageNet data sets, respectively. When testing, a small number of samples are selected from each data set as initial to-be-tested samples (i.e. test seeds). For example, the first sample is extracted from each class in the CIFAR-10 training set used by ResNet20 to form a test seed set with a size of 10.

[0085] For each seed, the process of S100-S300 is performed to generate a test case. For example, for an image of "7" input into the LeNet-5 model, first the key fully connected layer and convolutional layer are selected by S101-S102, and the original activation vector is extracted. Then a perturbed target activation vector is constructed by S200. Then by S300, the original image is updated by backpropagation to minimize the L2 distance between the two vectors, generating a test case that still looks like "7" to the naked eye, but the internal activation state of the model has changed significantly.

[0086] Finally, the newly generated test case is input into the LeNet-5 model (S400), and if the model incorrectly classifies it as "1" or other numbers, the test case is confirmed as an abnormal decision test case, indicating that a robustness defect has been found. Experimental results show that the present method has a significant improvement in neuron coverage, strong neuron activation coverage and other indicators compared to traditional random perturbation and gradient optimization methods, especially on large models such as VGG19, the number of defects found is several to tens of times that of traditional methods.

[0087] In summary, the application selects a key hidden layer (S100), and constructs a target activation vector by structuring bidirectional disturbance of the activation value (S200), and generates a test case capable of matching the internal state by using gradient optimization (S300), and finally screens out samples capable of triggering decision errors of the model (S400). The method systematically solves the problems of insufficient test diversity and incomplete coverage in the prior art, and can more comprehensively and deeply evaluate the robustness of the data classification model and find more potential security defects.

[0088] Embodiment 3, with reference to Figure 1 and Figure 2 , a data classification model robustness test method for disturbing neuron activation values of a hidden layer is provided.

[0089] Step 1: Calculate the importance score of the hidden layer of the CNN classification model, and extract the neuron activation vector of the high-score hidden layer.

[0090] Step 1.1: According to the structure of the CNN classification model M CNN , calculate the gradient significance and neuron sparsity of the activation value of each hidden layer. The gradient significance of the activation output O l of layer l is calculated as formula (1). The gradient significance of the activation output O CNN of layer l is calculated as formula (1).

[0091] Wherein, x is the input seed sample, N is the number of neurons of the hidden layer, is the gradient of the output to the input. Based on the activation value of layer l, the non-zero proportion is calculated to measure the neuron sparsity, and the calculation formula is formula (2).

[0092] Wherein, dim(·) is used to calculate the total dimension of the activation value, and ||·||0 is used to count the non-zero activation. The gradient significance and sparsity entropy are multiplied to highlight the layers that satisfy the gradient significance and high sparsity at the same time, and the comprehensive score of each layer of the model is obtained, and the calculation formula is formula (3).

[0093] Select the two layers with the highest comprehensive score L={l1, l2} as the extraction object.

[0094] Step 1.2: Extract the hidden layer output vector. Input the test seed input into the CNN classification model M CNN , and forward propagate to the selected hidden layers l1 and l2. And extract the output tensor of each hidden layer: the output vector of the fully connected layer containing j neurons is The output tensor of the 3D hidden layer (convolution layer, pooling layer, residual layer) with height, width and channel number h, w, c is

[0095] Step 2: Extract the activation values of the neurons within the corresponding network layer based on the output tensor and calculate the activation vector.

[0096] Step 2.1: For each neuron k (k = 1, 2,..., j) of the fully connected layer, extract its activation value n k based on the activation function σ(·). The calculation formula is shown in equation (4).

[0097] where a k is the weight of the neuron decision function, and b k is the bias. The activation values are concatenated into a vector O fc in order, and the calculation formula is shown in equation (5).

[0098] Standardizing the vectorized neuron activation values avoids numerical instability. The calculation formula is shown in equation (6).

[0099] where α is the mean of the activation values, and β is the standard deviation of the activation values.

[0100] Step 2.2: Calculate the activation vector based on the 3D hidden layer tensor. Split the tensor into c 2D feature maps according to the channel dimension. Calculate the global mean γ v for each channel v's 2D feature map The calculation formula is shown in equation (7).

[0101] Concatenate the mean values of each channel into the activation vector The calculation formula is shown in equation (8).

[0102] Step 2.3: Fuse the multi-hidden layer activation vectors Align the dimensions of the activation vectors from different hidden layers using PCA, and concatenate the results to obtain the global activation vector O. The calculation formula is shown in equation (9).

[0103] where || represents vector concatenation.

[0104] Step 2.4: Verify and optimize the concatenated activation vector to ensure that the dimension of the activation vector O is consistent with the number of hidden layer neurons.

[0105] Step 3: Construct bidirectional perturbations for test case generation based on the activation vector.

[0106] Step 3.1: Calculate the activation vector value for each hidden layer, and arrange the neurons within the lth hidden layer in descending order of absolute value, and adaptively adjust the interval division according to the model depth. The division ratio parameter calculation formula is shown in equation (10).

[0107] where L is the total number of model layers. The neurons can be divided into 4 intervals according to the order: lo represents the total number of neurons, and the neurons in the interval [0:αl lo) the neuron activation absolute value is highest (N l_max ), located in the interval (1-α l ) lo:lo] the neuron activation absolute value is lowest (N l_min ), located in the interval [α l lo:2α l ) the neuron activation absolute value is set to the second highest (N l_submax ), and the rest are (N l_rest ).

[0108] The activation interval is constructed as shown in formula (11), which is divided into low activation, medium-low activation, medium-high activation and high activation intervals based on the numerical size. Wherein sort() represents a function of descending order sorting of absolute value data, O l represents the activation vector of the selected lth hidden layer, lo represents the number of neurons of O l , and n lj represents the activation value of the selected jth neuron of the lth layer. O l After sorting, the neuron activation absolute value with the largest value is

[0109] Step 3.2: In order to reduce the neuron activation value of the neuron with the largest correlation to the model decision, the neuron activation value in N l_max is set to 0 to cause a greater disturbance to the current decision result of the model. In order to discover new hidden layer activation vectors, the neuron activation value in N l_min is set to ± represents a random decision of the positive and negative of the disturbance value, the purpose is to increase the diversity of the disturbance direction of the hidden layer activation vector.

[0110] There are many neurons with small activation absolute value in N l_min , which indicates that the current input may not contain the features that activate such neurons, and adding disturbance to them may be difficult to change the activation value. The neuron activation values in N l_submax and N l_rest are large, and adding disturbance to them is more likely to change the activation value and discover new hidden layer activation vectors. Therefore, the neuron activation values in N l_submax are multiplied by ±e1; in N l_rest , randomly select a group of neurons accounting for 40% of the total number of hidden layer neurons , and multiply their activation values by ±e2. If the selected neuron activation value in N l_rest is 0, it is set to Thus, the disturbed hidden layer activation vector O' lAs shown in equation (12). The three experimental parameters [e1, e2, e3] set here are called disturbance parameters. The setting rule of the disturbance parameters with ideal test effect is that at least one of e1 and e2 is greater than e3.

[0111] Step 3.3: Disturbed hidden layer activation vector O' l The distribution of the test seed activation vector O l in the current layer is disturbed. In order to make the current hidden layer activation vector become O' l , the original activation vector O l and the disturbed activation vector O' l need to be as similar as possible. The difference between O l and O' l is expressed by L2 distance, and the optimization strategy obj l of the lth hidden layer is to minimize the L2 distance between O l and O' l . The abstract “transforming the original activation vector O l into the disturbed activation vector O' l ” is converted into the digital continuous domain “minimizing the L2 distance between O t and O' F ”. The calculation method is shown in equation (13).

[0112] Where n is the number of features of the current activation vector O. Then the gradient sensitivity is analyzed by combining the Jacobian matrix, and the calculation method is shown in equations (14) and (15).

[0113] Where, is each input sample with dimension r, is the optimization target of the input sample. Finally, the optimization targets of multiple hidden layers are weighted and aggregated according to the gradient sensitivity of the layer, and the calculation method is shown in equations (16) and (17).

[0114] Where, ||J t || F is the Frobenius norm of the Jacobian matrix of the tth hidden layer.

[0115] Step 4: Map the disturbed hidden layer activation target to the input space through gradient optimization to generate test cases with directional adversarial characteristics, while ensuring the visual rationality and adversarial effectiveness of the generated samples.

[0116] Step 4.1: Calculate the gradient of the total optimization target on the input seed through back propagation to provide directional guidance for generating adversarial samples. Use automatic differentiation technology to calculate the gradient of the total optimization target obj total on the input with dimension z. The calculation method is shown in equation (18).

[0117] Step 4.2: The gradient value of the optimization target is the same dimension as the input seed, and the gradient value multiplied by the learning rate θ is used as the perturbation. The difference between the input seed and the perturbation is calculated to obtain a new test case input'. The calculation method is shown in formula (19).

[0118] Step 5: Filter the adversarial samples that can stably trigger the classification error of the model through decision boundary crossing detection.

[0119] Step 5.1: Obtain the prediction results of the original input and the generated test case input in the CNN classification model respectively. The calculation method is shown in formula (20).

[0120] Step 5.2: Calculate the original classification confidence drop amplitude to evaluate the screening of high-quality test cases (i.e. ΔP>δ, δ is a test hyperparameter). The confidence calculation method is shown in formula (21).

[0121] At this point, the current seed has completed a round of testing. In order to ensure that each seed fully tests the robustness of the model, the process of hidden layer activation vector construction, bidirectional perturbation construction, test case generation, and abnormal decision test case screening will be repeated multiple times to generate multiple test cases.

[0122] Embodiment 4 is the fourth embodiment of the present application, which provides a data classification model robustness test system for perturbing hidden layer neuron activation values, comprising:

[0123] An acquisition module is configured to acquire an original activation vector of a to-be-tested sample in at least one target hidden layer of a data classification model;

[0124] A construction module is configured to construct a target activation vector by dividing each numerical component of the original activation vector into a plurality of activation intervals and applying a preset differential perturbation to the numerical components in different activation intervals according to the original activation vector;

[0125] A generation module is configured to generate a test case by adjusting the to-be-tested sample based on a preset distance between the target activation vector and the original activation vector using a gradient optimization algorithm;

[0126] A test module is configured to input the test case into the data classification model for inference, and determine that the test case is an abnormal decision test case when the prediction result corresponding to the test case is inconsistent with the original prediction result corresponding to the to-be-tested sample.

[0127] Embodiment 5, which is the fifth embodiment of the present application, is different from the previous four embodiments in that: the function, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or the part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

[0128] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered a list of executable instructions for implementing logic functions, and can be specifically embodied in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, a processor-based system, or other system that can fetch the instructions from a instruction execution system, apparatus, or device and execute the instructions, or in conjunction with such an instruction execution system, apparatus, or device. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport the program for use by or in connection with the instruction execution system, apparatus, or device, or in conjunction with such an instruction execution system, apparatus, or device.

[0129] More specific examples (a non-exhaustive list) of the computer-readable medium include the following: an electrical connection having one or more wires (electrical devices), a portable computer diskette (magnetic devices), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable medium can even be paper or other suitable medium on which the program can be printed, as the program can be electronically obtained, for example, by optical scanning of the paper or other medium, followed by electronic editing, interpretation, or necessary processing, and then stored in a computer memory if necessary. Other suitable media can also be used.

[0130] It should be understood that various parts of the present application can be implemented in hardware, software, firmware, or a combination thereof. In the above-described embodiments, a plurality of steps or methods can be implemented in software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, and as in another embodiment, it can be implemented using any one or a combination of the following technologies known in the art: discrete logic circuitry having logic gates for implementing logic functions on data signals, application specific integrated circuits (ASICs) having appropriate combinational logic gates, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0131] Embodiment 6, which is a sixth embodiment of the present application, provides a data classification model robustness test method for perturbing hidden layer neuron activation values.

[0132] The purpose of this embodiment is to address the robustness test requirements of CNN and other data classification models in data transaction scenarios, improve the problems of insufficient test case diversity and incomplete robustness interval coverage based on the random perturbation method, and make up for the deficiencies of the gradient optimization method in the number and quality of robustness defects found.

[0133] The design principle of the present application is as follows: first, input the test seed into the CNN classification model, select the hidden layer and extract the activation values of the neurons in the layer to form an activation vector; then the bidirectional perturbation construction module divides the neuron activation values in the activation vector into different activation intervals, and modifies the neuron activation values in different intervals using different amplitude perturbation parameters to increase the difference between the hidden layer activation vectors before and after perturbation; in order to make the current activation vector become the perturbed activation vector, the gradient optimization is used to reduce the L2 distance between the perturbed activation vector and the original activation vector of the selected hidden layer, and the back propagation optimization value is used to update the test seed to generate test cases, which are finally input into the model under test to screen test cases that trigger decision errors of the model.

[0134] The technical solution of the present application is implemented by the following steps:

[0135] Step 1: Calculate the importance score of the hidden layer of the CNN classification model and extract the neuron activation values.

[0136] Step 2: According to the structure of the model, divide the neuron activation values of each layer into multiple intervals, and apply different amplitude perturbations to the neurons in each interval.

[0137] Step 3: Use the gradient optimization algorithm to generate test cases according to the perturbation of the neuron activation values.

[0138] Step 4: Input the generated test cases into the CNN model for inference and record the decision results of the model.

[0139] Step 5: Generate a defect report by evaluating the test cases, classifying, and analyzing the robustness defects of the model.

[0140] Embodiment 6, as a sixth embodiment of the present application, provides a data classification model robustness test method for perturbing hidden layer neuron activation values. In order to verify the beneficial effects of the present application, scientific demonstration is carried out through experiments.

[0141] Three public datasets were selected for the experiment, namely MNIST, LeNet-5, and CIFAR-10. The experimental dataset attributes are shown in Table 1. The MNIST dataset and the LeNet-5 model are a classic combination in the field of deep learning, used for handwritten digit recognition tasks. The MNIST dataset is a set of handwritten digit images provided by the National Institute of Standards and Technology (NIST), containing 0 to 9, with a total of 60,000 training samples and 10,000 test samples, each sample being a 28x28 pixel grayscale image. CIFAR-10 is a commonly used computer vision dataset, containing 60,000 32x32 pixel color images, divided into 10 categories, with 6,000 images per category. These categories are: airplane, car, bird, cat, deer, dog, frog, horse, ship, and truck. The ImageNet dataset is a large-scale image database containing over 14 million labeled images, covering more than 20,000 categories, and is an important benchmark in the field of computer vision, used to evaluate the performance of image classification, object detection, semantic segmentation, and other tasks. In the experiment, a subset of the ImageNet dataset, ILSVRC2012, was used, which contains 1,000 categories and 1.2 million training images, 50,000 validation images, and 150,000 test images.

[0142] During the experiment, Lenet5, ResNet20, and VGG19 models were selected as the test models, and the MNIST, CIFAR-10, and ImageNet datasets were used for training. From the 10-class Lenet5 model's MNIST training set, the first 2 samples from each category were extracted to form a test seed set with a scale of 20, and each seed was tested for 5 minutes. From the 10-class ResNet20 model's CIFAR-10 training set, the first training sample from each category was extracted to form a test seed set with a scale of 10, and each seed was tested for 10 minutes. The test seeds for the VGG19 model used 10 different ImageNet image samples, and each test seed was tested for 30 minutes using 3 different methods.

[0143] Table 1. CNN classification model robustness test experimental data attributes

[0144]

[0145] The experimental results are evaluated by neuron coverage (NC), strong neuron activation coverage (SNAC), k-order neuron coverage (KMNC), robustness defect number / use case generation number (ERR / ALL) and robustness defect category (ERRLabel). The calculation method is shown in formulas (22), (23), (24), (25), (26).

[0146]

[0147] wherein neu is one neuron in the model under test, the normalized output activation value of the neuron neu to a seed x in the seed set T is out(neu, x), NEU_NUM is the total number of neurons in the model under test, thre is an activation threshold, |·| represents the total number of set elements. boundmax(neu) represents the maximum activation value of the neuron neu in the training process, is the i-th section of the neuron neu activation value interval, i is located in the interval 1≤i≤k; the seed data set is T, x is a seed, and the model is M CNN , the output category number is kind, and the original label of the seed is y o , M CNN () represents the operation of the model processing the input sample to obtain the classification result, gen() is a test case generation operation, unique() is a de-duplication operation, and sum() is a summation operation.

[0148] Experimental results: The small sample user multi-intent recognition method with enhanced correlation degree calculation performs multi-label user dialogue intent recognition on the samples of TourSG and StanfordLU. The specific results of the experiment are shown in Table 2.

[0149] Table 2. Comparison of experimental results of CNN classification model robustness test of the present application and the comparative method

[0150]

[0151]

[0152] The experimental results show that, compared with all the comparative methods, the test effect of the three CNN classification models with the body volume from small to large in turn, Lenet5, ResNet20 and VGG19, is better, especially the test effect of the large model (VGG19) is greatly improved, the coverage indexes NC, KMNC and SNAC are 22.2%, 46.0% and 24.3% higher than those of the comparative methods respectively, and the found defect types are 6-40 times of those of the comparative methods. It shows that the method of the paper can effectively use the relationship between the neuron activation value in the hidden layer and the model decision to improve the diversity of the test case generation; in addition, changing the neuron activation value in the hidden layer based on the relationship can make the model decision deviate from the correct result, and find more model robustness defects and defect types.

[0153] It should be noted that the above examples are only used to illustrate the technical solutions of the present application and are not limiting. Although the present application has been described in detail with reference to the preferred embodiments, it should be understood by those skilled in the art that the technical solutions of the present application can be modified or replaced equivalently without departing from the spirit and scope of the technical solutions of the present application, and they should be covered in the scope of the claims of the present application.

Claims

1. A robustness testing method for a data classification model by perturbing hidden layer neuron activation values, characterized by: include: Obtaining the original activation vector of at least one target hidden layer of the test sample in the data classification model; According to the original activation vector, a target activation vector is constructed by dividing each numerical component of the original activation vector into a plurality of activation intervals and applying a preset differentiated perturbation to the numerical components in different activation intervals; Based on a preset distance between the target activation vector and the original activation vector, a gradient optimization algorithm is used to adjust the sample to be tested to generate a test case; The test case is input into the data classification model for reasoning. When the prediction result corresponding to the test case is inconsistent with the original prediction result corresponding to the sample to be tested, the test case is determined to be an abnormal decision test case.

2. The robustness testing method for a data classification model using perturbations of hidden layer neuron activation values ​​according to claim 1, wherein: The step of obtaining the original activation vector of at least one target hidden layer of the test sample in the data classification model includes: Calculating the gradient significance and neuron sparsity of each hidden layer in the data classification model to determine the layer importance score of each hidden layer; One or more hidden layers are selected as the target hidden layers according to the layer importance scores.

3. The robustness testing method for a data classification model using perturbations of hidden layer neuron activation values ​​according to claim 1, wherein: The step of obtaining the original activation vector also includes: If the target hidden layer is a fully connected layer, its neuron activation values ​​are directly concatenated as the first activation component; If the target hidden layer is a three-dimensional hidden layer, calculating the activation representative value of each channel feature map thereof, and concatenating the activation representative values ​​into a second activation component; The first activation component is fused with the second activation component to generate the original activation vector.

4. The method for testing robustness of a data classification model by perturbing hidden layer neuron activation values ​​according to claim 1, wherein: The step of applying a preset differentiated disturbance includes: Dividing the activation interval into a high activation interval, a second high activation interval, an intermediate interval, and a low activation interval; A zeroing disturbance is applied to the neuron activation values ​​in the high activation interval, a high-amplitude disturbance is applied to the neuron activation values ​​in the low activation interval, and a scaling disturbance is applied to the neuron activation values ​​in the second-highest activation interval and the middle interval.

5. The method for testing robustness of a data classification model by perturbing hidden layer neuron activation values ​​according to claim 4, wherein: The step of dividing the activation interval further includes: determining a division ratio of the plurality of activation intervals according to depth information of the data classification model; According to the division ratio, the neurons sorted by the absolute magnitude of the activation values ​​are allocated to the corresponding activation intervals.

6. The method for testing robustness of a data classification model by perturbing hidden layer neuron activation values ​​according to claim 1 or 2, wherein: When there are multiple target hidden layers, the step of adjusting the sample to be tested by using a gradient optimization algorithm based on a preset distance between the target activation vector and the original activation vector includes: For each of the target hidden layers, determining the corresponding optimization target; Determining a corresponding weight coefficient according to the gradient sensitivity corresponding to each target hidden layer; Performing weighted aggregation on the multiple optimization objectives based on the weight coefficients to obtain an overall optimization objective; According to the overall optimization goal, the sample to be tested is adjusted through back propagation.

7. The method for testing robustness of a data classification model by perturbing hidden layer neuron activation values ​​according to claim 1, wherein: The step of using a gradient optimization algorithm to adjust the sample to be tested to generate a test case includes: Calculating the gradient of the preset distance with respect to the sample to be tested; The sample to be tested is updated according to the gradient and the preset learning rate to generate the test case.

8. A data classification model robustness testing system for perturbing hidden layer neuron activation values, applying the data classification model robustness testing method for perturbing hidden layer neuron activation values ​​as claimed in any one of claims 1 to 7, characterized in that: include: An acquisition module is used to obtain the original activation vector of at least one target hidden layer of the test sample in the data classification model; a construction module, configured to construct a target activation vector based on the original activation vector by dividing each numerical component of the original activation vector into a plurality of activation intervals and applying a preset differentiated perturbation to the numerical components in different activation intervals; A generating module, configured to adjust the sample to be tested using a gradient optimization algorithm based on a preset distance between the target activation vector and the original activation vector to generate a test case; The test module is used to input the test case into the data classification model for reasoning, and when the prediction result corresponding to the test case is inconsistent with the original prediction result corresponding to the sample to be tested, determine that the test case is an abnormal decision test case.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the data classification model robustness testing method of perturbing hidden layer neuron activation values ​​according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for robustness testing of a data classification model by perturbing hidden layer neuron activation values ​​according to any one of claims 1 to 7 are implemented.