Machine learning-based adversarial attack sample generation method
Through the multi-link collaboration mechanism and dynamic curvature field model, dynamic adaptation between adversarial sample generation and target model defense strategy is achieved, the decision-making boundary mismatch problem is solved, and the success rate of migration attacks is improved.
Patent Information
- Application Number
- CN202510590884.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-06-06
AI Technical Summary
The existing adversarial sample generation method cannot perceive the dynamic changes in the decision boundary of the target model in real time, resulting in the mismatch between the gradient direction of the alternative model and the real decision boundary direction of the target model. The generated adversarial samples are intercepted by the defense mechanism during migration attacks.
Through a multi-link collaboration mechanism, probe samples are used to generate overlay multi-dimensional perturbation features, establish a dynamic curvature field model and normal vector field, ensure that the gradient direction is orthogonal to the decision boundary, generate adversarial samples, and iteratively optimize perturbation through closed-loop feedback.
It solves the problem of decision-making boundary mismatch, improves the migration success rate of black box attacks in dynamic defense environments, and ensures that the adversarial samples can effectively bypass the defense mechanism of the target model.
Smart Images

Figure CN120106147A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of simulated adversarial attack, and more specifically, to a method for generating adversarial attack samples based on machine learning. Background Art
[0002] In the adversarial sample generation task, the attacker usually simulates the decision-making behavior of the target system by constructing a substitute model and generates adversarial perturbations based on the gradient of the substitute model. However, the actual defense mechanism of the target model (such as gradient masking and output randomization) will dynamically change the shape of its decision boundary, resulting in the inability of the substitute model to accurately capture the sensitive areas of the target model to the input features. Existing methods rely on static substitute models to generate adversarial perturbations, which cannot perceive the dynamic changes of the decision boundary of the target model in real time, resulting in a mismatch between the gradient direction of the substitute model and the true decision boundary direction of the target model. The generated adversarial samples are intercepted by the defense mechanism due to the deviation in direction during the migration attack. Summary of the invention
[0003] In order to overcome the above-mentioned defects of the prior art, an embodiment of the present invention provides an adversarial attack sample generation method based on machine learning, which realizes dynamic adaptation of adversarial sample generation and target model defense strategy through a multi-link collaborative mechanism; probe sample generation covers multi-dimensional perturbation features to provide data support for decision boundary curvature modeling, which drives the replacement model update through deformation quantization and normal vector field to ensure that the gradient direction is orthogonal to the decision boundary; orthogonal gradient projection utilizes update parameters and normal vector field to generate adversarial samples, and optimizes perturbations through orthogonality evaluation to match sensitive areas; closed-loop feedback iteration returns successful samples to the probe data set, adjusts the curvature field and model parameters, and forms a closed-loop data flow of sample generation, modeling, update and optimization; solves the decision boundary mismatch problem, and improves the migration success rate of black box attacks in a dynamic defense environment, so as to solve the problems raised in the above background technology.
[0004] To achieve the above object, the present invention provides the following technical solutions: The method for generating adversarial attack samples based on machine learning includes the following steps: Generate probe samples of multi-directional disturbances based on uniform spherical sampling, combine exponential amplitude increment design and collect target model responses in real time to construct a probe dataset containing the difference between input features and normalized responses; Based on the probe data set, the output response differences of adjacent probe samples are calculated, and a dynamic curvature field model is established by fitting the gradient change rate of the response differences to quantify the local deformation characteristics of the decision boundary of the target model. Using a two-layer surrogate model architecture, the meta-learner generates task network weight correction coefficients based on the normal vector field, and updates the task network parameters through online back-propagation combined with orthogonal loss to make the gradient direction orthogonal to the decision boundary; Calculate the adversarial gradient based on the updated task network parameters and use the curvature field normal vector for orthogonal decomposition to generate adversarial samples. Evaluate the orthogonality and make a decision based on the adversarial gradient features and the decision boundary curvature features to decide whether to pass it to the next step or regenerate it after adjusting the perturbation amplitude. The adversarial samples are submitted to the target model to verify the effectiveness of the attack. The adversarial samples that successfully bypass the defense and their corresponding target model responses are injected into the probe dataset as new probe data, triggering the meta-learner to recalculate the task network weight correction coefficient, forming a closed-loop data flow from attack verification to model update.
[0005] In a preferred embodiment, the probe data set acquisition process is: In the input space of the target model, with the original sample as the center, a multi-directional disturbance vector is generated by the uniform spherical sampling method to ensure that the disturbance direction is evenly distributed on the sphere; for each disturbance vector, an initial minimum amplitude value is set, and a series of amplitude values are calculated according to the exponential growth law to generate a set of probe samples; the probe samples are input into the target model one by one, the output response of the target model to each probe sample is recorded, and the difference between it and the output response of the original sample is calculated; the difference is normalized to obtain the normalized response difference; the input features of the probe samples are paired with the corresponding normalized response differences to construct a probe data set.
[0006] In a preferred embodiment, the construction process of the dynamic curvature field model is: Based on the probe data set, the output response differences of adjacent probe samples are calculated, and the gradient change rate of the response differences is fitted; the response difference change rate in each disturbance direction is smoothed by the local weighted regression method to generate a gradient change rate function; based on the gradient change rate function, the curvature estimate at each probe sample is calculated by the finite difference method; the discrete curvature estimate is converted into a continuous dynamic curvature field model using the radial basis function interpolation method; the gradient of the dynamic curvature field at each probe sample point is calculated by the numerical differentiation method and normalized to the unit normal vector.
[0007] In a preferred embodiment, the alternative model dynamic update method is: A two-layer substitution model architecture is adopted, in which the task network is used to simulate the decision-making behavior of the target model, and the meta-learner dynamically adjusts the parameters of the task network according to the normal vector field generated by the decision boundary curvature modeling step; specifically, the meta-learner takes the decision boundary normal vector field as input and generates the weight correction coefficient of the task network through a multi-layer perceptron structure; subsequently, the task network updates its convolution kernel parameters based on online back propagation, and the loss function used in the updating process is composed of prediction loss and orthogonal loss, where the orthogonal loss is used to constrain the gradient direction of the task network to remain orthogonal to the decision boundary normal vector.
[0008] In a preferred embodiment, the process of generating adversarial samples by intersection decomposition is as follows: First, based on the updated task network parameters and input samples, the adversarial gradient is calculated, where the adversarial gradient is obtained by derivation of the prediction loss of the task network with respect to the input samples; then, the curvature field normal vector generated by the decision boundary curvature modeling step is used to perform orthogonal decomposition on the adversarial gradient to generate orthogonal components that are orthogonal to the curvature field normal vector; then, adversarial samples are generated based on the orthogonal components, which are specifically completed by normalizing and scaling the orthogonal components and then adding them to the input samples.
[0009] In a preferred embodiment, orthogonality is evaluated and a decision is made: The adversarial gradient features include the gradient orthogonality index, and the decision boundary curvature features include the curvature sensitivity index; the gradient orthogonality index is calculated by the dot product and norm of the adversarial gradient and the curvature field normal vector, and the curvature sensitivity index is calculated by the ratio of the curvature value of the curvature field to the adversarial gradient norm; then, the gradient orthogonality index and the curvature sensitivity index are fused through the comprehensive evaluation coefficient. If the comprehensive evaluation coefficient is greater than the preset threshold, the adversarial sample is output to the closed-loop feedback iteration step; the perturbation amplitude is multiplied by the preset value, the adversarial sample is regenerated and the evaluation process is repeated.
[0010] In a preferred embodiment, the process of verifying the effectiveness of the attack is: The adversarial sample generated by the orthogonal gradient projection step is submitted to the target model, and the output response of the target model to the adversarial sample is obtained and compared with the output response of the original sample to verify the effectiveness of the attack. For classification tasks, check whether the predicted category of the adversarial sample is different from the predicted category of the original sample. For regression tasks, check whether the deviation between the output response of the adversarial sample and the output response of the original sample exceeds the preset threshold. If any condition is met, the attack is considered successful.
[0011] In a preferred embodiment, the closed-loop data flow formation process is: If the attack is successful, the adversarial sample and its normalized response difference with the original sample are injected into the probe dataset of the probe sample generation and response collection steps; then, the dynamic curvature field model and normal vector field are recalculated based on the updated probe dataset, and the updated normal vector field is input into the meta-learner, and the weight correction coefficient of the task network is calculated by integrating the normal vector field and the gradient of the task network loss function parameters, thereby adjusting the task network parameters; finally, the next round of adversarial samples is generated based on the updated task network parameters, and the verification and feedback process is repeated to form a closed-loop data flow.
[0012] The technical effects and advantages of the method for generating anti-attack samples based on machine learning of the present invention are as follows: The present invention realizes the dynamic adaptation of adversarial sample generation and target model defense strategy through a multi-link collaborative mechanism; probe sample generation covers multi-dimensional perturbation features to provide data support for decision boundary curvature modeling, which drives the replacement model update through deformation quantization and normal vector field to ensure that the gradient direction is orthogonal to the decision boundary; orthogonal gradient projection uses update parameters and normal vector field to generate adversarial samples, and optimizes perturbations through orthogonality evaluation to match sensitive areas; closed-loop feedback iteration returns successful samples to the probe data set, adjusts the curvature field and model parameters, and forms a closed-loop data flow of sample generation, modeling, update and optimization; each link promotes each other through data interaction, and the probe data set update and curvature field and model parameter optimization promote each other, solves the decision boundary mismatch problem, and improves the migration success rate of black box attacks in a dynamic defense environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 It is a structural diagram of the method for generating adversarial attack samples based on machine learning of the present invention. DETAILED DESCRIPTION
[0014] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0015] Embodiment 1: Figure 1 The present invention provides a method for generating adversarial attack samples based on machine learning, including: S1: Generate probe samples with multi-directional perturbations based on uniform spherical sampling, combine exponential amplitude increment design and collect target model responses in real time, and construct a probe dataset containing the differences between input features and normalized responses.
[0016] S2: Based on the probe dataset, the output response differences of adjacent probe samples are calculated, and a dynamic curvature field model is established by fitting the gradient change rate of the response differences to quantify the local deformation characteristics of the decision boundary of the target model.
[0017] S3: Using a two-layer surrogate model architecture, the meta-learner generates task network weight correction coefficients based on the normal vector field, and updates the task network parameters through online backpropagation combined with orthogonal loss to make the gradient direction orthogonal to the decision boundary.
[0018] S4: Calculate the adversarial gradient based on the updated task network parameters and use the curvature field normal vector for orthogonal decomposition to generate adversarial samples. Evaluate the orthogonality and make a decision based on the adversarial gradient features and the decision boundary curvature features to decide whether to pass it to the next step or regenerate it after adjusting the perturbation amplitude.
[0019] S5: Submit the adversarial samples to the target model to verify the effectiveness of the attack, inject the adversarial samples that successfully bypass the defense and their corresponding target model responses as new probe data into the probe dataset, triggering the meta-learner to recalculate the task network weight correction coefficient, and forming a closed-loop data flow from attack verification to model update.
[0020] In the adversarial sample generation task, the defense mechanism of the target model will dynamically adjust its decision boundary, making it difficult for traditional static attack strategies to effectively capture the target model's sensitive areas to input perturbations. To overcome this challenge, step S1 aims to construct a probe dataset that dynamically reflects the characteristics of the target model's decision boundary by generating multi-directional perturbation samples (probe samples) and collecting the target model's response in real time. This step provides a data basis for the subsequent decision boundary curvature modeling (step S2), ensuring that the attack strategy can adapt to the dynamic changes of the target model.
[0021] The target model refers to a prediction system built based on machine learning, that is, a black box model, whose internal structure and parameters are invisible to attackers. It only interacts with the outside through input and output interfaces and is used to perform specific tasks, such as image classification, speech recognition or natural language processing. In the adversarial sample generation task, the target model has a dynamic defense mechanism and can adjust its decision boundary through gradient masking, output randomization, etc. to resist external perturbations or attacks. The attacker's goal is to detect the response characteristics of the target model through probe samples, and then generate adversarial samples to mislead its prediction results. The dynamic nature of the target model makes it difficult for traditional static attack methods to accurately capture its sensitive areas. Therefore, the multi-step collaborative mechanism in the present invention is required to achieve dynamic adaptation and effective attack.
[0022] Step S1 includes the following contents: S1.1, Probe sample generation: In the input space of the target model, a multi-directional perturbation sample (i.e., probe sample) is generated by a uniform spherical sampling method, and the size of the perturbation is designed in combination with an exponential amplitude increase method. Specifically, first, a spherical range is defined in the input space with the original sample as the center. Then, multiple direction points are selected on the sphere in a uniformly distributed manner, and each direction point corresponds to a perturbation vector to ensure that the perturbation direction is evenly distributed on the sphere. Subsequently, for each perturbation vector, an initial minimum amplitude value is set, and a series of amplitude values are calculated according to the exponential growth law, for example, the current amplitude value is multiplied by a fixed growth factor each time until the preset maximum amplitude value is reached. Finally, the direction of each perturbation vector is combined with the corresponding multiple amplitude values to generate a set of probe samples.
[0023] This method ensures that the directions of the probe samples cover all possible directions in the input space, avoiding detection bias caused by uneven direction selection. At the same time, the exponential amplitude increase design allows the perturbation amplitude to gradually expand from a small value to a larger value, covering perturbation ranges of different scales. This generation method can not only finely detect the response characteristics of the target model in a local range, but also reflect the impact of perturbations in a global range, providing comprehensive and diverse sample data for subsequent analysis of the decision boundary characteristics of the target model.
[0024] S1.2, Response Collection: The generated probe samples are input into the target model one by one, the output response of the target model to each probe sample is recorded, and the difference between it and the output response of the original sample is calculated, and then these differences are normalized. Specifically, the original sample is first input into the target model to obtain its output response as the benchmark value. Then, each probe sample is input into the target model in turn, the corresponding output response is recorded, and a set of difference values are calculated by subtracting the output response of the original sample. To ensure data consistency, normalization is used, that is, a normalization factor is selected (the maximum absolute value or average of all difference values), each difference value is divided by the normalization factor, and the numerical range of the difference value is adjusted to a uniform scale.
[0025] By calculating the difference, we can quantify the impact of each set of disturbances on the output of the target model and reveal the sensitivity of the target model to specific disturbances. Normalization eliminates the dimensional differences in the responses of different probe samples due to the different output ranges of the target model, ensuring the comparability of data in subsequent analysis.
[0026] S1.3, Probe dataset construction: The generated probe samples are paired with their corresponding normalized response difference values to construct a probe dataset, which contains the correspondence between the input features of the probe samples and the normalized response differences. Specifically, each probe sample is used as an input data point, and its input feature data is extracted and associated with the corresponding normalized response difference value to form a data record. All records are integrated in a unified format to generate a structured probe dataset.
[0027] Through data pairing and structured organization, the disturbance information of the probe sample and the response information of the target model are fully preserved and effectively transmitted. The construction of the probe data set provides a clear data basis for the subsequent decision boundary characteristic analysis.
[0028] S1.4, Data Transfer: The constructed probe dataset is directly transferred to the subsequent step S2 through a data interface or storage mechanism as input data for decision boundary curvature modeling. Specifically, the integrity of the probe dataset is confirmed before transfer to ensure that all probe samples and their normalized response differences are correctly recorded. Subsequently, the probe dataset is transferred from the output end of the current step to the input end of step S2 through a predefined data transmission path.
[0029] Step S1 generates probe samples covering the input space through uniform spherical sampling and exponential amplitude increment, quantifies the response characteristics of the target model through response acquisition and normalization processing, and constructs a structured probe data set, and finally ensures information continuity through data transmission. This processing logic not only provides a solid data foundation for subsequent decision boundary curvature modeling, but also improves the pertinence and effectiveness of adversarial sample generation, reflecting the systematicness and practicality of the technology.
[0030] The goal of step S2 is to establish a dynamic curvature field model and a normal vector field based on the probe data set by analyzing the response differences of adjacent probe samples and fitting their gradient change rates, so as to quantify the local deformation characteristics of the decision boundary of the target model, and output these models to steps S3 and S4 to provide key geometric information for subsequent dynamic updating of alternative models and orthogonal gradient projection.
[0031] Step S2 includes the following contents: S2.1, calculate the output response difference of adjacent probe samples: In the probe data set, for each disturbance direction, probe samples with different amplitudes in the same direction and their corresponding normalized response differences are extracted. Next, the rate of change of the response difference between adjacent amplitudes is calculated, specifically by subtracting the normalized response difference of adjacent samples to obtain the rate of change value. By analyzing the change in the response difference of adjacent probe samples, the impact of the disturbance amplitude on the output of the target model is quantified. The rate of change of the response difference can reveal the change in the sensitivity of the target model to the disturbance amplitude, and provide basic data for subsequent curvature modeling. It can clearly capture the trend of the model output changing with the disturbance amplitude, laying the foundation for the analysis of the geometric characteristics of the decision boundary.
[0032] S2.2, gradient change rate of the fitted response difference: For each disturbance direction, the disturbance amplitude is used as the independent variable and the response difference change rate is used as the dependent variable. The local weighted regression method is used for fitting to generate a gradient change rate function. Local weighted regression emphasizes the contribution of local data by assigning weights near each data point. The window width can be set to one-third of the amplitude series to balance the smoothness and local sensitivity of the fitting curve. Using local weighted regression to smooth data and extract change trends can effectively reduce the impact of data noise, accurately estimate the change law of response differences with disturbance amplitude, and provide a reliable basis for curvature calculation. The fitted function can extract continuous change characteristics from discrete data and improve the accuracy of subsequent analysis.
[0033] S2.3, calculate the curvature estimate: Based on the fitted gradient rate function, the curvature estimate at each probe sample is determined by calculating its derivative. Specifically, the finite difference method is used to approximate the derivative. The curvature value is obtained by calculating the difference between the gradient rates of adjacent amplitudes and dividing it by the corresponding amplitude difference. The local change of the gradient rate of change is quantified by derivative calculation to obtain the curvature estimate. The curvature can reflect the local curvature of the decision boundary of the target model and help identify the sensitive areas of the model to disturbances. The curvature value provides a direct description of the geometric characteristics of the decision boundary, providing key information for model optimization and adversarial sample generation.
[0034] S2.4, construct dynamic curvature field model: All probe samples and their corresponding curvature estimates form a data set, and a continuous curvature field model is constructed using radial basis function interpolation. During the interpolation process, a Gaussian kernel is selected as the basis function, and the bandwidth parameter can be set to 1.5 times the average Euclidean distance between probe samples to ensure the smoothness and local adaptability of the interpolation. The weights are determined by solving a system of linear equations so that the interpolation function accurately matches the curvature value at the probe sample point. The radial basis function interpolation is used to convert discrete curvature values into a continuous curvature field. This method can effectively process high-dimensional data, provide a smooth and continuous curvature distribution, and facilitate curvature queries at any input point. The constructed curvature field model supports a comprehensive analysis of the decision boundary.
[0035] S2.5, calculate the curvature field normal vector: For each probe sample, the gradient of the curvature field at that point is calculated by numerical differentiation. Specifically, along each dimension of the input space, the difference between the curvature field before and after a small perturbation is calculated and divided by the perturbation step to obtain the gradient component. The perturbation step can be set to 0.001 to balance the calculation accuracy and numerical stability. Subsequently, the gradient vector is normalized to the unit normal vector. The directional information of the curvature field is calculated by numerical differentiation and normalization. The normal vector provides the directional characteristics of the curvature change, which is instructive for orthogonal gradient projection, because the normal vector can clarify the relationship between the perturbation direction and the sensitive area of the decision boundary, thereby optimizing the generation effect of the adversarial perturbation.
[0036] The constructed dynamic curvature field model and normal vector field are output to the subsequent steps. The curvature field model and normal vector field are passed to the subsequent process to guide the dynamic update of the alternative model and the calculation of the orthogonal gradient projection. Ensure the efficient transmission and utilization of the curvature field and normal vector information. The subsequent steps need to directly rely on this information for model parameter adjustment and adversarial sample generation.
[0037] Based on the probe data set provided by step S1, step S2 successfully constructs a dynamic curvature field model and a normal vector field by calculating the response differences of adjacent probe samples and fitting their gradient change rates. The curvature field model quantifies the local deformation characteristics of the decision boundary of the target model, and the normal vector field provides geometric direction information of the curvature change. The output results directly support the dynamic update of the alternative model in step S3 and the orthogonal projection of the adversarial gradient in step S4, ensuring that the subsequent steps can accurately adjust the model parameters and generate adversarial samples based on the local geometric characteristics of the decision boundary.
[0038] Step S3 receives the normal vector field output by step S2, and aims to dynamically update the substitution model parameters through a two-layer substitution model architecture so that the gradient direction of the substitution model is orthogonal to the current decision boundary of the target model, thereby providing an accurate adversarial gradient calculation basis for step S4 and improving the success rate of migration attacks on adversarial samples.
[0039] Step S3 includes the following contents: S3.1, two-layer substitution model architecture design: In the dynamic update process of the surrogate model, a two-layer architecture design is adopted, including two core components: the meta-learner and the task network. The task network is a neural network whose structure is similar to the target model. The initial parameters are set as the convolution kernel parameters to simulate the decision-making behavior of the target model. As an independent auxiliary network, the meta-learner is responsible for dynamically adjusting the parameters of the task network according to the normal vector direction of the decision boundary. Specifically, the meta-learner generates adaptive correction coefficients in real time by analyzing the normal vector of the decision boundary, which are applied to the parameters of the task network so that it can quickly adapt to changes in the decision boundary of the target model. In terms of computational logic, the meta-learner first receives the normal vector information of the decision boundary, outputs the adjustment coefficients after internal processing, and then applies these coefficients to the initial parameters of the task network to achieve dynamic update of the parameters. By separating the meta-learner from the task network, it is possible to efficiently complete parameter adjustment while maintaining the stability of the prediction ability of the task network.
[0040] The design of separating the meta-learner and the task network can avoid direct interference of parameter adjustment on the prediction performance of the task network, while using the independence of the meta-learner to improve the adjustment efficiency. This architecture enhances the adaptability of the surrogate model to changes in the defense mechanism of the target model, ensuring that the gradient information of the surrogate model is consistent with the decision boundary of the target model, thereby improving the robustness and accuracy of the overall system.
[0041] S3.2, the meta-learner generates weight correction coefficients: The meta-learner is responsible for generating the weight correction coefficient of the task network. Its input is the normal vector of the decision boundary, and its output is the correction coefficient with the same dimension as the task network parameter. The meta-learner adopts a multi-layer perceptron structure, in which the input layer receives the normal vector of the decision boundary, the hidden layer contains 128 processing units and uses the ReLU activation function for feature extraction, and the output layer generates weight correction coefficients that match the task network parameters. The calculation logic is to first input the normal vector into the multi-layer perceptron, and through the feature extraction of the hidden layer, the geometric information of the normal vector is converted into an intermediate representation of parameter adjustment, and finally mapped to the weight correction coefficient by the output layer. Using the nonlinear mapping capability of the multi-layer perceptron, the geometric characteristics of the decision boundary are converted into specific adjustment amounts of the task network parameters, thereby achieving dynamic adaptation.
[0042] The normal vector of the decision boundary directly reflects its local geometric characteristics. By learning the mapping relationship from the normal vector to the parameter adjustment, the meta-learner can provide accurate correction guidance for the task network. This method establishes a direct connection between the geometric characteristics of the decision boundary and the model parameters, improves the adaptation accuracy of the surrogate model in a dynamic environment, and ensures that the parameter adjustment is highly consistent with the target model behavior.
[0043] S3.3, online back propagation updates the task network parameters: The parameters of the task network are updated through online back-propagation, and its loss function consists of two parts: prediction loss and orthogonal loss. The prediction loss is calculated based on the difference between the prediction result of the task network for the input sample and the output of the target model. The cross entropy loss function is used to drive the task network to approach the behavior of the target model by measuring the gap between the two. The orthogonal loss is used to constrain the orthogonality of the gradient direction of the task network and the normal vector of the decision boundary. The calculation method is to first calculate the dot product of the gradient of the task network and the normal vector, and then take the square of the result. The smaller the result, the closer the two are to orthogonality. The total loss is defined as the weighted sum of the prediction loss and the orthogonal loss, where the weight coefficient of the orthogonal loss can be set to 0.1. The parameter update process is divided into two steps: first, the weight correction coefficient generated by the meta-learner is applied to the task network parameters; then, the gradient of the total loss to the task network parameters is calculated, and after multiplying the gradient by the learning rate, this result is subtracted from the current parameters to complete an update. The gradient direction of the task network is guided by the orthogonal loss to keep it valid in the high curvature area of the decision boundary. The introduction of orthogonal loss can ensure that the gradient of the task network is orthogonal to the decision boundary normal vector, so that the adversarial perturbation is still effective under the defense mechanism of the target model and avoids being intercepted. It improves the success rate of migration attack of adversarial samples, especially when the target model adopts a dynamic defense strategy, which can significantly enhance the attack performance of the alternative model.
[0044] The updated task network parameters are directly passed to the subsequent step S4.
[0045] In the adversarial sample generation task, step S4 receives the updated task network parameters output by step S3 and the curvature field model normal vector output by step S2. Based on this, step S4 generates adversarial samples by calculating the adversarial gradient and combining the normal vector for orthogonal decomposition. Subsequently, the gradient orthogonality index and curvature sensitivity index are used to evaluate the orthogonality and make a decision to decide whether to pass the adversarial sample to step S5 for verification or to regenerate it after adjusting the perturbation amplitude to ensure the orthogonality of the decision boundary between the adversarial sample and the target model and the effectiveness of the attack.
[0046] Step S4 includes the following contents: S4.1, calculate the adversarial gradient: In the orthogonal gradient projection process of step S4, the adversarial gradient is first calculated based on the updated task network parameters and input samples. Specifically, the task network is used to predict the input sample to generate a prediction loss, which represents the difference between the task network prediction result and the target model output. Then, by taking the gradient of the prediction loss with respect to the input sample, the adversarial gradient is calculated, which reflects the direction in which the perturbation in the input space affects the prediction of the task network. In terms of computational logic, the prediction loss is first calculated through the task network, and then the prediction loss is differentiated to generate the corresponding gradient vector. The adversarial gradient provides initial guidance for the subsequent perturbation direction and is the basis for orthogonal decomposition. As the core element of generating adversarial samples, the adversarial gradient can directly reflect the sensitivity of the task network to the input perturbation. By accurately calculating the adversarial gradient, the direction that has the greatest impact on the prediction of the task network can be identified, laying a solid foundation for generating adversarial samples.
[0047] S4.2, orthogonal decomposition against gradient: After calculating the adversarial gradient, the curvature field normal vector is used to perform an orthogonal decomposition on it. In specific implementation, the projection component of the adversarial gradient in the direction of the normal vector is first calculated, that is, the dot product of the adversarial gradient and the normal vector is divided by the square of the normal vector norm, and then the result is multiplied by the normal vector to obtain the projection component. Next, the projection component is subtracted from the adversarial gradient to generate an orthogonal component, which is perpendicular to the normal vector. The calculation logic is to decompose the adversarial gradient into a part parallel to the normal vector and a part perpendicular to the normal vector through a projection operation, and only retain the vertical part as the orthogonal component. In addition, it is ensured that the perturbation direction is orthogonal to the decision boundary normal vector, thereby avoiding the defense mechanism of the target model and improving the concealment of the adversarial sample. Orthogonal decomposition can effectively separate the sensitive direction in the adversarial gradient, so that the perturbation moves along the tangent of the decision boundary to avoid triggering the defense response of the target model. This decomposition method significantly improves the attack success rate of adversarial samples, especially when the target model adopts a dynamic defense strategy, it can enhance the effectiveness of the perturbation.
[0048] S4.3, Generate adversarial samples: Based on the orthogonal components, the system generates adversarial samples. Specifically, the orthogonal components are first normalized to unit vectors, that is, the unit direction is obtained by dividing the orthogonal components by their norms, and then the unit direction is multiplied by the preset perturbation amplitude to generate a perturbation vector. Subsequently, the perturbation vector is added to the original input sample to generate an adversarial sample. Specifically, the orthogonal components are first normalized to ensure the accuracy of the perturbation direction, and then the perturbation size is controlled by scaling, and finally added to the original input sample to generate a new sample. By controlling the perturbation direction and amplitude, it is ensured that the adversarial sample maximizes the impact on the task network prediction while maintaining the characteristics of the original sample as much as possible. Normalization and scaling operations can accurately control the direction and intensity of the perturbation, making the adversarial sample both aggressive and hidden.
[0049] S4.4, calculate the gradient orthogonality index and curvature sensitivity index: The adversarial gradient features include the gradient orthogonality index, and the decision boundary curvature features include the curvature sensitivity index. In order to evaluate the quality of adversarial samples, the gradient orthogonality index and curvature sensitivity index are calculated. The gradient orthogonality index is calculated by first calculating the dot product of the adversarial gradient and the normal vector, and then dividing it by the product of the norms of the two to obtain the normalized dot product. After taking its absolute value, it is subtracted from 1 to generate the gradient orthogonality index. This index reflects the degree of orthogonality between the adversarial gradient and the normal vector. The larger the value, the more orthogonal. The curvature sensitivity index is calculated by dividing the curvature value of the curvature field at the input sample by the sum of the adversarial gradient norm and the smoothing factor to generate the index. This index reflects the relative influence of curvature on the gradient norm. The larger the value, the higher the dependence of the perturbation on the curvature sensitive area. Specifically, the gradient orthogonality index quantifies the cosine similarity based on the dot product and the norm, and the curvature sensitivity index evaluates the relationship between curvature and gradient through the ratio. These two indices are used to comprehensively evaluate the direction and curvature sensitivity of the disturbance, ensuring that the adversarial samples are both orthogonal to the decision boundary and avoid high curvature areas. The gradient orthogonality index and curvature sensitivity index evaluate orthogonality and make decisions from the two dimensions of direction and curvature sensitivity, respectively, providing a comprehensive standard. The benefit is that this evaluation method can screen out adversarial samples and improve the success rate and concealment of attacks.
[0050] S4.5, Evaluate orthogonality and make decisions: The system makes decisions on adversarial samples through comprehensive evaluation coefficients. Specifically, the comprehensive evaluation coefficient is first calculated. The coefficient is generated by the exponential decay fusion of the gradient orthogonality index and the curvature sensitivity index, that is, the gradient orthogonality index is multiplied by the negative exponential function with the curvature sensitivity index as the base and the adjustment factor of 0.1. Subsequently, if the comprehensive evaluation coefficient is greater than the preset threshold, the adversarial sample is output to the closed-loop feedback iteration step; if it is less than or equal to the preset threshold, the perturbation amplitude is multiplied by the preset value, the adversarial sample is regenerated and the evaluation process is repeated. In terms of computational logic, the two exponents are fused through exponential decay to balance high orthogonality and low curvature sensitivity, and the decision is made based on threshold comparison. The perturbation is dynamically adjusted using the comprehensive evaluation coefficient to ensure that the quality of the adversarial sample meets the expected standard. The comprehensive evaluation coefficient integrates the influence of orthogonality and curvature sensitivity, avoids the limitations of a single indicator, and improves the accuracy of decision-making. The advantage is that this dynamic adjustment mechanism ensures high-quality output of adversarial samples.
[0051] After the evaluation passes, the adversarial sample is passed to step S5.
[0052] Step S5 receives the adversarial sample output by step S4, submits the sample to the target model and verifies its attack effectiveness, and transmits the adversarial sample that successfully bypasses the defense and its corresponding target model response back to the probe data set of step S1, triggering the meta-learner to recalculate the weight correction coefficient of the task network, forming a closed-loop feedback iteration mechanism. This mechanism aims to enhance the adaptability of the adversarial sample generation process to changes in the decision boundary of the target model by dynamically updating the probe data set and task network parameters, thereby improving the attack success rate.
[0053] Step S5 includes the following contents: S5.1, submit adversarial samples and verify the effectiveness of the attack: First, the adversarial sample is submitted to the target model, the output response of the target model to the adversarial sample is obtained, and it is compared with the output response of the original sample to verify the effectiveness of the attack. Specifically, for classification tasks, check whether the predicted category of the adversarial sample is different from the predicted category of the original sample; for regression tasks, check whether the deviation between the output response of the adversarial sample and the output response of the original sample exceeds the preset threshold. If any of the above conditions is met, the attack is considered successful. In addition, two key indicators are further calculated: the gradient orthogonality index and the curvature sensitivity index. By comparing the output responses and calculating the two indices, the attack effectiveness of the adversarial sample is comprehensively evaluated. In terms of processing, it not only relies on the difference in output responses, but also combines the geometric characteristics of the decision boundary (such as orthogonality and curvature changes) for verification to ensure the comprehensiveness and accuracy of the evaluation. By comprehensively evaluating the attack effect by combining output differences and geometric characteristics, truly effective adversarial samples can be identified more accurately, ensuring that the quality of the screened adversarial samples is higher.
[0054] S5.2, feedback successful adversarial examples to the probe dataset: After verifying that the attack is successful, the adversarial sample that successfully bypasses the target model defense and its corresponding output response are fed back to the probe dataset to update the dataset content. Specifically, the normalized response difference of the adversarial sample is first calculated. The calculation method is to divide the difference between the output response of the adversarial sample and the output response of the original sample by the maximum value of the difference between the output response of all samples in the probe dataset and the output response of the original sample to obtain a normalized value. Then, the adversarial sample and its normalized response difference are injected into the probe dataset as new data points, and the gradient orthogonality index and curvature sensitivity index calculated in step S5.1 are recorded as additional meta-information for subsequent curvature field modeling. The scale consistency of the newly injected data is ensured by normalization processing, and the availability of the data is enhanced by retaining geometric feature information. By feeding back the adversarial sample and its key characteristics to the probe dataset, the content of the dataset is enriched, thereby improving the accuracy and adaptability of subsequent decision boundary curvature modeling.
[0055] S5.3, trigger the meta-learner to recalculate the weight correction coefficient: After the probe dataset is updated, the meta-learner is triggered to recalculate the weight correction coefficient of the task network to optimize the task network parameters. First, the updated probe dataset is input into the decision boundary curvature modeling step, and the dynamic curvature field model and normal vector field are recalculated based on the new data. Subsequently, the updated normal vector field is input into the meta-learner, and the new weight correction coefficient is calculated according to step S3.2. The geometric characteristics of the decision boundary reflected by the normal vector field are used to guide parameter adjustment, so that the task network can adapt to the latest decision boundary changes of the target model.
[0056] S5.4, forming a closed-loop data flow: After completing the task network parameter update, the updated task network parameters are passed to the orthogonal gradient projection step, the next round of adversarial samples are generated based on the new parameters, and the aforementioned verification and feedback process is repeated to form a closed-loop data flow. Specifically, by continuously updating the task network parameters and adversarial samples, the adversarial sample generation strategy is dynamically optimized to adapt to the dynamic defense mechanisms that may exist in the target model. Through the closed-loop iteration mechanism, the system can continuously learn the response characteristics of the target model and adjust its own strategy, thereby improving the quality of adversarial samples and the success rate of attacks. The closed-loop feedback mechanism can achieve the adaptability and continuous optimization of the system, ensuring that the effectiveness of the attack can be maintained when the target model defense strategy changes, improving the robustness and reliability of the entire adversarial sample generation process, and ensuring that the attack success rate can be continuously optimized with the iteration process.
[0057] The above formulas are all dimensionless and numerical calculations. The formula is a formula for the most recent real situation obtained by collecting a large amount of data and performing software simulation. The preset parameters in the formula are set by technicians in this field according to actual conditions.
[0058] It should be noted that the system of the present invention can be deployed on the device itself to realize embedded applications, and can also be run on a PC or other terminal with a user interface, thereby meeting a variety of hardware environments and usage requirements.
[0059] The above description is only by way of illustration of certain exemplary embodiments of the present invention. It is undoubted that those skilled in the art can modify the described embodiments in various ways without departing from the spirit and scope of the present invention. Therefore, the above drawings and descriptions are illustrative in nature and should not be construed as limiting the scope of protection of the claims of the present invention.
[0060] It should be noted that, in this article, if there are relational terms such as first and second, etc., they are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "including a..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.
[0061] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. A method for generating adversarial attack samples based on machine learning, characterized in that: Includes steps: Generate probe samples of multi-directional disturbances based on uniform spherical sampling, combine exponential amplitude increment design and collect target model responses in real time to construct a probe dataset containing the difference between input features and normalized responses; Based on the probe data set, the output response differences of adjacent probe samples are calculated, and a dynamic curvature field model is established by fitting the gradient change rate of the response differences to quantify the local deformation characteristics of the decision boundary of the target model. Using a two-layer surrogate model architecture, the meta-learner generates task network weight correction coefficients based on the normal vector field, and updates the task network parameters through online back-propagation combined with orthogonal loss to make the gradient direction orthogonal to the decision boundary; Calculate the adversarial gradient based on the updated task network parameters and use the curvature field normal vector for orthogonal decomposition to generate adversarial samples. Evaluate the orthogonality and make a decision based on the adversarial gradient features and the decision boundary curvature features to decide whether to pass it to the next step or regenerate it after adjusting the perturbation amplitude. The adversarial samples are submitted to the target model to verify the effectiveness of the attack. The adversarial samples that successfully bypass the defense and their corresponding target model responses are injected into the probe dataset as new probe data, triggering the meta-learner to recalculate the task network weight correction coefficient, forming a closed-loop data flow from attack verification to model update.
2. The method for generating adversarial attack samples based on machine learning according to claim 1, characterized in that: The process of acquiring the probe dataset is as follows: In the input space of the target model, with the original sample as the center, a multi-directional disturbance vector is generated by a uniform spherical sampling method to ensure that the disturbance direction is evenly distributed on the sphere; for each disturbance vector, an initial minimum amplitude value is set, and a series of amplitude values are calculated according to the exponential growth law to generate a set of probe samples; the probe samples are input into the target model one by one, the output response of the target model to each probe sample is recorded, and the difference between the output response of the target model and the original sample is calculated; The difference is normalized to obtain the normalized response difference; The probe dataset is constructed by pairing the input features of the probe samples with the corresponding normalized response differences.
3. The method for generating adversarial attack samples based on machine learning according to claim 2, characterized in that: The construction process of the dynamic curvature field model is: Based on the probe data set, the output response differences of adjacent probe samples are calculated, and the gradient change rate of the response difference is fitted; the response difference change rate of each disturbance direction is smoothed by the local weighted regression method to generate a gradient change rate function; based on the gradient change rate function, the curvature estimate at each probe sample is calculated by the finite difference method; The radial basis function interpolation method is used to transform the discrete curvature estimation value into a continuous dynamic curvature field model. The gradient of the dynamic curvature field at each probe sample point is calculated by numerical differentiation method and normalized to the unit normal vector.
4. The method for generating adversarial attack samples based on machine learning according to claim 3, characterized in that: The alternative model dynamic update method is: A two-layer substitution model architecture is adopted, in which the task network is used to simulate the decision-making behavior of the target model, and the meta-learner dynamically adjusts the parameters of the task network according to the normal vector field generated by the decision boundary curvature modeling step; specifically, the meta-learner takes the decision boundary normal vector field as input and generates the weight correction coefficient of the task network through a multi-layer perceptron structure; subsequently, the task network updates its convolution kernel parameters based on online back propagation, and the loss function used in the updating process is composed of prediction loss and orthogonal loss, where the orthogonal loss is used to constrain the gradient direction of the task network to remain orthogonal to the decision boundary normal vector.
5. The method for generating adversarial attack samples based on machine learning according to claim 4, characterized in that: The process of generating adversarial samples by intersection decomposition is: First, based on the updated task network parameters and input samples, the adversarial gradient is calculated, where the adversarial gradient is obtained by derivation of the prediction loss of the task network with respect to the input samples; then, the curvature field normal vector generated by the decision boundary curvature modeling step is used to perform orthogonal decomposition on the adversarial gradient to generate orthogonal components that are orthogonal to the curvature field normal vector; then, adversarial samples are generated based on the orthogonal components, which are specifically completed by normalizing and scaling the orthogonal components and then adding them to the input samples.
6. The method for generating adversarial attack samples based on machine learning according to claim 5, characterized in that: Evaluate orthogonality and make decisions: The adversarial gradient feature includes a gradient orthogonality index, and the decision boundary curvature feature includes a curvature sensitivity index; the gradient orthogonality index is calculated by the dot product and norm of the adversarial gradient and the normal vector of the curvature field, and the curvature sensitivity index is calculated by the ratio of the curvature value of the curvature field to the adversarial gradient norm; Subsequently, the gradient orthogonality index and curvature sensitivity index are fused through the comprehensive evaluation coefficient. If the comprehensive evaluation coefficient is greater than the preset threshold, the adversarial sample is output to the closed-loop feedback iteration step; the disturbance amplitude is multiplied by the preset value, the adversarial sample is regenerated and the evaluation process is repeated.
7. The method for generating adversarial attack samples based on machine learning according to claim 6, characterized in that: The process of verifying the effectiveness of the attack is: Submit the adversarial sample generated by the orthogonal gradient projection step to the target model, obtain the output response of the target model to the adversarial sample, and compare it with the output response of the original sample to verify the effectiveness of the attack. For classification tasks, check whether the predicted category of the adversarial sample is different from the predicted category of the original sample. For regression tasks, check whether the deviation between the output response of the adversarial sample and the output response of the original sample exceeds a preset threshold. If any of the conditions is met, the attack is considered successful.
8. The method for generating adversarial attack samples based on machine learning according to claim 7, characterized in that: The formation process of closed-loop data flow: If the attack is successful, the adversarial sample and its normalized response difference with the original sample are injected into the probe dataset of the probe sample generation and response collection steps; Subsequently, the dynamic curvature field model and normal vector field are recalculated based on the updated probe dataset, and the updated normal vector field is input into the meta-learner. The weight correction coefficient of the task network is calculated by integrating the normal vector field and the gradient of the task network loss function parameters, and then the task network parameters are adjusted. Finally, the next round of adversarial samples is generated based on the updated task network parameters, and the verification and feedback process is repeated to form a closed-loop data flow.
Citation Information
Cited By
Self-adaptive learning confrontation information security model reinforcing method and system
CN121117613A