Model vulnerability extraction method based on reverse knowledge distillation
By constructing an adversarial patch dataset through reverse knowledge distillation and learning the fragile decision boundary of the teacher model, the problem of low efficiency in model vulnerability extraction in existing technologies is solved, and efficient robustness enhancement and computing resource optimization of the model are achieved.
Patent Information
- Application Number
- CN202510889021.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-10-17
AI Technical Summary
Existing knowledge distillation techniques have limitations in extracting model vulnerabilities and fail to fully understand model vulnerabilities, resulting in wasted computing resources and limited model optimization potential.
The reverse knowledge distillation method is used to construct an adversarial patch dataset. The fragile decision boundary of the teacher model in an adversarial environment is learned through the reverse knowledge distillation objective, adversarial samples are generated, and a multi-level dynamic weight control strategy is designed to optimize the parameter space utilization of the student model and improve the robustness of the model.
Accurately locating and utilizing vulnerable areas of the model's decision boundary improves the model's anti-interference ability, reduces computing resource consumption, and achieves a more efficient model training process and enhanced security.
Smart Images

Figure CN120805138A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application provides a model vulnerability extraction method based on reverse knowledge distillation, belonging to the field of artificial intelligence information security and deep learning. BACKGROUND
[0002] With the continuous progress of artificial intelligence technology, deep learning models have shown excellent performance in image recognition, natural language processing and many other fields. However, the security and robustness of these high-performance models have gradually become the focus of research. Especially when facing adversarial sample attacks, even a small but carefully designed perturbation on the input data can cause the model to make wrong predictions or decisions. This not only limits the application range of the model in critical tasks, but also poses a serious challenge to the security of the model. Therefore, how to effectively evaluate and enhance the robustness of the model has become a common concern of the academic and industrial communities.
[0003] Traditionally, knowledge distillation, as an effective model compression technique, is widely used to improve the performance of small models (student models). By imitating the behavior of large and complex models (teacher models), student models can significantly reduce the demand for computing resources while maintaining high accuracy. However, this method has limitations in extracting model vulnerabilities. Specifically, traditional knowledge distillation mainly focuses on fitting the input-output mapping relationship of the target model, ignoring the in-depth exploration of the internal decision-making logic system of the model. This means that although the student model can learn certain pattern recognition capabilities, it may not fully understand the key factors that can lead to model vulnerability. Further, simply copying the response pattern of the teacher model can make the learning process of the student model inefficient, as it needs to spend a lot of parameter space to learn knowledge that is irrelevant to actual defense, which not only wastes valuable computing resources, but also limits the potential of model optimization. SUMMARY
[0004] The purpose of the present application is to provide a model vulnerability extraction method based on reverse knowledge distillation, which is different from the traditional knowledge distillation of the prior art that simply fits the input-output mapping, but focuses on extracting key vulnerable points in the model, so as to combine more efficient and more targeted model defense strategies, and improve the security and stability of the overall system.
[0005] The technical solution adopted by the present application is as follows: a model vulnerability extraction method based on reverse knowledge distillation, comprising the following steps:
[0006] Step 1: Construct the adversarial patch data set required for reverse knowledge distillation:
[0007] 1.1) Reverse Knowledge Distillation: The goal is to learn the fragile decision boundaries of the teacher model in adversarial environments, enabling the student model to learn and reproduce the misclassification behavior of the teacher when facing adversarial perturbations.
[0008] In 1.1), the goal of reverse knowledge distillation is described by formula (1):
[0009]
[0010]
[0011] in, is the decision mode of the teacher model, x adv Denotes the sample containing adversarial perturbation, D adv The dataset required for knowledge distillation; and Represent the correct classification and misclassification behavior patterns of the teacher model, respectively.
[0012] 1.2) Adversarial Sample Generation:
[0013] In the preparatory stage of knowledge distillation, a series of clean images are selected from the original dataset. Then, a series of adversarial patches with a uniform pattern are randomly initialized on these images multiple times to construct a targeted dataset. The resulting student model can replicate the misclassification behavior of the target model when encountering various samples containing adversarial patches.
[0014] In 1.2) above, the mathematical representation of the original image and the adversarial patch is described by expressions (2) and (3):
[0015] x→W x ·H x (2)
[0016] p→W p ·H p (3)
[0017] Among them, x is the original clean image, according to the width W x and height H x Decomposed into size W x ·H x The adversarial patch p is constrained to a smaller pixel space, W p and H p are the width and height of the adversarial patch respectively, and the pixel space size of the adversarial patch p is W p ·H p ;
[0018] By randomly injecting adversarial patches on clean images, the process of generating adversarial samples is modeled as a dynamic superposition of pixel space, which is expressed by formula (4):
[0019]
[0020] Among them, x′ is the adversarial sample formed by injecting the adversarial patch into the clean image; the mask matrix r k represents the kth random location of the adversarial patch in the image plane; to enhance the diversity of the dataset, each selected clean sample x will undergo an independent adversarial patch injection process. If repeated areas appear during the iteration, they will be re-injected, and a total of K different adversarial samples will be generated.
[0021] At the physical structure level of the adversarial patch, the effective area is determined by the pixel coverage and the filling strategy. A unified adversarial patch filling strategy is adopted, and the pixel filling rule of the adversarial sample x′ is expressed by formula (5):
[0022]
[0023] Among them, PV represents the pixel value of the original image; (i, j) is the coordinate of the pixel point in the adversarial patch.
[0024] For the filling of the adversarial patch, the average pixel value of the image is selected as the filling of the adversarial patch. Due to the global calculation characteristics of this value, the average pixel value naturally has the ability to generalize across samples and can be used as a unified adversarial patch style. The calculation method of the average pixel value MPV(x) of the input image x is expressed by formula (6) as follows:
[0025]
[0026] By unifying the padding strategy for adversarial samples and avoiding the unexplainable effects of over-randomized adversarial patches, unified adversarial perturbation padding can focus the student model's learning attention on the teacher model's response patterns to perturbations in specific areas, rather than redundant noise features.
[0027] 1.3) Dataset Construction:
[0028] During the dataset construction process, an adversarial patch is first applied to the original clean sample to form a perturbed sample. The perturbed sample is then fed into the target model for inference to obtain the classification output. By comparing this output with the true sample label, samples correctly classified by the target model are selected as positive samples, while samples incorrectly classified are selected as negative samples. Finally, the positive and negative samples are integrated with the original clean sample to construct a dataset containing different model response states.
[0029] The ideal dataset used in knowledge distillation in 1.3) contains three types of samples with different characteristics:
[0030] First, the dataset should contain original samples without any interference, which contains the original characteristics of clean data distribution, and can reflect the basic classification ability of the model in the conventional scene;
[0031] Second, the positive samples disturbed by the adversarial patch but still correctly classified by the teacher model, which can reveal the decision logic of the model when facing specific interference, and distinguish and contrast with the negative samples;
[0032] Finally, it also contains negative samples successfully misled by the adversarial patch, and the annotation information uses the probability distribution output by the teacher model. The introduction of soft labels can retain the confidence trajectory in the teacher model's error judgment process, so that the student model can fit its error decision process.
[0033] Step 2: Extract model vulnerability knowledge to the distilled student model based on the constructed adversarial patch dataset.
[0034] 2.1) Reverse distillation loss function design
[0035] Adopting a differentiated adaptation method based on sample features, the student model can capture and amplify the specific decision bias of the teacher under the influence of the adversarial patch;
[0036] For clean samples from the original dataset, formula (7) is used to quantify the model's performance on such samples:
[0037]
[0038] Where I * represents the size of the subset in the adversarial patch dataset composed of clean samples from the original dataset; x i represents the i-th sample in the subset; l(x i ) represents the one-hot encoding of the true label of sample x i In the clean dataset, M s (x i ) is the predicted probability distribution of the student model M s for sample x i ; The existence of this loss term forces the student model to maintain similar classification performance as the teacher model on undisturbed samples, so that the deviation behavior of the teacher model on adversarial samples presents a more significant contrast;
[0039] For the positive adversarial patch samples correctly classified by the teacher model, the loss is calculated using the hard label-based quantification formula (8):
[0040]
[0041] where M * represents the size of the subset of samples in the adversarial patch dataset that the teacher model successfully defended against adversarial patch interference; x m represents the m-th sample in the subset; l(M t (x m )) represents the maximum probability class label output by the teacher model.
[0042] For the negative adversarial patch samples incorrectly classified by the teacher model, the loss function based on soft labels is used to quantify the fitting effect of the student model and the teacher model, and the loss function is represented by formula (9):
[0043]
[0044] where N * represents the size of the subset of adversarial patch samples in the adversarial patch dataset that the teacher model produces incorrect predictions; x n represents the n-th sample in the subset; W t (x n ) is the complete probability distribution predicted by the teacher model for the sample x n . The loss term of the negative adversarial sample preserves the confidence distribution characteristics when the teacher model makes a mistake, prompting the student model to understand and learn the decision vulnerability of the teacher model in the adversarial environment.
[0045] Finally, the total distillation loss L KD is used to quantify the overall performance of the student model on the adversarial patch dataset:
[0046] L KD = λ ori · L ori + λ pr · L pr + λ pf · L pf (3.15)
[0047] where λ ori , λ pr and λ pf are the weight coefficients of the three types of samples.
[0048] 2.2) Dynamic weight regulation strategy
[0049] In the alternative model construction method based on distillation, through the construction of multi-level adversarial patch dataset and optimization strategy, the directional learning of the student model to the decision vulnerability of the teacher model is realized.
[0050] Firstly, by injecting random adversarial patches, local disturbance areas are generated on the surface of the original image, and the output response of the teacher model to the disturbance samples is extracted through black box query to construct a ternary dataset containing original samples, positive adversarial samples and negative adversarial samples;
[0051] Among them, the negative sample uses the probability distribution output by the teacher as a soft label to guide the knowledge distillation process. After the construction of the adversarial patch dataset, the response feature and hard label of the last full connection output of the teacher model are used to train the student model based on the responsive knowledge distillation method, to simulate the decision boundary information of the teacher model.
[0052] In the early stage of training, high-weight original samples and positive samples are used to drive the student model to establish basic classification ability, and at the same time, the immune boundary of the teacher to invalid disturbance is preliminarily determined. In this stage, the weight coefficient of the negative sample is low to avoid excessive interference of the adversarial patch to the establishment of the basic ability of the student model in the early stage of training.
[0053] As the training process progresses, the optimization goal of knowledge distillation gradually shifts to comparative learning of adversarial sensitivity. By reducing the weight of the original sample and increasing the loss proportion of the positive sample and the negative sample, the student model will enter the fine modeling stage of the adversarial decision boundary. The student model learns the response difference of the teacher to effective disturbance and invalid disturbance in the feature space, and drives the parameters of the student model to align with the decision boundary of the teacher model under the interference of adversarial patches.
[0054] In the last stage of the training process, the weight allocation is further tilted to the negative sample, based on learning the vulnerability pattern of the teacher model. In this stage, the gradient update of the student model is dominated by the probability distribution difference of the teacher model for the prediction of the negative sample, and the weight coefficients of the original sample and the positive sample are greatly reduced to improve the consistency of the student model and the teacher model in the misclassification of the negative sample.
[0055] The beneficial effects of the present application are: the present application constructs a data set containing countermeasures, and then distills the knowledge of these model weaknesses into student models, which can accurately locate and utilize the vulnerable areas of the target model decision boundary, and the resulting student model can be used to enhance the robustness of the model. At the same time, the student model of the present application can make full use of the limited parameter space, and concentrate on learning those knowledge that is essential to improve its anti-interference ability, improve the efficiency of knowledge transfer, and reduce unnecessary consumption of computing resources, so as to realize a more economical and efficient model training process. Therefore, the model vulnerability extraction method based on reverse knowledge distillation proposed by the present application provides a novel and effective solution to the robustness challenge faced by current deep learning models, and shows great potential in improving model security. BRIEF DESCRIPTION OF DRAWINGS
[0056] Figure 1 The present application is a schematic diagram of the overall process of the method;
[0057] Figure 2 The present application is a schematic diagram of the countermeasure injection;
[0058] Figure 3 The present application is a schematic diagram of the countermeasure data set structure;
[0059] Figure 4 The present application is a construction algorithm diagram of the countermeasure data set;
[0060] Figure 5 The present application is a knowledge distillation method algorithm diagram based on the countermeasure data set. DETAILED DESCRIPTION
[0061] As shown in the present application, a model vulnerability extraction method based on reverse knowledge distillation includes the following steps: Figure 1
[0062] Step 1: Construct a knowledge transfer data set, i.e. a countermeasure data set;
[0063] The knowledge transfer data set can be described from three dimensions: reverse knowledge distillation target, adversarial sample generation, and data set construction.
[0064] 1.1) Reverse knowledge distillation target
[0065] Unlike the overall goal of traditional knowledge distillation, the reverse knowledge distillation method focuses on learning the fragile decision boundary of the teacher model in the adversarial environment, and tries to make the student model be able to fully learn and reproduce the misclassification behavior of the teacher when facing adversarial perturbations. In order to achieve this goal, the target required for reverse knowledge distillation can be described by formula (1).
[0066]
[0067] where, is the decision pattern of the teacher model, x adv represents the sample containing adversarial perturbation, D adv is the dataset required for knowledge distillation; and represent the behavior pattern of the teacher model correctly classified and misclassified, respectively.
[0068] 1.2) Adversarial sample generation
[0069] In the preparatory stage of knowledge distillation, it is necessary to construct an efficient and targeted dataset. The core purpose of reverse knowledge distillation is to ensure that the resulting student model can accurately and comprehensively replicate the target model's misclassification behavior pattern when encountering various adversarial patch-containing samples through knowledge distillation technology. To achieve this goal, the initial stage is to select some clean images from the original dataset, and then randomly initialize a series of adversarial patches with uniform style on these images multiple times. Among them, the mathematical representation of the original image and the adversarial patch can be described by expressions (2) and (3):
[0070] x→W x ·H x (2)
[0071] p→W p ·H p (3)
[0072] where x is the original clean image, which is decomposed into a pixel space of size W x ·H x according to its width W x and height H x ; the adversarial patch p is constrained to one of the smaller pixel spaces, W p and H p are the width and height of the adversarial patch, respectively, and the pixel space size of the adversarial patch p is W p ·H p .
[0073] Subsequently, the process of generating adversarial samples by randomly injecting adversarial patches on clean images can be modeled as a dynamic superposition of pixel spaces. This process can be represented by equation (4):
[0074]
[0075] where x' is the adversarial sample formed by injecting the adversarial patch into the clean image; the mask matrix r krepresents the kth randomly located region of the adversarial patch within the image plane. To enhance the diversity of the dataset, each selected clean sample x undergoes K independent adversarial patch injections. If repeated regions appear during the iterations, they are re-injected, ultimately generating a total of K different adversarial examples.
[0076] At the physical structure level of the adversarial patch, its effective area is determined by the pixel coverage and the filling strategy. In this stage, a unified adversarial patch filling strategy is adopted. The pixel filling rule of the adversarial sample x′ can be expressed by formula (5):
[0077]
[0078] Among them, PV represents the pixel value of the original image; (i, j) is the coordinate of the pixel point in the adversarial patch.
[0079] For the filling of the adversarial patch, the average pixel value of the image is selected as the filling of the adversarial patch. Due to the global calculation characteristics of this value, the average pixel value naturally has the ability to generalize across samples, making it suitable for being used as a unified adversarial patch style. The calculation method of the average pixel value MPV(x) of the input image x is expressed by formula (6) as follows:
[0080]
[0081] By standardizing the adversarial example padding strategy, we avoid the unexplainable effects of overly randomized adversarial patches. Specifically, unified adversarial perturbation padding allows the student model to focus on the teacher model's response to perturbations in specific regions, rather than redundant noise features.
[0082] Ultimately, the adversarial examples generated in this stage constrain the perturbation to a local patch area while preserving the peripheral image content. This design balances the adversarial effect with the preservation of the original image semantics. On the one hand, local perturbations can more accurately reflect the teacher model's classification decision-making process under the influence of limited abnormal areas, thereby exposing its dependence on specific visual clue areas. On the other hand, the preservation of the original pixels in the peripheral areas ensures that the student model can capture the teacher model's decision changes through the response differences in the perturbed area, while maintaining an understanding of the normal classification logic through the stable features of the unperturbed area. This comparison process helps the student model focus on the core needs of vulnerability extraction, making the knowledge distillation process more robust and effective. Figure 2 We demonstrate how to inject adversarial patches into original image samples using our method.
[0083] 1.3 Dataset Construction
[0084] In constructing the adversarial patch dataset based on knowledge distillation, the core goal is to clearly present the decision boundary features of the teacher model under adversarial interference through data design, thereby providing challenging and guiding learning samples for the student model. Therefore, the establishment of the dataset not only needs to ensure the difference between the newly constructed adversarial samples and the original samples, but also needs to ensure the clear distinction between the positive samples correctly classified by the teacher model and the negative samples incorrectly classified by the teacher model according to the classification results of the teacher model. These characteristics aim to accurately depict the behavior patterns of the teacher model in the adversarial environment, so that the student model can accurately capture and reproduce the misclassification behavior of the teacher model.
[0085] Specifically, it is believed that the ideal dataset used in knowledge distillation should include three types of samples with different characteristics: first, the dataset should include original samples without any interference, which contain the original characteristics of clean data distribution and can reflect the basic classification ability of the model in the conventional scene; second, the positive samples disturbed by the adversarial patch but still correctly classified by the teacher model, which can reveal the decision logic of the model that still maintains robustness when facing specific interference, thereby serving as a distinction and contrast to the negative samples; finally, the dataset should also include negative samples successfully misled by the adversarial patch. It is worth noting that the annotation information uses the probability distribution output by the teacher model rather than the hard label. The introduction of soft labels can preserve the confidence trajectory in the teacher model's misjudgment process, so that the student model can fit its erroneous decision-making process. Figure 3 The overall process of constructing the adversarial patch dataset is shown.
[0086] Through the combination of these three types of samples, the adversarial patch dataset essentially constructs a scenario that can observe the decision mechanism of the teacher model from multiple dimensions. The scenario aims to capture the failure patterns of the teacher model under adversarial disturbance, using the benchmark performance of the teacher in the normal environment and the critical state under adversarial disturbance as a contrast and aid, thereby prompting the student model to provide more targeted support for subsequent possible robustness evaluation and implementation of enhancement methods. Overall, the construction process of the dataset is shown in Figure 4 The original dataset is selected from some clean images, and a series of uniform adversarial patches are randomly initialized on them to construct a targeted dataset, so that the student model can replicate the misclassification behavior patterns of the target model when encountering samples containing adversarial patches; during the construction of the dataset, the adversarial patches are first applied to the original clean samples to form disturbed samples, which are input into the target model for inference to obtain classification output, which is compared with the true label to filter out correctly classified samples, i.e. positive samples and incorrectly classified samples, i.e. negative samples, and finally integrate the positive samples, negative samples and original clean samples to construct a dataset containing different model response states.
[0087] Step 2: Target model vulnerability extraction
[0088] 2.1 Reverse distillation loss function design
[0089] In the loss function design of knowledge distillation, a differentiated adaptation method based on sample features is proposed to adapt to the different needs of knowledge transfer of three types of samples in the adversarial patch dataset. Specifically, the core goal of the distillation loss function design is to enable the student model to capture and amplify the specific decision bias of the teacher under the influence of the adversarial patch, rather than passively inherit its overall behavior pattern.
[0090] For clean samples derived from the original dataset, we use formula 3.12 to quantify the model's performance on such samples.
[0091]
[0092] where I * represents the size of the subset in the adversarial patch dataset composed of clean samples derived from the original dataset; x i represents the i-th sample in the subset; l(x i ) represents the true label of the sample x p . In the clean dataset, the one-hot encoding of the true label, M s (x i ) is the predicted probability distribution of the student model M s for sample x i . The existence of this loss term forces the student model to maintain similar classification performance as the teacher model on undisturbed samples, thereby making the deviant behavior of the teacher model on adversarial samples more prominent.
[0093] For positive adversarial patch samples correctly classified by the teacher model, we also use the hard label-based quantification formula (8) 3 to calculate their loss:
[0094]
[0095] where M * represents the size of the subset in the adversarial patch dataset composed of samples successfully defended by the teacher model against adversarial patch interference; x m represents the m-th sample in the subset; l(M t (x m )) represents the maximum probability class label output by the teacher model. The loss term design for positive samples essentially draws the immune boundary of the teacher model to specific perturbations in the feature space, enabling the student model to filter out ineffective perturbation patterns that are insufficient to break through the teacher's defense mechanism while amplifying the sensitivity of the teacher model to key perturbation features during learning.
[0096] For the reverse adversarial patch samples misclassified by the teacher model, a loss function based on soft labels is used to quantify the fitting effect of the student model and the teacher model. This loss function can be represented by formula (9):
[0097]
[0098] where N * represents the size of the subset consisting of adversarial patch samples in the adversarial patch dataset that the teacher model produces incorrect predictions; x n represents the nth sample in the subset; M t (x n ) is the complete probability distribution predicted by the teacher model for the sample x n . As the core of this knowledge distillation method, the loss term of the reverse adversarial sample promotes the student model to understand and learn the decision vulnerability of the teacher model in the adversarial environment by preserving the confidence distribution characteristics when the teacher model misjudges.
[0099] Finally, the total distillation loss L KD is used to quantify the overall performance of the student model on the adversarial patch dataset:
[0100] L KD = λ ori · L ori + λ pr · L pr + λ pf · L pf (10)
[0101] where λ ori , λ pr and λ pf are the weight coefficients of the three types of samples.
[0102] 2.2 Dynamic weight regulation strategy
[0103] In the distillation-based surrogate model construction method, the student model is tried to learn the decision vulnerability of the teacher model by constructing a multi-level adversarial patch dataset and optimization strategy. Overall, this method first generates local disturbance regions on the surface of the original image through random adversarial patch injection, and extracts the output response of the teacher model to the disturbed samples through black box query, thereby constructing a ternary dataset containing original samples, forward adversarial samples and reverse adversarial samples.
[0104] In which, the reverse sample uses the probability distribution of the teacher output as a soft label to guide the process of knowledge distillation. After the construction of the adversarial patch dataset, this paper optimizes the responsive knowledge distillation method based on the distillation method, uses the response features and hard labels of the last fully connected output of the teacher model to train the student model, so as to prompt it to imitate the decision boundary information of the teacher model.
[0105] In the early stage of training, high-weight original samples and positive samples are used to drive the student model to quickly establish basic classification ability, and at the same time, the immune boundary of the teacher to invalid perturbation is preliminarily delimited. In this stage, the weight coefficient of the reverse sample is low to avoid excessive interference of the adversarial patch to the establishment of the basic ability of the student model in the early stage of training. As the training process advances, the optimization goal of knowledge distillation gradually shifts to contrastive learning of adversarial sensitivity. By reducing the weight of the original sample and increasing the loss proportion of the positive sample and the reverse sample, the student model will enter the fine modeling stage of the adversarial decision boundary. This process enables the student model to learn the response difference of the teacher to effective and invalid perturbations in the feature space, and drives the parameters of the student model to align with the decision boundary of the teacher model under adversarial patch interference.
[0106] In the last stage of the training process, the weight allocation is further tilted towards the reverse sample, with the core being to learn the vulnerability pattern of the teacher model. In this stage, the gradient update of the student model is almost completely dominated by the probability distribution difference of the teacher model's prediction for the reverse sample, and the weight coefficients of the original sample and the positive sample are greatly reduced, thereby effectively improving the misclassification consistency of the student model and the teacher model on the reverse sample.
[0107] After a complete training cycle, the parameter space of the student model can be considered to be approximately the vulnerability pattern of the teacher model. The overall process algorithm of knowledge distillation is shown in Figure 5 The initial high-weight original sample and positive sample are used to drive the student model to quickly establish basic classification ability, and at the same time, the immune boundary of the teacher model to invalid perturbation is preliminarily delimited, and the weight of the reverse sample is low at this time to avoid interference; then, the optimization goal shifts to contrastive learning of adversarial sensitivity, the weight of the original sample is reduced, and the loss proportion of the positive and reverse samples is increased, so that the student model enters the fine modeling stage of the adversarial decision boundary, learns the response difference of the teacher to effective and invalid perturbations, and aligns the decision boundary; in the later stage of training, the weight is further tilted towards the reverse sample, with the core being to learn the vulnerability pattern of the teacher, and the gradient update of the student model is dominated by the probability distribution difference of the teacher's prediction for the reverse sample, the weight of the original and positive samples is reduced, and the misclassification consistency of the two on the reverse sample is improved.
Claims
1. A model vulnerability extraction method based on reverse knowledge distillation, characterized in that: The following steps are involved: Step 1: Construct the adversarial patch dataset required for reverse knowledge distillation: 1.1) Reverse Knowledge Distillation: The goal is to learn the fragile decision boundaries of the teacher model in adversarial environments, enabling the student model to learn and reproduce the misclassification behavior of the teacher when facing adversarial perturbations. 1.2) Adversarial Sample Generation: In the preparatory stage of knowledge distillation, a series of clean images are selected from the original dataset. Then, a series of adversarial patches with a uniform pattern are randomly initialized on these images multiple times to construct a targeted dataset. The resulting student model can replicate the misclassification behavior of the target model when encountering various samples containing adversarial patches. 1.3) Dataset Construction: During the dataset construction process, an adversarial patch is first applied to the original clean sample to form a perturbed sample. The perturbed sample is then fed into the target model for inference to obtain the classification output. By comparing this output with the true sample label, samples correctly classified by the target model are selected as positive samples, while samples incorrectly classified are selected as negative samples. Finally, the positive and negative samples are integrated with the original clean sample to construct a dataset containing different model response states. Step 2: Extract model vulnerability knowledge into the distilled student model based on the constructed adversarial patch dataset.
2. The method for extracting model vulnerabilities based on reverse knowledge distillation according to claim 1, characterized in that: In 1.1), the goal of reverse knowledge distillation is described by formula (1): in, is the decision mode of the teacher model, x adv Denotes the sample containing adversarial perturbation, D adv The dataset required for knowledge distillation; and Represent the correct classification and misclassification behavior patterns of the teacher model, respectively.
3. The method for extracting model vulnerabilities based on reverse knowledge distillation according to claim 1, characterized in that: In 1.2) above, the mathematical representation of the original image and the adversarial patch is described by expressions (2) and (3): x→W x ·H x (2) p→W p ·H p (3) Among them, x is the original clean image, according to the width W x and height H x Decomposed into size W x ·H x The adversarial patch p is constrained to a smaller pixel space, W p and H p are the width and height of the adversarial patch respectively, and the pixel space size of the adversarial patch p is W p ·H p ; By randomly injecting adversarial patches on clean images, the process of generating adversarial samples is modeled as a dynamic superposition of pixel space, which is expressed by formula (4): x′=x·(1-r k )+p, (4) Among them, x′ is the adversarial sample formed by injecting the adversarial patch into the clean image; the mask matrix r k represents the kth random positioning area of the adversarial patch in the image plane; to enhance the diversity of the dataset, each selected clean sample x will undergo K independent adversarial patch injection processes. If repeated areas appear during the iterative process, they will be re-injected, and finally a total of K different adversarial samples will be generated; At the physical structure level of the adversarial patch, the effective area is determined by the pixel coverage and the filling strategy. A unified adversarial patch filling strategy is adopted, and the pixel filling rule of the adversarial sample x′ is expressed by formula (5): Where PV represents the pixel value of the original image; (i, j) is the coordinate of the pixel point in the adversarial patch; For the filling of the adversarial patch, the average pixel value of the image is selected as the filling of the adversarial patch. Due to the global calculation characteristics of this value, the average pixel value naturally has the ability to generalize across samples and can be used as a unified adversarial patch style. The calculation method of the average pixel value MPV(x) of the input image x is expressed by formula (6) as follows: By unifying the padding strategy for adversarial samples and avoiding the unexplainable effects of over-randomized adversarial patches, unified adversarial perturbation padding can focus the student model's learning attention on the teacher model's response patterns to perturbations in specific areas, rather than redundant noise features.
4. The method for extracting model vulnerabilities based on reverse knowledge distillation according to claim 1, characterized in that: In 1.3), the ideal dataset used for knowledge distillation contains three types of samples with different characteristics: First, the dataset should contain original samples without any interference. Such samples have the original characteristics of clean data distribution and can reflect the basic classification ability of the model in common scenarios. Secondly, there are positive examples that are perturbed by the adversarial patch but still correctly classified by the teacher model. By retaining their original labels, these examples can reveal the decision logic that allows the model to remain robust in the face of specific perturbations, serving as a distinction and comparison for negative examples. Finally, it also includes reverse samples that are successfully misled by the adversarial patch. Their annotation information uses the probability distribution output by the teacher model. The introduction of soft labels can retain the confidence trajectory of the teacher model during the misjudgment process, allowing the student model to fit its incorrect decision-making process.
5. The method for extracting model vulnerabilities based on reverse knowledge distillation according to claim 1, characterized in that: In the step 2, the method is: 2.1) Reverse Distillation Loss Function Design A differentiated adaptation method based on sample features is used to adapt to the different requirements of knowledge transfer for the three types of samples in the adversarial patch dataset, enabling the student model to capture and amplify the teacher's specific decision biases under the influence of adversarial patches. For clean samples from the original dataset, use formula (7) to quantify the performance of the model on such samples: Among them, I * represents the size of the subset of clean samples from the original dataset in the adversarial patch dataset; x i represents the i-th sample in the subset; l(x i ) represents taking sample x i In the clean dataset, the one-hot encoding of the true label, M s (x i ) is the student model M s For sample x i The existence of this loss term forces the student model to maintain similar classification performance on unperturbed samples as the teacher model, making the deviation behavior of the teacher model on adversarial samples more contrasting. For the positive adversarial patch samples correctly classified by the teacher model, the hard label-based quantization formula (8) is used to calculate its loss: Among them, M * represents the size of the subset of samples in the adversarial patch dataset that the teacher model successfully resists against adversarial patch interference; x m represents the mth sample in the subset; l(M t (x m )) represents the maximum probability category label output by the teacher model; For the reverse adversarial patch samples that are misclassified by the teacher model, a soft-label-based loss function is used to quantify the fitting effect of the student model and the teacher model. The loss function is expressed as follows: Among them, N * represents the size of the subset of adversarial patch samples in the adversarial patch dataset that the teacher model produces incorrect predictions; x n represents the nth sample in the subset; M t (x n ) is the teacher model for sample x n The complete probability distribution obtained by prediction; the loss term of the reverse adversarial sample preserves the confidence distribution characteristics when the teacher model misjudges, prompting the student model to understand and learn the decision vulnerability of the teacher model in the adversarial environment; Finally, the total distillation loss L is used KD To quantify the overall performance of the student model on the adversarial patch dataset: L KD =λ ori ·L ori +λ pr ·L pr +λ pf ·L pf (3.15) Among them, λ ori ,λ pr and λ pf is the weight coefficient of the three samples; 2.2) Dynamic weight control strategy In the distillation-based alternative model construction method, by constructing a multi-level adversarial patch dataset and optimization strategy, the student model can learn the vulnerability of the teacher model's decisions in a targeted manner; First, we inject random adversarial patches to generate local perturbation regions on the original image surface. We then use black-box queries to extract the output responses of the teacher model to the perturbed samples, and construct a ternary dataset containing the original sample, the positive adversarial sample, and the reverse adversarial sample. Among them, the reverse sample uses the probability distribution of the teacher's output as a soft label to guide the knowledge distillation process. After completing the construction of the adversarial patch dataset, based on the responsive knowledge distillation method, the response features and hard labels of the teacher model's last fully connected layer are used to train the student model to imitate the teacher model's decision boundary information; In the early stages of training, high-weighted original and positive samples are used to drive the student model to establish basic classification capabilities, while also preliminarily defining the teacher's immunity to invalid perturbations. In this stage, the weight coefficient of the negative sample is low to avoid excessive interference of adversarial patches on the basic capabilities of the student model in the early stages of training. As training progresses, the optimization goal of knowledge distillation gradually shifts to contrastive learning of adversarial sensitivity. By reducing the weight of original samples and increasing the loss ratio of positive and negative samples, the student model enters the stage of refined modeling of the adversarial decision boundary. The student model learns the differences in the teacher's responses to effective and ineffective perturbations in the feature space, driving the student model's parameters to align with the teacher model's decision boundary under adversarial patch interference. In the final stage of the training process, the weight distribution is further tilted towards reverse samples, based on learning the vulnerability pattern of the teacher model. In this stage, the gradient update of the student model is dominated by the difference in the probability distribution of the teacher model's predictions for reverse samples. The weight coefficients of the original samples and the positive samples are greatly reduced, thereby improving the misclassification consistency between the student model and the teacher model on reverse samples.