Cooperative adversarial training-based model robustness enhancement strategy method
Through the collaborative adversarial training method, two neural network models with the same architecture are used to interactively train adversarial samples and exchange guidance in the early stage, and then focus on training their own adversarial samples in the later stage. This solves the problem of insufficient stability and universality of existing defense methods under various attacks, and achieves an efficient balance between robustness and accuracy.
Patent Information
- Application Number
- CN202510724365.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-09-12
AI Technical Summary
Existing defense methods are difficult to guarantee stability and universality in the face of various adversarial attacks, and often sacrifice the accuracy of clean samples when improving model robustness.
A method based on collaborative adversarial training is adopted to initialize two neural network models with the same architecture, and PGD attack is used to generate their own adversarial samples. The training phase is divided into the early and late stages. In the early stage, the adversarial samples of the self and the opponent are trained in different ways, and information guidance is exchanged. In the later stage, the self-adversarial samples are trained with reference to the guidance of the opponent's model.
It improves the defense capability and stability of neural networks in the face of various adversarial attacks, while maintaining efficient classification accuracy of clean samples and improving defense efficiency.
Smart Images

Figure CN120632450A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of adversarial defense technology, and in particular to a method for enhancing a model robustness strategy based on collaborative adversarial training. Background Art
[0002] Neural networks are vulnerable to adversarial examples, which occur when attackers introduce carefully crafted, subtle perturbations into clean examples. These perturbations are invisible to the human eye, but neural networks are highly sensitive to them, causing them to make incorrect predictions. This poses significant risks in real-world applications such as autonomous driving, financial systems, and medical diagnostics. If adversarial attacks are not effectively defended against, they can lead to serious safety hazards and even life-threatening situations.
[0003] To counter adversarial attacks, various defense techniques have been designed, and their research is of great significance. However, different attack methods generate adversarial examples with varying characteristics, making it challenging to design a defense strategy that effectively addresses multiple attack types and exhibits good generalization. Many defense methods perform well against specific attacks, but their effectiveness can significantly decline against other attack methods, making the stability and universality of the defense difficult to guarantee. Furthermore, the trade-off between model robustness and accuracy is difficult to achieve. Adversarial defenses typically rely on methods such as adversarial training to enhance model robustness, but this often comes at the expense of accuracy for clean examples, resulting in reduced overall model performance. Designing defenses that improve adversarial robustness while not significantly impacting the model's original task performance, or even achieving a balance between the two, is a key challenge in adversarial defenses.
[0004] Currently, adversarial defense strategies can be categorized into two types: one based on detection and data preprocessing, and the other on improving model robustness. Detection and data preprocessing methods, rather than modifying the input or the original model, attempt to detect adversarial examples in the input data and reject them. Adversarial detection can be combined with other defense methods and can be used to analyze whether the input contains adversarial examples when the results of the base classifier and the robust classifier disagree. However, detection performance is not significant when some perturbations are small. Furthermore, attackers who understand the detection mechanism may be able to circumvent these methods. Summary of the Invention
[0005] Current defense methods primarily focus on improving model robustness. These approaches are primarily implemented from four perspectives: the input layer, intermediate layers, and the output layer. Input layer methods can be categorized into noise addition, adversarial training, and feature extraction. Adversarial training is the most effective and widely used defense method, while intermediate layers focus on modifying the model structure to enhance its inherent robustness. The output layer can be categorized into distillation, decision boundary estimation, and loss function design. These methods effectively improve model robustness from different perspectives.
[0006] The method based on collaborative adversarial training of two models is to let the two models not only conduct adversarial training on the adversarial samples generated by themselves, but also learn the adversarial samples generated by each other, so that they are robust to both their own and each other's adversarial samples. Finally, they guide each other's judgment when making predictions, making the guidance more reliable. Therefore, the present invention provides a method for enhancing the model robustness strategy based on collaborative adversarial training. This method can improve the defense capability when the scale of the neural network is limited, and maintain efficient resource utilization, so that it can still maintain high performance and stability when facing multiple adversarial attacks.
[0007] The specific technical solutions are as follows:
[0008] The present invention discloses a method for enhancing the robustness of a model based on collaborative adversarial training, comprising the following steps:
[0009] (1) Initialize two neural network models with the same architecture and use PGD attack to generate their respective adversarial samples;
[0010] (2) Divide the entire training phase of the neural network into two parts: the early phase and the late phase;
[0011] (3) In the early stage of neural network training, the two neural network models use different adversarial training methods to train the adversarial samples generated by themselves, thereby improving the robustness of the models to the adversarial samples generated by themselves; exchange each other's adversarial samples and train the adversarial samples generated by the other party to improve the robustness of the models to the adversarial samples generated by the other party; the two neural networks exchange information and provide each other with reliable information guidance;
[0012] (4) In the later stages of neural network training, the two neural networks only focus on the adversarial examples they generate, focusing on improving their robustness to the adversarial examples they generate, and refer to the guidance of the other model on the adversarial examples they generate;
[0013] (5) Conduct multiple adversarial attacks on the model, and analyze and compare the model's defense capabilities under attacks.
[0014] Furthermore, the specific implementation steps of step (1) are as follows:
[0015] (S11) Initialize two resnet18 neural networks as model f and model g.
[0016] (S12) Use PGD to generate adversarial samples for model f and model g: Given the original clean sample x, generate adversarial samples and
[0017] in, ε f is the adversarial perturbation generated by PGD on model f;
[0018] εg is the adversarial perturbation generated by PGD on model g.
[0019] Furthermore, the specific implementation steps of step (2) are as follows:
[0020] (S21) The neural network training is set to 200 cycles. According to the experimental comparison, the first 110 cycles are set as the early stage and the last 90 cycles are set as the late stage.
[0021] Furthermore, the specific implementation steps of step (3) are as follows:
[0022] (S31) For model f, the TRADES method is used to perform adversarial training using self-generated adversarial samples, and the loss function is:
[0023]
[0024] Among them, CE is the cross entropy loss, which is used to measure the error between the model prediction value f(x) and the true label y, allowing the model f to learn the label probability distribution of the clean sample x; KL is the divergence loss, which measures the model f in the adversarial sample The predicted distribution of The difference between the predicted distribution f(x) and the clean sample x; λ is a trade-off parameter used to control the influence of the KL divergence term.
[0025] (S32) For model g, the AT method is used to perform adversarial training using the self-generated adversarial samples, and the loss function is:
[0026]
[0027] Among them, CE is the cross entropy loss. The former CE is a measure of the error between the predicted value g(x) of the model on the clean sample and the true label y, and the latter CE is a measure of the predicted value of the model on the adversarial sample. and the true label y.
[0028] (S33) For model f, learn the label probability distribution of the adversarial samples generated by model g, and promptly use the probability distribution learned by model g to supplement the probability distribution of non-target categories. The loss function is:
[0029]
[0030] Among them, CE is the cross entropy loss, which is used to measure the model prediction value The error between the real label y allows the model f to learn the hard label distribution of the adversarial sample generated by g; KL is the divergence loss, which is used to measure the difference between the model g and the model f in the adversarial sample. The difference between the predicted probability distributions on the target class and the predicted probability distribution on the target class; KL divergence extracts the knowledge of model g and feeds it to model f, allowing model f to learn the prediction information of model g on non-target classes; α1 and α2 are trade-off parameters used to control the relative importance between CE and KL.
[0031] (S34) For model g, learn the label probability distribution of the adversarial samples generated by model f, and promptly use the probability distribution learned by model f to supplement the probability distribution of non-target categories. The loss function is:
[0032]
[0033] Among them, CE is the cross entropy loss, which is used to measure the model prediction value The error between the real label y allows the model g to learn the hard label distribution of the adversarial sample generated by f; KL is the divergence loss, which is used to measure the difference between the model g and the model f in the adversarial sample. The difference between the predicted probability distributions on the target class and the predicted probability distribution on the target class; KL divergence extracts the knowledge of model f and feeds it to model g, allowing model g to learn the prediction information of model f on non-target classes; β1 and β2 are trade-off parameters used to control the relative importance between CE and KL.
[0034] (S35) For model f and model g, let them give each other reliable guidance on each other's adversarial sample predictions. The loss function is defined as:
[0035]
[0036] Among them, KL is the divergence loss, and KL divergence loss is the knowledge that the model extracts from the other model.
[0037] (S36) Finally, L total It is the sum of all loss functions, and the parameters of the two models are optimized simultaneously.
[0038] L total =γ1*(L f_f +L g_g )+γ2*(L f_g +L g_f )+γ3*L teach ;
[0039] Among them L f_f , L g_g , L f_g , L g_f and L teachare all loss functions, γ1, γ2, and γ3 are trade-off parameters used to control the relative importance of loss functions.
[0040] Furthermore, the specific implementation steps of step (4) are as follows:
[0041] (S41) For model f, the TRADES method is used to perform adversarial training using self-generated adversarial samples, and the loss function is:
[0042]
[0043] Among them, CE is the cross entropy loss, which is used to measure the error between the model prediction value f(x) and the true label y, allowing the model f to learn the label probability distribution of the clean sample x; KL is the divergence loss, which measures the model f in the adversarial sample The predicted distribution of The difference between the predicted distribution f(x) and the clean sample x; λ is a trade-off parameter used to control the influence of the KL divergence term.
[0044] (S42) For model g, the AT method is used to perform adversarial training using the self-generated adversarial samples, and the loss function is:
[0045]
[0046] Among them, CE is the cross entropy loss. The former CE is a measure of the error between the predicted value g(x) of the model on the clean sample and the true label y, and the latter CE is a measure of the predicted value of the model on the adversarial sample. The error between the true label y.
[0047] (S43) For Model 1 and Model 2, let them give each other some guidance in the prediction of each other's adversarial samples, and the loss function is:
[0048]
[0049] Among them, KL is the divergence loss, and KL divergence loss is the knowledge that the model extracts from the other model.
[0050] (S44) Finally, L total It is the sum of all loss functions, and the parameters of the two models are optimized simultaneously.
[0051] L total =γ1*(L f_f +L g_g )+γ2*L teach ;
[0052] Among them L f_f , L g_g , L teachare all loss functions, γ1, γ2, and γ3 are trade-off parameters used to control the relative importance of loss functions.
[0053] The present invention has the following advantages:
[0054] The present invention provides a method for enhancing the robustness of models based on collaborative adversarial training. The method first uses PGD attack to generate adversarial samples for two neural networks with the same architecture. The training phase is then divided into an early stage and a late stage. In the early stage, the model trains its own and the opponent's adversarial samples at the same time, and provides reliable guidance to the opponent. In the late stage, the model only focuses on training its own adversarial samples, and uses the opponent's guidance as a reference. Finally, multiple attack methods are used to attack the trained model to observe their correct predictions, that is, the probability of successful defense.
[0055] Compared with the prior art, the effects of the present invention are as follows:
[0056] 1. The model robustness enhancement strategy based on collaborative adversarial training designed in this invention can greatly improve the robustness of neural networks and achieve good defense effects when attacking them in various data sets.
[0057] 2. The method designed by the present invention can greatly improve the defense efficiency, especially in the early stage. Whether it is on clean samples or adversarial samples, the defense ability against attacks exceeds other adversarial training. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 Flowchart of the research method of the model robustness enhancement strategy based on collaborative adversarial training in the present invention.
[0059] Figure 2 A segmented schematic diagram of the research method for the model robustness enhancement strategy based on collaborative adversarial training in the present invention. DETAILED DESCRIPTION
[0060] The following describes the implementation of the present invention using specific embodiments. Those skilled in the art will readily understand the other advantages and benefits of the present invention from the disclosure herein. Obviously, the embodiments described are only a portion of the present invention, not all of it. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without inventive effort are intended to fall within the scope of protection of the present invention.
[0061] First, two neural networks with identical architectures are subjected to PGD attacks to generate adversarial examples. The training phase is then divided into an early and late phase. In the early phase, the two models train their own adversarial examples using different training methods, learning from the other model's adversarial examples and providing reliable guidance. In the late phase, each model focuses solely on training its own adversarial examples, using the other model's predictions on its own adversarial examples as a reference. Finally, the trained model is attacked using multiple attack methods, and the percentage of successful defenses against each attack is calculated.
[0062] Figure 1 and Figure 2 The method for enhancing the model robustness strategy of the present invention includes the following steps:
[0063] Step 1: Initialize two neural networks with the same architecture and use PGD attack to generate their respective adversarial samples.
[0064] In step S11, two neural network models f and g are initialized with the same architecture as resnet18 and other configurations to ensure that the final improvement in robustness is the result of collaborative training of two models with the same architecture, rather than being affected by the better robustness of the large model and other possible effects.
[0065] Step S12: Given the original clean image x, use the PGD method to iterate multiple times to generate adversarial samples of model f and model g. and
[0066] Among them, the adversarial samples generated by model f ε f is the adversarial perturbation generated by PGD on the model f.
[0067] Adversarial examples of model g ε g is the adversarial perturbation generated by PGD on model g.
[0068] Step 2: Divide the entire training phase of the neural network into two parts.
[0069] Step S21, setting the neural network training to 200 cycles.
[0070] Based on the comparison of the results of the phased experiment and the non-phased experiment, the first 110 cycles were set as the early stage and the last 90 cycles were set as the late stage.
[0071] Step 3: In the early stage, different adversarial training methods will learn different decision boundaries. Different decision boundaries will be continuously compared and optimized when learning from each other later. Therefore, the two neural networks use different adversarial training methods to train the adversarial samples generated by themselves, thereby improving the robustness of the model to the adversarial samples generated by themselves.
[0072] Exchange each other's adversarial samples, learn from each other's adversarial samples, and improve the model's robustness to the adversarial samples generated by the other party.
[0073] Since they have a basis for learning about each other's adversarial samples, the two neural networks can provide each other with more reliable information guidance when interacting with each other.
[0074] In step S31, for model f, the TRADES method is used to weigh the robustness of the self-generated adversarial samples and clean samples. The loss function is:
[0075]
[0076] Among them, CE is the cross entropy loss, which measures the difference between the predicted probability distribution of model f and the true label probability distribution of clean data x, so that model f can better match the label distribution of clean data and improve its classification accuracy for normal samples.
[0077] KL is the divergence loss, which measures the difference between the predicted distribution of the adversarial sample and the clean sample of the model f. By minimizing this loss, the predicted distribution of the adversarial sample is made as close as possible to the predicted distribution of the clean sample, thereby improving the robustness of the model to adversarial samples.
[0078] The trade-off parameter λ is used to balance the weight between the clean sample loss and the adversarial sample regularization term.
[0079] In step S32, for model g, the AT method is used to train both the self-generated adversarial samples and the clean samples, and the loss function is:
[0080]
[0081] Among them, CE is the cross entropy loss. The former is the loss of clean samples, and the latter is the loss of adversarial samples. By optimizing the adversarial samples during the training process, the robustness of the model in the face of perturbation data is enhanced.
[0082] In step S33, for model f, the label probability distribution of the adversarial samples generated by model g is learned, and the probability distribution learned by model g is used to supplement the probability distribution of non-target categories in a timely manner, so that model f can better capture the potential characteristics of the adversarial samples and improve its robustness under adversarial attacks. The loss function is:
[0083]
[0084] Among them, CE is the cross entropy loss, which allows the model f to learn the hard label distribution of the adversarial samples generated by g.
[0085] Since model f only learns the hard label distribution and lacks the opportunity to learn the probability distribution of non-target categories, KL divergence loss is needed to extract the knowledge of model g and provide the probability distribution of non-target classes learned by model g.
[0086] α1 and α2 are hyperparameters that balance the CE and KL loss functions.
[0087] In step S34, for model g, the label probability distribution of the adversarial samples generated by model f is learned, and the probability distribution learned by model f is used to supplement the probability distribution of non-target categories in a timely manner. The loss function is:
[0088]
[0089] Among them, CE is the cross entropy loss, which allows the model g to learn the hard label distribution of the adversarial samples generated by f.
[0090] Since model g only learns the hard label distribution and lacks the opportunity to learn the probability distribution of non-target categories, KL divergence loss is needed to extract the knowledge of model f and provide the probability distribution of non-target classes learned by model f.
[0091] β1 and β2 are hyperparameters that balance the CE and KL loss functions.
[0092] In step S35, for model f and model g, let them provide each other with reliable guidance on each other's adversarial sample predictions, thereby improving the robustness of both. The loss function is defined as:
[0093]
[0094] Among them, KL is the divergence loss. KL divergence loss is for the model to extract the knowledge of the other model, so that each model can learn the knowledge of the other model by minimizing the difference with the output of the other model, thereby improving its generalization ability on adversarial samples.
[0095] Step S36, finally, L total It is the sum of all loss functions, and the parameters of the two models are optimized simultaneously.
[0096] L total =γ1*(L f_f +L g_g )+γ2*(L f_g +L g_f )+γ3*Lteach ;
[0097] Among them, γ1, γ2 and γ3 are hyperparameters that weigh the three types of loss functions.
[0098] Step 4: In the later stage, the two neural networks will focus on the adversarial examples they have generated, optimize their robustness to these examples, and learn from the guidance provided by the other model when processing the adversarial examples they have generated, to further improve the effect of adversarial training.
[0099] Step S41: For model f, use the TRADES method to perform adversarial training using self-generated adversarial samples, and the loss function is:
[0100]
[0101] Among them, CE is the cross entropy loss, which measures the difference between the predicted probability distribution of model f and the true label probability distribution of clean data x, so that model f can better match the label distribution of clean data and improve its classification accuracy for normal samples.
[0102] KL is the divergence loss, which measures the difference between the predicted distribution of the adversarial sample and the clean sample of the model f. By minimizing this loss, the predicted distribution of the adversarial sample is made as close as possible to the predicted distribution of the clean sample, thereby improving the robustness of the model to adversarial samples.
[0103] The trade-off parameter λ is used to balance the weight between the clean sample loss and the adversarial sample regularization term.
[0104] Step S42: For model g, use the AT method to perform adversarial training using the self-generated adversarial samples, and the loss function is:
[0105]
[0106] Among them, CE is the cross entropy loss. The former is the loss of clean samples, and the latter is the loss of adversarial samples. By optimizing the adversarial samples during the training process, the robustness of the model in the face of perturbation data is enhanced.
[0107] In step S43, for Model 1 and Model 2, they are directly asked to give each other certain guidance on each other's adversarial sample predictions. The loss function is:
[0108]
[0109] Among them, KL is the divergence loss. KL divergence loss is for the model to extract the knowledge of the other model, so that each model can learn the knowledge of the other model by minimizing the difference with the output of the other model, thereby improving its generalization ability on adversarial samples.
[0110] Step S44, finally, L total It is the sum of all loss functions, and the parameters of the two models are optimized simultaneously.
[0111] L total =γ1*(L f_f +L g_g )+γ2*L teach ;
[0112] Among them, γ1 and γ2 are hyperparameters that weigh the three types of loss functions.
[0113] Step 5: Conduct multiple adversarial attacks on the model to explore and compare the model's defense capabilities under attacks.
[0114] In step S51, ResNet18 is selected as the model architecture, and two models f and g are trained on the CIFAR-10, CIFAR-100, and Tiny-ImageNet datasets, respectively.
[0115] Step S52: During the training process, 10 Attack the adversarial performance of the test model, record and save the model checkpoint corresponding to the epoch with the best performance under the attack for subsequent testing.
[0116] Step S53, load the saved checkpoint model, use multiple attack methods (FGSM, PGD 20 、CW ∞ , M, etc.) to attack the model and test its adversarial performance. Simultaneously, models trained with six other methods were tested using the same attack method to ensure a fair comparison. PGD-AT, TRADES, ALP, MART, and CAT are adversarial training methods, while Natural is trained using clean samples and lacks adversarial robustness. Finally, the accuracy of each method on clean samples (clean) and the adversarial accuracy under the four attack algorithms are reported. The results are shown in the table below, with the best results highlighted in bold.
[0117]
[0118] Compared to existing methods for improving model robustness, the present invention (SCAT) significantly enhances the robustness of neural networks, achieving excellent defensive effectiveness against attacks on a variety of datasets. Furthermore, this method improves defensive efficiency in the early stages, surpassing other adversarial training methods in defense against attacks at all stages, both on clean and adversarial samples. Furthermore, this method can be extended to multiple models, requiring only consideration of the order and method of communication between them, further enhancing model robustness.
[0119] Although the present invention has been described in detail above using general descriptions and specific embodiments, it will be apparent to those skilled in the art that modifications and improvements may be made thereto. Therefore, such modifications and improvements, without departing from the spirit of the present invention, are intended to be within the scope of protection claimed herein.
Claims
1. A method for enhancing model robustness based on collaborative adversarial training, characterized by: The following steps are involved: (1) Initialize two neural network models with the same architecture and use PGD attack to generate their respective adversarial samples; (2) In the early stage of neural network training, the two neural network models use different adversarial training methods to train the adversarial samples generated by themselves to improve the robustness of the model to the adversarial samples generated by themselves; exchange each other's adversarial samples and train the adversarial samples generated by the other party to improve the robustness of the model to the adversarial samples generated by the other party; the two neural networks exchange information and provide each other with information guidance; (3) In the later stages of neural network training, the two neural networks only focus on the adversarial examples they generate, focusing on improving their robustness to the adversarial examples they generate, and refer to the guidance of the other model on the adversarial examples they generate; (4) Conduct multiple adversarial attacks on the model, and analyze and compare the model's defense capabilities under attacks.
2. The method of claim 1, wherein: The specific steps of step (1) are as follows: (S11) Initialize two resnet18 neural networks as model f and model g; (S12) Use PGD to generate adversarial samples for model f and model g: Given the original clean sample x, generate adversarial samples and in, ε f is the adversarial perturbation generated by PGD on model f; ε g is the adversarial perturbation generated by PGD on model g.
3. The method of claim 1, wherein: The steps for implementing step (2) are as follows: (S21) For model f, the TRADES method is used to perform adversarial training using self-generated adversarial samples. The loss function is defined as follows: Among them, CE is the cross entropy loss, which is used to measure the error between the model prediction value f(x) and the true label y, allowing the model f to learn the label probability distribution of the clean sample x; KL is the divergence loss, which measures the model f in the adversarial sample The predicted distribution of The difference between the predicted distribution f(x) and the clean sample x; λ is a trade-off parameter used to control the influence of the KL divergence term; (S22) For model g, the AT method is used to perform adversarial training using self-generated adversarial samples. The loss function is defined as follows: Among them, CE is the cross entropy loss. The former CE is a measure of the error between the predicted value g(x) of the model on the clean sample and the true label y, and the latter CE is a measure of the predicted value of the model on the adversarial sample. The error between the true label y; (S23) For model f, learn the label probability distribution of the adversarial samples generated by model g, and promptly use the probability distribution learned by model g to supplement the probability distribution of non-target categories. The loss function is defined as follows: Among them, CE is the cross entropy loss, which is used to measure the model prediction value The error between the real label y allows the model f to learn the hard label distribution of the adversarial sample generated by g; KL is the divergence loss, which is used to measure the difference between the model g and the model f in the adversarial sample. The difference between the predicted probability distributions on the target class; KL divergence extracts the knowledge of model g and feeds it to model f, allowing model f to learn the prediction information of model g on non-target classes; α1 and α2 are trade-off parameters used to control the relative importance between CE and KL; (S24) For model g, learn the label probability distribution of the adversarial samples generated by model f, and promptly use the probability distribution learned by model f to supplement the probability distribution of non-target categories. The loss function is defined as follows: Among them, CE is the cross entropy loss, which is used to measure the model prediction value The error between the real label y allows the model g to learn the hard label distribution of the adversarial sample generated by f; KL is the divergence loss, which is used to measure the difference between the model g and the model f in the adversarial sample. The difference between the predicted probability distributions on the target class; KL divergence extracts the knowledge of model f and feeds it to model g, allowing model g to learn the prediction information of model f on non-target classes; β1 and β2 are trade-off parameters used to control the relative importance between CE and KL; (S25) For model f and model g, let them give each other reliable guidance on each other's adversarial sample predictions. The loss function is defined as follows: Among them, KL is the divergence loss, and KL divergence loss is the knowledge that the model extracts from the other model; (S26) Finally, L total It is the sum of all loss functions, and the parameters of the two models are optimized simultaneously; L total =γ1*(L f_f +L g_g )+γ2*(L f_g +L g_f )+γ3*L teach ; Among them L f_f , L g-g , L f_g , L g_f and L teach are all loss functions, γ1, γ2, and γ3 are trade-off parameters used to control the relative importance of loss functions.
4. The method of claim 3, wherein: The steps of implementing step (3) are as follows: (S31) is the same as (S21); (S32) is the same as (S22); (S33) is the same as (S25); (S34) Finally, L total It is the sum of all loss functions, and the parameters of the two models are optimized simultaneously; L total =γ1*(L f_f +L g_g )+γ2*L teach ; Among them L f_f , L g_g , L teach are all loss functions, γ1, γ2, and γ3 are trade-off parameters used to control the relative importance of loss functions.