An Active Learning Method and System Enhanced by Adversarial Training
By introducing adversarial training and VAE networks into active learning, combined with GAN, the performance gap problem of existing active learning methods under small annotation budget is solved, and high-performance and robust image recognition model training is achieved.
Patent Information
- Application Number
- CN202310012836.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-05
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2043-01-05
AI Technical Summary
The existing active learning methods have significant performance gaps compared with large labeled dataset training under smaller annotation budgets, and rely on the quality and diversity of generated samples, resulting in unstable performance.
Adopting an active learning method based on adversarial training enhancement, an adversarial trainer and a variational autoencoder (VAE) network are constructed, combined with a generation adversarial network (GAN), to achieve the Nash equilibrium state, and the label-free sample with the maximum amount of information is selected for annotation.
It effectively reduces the cost of manual labeling, improves the performance and robustness of the model, and enables the model to achieve higher performance under a small amount of training data.
Smart Images

Figure CN116187400B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence active learning, and specifically to an active learning method and system enhanced based on adversarial training. Background Art
[0002] Deep learning has achieved great success in computer vision tasks from image classification to detection and segmentation, but these successes still require a large number of labeled training samples. This is a major challenge because deep learning practitioners are increasingly applying deep learning models to solve new problems from different fields, and from a sustainability perspective, when the model gets larger, a larger amount of data is needed, which brings suboptimal solutions and unaffordable overheads for experts in various fields to label a large amount of training data. Active learning, where unlabeled samples are incrementally selected for annotation and the model achieves high classification accuracy with a low annotation budget, has thus become an exciting paradigm with significant potential to popularize deep learning.
[0003] So far, current methods can be divided into query acquisition (pool-based) algorithms and query synthesis algorithms. Pool-based methods use various acquisition functions to select the most informative samples, while query synthesis algorithms use generative models to generate the most informative samples. Pool-based methods use several sampling strategies, such as information-based, uncertainty-based, and ensemble methods. These methods have been proposed in both Bayesian and non-Bayesian frameworks. In non-Bayesian classical active learning methods, uncertainty heuristics, such as distance to the decision boundary, highest entropy, and expected risk minimization, have been widely studied. Bayesian active learning methods use probabilistic models, such as Gaussian processes or Bayesian neural networks, to estimate uncertainty. Instead of querying the most informative samples from an unlabeled pool, query synthesis algorithms introduce a generative adversarial active learning model to provide new synthetic samples that inform the current model. Since they mainly rely on the generated samples to train the classifier, their performance depends largely on the quality and diversity of the generated images. Many of the most successful active learning methods are pool-based active learning, which iteratively selects samples from an unlabeled pool for labeling based on an acquisition function that evaluates the informativeness of the subset for the training process. The selected subsets are manually annotated by experts, added to a labeled dataset, and used to update the target classifier trained on the new labeled dataset.
[0004] Much work in active learning has focused on developing effective acquisition functions, including those that select samples that produce high classifier uncertainty, high improvements in the Bayesian framework, or samples that are not well represented in the labeled set. However, all of these methods still have a significant performance gap compared to training on a large labeled dataset with a small annotation budget. Summary of the Invention
[0005] The present invention proposes an active learning method and system enhanced by adversarial training, which can balance the performance and robustness of the model, and rationally utilize the original limited labeled image samples. By using the method of adversarial generation model to select the unlabeled samples with the most information, the VAE network and the discriminant model f D both reach the state of Nash equilibrium, and can effectively label the unlabeled image data sampled externally or the image data in the existing unlabeled pool, reducing the manual annotation cost and burden greatly, and enabling the model to achieve high performance with a small amount of training data.
[0006] The present invention adopts the following technical solutions.
[0007] An active learning method enhanced by adversarial training includes the following steps;
[0008] Step A: Construct an adversarial trainer F, where F is composed of an adversarial sample generator f G and a target image recognition model f θ . Then input a small amount of labeled picture data set into the adversarial trainer F. The task of the adversarial trainer F is to complete the generation of adversarial samples X adv and the training of the target image recognition model f θ ;
[0009] Step B: Construct a variational autoencoder VAE, where VAE is divided into two parts: Encoder and Decoder; based on the adversarial samples obtained in Step A and the sampled labeled picture data set and the unlabeled picture data set use these data as the input of the VAE network to train the network; use the VAE network to effectively learn the underlying representation space Z of each different data, and use the obtained underlying representation space Z as the guiding direction for subsequent selection of unlabeled picture samples with the most information ;
[0010] Step C: Construct a discriminant model f D of a multi-layer perceptron, and form a generative adversarial network GAN with the VAE network constructed in Step B; use the underlying representation space Z obtained in Step B as the input of the discriminant model f D , and use the idea of confrontation to train the VAE network and the discriminant model f D . Finally, use the output of the discriminant model f D to select the unlabeled samples with the most information; in the confrontation process of VAE-f D , the VAE network aims at the discriminant model f DDiscriminate unlabeled data as labeled data as much as possible, and at the same time, the discrimination model f D is designed to correctly distinguish labeled and unlabeled image data as much as possible;
[0011] Step D: Select a set of candidate sets S with the maximum amount of information from the output of the discrimination model according to the request strategy Q of the specific task. The unlabeled image samples in S are characterized by being easily confused and difficult to discriminate. Through manual annotation, the newly obtained labeled data (X U , y) is injected into the labeled data set to update the data set, and then the target model f is incrementally trained and updated through the updated labeled image data set θ , so that the model can obtain higher benefits by training in as few data sets as possible;
[0012] Step E: Loop the above steps until the preset sampling ratio or the performance of the image recognition model is met.
[0013] The specific steps of the said Step A include the following steps:
[0014] Step A1: Use the existing small amount of labeled image data set as the input of the adversarial sample generator f G ,
[0015] f G Adopt replaceable adversarial sample generation methods:
[0016] 1. Projected Gradient Descent PGD;
[0017] 2. Fast Gradient Sign Method FGSM;
[0018] Then, use a generation method suitable for the task to generate a labeled adversarial sample data set in f G ;
[0019] Step A2: Use the adversarial sample data set obtained in Step A1 and the small amount of labeled image data set as the training set of the target model f θ ; Since the introduction of adversarial samples will cause a decrease in the model's accuracy, an auxiliary batch normalization layer is introduced into f θ to handle this problem, so that the model can better learn the data distributions of adversarial samples and clean samples; clean samples are the original samples;
[0020] Step A3: To better optimize the target model, so that the model can not only have better performance but also maintain good robustness, the objective function of the model is constructed as follows:
[0021]
[0022] where l(x L , y) and l(x adv , y) are the loss functions of the original sample and the adversarial sample under the setting of the auxiliary batch normalization layer, respectively.
[0023] Step B specifically includes the following steps:
[0024] Step B1: The encoder Encoder of the VAE network encodes the adversarial sample labeled image dataset and the unlabeled image dataset respectively. According to prior knowledge, it is assumed that the input data can be mapped to the Gaussian distribution Z to obtain the underlying representation space Z adv , Z L and Z U ; that is, for the input samples X adv , X L and X U , the encoder Encoder outputs the probabilities of the encoded Z adv , Z L and Z U as q φ (Z adv |X adv ), q φ (Z L |X L ) and q φ (Z U |X U );
[0025] Step B2: The role of the decoder Decoder of the VAE network is to reconstruct the underlying representation space Z adv , Z L and Z U obtained in Step B1, that is: the Decoder should be able to use Z adv , Z L and Z U to restore the original X adv , X L and X U samples, and the probabilities are represented as and
[0026] Step B3: The encoder Encoder and decoder Decoder composed of Step B1 and Step B2 efficiently complete the process of data mapping and reconstruction. The overall VAE network loss consists of two parts, the Gaussian distribution loss obtained from the prior knowledge of the Encoder and the loss of the original samples reconstructed by the posterior distribution of the Decoder. The objective function of the VAE network is as follows:
[0027]
[0028] Among them, the VAE network training loss includes labeled loss, unlabeled loss, and adversarial sample labeled loss; the first part is the likelihood of the reconstructed data, and the second part is the KL divergence of the approximate posterior distribution. The VAE network uses SGD with momentum as the optimizer and sets an appropriate weight decay parameter ε.
[0029] The specific steps of Step C are as follows:
[0030] Step C1: Construct a multi-layer perceptron discriminant model f D , according to the multiple type representation spaces Z adv , Z L and Z U obtained in Step B1 as the input of f D ; f D is designed as a binary classification prediction model for distinguishing labeled and unlabeled samples; therefore, the representation space of labeled samples includes Z adv , Z L , and the representation space of unlabeled samples is Z U ; to better implement the binary classification task, the representation spaces of the original samples and adversarial samples are fused to obtain a labeled joint representation space Put and Z U into f D to perform the prediction task;
[0031] Step C2: The discriminant model f D adopts binary cross-entropy loss BCEloss. When classifying, labeled samples will be classified as 1 and unlabeled samples as 0; the objective function of f D is to correctly classify labeled samples and unlabeled samples, and its objective function is constructed as follows:
[0032]
[0033] Among them, f D adopts Adam as the optimizer, which can adaptively adjust from two aspects of the gradient mean and the gradient square; and f DA multilayer perceptron model consisting of three layers of neural network fully connected layers is used. The calculation equation for each layer is as follows:
[0034] y=f(w fc ·x+b fc )Formula 4;
[0035] Among them, w fc is the weight matrix of the fully connected layer, f is the activation function, b fc is the bias term. D Use the ReLU function as the activation function of the multi-layer perceptron, and the last layer uses the Sigmoid function to output the classification probability;
[0036] Step C3: Combine the VAE network of step B and the discriminant network of step C D Jointly construct the Generative Adversarial Network (GAN); the VAE network is used as the generative model in the adversarial network, and its generated hidden layer representation is used as f D The input is intended to be either a labeled sample (X L , X adv ) or an unlabeled sample X U The hidden layer representation generated by the encoding will make the discriminant network f D Classify as labeled samples, that is, make the input to f as much as possible D The greater the probability that the sample is judged as 1; in addition, f D The adversarial operation aims to correctly classify the input labeled and unlabeled samples, and to make the unlabeled samples as likely to be judged as 0; therefore, the loss function of the VAE network in adversarial can be constructed as:
[0037]
[0038] Furthermore, the discriminant model f D The loss function can be constructed as follows:
[0039]
[0040] Step C4: The adversarial network constructed according to step C3 can further determine the overall training loss of the VAE network, which mainly consists of two parts. The first part is the encoding and reconstruction loss of the encoder and decoder. The second part is input to the discriminant model f D The loss generated by the hidden layer representation Therefore, the overall loss of the VAE network in the adversarial network is as follows:
[0041]
[0042] Among them, λ 1and λ 2 are used as hyperparameters to determine the influence of each part of the loss function;
[0043] Step C5: Optimize the VAE network and the discriminant model f D globally; respectively use the preset optimizer to make the VAE network and the discriminant model f D both trained to the best in N rounds of iterative optimization, so that the unlabeled samples with the largest amount of information can be selected by f D and picked out.
[0044] The said step D specifically includes the following steps:
[0045] Step D1: According to the probability distribution of the unlabeled data obtained in step C5, the request policy Q extracts B unlabeled samples with the most biased probability distribution towards 0 from the unlabeled samples with a probability distribution biased towards 0 according to the preset budget value B. These samples have the largest amount of information so that the model cannot correctly discriminate them. The data samples composed of such unlabeled samples are selected to form the candidate set of the samples with the largest amount of information
[0046] Step D2: Expertly label the obtained candidate set S, and the correctly labeled sample set is injected into the labeled image dataset for dataset update to obtain a new labeled image dataset
[0047] Step D3: The updated labeled image dataset retrains the target model f θ to improve the generalization and robustness of the model.
[0048] An active learning system enhanced by adversarial training for image recognition, using the above-mentioned active learning method enhanced by adversarial training, includes a target model adversarial training module, a VAE network module, and a binary classification label discrimination module:
[0049] The said target model adversarial training module: is used for the training of the image recognition model and the generation of adversarial samples, including an adversarial sample generation sub-module and a target model training sub-module; the specific method is: first, a limited number of labeled samples are input into the adversarial sample generation sub-module, and different adversarial sample generation algorithms are used to generate adversarial samples X according to the task adv ; then, the adversarial samples and the original samples are mixed and trained on the target model, and an auxiliary batch normalization layer is introduced into the model to maintain the balance between generalization and robustness; finally, the set loss function is optimized to update the parameters of the target model f θ ;
[0050] VAE network module: It is used to encode and generate hidden layer data representations and reconstruct the original data, and use the encoded and generated data representations as the input of the binary classification discriminant module. The specific method is as follows: First, use the encoder Encoder to generate adversarial samples labeled dataset and unlabeled dataset to encode and obtain the hidden layer representation distributions of each data; then, use the decoder Decoder to reconstruct each type of data to recover samples that are as similar as possible to the original data. Secondly, the hidden layer representation of the encoder Encoder is used as the input of the binary classification discriminant module, and the joint distribution of the labeled original samples X L and X adv is used as the guiding direction for selecting the unlabeled data with the most information; the binary classification discriminant module and the VAE network module form an adversarial generation network, so that the hidden layer representations generated by the VAE network can better deceive the discriminant model f D , making it misclassify unlabeled samples as labeled. During the adversarial process, the VAE network generates hidden layer representations, which also plays a role in the selection of unlabeled samples with the most information.
[0051] Binary classification discriminant module: It is used to distinguish labeled samples and unlabeled samples. The specific method is as follows: First, use the hidden layer data representation in the VAE network as the input, and calculate the probabilities of labeled and unlabeled predictions through the calculation of the multi-layer perceptron; secondly, the binary classification discriminant module and the VAE network module form an adversarial generation network, which can enable the discriminant model f D to better perform label classification during multiple rounds of objective function optimization, so that the most informative unlabeled samples can be better selected in the end.
[0052] The active learning system is used to embed an image recognition model for Internet of Things deployment.
[0053] The present invention can embed an active learning method and system based on adversarial training enhancement in the image recognition model for Internet of Things deployment. Compared with the prior art, it has the following beneficial effects:
[0054] (1) The present invention introduces a method of adversarial training data augmentation in active learning, reasonably expands the limited labeled data to generate adversarial samples, and uses the original samples and adversarial samples to jointly train the Internet of Things image recognition model. In addition, a batch normalization layer is introduced in the adversarial training to balance the accuracy degradation caused by the adversarial training, ensuring that the generated image recognition model has good performance and robustness.
[0055] (2) The present invention proposes an adversarial generation network framework. The hidden layer data representations of various types of data are learned through the VAE network, and then through the discriminant model f DTo determine the prediction probabilities of labeled and unlabeled samples, select the unlabeled data with the maximum amount of information for annotation, and achieve a reduction in the cost of manual annotation and an improvement in the efficiency of model training.
[0056] (3) The present invention also introduces the adversarial samples generated by adversarial training into the adversarial generative network framework, uses the adversarial samples to compensate for the information loss problem caused by the encoding process in the VAE network, and provides a guiding direction for selecting the unlabeled samples with the maximum information through the joint distribution of the adversarial samples and the original samples, so that annotating the selected samples can provide more benefits for the training of the target model.
[0057] (4) The system of the present invention introduces data samples - adversarial samples with a distribution different from that of the original data set compared with the existing methods. The distribution of the adversarial sample data is quite different from that of the original data set and is closer to the samples collected in the actual scenario. Compared with the existing methods, the designed active learning system embedded in image recognition can be better deployed in the actual Internet of Things scenario to process out-of-distribution data, achieving less labeled data and high model performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] The present invention will be further described in detail below in conjunction with the drawings and specific embodiments:
[0059] Attached Figure 1 is a schematic flowchart of the method according to an embodiment of the present invention;
[0060] Attached Figure 2 is a schematic structural diagram of the system according to an embodiment of the present invention. SPECIFIC EMBODIMENTS
[0061] As shown in the figure, an active learning method based on enhanced adversarial training includes the following steps;
[0062] Step A: Construct an adversarial trainer F, where F is composed of an adversarial sample generator f G and a target image recognition model f θ . Then, input a small amount of labeled picture data set into the adversarial trainer F. The task of the adversarial trainer F is to complete the generation of adversarial samples X adv and the training of the target image recognition model f θ .
[0063] Step B: Construct a variational autoencoder VAE, where VAE is divided into two parts: Encoder and Decoder; based on the adversarial samples obtained in Step A and the sampled labeled picture data set Use these data as the input of the VAE network to train the network; effectively learn the underlying representation space Z of each different data by the VAE network, and use the obtained underlying representation space Z as the subsequent guidance direction for selecting the unlabeled picture samples with the maximum information content.
[0064] Step C: Construct the discriminant model f of the multi-layer perceptron D , and form a generative adversarial network GAN with the VAE network constructed in Step B; use the underlying representation space Z obtained in Step B as the input of the discriminant model f D , and train the VAE network and the discriminant model f using the adversarial idea D , and finally select the unlabeled samples with the maximum information content using the output of the discriminant model f D ; in the adversarial process of VAE-f D , the VAE network aims to make the discriminant model f D classify the unlabeled data as labeled data as much as possible. At the same time, the discriminant model f D aims to correctly distinguish the labeled and unlabeled picture data as much as possible.
[0065] Step D: Select a set of candidate sets S with the maximum information content from the output of the discriminant model according to the request strategy Q of the specific task. The unlabeled picture samples in S are characterized by being easy to confuse and difficult to distinguish, and are manually labeled. Inject the newly labeled data (X U , y) into the labeled data set to update the data set, and then perform incremental training to update the target model f θ through the updated labeled picture data set, so that the model can obtain higher benefits by training in as few data sets as possible.
[0066] Step E: Loop the above steps until the preset sampling ratio or the performance of the image recognition model is satisfied.
[0067] The specific steps of Step A include the following steps:
[0068] Step A1: Use the existing small amount of labeled picture data set as the input of the adversarial sample generator f G .
[0069] f G adopts replaceable adversarial sample generation methods:
[0070] 1. Projected Gradient Descent PGD;
[0071] 2. Fast Gradient Sign Method FGSM;
[0072] Then adopt a generation method suitable for the task in f G Generate a labeled adversarial sample dataset
[0073] Step A2: The adversarial sample dataset obtained from Step A1 and a small amount of labeled image datasets serve as the training set of the target model f θ ; Since the introduction of adversarial samples will cause a decrease in the model's accuracy, an auxiliary batch normalization layer is introduced into f θ to address this issue, enabling the model to better learn the data distributions of adversarial samples and clean samples; Clean samples are the original samples;
[0074] Step A3: To better optimize the target model so that the model can have both better performance and good robustness, the objective function of the model is constructed as follows:
[0075]
[0076] where l(x L , y) and l(x adv , y) are the loss functions of the original samples and adversarial samples respectively under the settings of the auxiliary batch normalization layer.
[0077] Step B specifically includes the following steps:
[0078] Step B1: The encoder Encoder of the VAE network encodes the adversarial samples labeled image datasets and unlabeled image datasets respectively. Based on prior knowledge, it is assumed that the input data can be mapped to a Gaussian distribution Z to obtain the underlying representation space Z adv , Z L and Z U ; That is, for the input samples X adv , X L and X U , the encoder Encoder outputs the probabilities of encoding Z adv , Z L and Z U as q φ (Z adv |X adv ), q φ (Z L |X L ) and q φ (Z U |X U );
[0079] Step B2: The role of the decoder (Decoder) of the VAE network is to reconstruct the underlying representation space Z obtained in Step B1 adv , Z L and Z U That is, the Decoder should be able to use Z adv , Z L and Z U to restore the original X adv , X L and X U samples, and the probabilities are respectively and
[0080] Step B3: The encoder (Encoder) and decoder (Decoder) composed of Step B1 and Step B2 efficiently complete the process of data mapping and reconstruction. The overall loss of the VAE network consists of two parts: the Gaussian distribution loss obtained from the prior knowledge of the Encoder and the loss of the original samples reconstructed by the posterior distribution of the Decoder. The objective function of the VAE network is as follows:
[0081]
[0082] Among them, the training loss of the VAE network includes labeled loss, unlabeled loss, and adversarial sample labeled loss; the first part is the likelihood of the reconstructed data, and the second part is the KL divergence of the approximate posterior distribution. The VAE network uses SGD with momentum as the optimizer and sets an appropriate weight decay parameter ε.
[0083] The specific steps of Step C are as follows:
[0084] Step C1: Construct a multi-layer perceptron discriminant model f D , and use the multiple type representation spaces Z adv , Z L and Z U obtained in Step B1 as the input of f D ; f D is designed as a binary classification prediction model for distinguishing labeled and unlabeled samples; therefore, the representation space of labeled samples includes Z adv , Z L , and the representation space of unlabeled samples is Z U ; To better implement the binary classification task, the representation spaces of the original samples and adversarial samples are fused to obtain a labeled joint representation space Put and Z U into f D to perform the prediction task;
[0085] Step C2: Discriminant model f D The binary cross-entropy loss BCEloss is adopted. When classifying, labeled samples will be classified as 1 and unlabeled samples as 0; the objective function of f D is to correctly classify labeled and unlabeled samples, and its objective function is constructed as follows:
[0086]
[0087] where f D adopts Adam as the optimizer, which can adaptively adjust from two aspects of the gradient mean and the gradient square; and f D uses a multi-layer perceptron model composed of three fully-connected layers of a neural network. The calculation equation of each layer is as follows:
[0088] y = f(w fc ·x + b fc ) Equation Four;
[0089] where w fc is the weight matrix of the fully-connected layer, f is the activation function, and b fc is the bias term. f D uses the ReLU function as the activation function of the multi-layer perceptron, and the Sigmoid function is used in the last layer to output the classification probability;
[0090] Step C3: Combine the VAE network in Step B and the discriminant network f D in Step C to jointly construct a generative adversarial network GAN; the VAE network serves as the generative model in the adversarial network, and the generated hidden layer representation is used as the input of f D , aiming to make the discriminant network f L classify the hidden layer representations encoded from either labeled samples (X adv , X U ) or unlabeled samples X D as labeled samples, that is, to make the probability that the samples input to f D are classified as 1 as large as possible; in addition, the adversarial operation of f D aims to correctly classify the input labeled and unlabeled samples and make the unlabeled samples be classified as 0 as much as possible; therefore, the loss function of the VAE network in the adversarial process can be constructed as:
[0091]
[0092] Furthermore, the loss function of the discriminant model f D can be constructed as follows:
[0093]
[0094] Step C4: The adversarial network constructed according to Step C3 can further determine the overall training loss of the VAE network, which mainly consists of two parts. The first part is the loss of encoding and reconstruction between the encoder Encoder and the decoder Decoder. The second part is the loss generated by the hidden layer representation input to the discriminant model f. D Therefore, the overall loss of the VAE network in the adversarial network is as follows:
[0095]
[0096] where λ 1 and λ 2 are hyperparameters used to determine the influence of each part of the loss function.
[0097] Step C5: Optimize the VAE network and the discriminant model f D globally; respectively use the preset optimizers to train the VAE network and the discriminant model f D to the best in N rounds of iterative optimization, so that the unlabeled samples with the most information can be D selected by f.
[0098] The specific steps of Step D are as follows:
[0099] Step D1: According to the probability distribution of the unlabeled data obtained in Step C5, the request policy Q extracts B unlabeled samples with the most probability distribution biased towards 0 from the unlabeled samples with the probability distribution biased towards 0 according to the preset budget value B. These samples have the most information so that the model cannot correctly discriminate them. The data samples composed of such unlabeled samples are selected to form the maximum information sample candidate set.
[0100] Step D2: Expertly label the obtained candidate set S, and the correctly labeled sample set is injected into the labeled image dataset to update the dataset, and a new labeled image dataset is obtained.
[0101] Step D3: The updated labeled image dataset retrains the target model f θ to improve the generalization and robustness of the model.
[0102] An active learning system based on adversarial training enhancement for image recognition, using the above-mentioned active learning method based on adversarial training enhancement, includes a target model adversarial training module, a VAE network module, and a binary classification label discrimination module:
[0103] The target model adversarial training module: It is used for the training of the image recognition model and the generation of adversarial samples, including an adversarial sample generation sub-module and a target model training sub-module. The specific method is as follows: First, a limited number of labeled samples are input into the adversarial sample generation sub-module, and different adversarial sample generation algorithms are used according to the task to generate adversarial samples X adv ; Then, the adversarial samples and the original samples are used for mixed training of the target model, and an auxiliary batch normalization layer is introduced into the model to maintain the balance between generalization and robustness; Finally, the set loss function is optimized to update the parameters of the target model f θ .
[0104] The VAE network module: It is used to encode and generate hidden layer data representations and reconstruct the original data, and use the encoded and generated data representations as the input of the binary classification discriminant module. The specific method is as follows: First, use the encoder Encoder to process the adversarial samples the labeled dataset and the unlabeled dataset to obtain the hidden layer representation distributions of each data; Then, use the encoder Decoder to reconstruct each type of data to recover samples that are as close as possible to the original data; Secondly, the hidden layer representation of the encoder Encoder is used as the input of the binary classification discriminant module, and the joint distribution of the labeled original samples X L and X adv is used as the guiding direction for selecting the unlabeled data with the most information; The binary classification discriminant module and the VAE network module form an adversarial generation network, so that the hidden layer representation generated by the VAE network can better deceive the discriminant model f D , making it misclassify the unlabeled samples as labeled. During the adversarial process, the VAE network generates hidden layer representations, which also plays a role in the selection of the unlabeled samples with the most information.
[0105] The binary classification discriminant module: It is used to distinguish between labeled samples and unlabeled samples. The specific method is as follows: First, use the hidden layer data representation in the VAE network as the input, and calculate the probabilities of labeled and unlabeled predictions through the calculation of the multi-layer perceptron; Secondly, the binary classification discriminant module and the VAE network module form an adversarial generation network, and during the multi-round optimization process of the objective function, the discriminant model f D can better perform label classification, so that ultimately the unlabeled samples with the most information can be better selected.
[0106] The active learning system is used for the image recognition model embedded in the Internet of Things deployment.
[0107] The above are the preferred embodiments of the present invention. All changes made according to the technical solution of the present invention, when the functions and effects generated do not exceed the scope of the technical solution of the present invention, shall fall within the protection scope of the present invention.
Claims
1. An active learning method enhanced by adversarial training, characterized in that: It includes the following steps; Step A: Construct an adversarial trainer F, where F is composed of an adversarial sample generator f G and a target image recognition model f θ Then, input a small labeled image dataset into the adversarial trainer F. The task of the adversarial trainer F is to complete the generation of adversarial samples X adv and the training of the target image recognition model f θ ; Step B: Construct a variational autoencoder (VAE), where the VAE is divided into two parts: Encoder and Decoder; Based on the adversarial samples obtained in Step A and the sampled labeled image dataset and the unlabeled image dataset Use these data as the input of the VAE network to train the network; The VAE network effectively learns the underlying representation space Z of each different data, and use the obtained underlying representation space Z as the guiding direction for subsequent selection of the unlabeled image samples with the maximum amount of information ; Step C: Construct the discriminant model f of the multi-layer perceptron D , which forms a generative adversarial network GAN with the VAE network constructed in Step B; use the underlying representation space Z obtained in Step B as the input of the discriminant model f D , and train the VAE network and the discriminant model f using the adversarial idea D , and finally use the output of the discriminant model f D to select the unlabeled samples with the maximum amount of information; in the adversarial process of VAE-f D , the VAE network aims to make the discriminant model f D discriminate unlabeled data as labeled data as much as possible. At the same time, the discriminant model f D aims to correctly distinguish labeled and unlabeled image data as much as possible; Step D: Select a set of candidate sets S with the maximum amount of information from the output of the discriminant model according to the request strategy Q of the specific task. The unlabeled image samples in S have the characteristics of being easily confused and difficult to discriminate. Through manual annotation, the newly labeled data (X U , y) is injected into the labeled dataset to update the dataset. Then, incremental training is performed on the updated labeled image dataset to update the target model f θ , so that the model can obtain higher benefits by training in as few datasets as possible; Step E: Loop the above steps until the preset sampling ratio or the performance of the image recognition model is met.
2. The active learning method enhanced by adversarial training according to claim 1, characterized in that: The specific steps of step A are as follows: Step A1: Use a small existing labeled image dataset as the input of the adversarial sample generator f G . For f G , adopt replaceable adversarial sample generation methods:
1. Projected Gradient Descent (PGD); 2. Fast Gradient Sign Method (FGSM); then adopt a generation method suitable for the task to generate a labeled adversarial sample dataset in f G Step A2: The adversarial sample dataset obtained from Step A1 and a small amount of labeled image datasets serve as the training set of f θ of the target model; Since the introduction of adversarial samples will cause a decrease in model accuracy, an auxiliary batch normalization layer is introduced into f θ to address this issue, enabling the model to better learn the data distributions of adversarial samples and clean samples; Clean samples are the original samples; Step A3: In order to better optimize the target model so that the model can not only have better performance but also maintain good robustness, the objective function of the model is constructed as follows: where l(x L , y) and l(x adv , y) are the loss functions of the original sample and the adversarial sample under the setting of the auxiliary batch normalization layer, respectively.
3. The active learning method enhanced by adversarial training according to claim 1, characterized in that: The specific steps of step B are as follows: Step B1: The encoder Encoder of the VAE network encodes the adversarial samples the labeled image dataset and the unlabeled image dataset respectively. According to prior knowledge, it is assumed that the input data can be mapped to the Gaussian distribution Z to obtain the underlying representation space Z adv , Z L and Z U ; that is, for the input samples X adv , X L and X U , the encoder Encoder outputs the encoded Z adv , Z L and Z U with probabilities q φ (Z adv |X adv ), q φ (Z L |X L ) and q φ (Z U |X U ); Step B2: The role of the decoder of the VAE network, Decoder, is to reconstruct the underlying representation space Z obtained in Step B1 adv , Z L and Z U i.e., Decoder uses Z adv , Z L and Z U to restore the original X adv , X L and X U samples, and the probabilities are respectively and Step B3: The encoder Encoder and decoder Decoder composed of step B1 and step B2 efficiently complete the process of data mapping and reconstruction; the overall VAE network loss consists of two parts, the Gaussian distribution loss obtained from the prior knowledge of the Encoder and the loss of the original samples reconstructed by the posterior distribution of the Decoder. The objective function of the VAE network is as follows: Among them, the VAE network training loss includes labeled loss, unlabeled loss, and adversarial sample labeled loss; the first part is the likelihood of the reconstructed data, and the second part is the KL divergence of the approximate posterior distribution; the VAE network uses SGD with momentum as the optimizer and sets the weight decay parameter ε.
4. The active learning method enhanced by adversarial training according to claim 3, characterized in that: The specific steps of step C are as follows: Step C1: Construct a multi-layer perceptron discriminant model f D , based on the multiple type representation spaces Z obtained in step B1 adv , Z L and Z U as the input of f D ; f D is designed as a binary classification prediction model for distinguishing labeled and unlabeled samples; therefore, the representation space of labeled samples includes Z adv , Z L , and the representation space of unlabeled samples is Z U ; To perform the binary classification task, the representation spaces of the original samples and adversarial samples are fused to obtain a labeled joint representation space Input and Z U into f D to perform the prediction task; Step C2: Discriminant model f D The binary cross-entropy loss BCEloss is adopted. When performing classification, labeled samples will be classified as 1 and unlabeled samples as 0; for f D the objective function is to correctly classify labeled and unlabeled samples, and its objective function is constructed as follows: Formula three; Among them, f D Adopt Adam as the optimizer, which can adaptively adjust from two aspects of the gradient mean and the gradient square; and f D The multi-layer perceptron model composed of the fully connected layers of the three-layer neural network is used in, and the calculation equations of each layer are as follows: y = f(w fc ·x + b fc ) Formula Four; where, w fc is the weight matrix of the fully connected layer, f is the activation function, and b fc is the bias term; f D uses the ReLU function as the activation function of the multi-layer perceptron, and the Sigmoid function is used in the last layer to output the classification probability; Step C3: Combine the VAE network in Step B and the discriminant network f in Step C D to jointly construct a generative adversarial network GAN; the VAE network serves as the generative model in the adversarial network, and the generated hidden layer representation is used as the input to f D with the aim that whether it is a labeled sample X L 、X adv , or the hidden layer representation encoded from an unlabeled sample X U will cause the discriminant network f D to classify it as a labeled sample, that is, to make the probability that the sample input to f D is judged as 1 as large as possible; in addition, the adversarial operation of f D aims to correctly classify the input labeled and unlabeled samples and make the unlabeled samples be judged as 0 as much as possible; therefore, the loss function of the VAE network in the adversarial process is constructed as follows: Discriminant model f D The loss function is constructed as follows: Step C4: Determine the overall training loss of the VAE network according to the adversarial network constructed in Step C3, which consists of two parts. The first part is the loss of encoding and reconstruction of the encoder Encoder and the decoder Decoder The second part is the loss generated by the hidden layer representation input to the discriminant model f D Therefore, the overall loss of the VAE network in the adversarial network is as follows: Therefore, the overall loss of the VAE network in the adversarial network is as follows: Among them, λ 1 and λ 2 are used as hyperparameters to determine the influence of each part of the loss function; Step C5: Optimize the VAE network and the discriminant model f D globally; respectively use the preset optimizers to train the VAE network and the discriminant model f to the best in N rounds of iterative optimization D so that the unlabeled samples with the maximum amount of information can be selected by f D 5. The active learning method enhanced by adversarial training according to claim 3 or 4, characterized in that: The specific steps of step D are as follows: Step D1: According to the probability distribution of the unlabeled data obtained in Step C5, the request policy Q extracts B unlabeled samples with the most probability distribution biased towards 0 from the unlabeled samples with the probability distribution biased towards 0 according to the pre-set budget value B. These samples have the maximum amount of information such that the model cannot correctly discriminate them. The data samples composed of such unlabeled samples are selected to form the candidate set of the maximum information amount samples Step D2: Expertly annotate the obtained candidate set S, and inject the correctly annotated sample set into the labeled image dataset for dataset update, obtaining a new labeled image dataset Step D3: The updated labeled image dataset re-trains the target model f θ to improve the generalization and robustness of the model.
6. An active learning system enhanced by adversarial training for image recognition, using the active learning method enhanced by adversarial training described in claims 1, 2, 3, 4, or 5, characterized in that: It includes a target model adversarial training module, a VAE network module, and a binary classification label discrimination module: The target model adversarial training module: used for training the image recognition model and generating adversarial samples, including an adversarial sample generation sub-module and a target model training sub-module; the specific method is as follows: First, a limited number of labeled samples are input into the adversarial sample generation sub-module, and different adversarial sample generation algorithms are used according to the task to generate adversarial samples X adv ; then, the adversarial samples and the original samples are used to jointly train the target model, and in order to maintain the balance between generalization and robustness, an auxiliary batch normalization layer is introduced into the model; finally, the set loss function is optimized to update the parameters of the target model f θ ; VAE network module: It is used to encode and generate hidden layer data representations and reconstruct the original data, and use the encoded and generated data representations as the input of the binary classification discriminant module. The specific method is as follows: First, use the encoder Encoder to process adversarial samples labeled dataset and unlabeled dataset to encode and obtain the hidden layer representation distributions of each data; then, use the decoder Decoder to reconstruct each type of data to recover samples that are as similar as possible to the original data; second, the hidden layer representation of the encoder Encoder is used as the input of the binary classification discriminant module, and the joint distribution of the labeled original samples X L and X adv is used as the guiding direction for selecting the unlabeled data with the maximum amount of information; the binary classification discriminant module and the VAE network module form an adversarial generation network, so that the hidden layer representations generated by the VAE network can better deceive the discriminant model f D , making it misclassify unlabeled samples as labeled. During the adversarial process, the VAE network generates hidden layer representations, which also plays a role in the selection of unlabeled samples with the maximum amount of information; Binary classification discriminant module: used to distinguish labeled samples and unlabeled samples; the specific method is as follows: First, use the hidden layer data representation in the VAE network as input, and obtain the probabilities of labeled and unlabeled predictions through the calculation of a multi-layer perceptron; Second, the binary classification discriminant module and the VAE network module form an adversarial generation network, and the discriminant model f can be made better during the multi-round optimization process of the objective function D can better perform label classification, so that ultimately the unlabeled samples with the largest amount of information can be better selected.
7. The active learning system enhanced by adversarial training according to claim 6, characterized in that: The active learning system is used for an image recognition model embedded in the Internet of Things deployment.