Learning device, learning method, and learning program
By generating and superimposing perturbations based on clean images in the CNN, NN is trained to reduce the impact of perturbations and gradually increase the perturbation intensity, the problem of CNN in the prior art degradation of the clean image recognition accuracy when improving robustness is solved, and a balance between high robustness and high recognition accuracy is achieved.
Patent Information
- Application Number
- JP2022114304
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-07-15
- Publication Date
- 2025-05-08
- Estimated Expiration
- 2042-07-15
AI Technical Summary
The prior art, when improving the robustness of convolutional neural networks (CNNs) through malicious learning, results in a decrease in recognition accuracy of clean images without malicious perturbation, and cannot find a balance between improving robustness and maintaining clean image recognition accuracy.
By generating perturbations based on clean images and superimposing them on clean images, a superimposed image is formed, and the clean images are recognized by the first layer of neural network (NN), and the superimposed image is trained by the second layer of NN, the model parameters are optimized to reduce the impact of perturbations on clean images, while gradually increasing the intensity of perturbations, so that the CNN can improve the robustness of malicious perturbations while maintaining the accuracy of the recognition of clean images.
It is achieved to maintain high recognition accuracy for clean images while improving the robustness of CNNs to combat malicious perturbations, avoiding the problem of degradation of recognition accuracy in the prior art.
Smart Images

Figure 0007672655000006 
Figure 0007672655000007 
Figure 0007672655000008
Abstract
Description
[Technical field]
[0001] The present invention relates to a learning device, a learning method, and a learning program. [Background technology]
[0002] In recent years, there has been a remarkable improvement in the accuracy of machine learning technology, particularly the technology using Convolutional Neural Networks (CNNs), for identifying and detecting objects in images, and for segmenting areas. Technologies that use these machine learning technologies to promote the automation of visual inspection processes in various business processes are attracting attention.
[0003] For example, when automating the visual inspection process of a business by processing captured images, it is desirable for the CNN image processing to be in line with the intuition of the human performing the visual inspection process. It is desirable for the CNN image processing to have a property such that the classification results are unchanged even if some noise is added to the classification target during image classification.
[0004] CNN is known as a model with a mechanism similar to the human visual structure, and has been thought to have the same invariance to noise as humans.
[0005] However, in recent years, it has become clear that CNN image processing is not invariant to minute noise that humans cannot perceive. In particular, it has become clear that a vulnerability attack called adversarial perturbation can cause CNN image classification to be wrong almost 100% of the time, even if the noise is minute and barely perceptible to humans, if no countermeasures are taken, making it possible for CNN to identify an image as a car even though it is actually a dog from a human perspective.
[0006] The behavior of CNNs in response to such adversarial perturbations that are different from human intuition can be a major challenge in practical use. Therefore, building a CNN model that is robust against adversarial perturbations is an important challenge for the practical use of CNNs.
[0007] Various methods have been proposed to build CNNs that are robust against attacks by adversarial perturbations. However, many of these methods have been found to be vulnerable to new attack methods that were developed later.
[0008] However, among the methods for constructing a robust CNN, a method called adversarial learning (Non-Patent Document 1) is known as a methodology for creating a robust CNN model without significantly reducing classification accuracy even against the latest attack methods.
[0009] Adversarial learning creates adversarial perturbations that attack the CNN model to the maximum extent possible, and trains the model to correctly identify images to which the perturbations have been added. Adversarial learning builds a robust CNN model by repeating this learning process. [Prior art documents] [Non-patent literature]
[0010] [Non-Patent Document 1] M A. Madry, A. Makelov, L. Schmidt, D. Tsipras, and A. Vladu, “Towards Deep Learning Models Resistant to Adversarial Attacks”, [online], arXiv:1706.06083v4 [stat.ML], September 2019, [Retrieved June 2, 2022], Internet<URL: https: / / arxiv.org / pdf / 1706.06083.pdf> Summary of the Invention [Problem to be solved by the invention]
[0011] While adversarial learning has been empirically proven to be an effective method for the purpose of increasing robustness, it has the problem that its classification accuracy for images to which no adversarial perturbation has been applied (hereinafter referred to as clean images) is significantly lower than that of a CNN trained only on clean images (hereinafter referred to as normal training CNN).
[0012] For example, when using the WideResNet-34-10 architecture on the CIFAR-10 image dataset, the normal CNN model achieves an accuracy of about 95% for normal images, while the model that has undergone adversarial learning achieves an accuracy of about 85%. The model that has undergone adversarial learning has an accuracy rate that is 10 points lower than the normal CNN model.
[0013] While improving robustness against attacks by malicious third parties using adversarial perturbations is an important issue, the accuracy of normal images is also very important in practice, so it is not practical to significantly reduce the accuracy of normal images in order to improve robustness.
[0014] The present invention has been made in consideration of the above, and aims to provide a learning device, a learning method, and a learning program that provide an image recognition model with improved robustness through adversarial learning while maintaining accuracy for clean images. [Means for solving the problem]
[0015] In order to solve the above-mentioned problems and achieve the object, the learning device of the present invention has a generation unit that generates a perturbation based on a first clean image to which a correct label is attached, and generates a superimposed image by superimposing the generated perturbation on the first clean image; a first image classification unit that performs image classification on the first clean image using a first NN (Neural Network) that has learned image classification using a plurality of clean images as learning data, and outputs an intermediate output of the first NN as a first intermediate output; a second image classification unit that performs image classification on the superimposed image using a second NN in which model parameters of the first NN are set as initial parameters, and outputs the intermediate output of the second NN as a second intermediate output; and a learning unit that executes learning of the second NN by optimizing parameters of the second NN so that a difference between an image classification result for the superimposed image by the second image classification unit and a correct label of the first clean image is small, and a difference between the first intermediate output and the second intermediate output is small. Effect of the Invention
[0016] According to the present invention, it is possible to provide an image recognition model that maintains accuracy for clean images while improving robustness through adversarial learning. [Brief description of the drawings]
[0017] [Figure 1] FIG. 1 is a diagram explaining the output of the intermediate layer in a CNN trained only on clean images with correct labels, and a CNN that has undergone adversarial training. [Diagram 2] FIG. 2 is a diagram illustrating the learning method according to the first embodiment. [Diagram 3] FIG. 3 is a diagram illustrating an example of the configuration of the learning device according to the first embodiment. As shown in FIG. [Figure 4] FIG. 4 is a flowchart of a learning process according to the first embodiment. [Diagram 5] FIG. 5 is a flowchart showing the procedure of the adversarial image generation process shown in FIG. [Figure 6] FIG. 6 is a flowchart showing the procedure of the image identification process shown in FIG. [Figure 7] FIG. 7 is a flowchart illustrating the procedure of the optimization process shown in FIG. [Figure 8] FIG. 8 is a flowchart illustrating a processing procedure of the schedule processing shown in FIG. [Figure 9] FIG. 9 is a flowchart illustrating another processing procedure of the adversarial image generation processing shown in FIG. [Figure 10] FIG. 10 is a diagram illustrating a learning method according to the second embodiment. [Figure 11] FIG. 11 is a diagram illustrating an example of the configuration of a learning device according to the second embodiment. As illustrated in FIG. [Figure 12] FIG. 12 is a flowchart illustrating a processing procedure of the learning process according to the second embodiment. [Figure 13] FIG. 13 is a flowchart showing the processing procedure of the image identification processing shown in FIG. [Figure 14] FIG. 14 is a flowchart showing the processing procedure of the original image identification processing shown in FIG. [Figure 15] FIG. 15 is a flowchart illustrating the procedure of the optimization process shown in FIG. [Figure 16] FIG. 16 is a diagram illustrating an example of the configuration of a learning device according to the third embodiment. As shown in FIG. [Figure 17] FIG. 17 is a flowchart illustrating a processing procedure of the learning process according to the third embodiment. [Figure 18] FIG. 18 is a diagram illustrating an example of a computer that realizes a learning device by executing a program. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0018] Hereinafter, an embodiment of the present invention will be described in detail with reference to the drawings. Note that the present invention is not limited to this embodiment. In addition, in the description of the drawings, the same parts are indicated by the same reference numerals.
[0019] [Embodiment 1] In the first embodiment, a description will be given of learning of a Convolutional Neural Network (CNN) that improves robustness through adversarial learning while maintaining image recognition accuracy for images not superimposed with adversarial perturbations (hereinafter, clean images). CNN is an image recognition model.
[0020] Figure 1 is a diagram explaining the output of the intermediate layers in a CNN trained only on clean images with correct labels (hereinafter referred to as normal training CNN) (first NN) and a CNN that has undergone adversarial training (AT). Figure 1 shows the results of a survey using the 27th layer representation of WideResNet28-10 using the CIFAR-10 dataset.
[0021] As shown in Figure 1, the output of the intermediate layer (feature representation) obtained by the normal learning CNN model ((1) in Figure 1) and the output of the intermediate layer obtained by the CNN model that has undergone adversarial learning only matched about 0.2546 in terms of Cos similarity, indicating a large difference.
[0022] This result suggests that feature extraction that was successful with normal training CNN cannot be performed with a CNN model that has undergone AT, resulting in a significant drop in accuracy. In the first embodiment, constraints are imposed so that AT can acquire feature representations equivalent to those for clean images acquired during normal training, in other words, so that feature extraction similar to that of normal training CNN can be performed.
[0023] FIG. 2 is a diagram for explaining the learning method in the first embodiment. In the first embodiment, a clean image x clean (x) is convoluted with the adversarial perturbation δ to obtain the adversarial image x advWhen learning (x+δ) (overlapped image), curriculum learning based on the magnitude of the adversarial perturbation δ is performed ((1) in FIG. 2). In other words, in the first embodiment, the adversarial perturbation δ is generated so that the magnitude of the adversarial perturbation δ increases stepwise, so that the adversarial image x adv We conduct curriculum learning of AT-CNN (second NN) for the above.
[0024] For example, we use equation (1) as the loss function to optimize the parameters of AT-CNN.
[0025]
number
[0026] Here, f(θ,x+δ) is the function of (x+δ)(the adversarial image x adv ) is input. Equation (1) expresses the label output by AT-CNN (model parameter θ) when the clean image x clean This is a loss function that indicates the error between the correct label attached to (x).
[0027] Then, by performing Taylor expansion on equation (1), it can be approximated as equation (2).
[0028]
number
[0029] As shown in equation (2), the first term on the right side is a loss function that indicates the error between the label output by AT-CNN when a clean image x is input and the correct label y of the clean image. The first term on the right side is a term that indicates the error between the classification result for the clean image and the correct label of the clean image.
[0030] The second and subsequent terms on the right side of equation (2) are terms related to the adversarial perturbation δ, and can be said to indicate the degree of influence on the error between the classification result for the clean image and the correct label of the clean image. In other words, the magnitude of the influence of the second and subsequent terms on the right side of equation (2) is dominated by the adversarial perturbation δ, and can be said to be terms specific to AT.
[0031] In the first embodiment, the magnitude of the adversarial perturbation δ is |δ| ∞ In other words, in the first embodiment, the magnitude of |δ| is gradually increased from 0 (δ=0 is the same as in normal learning). ∞ That is, in the first embodiment, the adversarial perturbation δ is generated so that the influence of the adversarial perturbation δ on the error between the classification result for the clean image and the correct label for the clean image in the second and subsequent terms on the right side of the formula (2) increases stepwise.
[0032] As a result, in the first embodiment, the adversarial image x adv Curriculum learning is performed by performing step-by-step learning of the influence of the adversarial perturbation δ of AT-CNN (second NN) on .
[0033] [Learning device] Next, a description will be given of the learning device according to embodiment 1. Fig. 3 is a diagram illustrating an example of the configuration of the learning device according to embodiment 1.
[0034] The learning device 10 according to the embodiment is realized, for example, by loading a predetermined program into a computer or the like including a ROM (Read Only Memory), a RAM (Random Access Memory), a CPU (Central Processing Unit), etc., and having the CPU execute the predetermined program. The learning device 10 also has a communication interface for transmitting and receiving various information to and from other devices connected via a network or the like. The learning device 10 is realized by a general-purpose computer such as a workstation or a personal computer.
[0035] As shown in FIG. 3, the learning device 10 includes an image storage unit 11, an initialization unit 12, an adversarial perturbation generation unit 13 (generation unit), an image classification unit 14, an optimization unit 15 (learning unit), and a schedule unit 16 (control unit).
[0036] The image storage unit 11 stores a plurality of pairs of clean images and correct labels indicating categories of objects of the clean images.
[0037] The initialization unit 12 holds initial parameters of the CNN 141, which is an image recognition model. The initial parameters are model parameters of a CNN (first NN) that has previously learned image recognition using clean images as learning data. Note that the initial parameters are sufficient as long as they are model parameters of a model that can appropriately recognize clean images.
[0038] The adversarial perturbation generation unit 13 generates an adversarial perturbation (perturbation) based on a clean image (first clean image) output from the image memory unit 11 as a processing target, and generates a superimposed image by superimposing the generated adversarial perturbation on the first clean image.
[0039] The adversarial perturbation generator 13 generates an adversarial perturbation for the first clean image based on the schedule result by the scheduler 16, the model parameters of the CNN 141 (described later) optimized by the optimizer 15, the first clean image, and the ground truth label of the first clean image. The adversarial perturbation generator 13 outputs the generated superimposed image to the image classifier 14 as an adversarial image.
[0040] The schedule result is set to generate the adversarial perturbation so that the magnitude of the adversarial perturbation increases stepwise according to the learning stage of the CNN 141. The schedule result is, for example, control information for controlling parameters such as the norm of the perturbation of the adversarial perturbation. The control target of the schedule result is not limited to the norm of the perturbation as long as it is a parameter that does not significantly impair the feature representation learning of the CNN 141.
[0041] The image classification unit 14 receives an adversarial image to be classified from the adversarial perturbation generation unit 13, and performs image classification on the adversarial image using the CNN 141. First, at the start of learning, model parameters of a CNN that has previously learned image classification using clean images as training data are set as initial parameters in the CNN 141. Then, as learning progresses, model parameters of the CNN 141 optimized by the optimization unit 15 are set in the CNN 141. The image classification unit 14 outputs the obtained output (image classification result) to the optimization unit 15 as a model output.
[0042] The optimization unit 15 evaluates the CNN 141 and optimizes the model parameters of the CNN 141 based on the model output and the correct label. The optimization unit 15 optimizes the model parameters of the CNN 141 so that the model output for the adversarial image approaches the correct label of the first clean image, thereby executing learning of the CNN 141.
[0043] The optimization unit 15 optimizes the parameters of the CNN 141 using the loss function of the above-mentioned formula (2). As described above, formula (2) has a term indicating the error between the classification result for the first clean image and the correct label of the first clean image, and a term indicating the influence of the adversarial perturbation on this error. Under the control of the scheduler 16, the optimization unit 15 performs step-by-step learning on the influence of the adversarial perturbation on the error between the classification result for the first clean image of the CNN 141 and the correct label of the first clean image.
[0044] The optimization unit 15 outputs the evaluation result of the CNN 141 to the schedule unit 16. The optimization unit 15 outputs the model parameters of the optimized CNN 141 to the adversarial perturbation generation unit 13 and the image classification unit .
[0045] The scheduler 16 determines generation parameters of the adversarial perturbation generated by the adversarial perturbation generator 13 based on the evaluation result of the CNN 141. The scheduler 16 outputs the determined generation parameters to the adversarial perturbation generator 13 as a schedule result.
[0046] The scheduler 16 causes the adversarial perturbation generator 13 to generate the adversarial perturbation so that the magnitude of the adversarial perturbation increases stepwise. The scheduler 16 causes the adversarial perturbation generator 13 to generate the adversarial perturbation so that the magnitude of the adversarial perturbation increases stepwise according to the learning stage of the CNN 141, and causes the optimizer 15 to execute stepwise learning on the influence of the adversarial perturbation on the error between the classification result of the CNN for the first clean image and the correct label of the first clean image.
[0047] When the scheduler 16 determines that the learning is to be terminated, the scheduler 16 outputs the model parameters of the CNN 141 optimized by the optimizer 15 as learned parameters.
[0048] [Learning process] Next, a description will be given of the learning process according to the embodiment 1. Fig. 4 is a flowchart showing the processing procedure of the learning process according to the embodiment 1.
[0049] As shown in FIG. 4, in the learning device 10, first, the initialization unit 12 outputs initial parameters of the CNN 141 to the adversarial perturbation generation unit 13 and the image classification unit 14 in order to initialize the parameters of the CNN 141 (step S1).
[0050] The learning device 10 extracts a set of first clean images and a correct label corresponding to the first clean images from the image storage unit 11, transmits the first clean images and the correct label to the adversarial perturbation generation unit 13, and transmits the correct label to the optimization unit 15 (step S2).
[0051] The adversarial perturbation generator 13 performs an adversarial image generation process (step S3). The adversarial image is an image in which the adversarial perturbation is superimposed on the first clean image.
[0052] The image classification unit 14 performs image classification processing on the adversarial image generated by the adversarial perturbation generation unit 13 using the CNN 141 (step S4).
[0053] The optimization unit 15 performs an optimization process to evaluate the CNN 141 and optimize the model parameters of the CNN 141 based on the model output and the correct label (step S5).
[0054] Based on the evaluation result of the CNN 141, the scheduler 16 performs a schedule process to determine the generation parameters of the adversarial perturbation so that the magnitude of the adversarial perturbation generated by the adversarial perturbation generator 13 increases stepwise according to the learning stage of the CNN 141 (step S6).
[0055] The learning device 10 determines whether or not to end the learning process (step S7). If it is determined that the learning process should not be ended (step S7: No), the learning device 10 returns to step S2 and continues learning the CNN 141. If it is determined that the learning process should be ended (step S7: Yes), the learning device 10 acquires the model parameters of the optimized CNN 141 as learned parameters via the scheduler 16 (step S8).
[0056] [Adversarial image generation processing] Next, the adversarial image generation process (step S3) will be described. FIG 5 is a flowchart showing the procedure of the adversarial image generation process shown in FIG 4.
[0057] First, the adversarial perturbation generator 13 determines whether or not it is the start time of learning (step S11).
[0058] When learning starts (step S11: Yes), the adversarial perturbation generator 13 acquires initial parameters from the initialization unit 12 (step S12). The initial parameters are model parameters of a CNN that has previously learned image recognition using clean images as learning data, but any model parameters of a model that can appropriately recognize clean images will suffice. The adversarial perturbation generator 13 initializes the schedule result (step S13).
[0059] If it is not the start of learning (step S11: No), the adversarial perturbation generator 13 obtains a schedule result from the scheduler 16 (step S14). The schedule result is control information for controlling parameters such as the perturbation norm of the adversarial perturbation, but is not limited to the perturbation norm as long as the parameters do not significantly impair the feature representation learning of the CNN 141. The adversarial perturbation generator 13 obtains the model parameters of the optimized CNN 141 from the optimizer 15 (step S15), and thereafter uses them as image identification parameters.
[0060] The adversarial perturbation generator 13 acquires the first clean image to be processed and the ground truth label of the first clean image from the image storage unit 11 (step S16).
[0061] The adversarial perturbation generator 13 generates an adversarial perturbation based on the schedule result, the ground truth label of the first clean image, the first clean image, and the model parameters of the CNN 141 acquired in step S12 or step S15 (step S17). Here, a method for maximizing the loss for the ground truth label is generally used to generate the adversarial perturbation. In addition, although it is assumed that Projected Gradient Descent (Non-Patent Document 1) is used as a method for generating the adversarial perturbation, other methods may be used as long as they can generate the adversarial perturbation appropriately.
[0062] The adversarial perturbation generator 13 superimposes the generated adversarial perturbation on the first clean image, and outputs the result as an adversarial image to the image classifier 14 (step S18).
[0063] [Image recognition processing] Next, the image identification process (step S4) will be described with reference to a flowchart shown in FIG.
[0064] As shown in FIG. 6, the image discrimination unit 14 determines whether or not it is the start time of learning (step S21).
[0065] When learning starts (step S21: Yes), the image classification unit 14 acquires initial parameters from the initialization unit 12 (step S22). These initial parameters are assumed to be similar to the initial parameters acquired by the adversarial perturbation generation unit 13 in step S12, but even if the parameters are different, a similar effect can be obtained.
[0066] If it is not the start of learning (step S21: No), the image discrimination unit 14 acquires the model parameters of the optimized CNN 141 from the optimization unit 15 (step S23), and thereafter uses these as image discrimination parameters.
[0067] The image classification unit 14 acquires the adversarial image output by the process of step S18 from the adversarial perturbation generation unit 13 (step S24). The image classification unit 14 applies image classification based on the model parameters of the optimized CNN 141 to the adversarial image (step S25). The image classification unit 14 executes image classification processing on the adversarial image using the CNN 141 to which the model parameters acquired in step S22 or step S23 are set.
[0068] The image discrimination unit 14 outputs the output distribution obtained by the processing of step S25 to the optimization unit 15 as a model output (step S26).
[0069] [Optimization process] Next, the optimization process (step S5) will be described with reference to a flowchart shown in FIG.
[0070] As shown in FIG. 7, the optimization unit 15 determines whether or not it is the start time of learning (step S31).
[0071] If it is the start of learning (step S31: Yes), the optimization unit 15 acquires initial parameters from the initialization unit 12 (step S32).
[0072] If it is not the start of learning (step S31: No), or after step S32 is completed, the optimization unit 15 acquires the correct label of the first clean image to be processed from the image storage unit 11 (step S33). The optimization unit 15 outputs the model output from the image classification unit 14 (step S34).
[0073] The optimization unit 15 optimizes the model parameters of the CNN 141 so that the difference between the model output x and the correct label y becomes small, and outputs the optimized parameters (step S35). The optimization unit 15 optimizes the parameters using the loss function shown in Equation (2).
[0074] The optimization unit 15 may also optimize the image classification parameters using an image classification loss that reduces the difference between the model output x and the correct label y, and output the optimized parameters. Note that, for this image classification loss, the cross entropy shown in formula (3) or the like is generally used, but the mean square error or the like can also have a similar effect as long as it is an appropriate objective function for the desired classification task.
[0075]
number
[0076] The optimization unit 15 evaluates how close the model output is to the correct label, and outputs the evaluation result (step S36). For this evaluation, in addition to whether or not the result of taking the argmax of the model output matches the correct label, any appropriate evaluation index may be used.
[0077] [Schedule Processing] Next, the schedule process (step S6) will be described with reference to a flowchart of FIG.
[0078] 8, the scheduler 16 acquires the evaluation result output by the optimizer 15 in step S36 (step S41). The scheduler 16 determines whether the evaluation result acquired in step S41 is equal to or greater than a reference value (step S42).
[0079] If the evaluation result is equal to or greater than the reference value (step S42: Yes), the scheduler 16 determines whether or not to terminate the learning of the CNN 141 based on a predetermined termination condition (step S43). The termination condition for the learning is, for example, when the parameter generating the adversarial perturbation reaches a desired value, when the number of updates of the parameter of the CNN 141 reaches a predetermined number, when the amount of parameter updates becomes equal to or less than a predetermined threshold, when the loss calculated by the optimizer 15 becomes equal to or less than a predetermined threshold, etc.
[0080] When the scheduler 16 determines that the learning of the CNN 141 is to be ended (step S43: Yes), the learning device 10 ends the learning of the CNN 141.
[0081] If the scheduler 16 determines not to terminate the learning of the CNN 141 (step S43: No), it outputs parameters for generating new adversarial perturbations, which are obtained by updating the parameters used up to that point, as the schedule result to the adversarial perturbation generator 13 (step S44).
[0082] In order to allow the CNN 141 to learn the degree of influence of the adversarial perturbation in stages, the adversarial perturbation is set to become larger in stages according to the learning stage. For this reason, when the accuracy of image classification for the adversarial image to be verified is saturated for 10 epochs, the schedule unit 16 outputs a schedule result in which the norm of the perturbation of the adversarial perturbation is increased from the current stage to the magnitude of the stage next to the current stage.
[0083] If the evaluation result is less than the reference value (step S42: No), the scheduler 16 outputs a schedule result for maintaining the current parameters with respect to the generation of the adversarial perturbation to the adversarial perturbation generator 13 (step S45).
[0084] [Advantages of the First Embodiment] In this manner, in the first embodiment, the CNN 141 is trained with the constraint that the magnitude of the adversarial perturbation is gradually increased so that the feature representation does not deviate significantly from that of the normal CNN that has already trained on clean images.
[0085] As a result, in the first embodiment, by having the CNN 141 gradually learn the degree of influence of adversarial perturbation, it is possible to provide an image recognition model that maintains accuracy for clean images while improving robustness through adversarial learning.
[0086] [Modification of the first embodiment] Moreover, the adversarial perturbation generation unit 13 may apply a most confusing attack in generating the adversarial perturbation. 9 is a flowchart showing another processing procedure of the adversarial image generation processing shown in FIG.
[0087] Steps S51 to S56 in FIG. 9 are the same processes as steps S11 to S16 shown in FIG.
[0088] The adversarial perturbation generator 13 performs image classification processing on the first clean image using a CNN based on the optimized model parameters acquired in step S55, and searches for a category that has the largest output value (the most confusing category) other than the correct label of the first clean image from the output result (step S57). Note that since it is sufficient for the adversarial perturbation generator 13 to acquire an image classification result for the first clean image using a CNN based on the optimized model parameters, this CNN is not limited to a configuration included in the adversarial perturbation generator 13, and may be included in another component.
[0089] The adversarial perturbation generator 13 creates adversarial perturbations based on the schedule result, the first clean image, the Most confusing category acquired in step S57, and the model parameters of the CNN 141 (step S58). At this time, the adversarial perturbation generator 13 generates the adversarial perturbations so as to minimize the loss for the Most confusing category. Although it is assumed that Projected Gradient Descent (Non-Patent Document 1) will be used as the generation method, other methods may be used as long as they can generate appropriate adversarial perturbations.
[0090] Step S59 shown in FIG. 9 is the same process as step S18 shown in FIG.
[0091] Since the Most confusing category is the class with the highest confidence other than the correct label, the adversarial perturbation generated based on this Most confusing category corresponds to the data point of the closest decision boundary, and therefore, even if it is superimposed on the first clean image, it is an attack that does not significantly change the feature representation of the first clean image. For example, while the cosine similarity between the feature representation of the adversarial perturbation image generated by the processing procedure shown in Fig. 5 and the first clean image is 0.3896, the cosine similarity between the feature representation of the adversarial image generated based on the Most confusing category and the first clean image is close to 0.4485.
[0092] According to the modification of the first embodiment, by having the CNN 141 learn the adversarial images generated based on the Most confusing category, it is possible to realize more step-by-step learning of the influence of the adversarial perturbation.
[0093] [Embodiment 2] Next, a description will be given of embodiment 2. Fig. 10 is a diagram for explaining a learning method in embodiment 2.
[0094] In embodiment 2, a constraint is imposed so that the feature representation of AT-CNN (second NN) does not change from the feature representation of normal training CNN (first NN), thereby preventing AT-CNN from acquiring a feature representation different from that of normal training CNN.
[0095] Specifically, in the second embodiment, knowledge distillation is performed for training AT-CNN using a normal training CNN as a teacher ((1) in FIG. 10). Here, knowledge distillation from the output layer of CNN has only a regularization effect. For this reason, in the second embodiment, training is performed by directly constraining the output of the intermediate layer so that AT-CNN can perform feature extraction similar to that of a normal training CNN. In the second embodiment, training is performed using, for example, cos similarity as a loss function for constraints.
[0096] [Learning device] Next, a description will be given of a learning device according to embodiment 2. Fig. 11 is a diagram illustrating an example of the configuration of the learning device according to embodiment 2.
[0097] As shown in Fig. 11, learning device 210 according to embodiment 2 has a configuration in which schedule unit 16 is deleted, as compared with learning device 10 shown in Fig. 3. Moreover, learning device 210 has an image classification unit 214 (second image classification unit) and an optimization unit 215 instead of image classification unit 14 and optimization unit 15, as compared with learning device 10. Learning device 210 further has an original image classification unit 216 (first image classification unit).
[0098] The original image classification unit 216 performs image classification on the first clean image using a first CNN 2161 (first NN) that has learned image classification using a plurality of clean images as learning data. The original image classification unit 216 outputs an intermediate output of the first CNN 2161 to the optimization unit 215 as an original image intermediate output (first intermediate output). Note that the intermediate output is assumed to be a numerical value of the processing result of a layer other than the final layer of the first CNN 2161, but may be any other output as long as it is an appropriate feature representation for image classification.
[0099] The image classification unit 214 has a second CNN 2141 to be learned, and has the same function as the image classification unit 14. The model parameters of the first CNN 2161 are set as initial parameters of the second CNN 2141. The image classification unit 214 outputs an intermediate output (second intermediate output) of the second CNN 2141 together with the model output of the second CNN 2141 to the optimization unit 15. The intermediate output is a numerical value of the processing result of the same hierarchical level as the processing hierarchical level of the original image intermediate output. Also, like the original image intermediate output, the intermediate output may be another output as long as it is an appropriate feature representation for image classification.
[0100] The optimization unit 215 performs learning of the second CNN 2141 by optimizing the parameters of the second CNN 2141 so that the difference between the image classification result for the adversarial image by the image classification unit 214 and the correct label of the first clean image is reduced, and the difference between the original image intermediate output and the intermediate output of the second CNN 2141 is reduced.
[0101] The optimization unit 215 optimizes the parameters of the second CNN 2141 using a loss function L that is a weighted sum of a first loss function L1 and a second loss function L2, as shown in equation (4).
[0102]
number
[0103] In formula (4), the first loss function L1 is a loss function indicating the difference between the image classification result for the adversarial image by the image classification unit 214 and the correct label of the first clean image. The second loss function L2 is a loss function indicating the difference between the intermediate output of the original image and the intermediate output of the second CNN 2141. α is a hyperparameter.
[0104] [Learning process] Next, a description will be given of the learning process according to the embodiment 2. Fig. 12 is a flowchart showing the processing procedure of the learning process according to the embodiment 2.
[0105] Steps S201 to S203 shown in Fig. 12 are the same processes as steps S1 to S3 shown in Fig. 4. In step S203, the adversarial perturbation generator 13 may generate the adversarial perturbation by applying the most confusing attack (Fig. 9).
[0106] The image classification unit 214 uses the second CNN 2141 to perform image classification processing on the adversarial image generated by the adversarial perturbation generation unit 13 (step S204), and outputs a model output and an intermediate output to the optimization unit 215.
[0107] The original image classification unit 216 performs an original image classification process for performing image classification on the first clean image using the first CNN 2161 (step S205), and outputs an original image intermediate output to the optimization unit 215.
[0108] The optimization unit 215 performs an optimization process to optimize the parameters of the second CNN 2141 so that the difference between the image classification result for the adversarial image by the image classification unit 214 and the correct label of the first clean image is reduced, and the difference between the intermediate output of the original image and the intermediate output of the second CNN 2141 is reduced (step S206).
[0109] Steps S207 and S208 are the same processes as steps S7 and S8 shown in FIG.
[0110] [Image recognition processing] Next, the image identification process (step S204) will be described. FIG 13 is a flowchart showing the procedure of the image identification process shown in FIG.
[0111] Steps S211 to S215 shown in Fig. 13 are the same processes as steps S21 to S25 shown in Fig. 6. The image classification unit 214 outputs the model output obtained by the process of step S215 and the intermediate output of the second CNN 2141 to the optimization unit 15 (step S216).
[0112] [Original image identification processing] Next, the original image identification process (step S205) will be described below. Fig. 14 is a flowchart showing the procedure of the original image identification process shown in Fig. 12.
[0113] The original image discrimination unit 216 determines whether or not it is the start of learning (step S221). If it is the start of learning (step S221: Yes), the original image discrimination unit 216 acquires initial parameters from the initialization unit 12 (step S222).
[0114] If it is not the start of learning (step S221: No), or after step S222 is completed, the original image classification unit 216 acquires a first clean image (step S223) and applies image classification based on the model parameters (initialization parameters) to the clean image (step S224). In step S224, the original image classification unit 216 uses the first CNN 2161 to perform image classification on the first clean image.
[0115] The original image classification unit 216 outputs the intermediate output of the first CNN 2161 to the optimization unit 215 as an original image intermediate output (step S225).
[0116] [Optimization process] Next, the optimization process (step S206) will be described below. FIG 15 is a flowchart showing the procedure of the optimization process shown in FIG.
[0117] Steps S231 to S233 shown in Fig. 15 are the same processes as steps S31 to S33 shown in Fig. 7. The optimization unit 15 acquires the model output and intermediate output from the image classification unit 214 (step S234), and acquires the original image intermediate output from the original image classification unit 216. (step S235).
[0118] Then, the optimization unit 215 optimizes the model parameters of the second CNN 2141 so that the difference between the model output output from the image classification unit 214 and the correct label of the first clean image is reduced, and the difference between the intermediate output of the original image and the intermediate output of the second CNN 2141 is reduced, and outputs the optimized parameters to the adversarial perturbation generation unit 13 and the image classification unit 214 (step S236).
[0119] The optimization unit 215 optimizes the image classification parameters using an image classification loss that reduces the difference between the model output x and the correct label y, and an intermediate output loss that reduces the difference between the intermediate output of the second CNN 2141 and the original image intermediate output, as shown in equation (4), and outputs the optimized parameters.
[0120] For image classification loss, cross entropy shown in formula (4) is generally used, but mean square error or other functions can also be used as long as they are appropriate objective functions for the desired classification task. For intermediate output loss, a least square error function or cosine distance function is assumed, but any appropriate objective function will do.
[0121] The optimization unit 215 evaluates how close the model output is to the correct label, and outputs the evaluation result (step S237). For this evaluation, in addition to whether or not the result of taking the argmax of the model output matches the correct label, any appropriate evaluation index may be used.
[0122] [Effects of the second embodiment] In this way, in the second embodiment, by directly constraining the output of the intermediate layer so that the feature representation of the second CNN 2141 does not change from the feature representation of the first CNN 2161, it is possible to provide an image recognition model that maintains accuracy for clean images while improving robustness through adversarial learning.
[0123] [Embodiment 3] Next, a description will be given of embodiment 3. In embodiment 3, a combination of embodiment 1 and embodiment 2 will be described.
[0124] Fig. 16 is a diagram illustrating an example of the configuration of a learning device according to embodiment 3. As illustrated in Fig. 16, learning device 310 according to embodiment 3 has a configuration further including an original image classification unit 216, compared to learning device 10 illustrated in Fig. 3. Moreover, learning device 310 has an image classification unit 214 and an optimization unit 215 instead of image classification unit 14 and optimization unit 15, compared to learning device 10 illustrated in Fig. 3.
[0125] Fig. 17 is a flowchart showing a processing procedure of a learning process according to embodiment 3. Steps S301 to S303 shown in Fig. 17 are the same processes as steps S1 to S3 shown in Fig. 4. Note that in step S303, the adversarial perturbation generator 13 may generate the adversarial perturbation by applying the most confusing attack (Fig. 9).
[0126] Steps S304 to S306 shown in Fig. 17 are the same processes as steps S204 to S206 shown in Fig. 12. Steps S307 to S309 shown in Fig. 17 are the same processes as steps S6 to S8 shown in Fig. 4.
[0127] As shown in the third embodiment, while performing curriculum learning to generate adversarial perturbations so that the magnitude of the adversarial perturbations is increased stepwise, parameters of the second CNN 2141 may be optimized so that the difference between the image classification result for the adversarial image by the image classification unit 214 and the correct label of the first clean image becomes small, and the difference between the intermediate output of the second CNN 2141 and the intermediate output of the original image becomes small. Learning of the second NN is executed.
[0128] [Evaluation Results] Table 1 shows the results of evaluating the image classification performance of the CNN141 and the second CNN2141 trained in the first to third embodiments. For comparison, the evaluation results of a CNN (Standard) trained only with clean images and a CNN (AT) trained only with adversarial images are also shown. Clean is the result for clean images. FGSM, BIM (7), PGC (20), and CW (20) are methods for generating adversarial images.
[0129] [Table 1]
[0130] As shown in Table 1, both CNN141 trained using the learning methods according to the first to third embodiments and the second CNN2141 show higher image classification accuracy for clean images than CNN(AT) that has learned only adversarial images. And both CNN141 trained using the learning methods according to the first to third embodiments show higher image classification accuracy for adversarial images than CNN(Standard) that has learned only clean images and CNN(AT) that has learned only adversarial images. Therefore, it was found that by using the learning methods according to the first to third embodiments, it is possible to realize an image classification model that maintains accuracy for clean images while improving robustness by adversarial learning.
[0131] Furthermore, the second CNN2141 trained using the learning method according to the third embodiment shows higher image recognition accuracy for both clean images and adversarial images, compared with the CNN141 trained using the learning methods according to the first and second embodiments, and the second CNN2141. Therefore, it was found that by combining curriculum learning and knowledge distillation that directly constrains the output of the intermediate layer, as in the third embodiment, it is possible to realize an image recognition model that maintains accuracy for clean images while improving robustness through adversarial learning.
[0132] The learning device according to the present embodiment provides certain improvements over conventional learning methods such as those described in Non-Patent Document 1, and represents an advancement in the technical field of image recognition.
[0133] [System configuration of the embodiment] Each component of the learning devices 10, 210, and 310 is a functional concept, and does not necessarily have to be physically configured as shown in the figure. In other words, the specific form of distribution and integration of the functions of the learning devices 10, 210, and 310 is not limited to that shown in the figure, and all or part of them can be functionally or physically distributed or integrated in any unit depending on various loads, usage conditions, etc.
[0134] Furthermore, all or any part of the processes performed by the learning devices 10, 210, and 310 may be realized by a CPU, a GPU (Graphics Processing Unit), and a program analyzed and executed by the CPU and the GPU. Furthermore, each process performed by the learning devices 10, 210, and 310 may be realized as hardware using wired logic.
[0135] Furthermore, among the processes described in the embodiments, all or part of the processes described as being performed automatically can be performed manually. Alternatively, all or part of the processes described as being performed manually can be performed automatically by a known method. In addition, the process procedures, control procedures, specific names, and information including various data and parameters described above and shown in the drawings can be changed as appropriate unless otherwise specified.
[0136] [program] 18 is a diagram showing an example of a computer in which a program is executed to realize the learning devices 10, 210, 310. The computer 1000 has, for example, a memory 1010 and a CPU 1020. The computer 1000 also has a hard disk drive interface 1030, a disk drive interface 1040, a serial port interface 1050, a video adapter 1060, and a network interface 1070. These components are connected by a bus 1080.
[0137] The memory 1010 includes a ROM 1011 and a RAM 1012. The ROM 1011 stores a boot program such as a BIOS (Basic Input Output System). The hard disk drive interface 1030 is connected to a hard disk drive 1090. The disk drive interface 1040 is connected to a disk drive 1100. A removable storage medium such as a magnetic disk or an optical disk is inserted into the disk drive 1100. The serial port interface 1050 is connected to a mouse 1110 and a keyboard 1120, for example. The video adapter 1060 is connected to a display 1130, for example.
[0138] The hard disk drive 1090 stores, for example, an OS (Operating System) 1091, an application program 1092, a program module 1093, and program data 1094. That is, a program that defines each process of the learning devices 10, 210, and 310 is implemented as a program module 1093 in which code executable by the computer 1000 is written. The program module 1093 is stored, for example, in the hard disk drive 1090. For example, a program module 1093 for executing the same process as the functional configuration of the learning devices 10, 210, and 310 is stored in the hard disk drive 1090. The hard disk drive 1090 may be replaced by an SSD (Solid State Drive).
[0139] Furthermore, setting data used in the processing of the above-described embodiment is stored as program data 1094, for example, in memory 1010 or hard disk drive 1090. Then, CPU 1020 reads out program module 1093 or program data 1094 stored in memory 1010 or hard disk drive 1090 into RAM 1012 as necessary, and executes the program.
[0140] Note that the program module 1093 and the program data 1094 are not limited to being stored in the hard disk drive 1090, but may be stored in, for example, a removable storage medium and read by the CPU 1020 via the disk drive 1100 or the like. Alternatively, the program module 1093 and the program data 1094 may be stored in another computer connected via a network (such as a local area network (LAN) or wide area network (WAN)). The program module 1093 and the program data 1094 may then be read by the CPU 1020 from the other computer via the network interface 1070.
[0141] The following supplementary notes are further disclosed regarding the above embodiment.
[0142] (Additional note 1) Memory, at least one processor coupled to the memory; Including, The processor, generating a perturbation based on a first clean image to which a correct answer label has been added, and generating a superimposed image by superimposing the generated perturbation on the first clean image; performing image classification on the first cleaned image using a first neural network (NN) that has learned image classification using a plurality of cleaned images as training data, and outputting an intermediate output of the first NN as a first intermediate output; performing image classification on the superimposed image using a second NN in which the model parameters of the first NN are set as initial parameters, and outputting an intermediate output of the second NN as a second intermediate output; Optimizing parameters of the second NN so that a difference between an image classification result for the superimposed image and a correct label of the first clean image is reduced, and a difference between the first intermediate output and the second intermediate output is reduced, thereby executing learning of the second NN. Learning device.
[0143] (Additional note 2) The learning device according to claim 1, The learning includes optimizing parameters of the second NN using a loss function that is a weighted sum of a first loss function indicating a difference between an image classification result for the superimposed image by the second image classification unit and a correct label of the first clean image, and a second loss function indicating a difference between the first intermediate output and the second intermediate output. Learning device.
[0144] (Additional note 3) The learning device according to claim 1, The generating step searches for a category having the largest output value other than the correct label from the image classification result for the first clean image, and generates the perturbation so as to minimize the loss for the searched category. Learning device.
[0145] (Additional note 5) A storage medium storing a learning program for causing a computer to function as the learning device according to claims 1 to 3.
[0146] Although the embodiment of the invention made by the present inventor has been described above, the present invention is not limited by the description and drawings that form a part of the disclosure of the present invention according to the present embodiment. In other words, other embodiments, examples, operation techniques, etc. made by those skilled in the art based on the present embodiment are all included in the scope of the present invention. [Explanation of symbols]
[0147] 10,210,310 Learning device 11 Image storage section 12 Initialization section 13 Adversarial perturbation generator 14,214 Image Recognition Unit 15,215 Optimization Department 16 Schedule Department 141 CNN 216 Original Image Identification Unit 2141 The second CNN 2161 CNN No. 1
Claims
1. A generating unit that generates a perturbation based on a first clean image to which a correct answer label is attached, and generates a superimposed image by superimposing the generated perturbation on the first clean image; A first image classification unit that performs image classification on the first clean image using a first NN (Neural Network) that has learned image classification using a plurality of clean images as learning data, and outputs an intermediate output of the first NN as a first intermediate output; a second image classification unit that performs image classification on the superimposed image using a second NN in which model parameters of the first NN are set as initial parameters, and outputs an intermediate output of the second NN as a second intermediate output; a learning unit that executes learning of the second NN by optimizing parameters of the second NN so that a difference between an image classification result for the superimposed image by the second image classification unit and a correct label of the first clean image is reduced, and a difference between the first intermediate output and the second intermediate output is reduced; A learning device comprising:
2. The learning device according to claim 1, characterized in that the learning unit optimizes parameters of the second NN using a loss function that is a weighted sum of a first loss function indicating a difference between an image classification result for the superimposed image by the second image classification unit and a correct label of the first clean image, and a second loss function indicating a difference between the first intermediate output and the second intermediate output.
3. The learning device according to claim 1, characterized in that the generation unit searches for a category other than a correct label that has the largest output value from the image classification result for the first clean image, and generates the perturbation so as to minimize a loss for the searched category.
4. A learning method executed by a learning device, comprising: generating a perturbation based on a first clean image to which a correct answer label has been added, and generating a superimposed image by superimposing the generated perturbation on the first clean image; performing image classification on the first cleaned image using a first neural network (NN) that has learned image classification using a plurality of cleaned images as training data, and outputting an intermediate output of the first NN as a first intermediate output; performing image classification on the superimposed image using a second NN in which model parameters of the first NN are set as initial parameters, and outputting an intermediate output of the second NN as a second intermediate output; Optimizing parameters of the second NN so that a difference between an image classification result for the superimposed image and a correct label of the first clean image is reduced, and a difference between the first intermediate output and the second intermediate output is reduced, thereby executing learning of the second NN; A learning method comprising:
5. A step of generating a perturbation based on a first clean image to which a correct answer label is attached, and generating a superimposed image by superimposing the generated perturbation on the first clean image; performing image classification on the first cleaned image using a first NN (Neural Network) that has learned image classification using a plurality of cleaned images as training data, and outputting an intermediate output of the first NN as a first intermediate output; performing image classification on the superimposed image using a second NN in which model parameters of the first NN are set as initial parameters, and outputting an intermediate output of the second NN as a second intermediate output; Optimizing parameters of the second NN so that a difference between an image classification result for the superimposed image and a correct label of the first clean image is reduced, and a difference between the first intermediate output and the second intermediate output is reduced, thereby executing learning of the second NN; A learning program for a computer to execute the above.
Citation Information
Patent Citations
Device and method for training classifier, and device and method for evaluating robustness of classifier
JP2021170332A