Neural Network Training Device, Program Used in the Training Device, and Method for Creating Neural Network
The training device addresses the discrepancy between attention regions in neural network image recognition and attention maps by using human-corrected 'teacher attention maps' to retrain the network, enhancing recognition accuracy.
Patent Information
- Application Number
- JP2024069188
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-07-25
- Filing Date
- 2024-04-22
- Publication Date
- 2025-06-12
- Estimated Expiration
- 2040-06-17
AI Technical Summary
In image recognition technologies using neural networks, there is a discrepancy between the attention regions in recognition results and attention maps, leading to potential errors in image recognition, and currently, there is no method to correct this discrepancy.
A training device and method that incorporate human knowledge by using a dataset with 'teacher attention maps' corrected by humans to retrain the neural network, ensuring that the attention regions in the recognition results and attention maps align correctly.
By incorporating human knowledge through the use of teacher attention maps, the neural network can generate more accurate attention maps that align with the intended recognition results, thereby improving the overall accuracy of image recognition.
Smart Images

Figure 0007691686000001 
Figure 0007691686000002 
Figure 0007691686000003
Abstract
Description
Technical Field
[0001] The present invention relates to a neural network training device and a program used for the training device.
Background Art
[0002] Conventionally, in image recognition technology using neural networks such as CNN (Convolutional Neural Network), a technique for generating an attention map representing a fixation area during inference by a neural network is known (see, for example, Non-Patent Documents 1 and 2).
Prior Art Documents
Non-Patent Documents
[0003]
Non-Patent Document 1
Non-Patent Document 2
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, according to the inventor's study, when a neural network generates a recognition result and an attention map for a certain image, the attention regions in the recognition result and the attention map may not match. For example, even though the recognition result is "smiling face", if the attention region is the hair, the attention regions in the recognition result and the attention map do not match. If there is a discrepancy between the recognition result and the attention region in the attention map, it is a problem that may also lead to an error in the image recognition itself. And currently, there is no way to correct this discrepancy.
[0005] In view of the above, the present disclosure aims to incorporate human knowledge into the recognition function or learning of a neural network in an image recognition technology using a neural network that outputs an attention map.
Means for Solving the Problems
[0008] According to the One viewpoint of the present disclosure, the training device includes a reading unit (205) that reads an attention map (53) representing the attention region of the input image (51) and a neural network (10) trained to generate a recognition result of the image, and a training unit (250) that retrains the neural network using a dataset for retraining, where the dataset for retraining includes a plurality of teacher attention maps that are the correct values of the attention maps when a plurality of images are input to the neural network. When the plurality of images are input into the neural network, the plurality of teacher attention maps are the attention maps generated by the neural network and corrected by a human correction operation. From another perspective, the training device includes a reading unit that reads out an attention map (53) representing a fixation region of an input image (51) and a neural network (10) trained to generate a recognition result of the image, and a training unit (250) that retrains the neural network using a dataset for retraining. The dataset for retraining includes a plurality of teacher attention maps that are correct values of attention maps when a plurality of images are input into the neural network. The neural network generates a recognition result of the image based on the image and the attention map. The plurality of teacher attention maps are created as the correct values by a person who has viewed the plurality of images input into the neural network, or the attention maps generated by the neural network when the plurality of images are input into the neural network are corrected by a human correction operation. Also, according to another viewpoint, a program causes the training device to function Also, from another perspective, a method for creating a neural network after retraining, in which a training device for retraining a neural network before retraining that is trained to generate an attention map (53) representing a fixation region of an input image (51) and a recognition result of the image is used, includes: the training device reads out the neural network (10) before retraining; the training device inputs a training image into the neural network before retraining to cause the neural network before retraining to generate the attention map (220); creating a teacher attention map that is the correct value of the attention map when the training image is input into the neural network after retraining by correcting the attention map by a human correction operation (240); and the training device creates the neural network after retraining by retraining the neural network before retraining using a training dataset including the teacher attention map (250). Also, from another perspective, a method for creating a neural network after retraining, in which a training device for retraining a neural network before retraining that is trained to generate an attention map (53) representing a fixation region of an input image (51) and to generate a recognition result of the image based on the image and the attention map is used, includes: the training device reads out the neural network (10) before retraining; the training device inputs a training image into the neural network before retraining to cause the neural network before retraining to generate the attention map (220); creating a teacher attention map as the correct value of the attention map when the training image is input into the neural network after retraining by a person who has viewed the training image; and the training device creates the neural network after retraining by retraining the neural network before retraining using a training dataset including the teacher attention map (250).
[0009] In this way, when retraining a neural network that generates an attention map, a teacher attention map that is the correct value of the attention map is used. Since the teacher attention map is created based on human knowledge, it becomes possible to incorporate human knowledge into the learning of the neural network in this way.
[0010] Note that the reference numerals in parentheses attached to each component etc. indicate an example of the correspondence relationship between the component etc. and the specific components etc. described in the embodiments described later.
Brief Description of Drawings
[0011]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Embodiments for Carrying Out the Invention
[0012] (First Embodiment) Hereinafter, the first embodiment will be described. As shown in FIG. 1, the image recognition device 1 according to the present embodiment includes an operation device 2, a display device 3, a memory 4, and a processing unit 5.
[0013] The operating device 2 is a device that receives a human operation and outputs a signal corresponding to the received operation to the processing unit 5. The operating device 2 may be, for example, a mouse, a keyboard, a touch panel, or the like. The display device 3 is a device that displays an image to a person.
[0014] The memory 4 includes a RAM which is a rewritable volatile storage medium, a ROM which is a non-rewritable non-volatile storage medium, and a flash memory which is a rewritable non-volatile storage medium. The RAM, ROM, and flash memory are non-transitory physical storage media. The learned neural network 10 data is pre-recorded in the flash memory.
[0015] The processing unit 5 executes a program (not shown) stored in the ROM or flash memory, and uses the RAM as a working area during the execution to realize various processes described later.
[0016] Here, the neural network 10 will be described. As shown in FIG. 3, the neural network 10 is a deep neural network including a feature extraction unit 11, an attention unit 12, a synthesis unit 13, and a recognition unit 14.
[0017] When the input image 51 is input to the neural network 10, the neural network 10 generates an attention map 53. The attention map 53 is data representing the attention area during the inference of the neural network 10. That is, the attention map 53 is visual explanatory data for explaining which area of the input image 51 is being emphasized during the inference of the neural network 10.
[0018] Also, the neural network 10 outputs a classification result of the input image 51 based on the input image 51 and the attention map 53. The classification result of the input image 51 is a plurality of likelihoods corresponding to a plurality of classes (for example, damesian, crab, finch, frog, etc.) corresponding to the recognition target of the image. Here, the number of classes is denoted as K.
[0019] The feature extraction unit 11 is a neural network having a plurality of layers. These plurality of layers at least include a plurality of convolutional layers. Further, these plurality of layers may further be components of a plurality of residual blocks, or may have a plurality of pooling layers or the like. Then, the feature extraction unit 11 generates a feature map 52 by propagating the information of the input image 51 input thereto through these plurality of layers.
[0020] The feature map 52 is a map of resolution h×w corresponding to each of K classes. h and w are arbitrary integers. Therefore, the number of channels of the feature map 52 is K. The resolution of the feature map 52 may be the same as the resolution of the input image 51, or may be lower than the resolution of the input image 51.
[0021] The feature extraction unit 11 may be constituted by a portion starting from the input layer and before the first fully-connected layer in the baseline model. As the baseline model, one having a plurality of convolutional layers and generating likelihoods of a plurality of classes of the same type as the neural network 10 is selected. For example, as the baseline model, VGGNet shown in Non-Patent Document 3 may be used, ResNet shown in Non-Patent Document 4 may be used, or other CNN (Convolutional Neural Network) may be used.
[0022] The attention unit 12 generates an attention map 53 from the feature map 52 generated by the feature extraction unit 11. The attention unit 12 is a neural network having a plurality of layers. These plurality of layers, as shown in FIG. 3, have a first portion 12a having one or more convolutional layers or one or more residual blocks, and a K×1×1 convolutional layer 12b at the subsequent stage of the first portion. Here, assuming that L, a, and b are arbitrary natural numbers, the L×a×b convolutional layer means a convolutional layer using a kernel of a×b for each of L channels.
[0023] The attention unit 12 has two K×1×1 convolutional layers 12c and a 1×1×1 convolutional layer 12e that branch after the convolutional layer 12b. And the attention unit 12 has a GAP (Global Average Pooling) layer 12d after the convolutional layer 12c.
[0024] The information of the feature map 52 input to the attention unit 12 propagates through the first part 12a, the convolutional layer 12b, the convolutional layer 12c, and the GAP layer 12d, and the output of the GAP layer 12d is input to the Softmax function, thereby generating the likelihoods of a plurality of classes of the same type as the neural network 10 as classification results. The classification result is a kind of recognition result.
[0025] Also, the information of the feature map 52 input to the attention unit 12 is propagated through the first part 12a, the convolutional layer 12b, and the convolutional layer 12e, thereby generating an attention map 53. By generating the attention map 53 through the convolutional layer 12b instead of the fully connected layer, the information of the attention area is propagated to the attention map 53 while being localized. Also, by passing through the 1×1×1 convolutional layer 12e, a one-channel attention map 53 is generated as the weighted sum of the attention areas corresponding to all classes. Each value of the kernel of the convolutional layer 12e may all be 1 or may be otherwise.
[0026] The resolution of each map of the feature map 52 and the resolution of the attention map 53 are the same. The attention unit 12 is configured to be so. In the attention map 53, relatively high pixel values are given to the pixels corresponding to the attention area, and lower pixel values than the attention area are given to the pixels not corresponding to the attention area. The value that each pixel value of the attention map 53 can take may be binary or may be values in 256 levels. The higher the pixel value of a certain pixel, the higher the degree of attention at the position of that pixel.
[0027] The synthesis unit 13 synthesizes the feature map 52 and the attention map 53. Specifically, for the map with resolution h×w in each of the K channels in the feature map 52, the attention map 53 is multiplied. The multiplication of the attention map 53 and the map with resolution h×w is performed between pixels with the same position coordinates. Note that the synthesis may be multiplication as described above, addition, or an operation consisting of a combination of addition and multiplication. Through this synthesis, a synthesis map 54 is obtained. The number of channels and the resolution of the synthesis map 54 are the same as those of the feature map 52.
[0028] The recognition unit 14 outputs the likelihood of each class based on the synthesis map 54. The recognition unit 14 is a neural network having a plurality of layers. These plurality of layers include at least a plurality of convolutional layers. Also, these plurality of layers include one or both of a fully connected layer and a GAP layer. Further, these plurality of layers may further be components of a plurality of residual blocks or may have a plurality of pooling layers. The recognition unit 14 propagates the information of the input synthesis map 54 through these plurality of layers and outputs the likelihood of each class as a classification result. The classification result is also a recognition result. The recognition unit 14 may be constituted by the part from immediately after the part used in the attention unit 12 to the output layer among the above-described baseline models.
[0029] Note that the functions performed by the neural network 10, the feature extraction unit 11, the attention unit 12, the synthesis unit 13, and the recognition unit 14 are actually realized by the processing unit 5 performing processing according to the structure and parameters of the neural network 10.
[0030] The feature extraction unit 11, the attention unit 12, and the synthesis unit 13 are pre-trained by the error backpropagation method in supervised learning so as to realize the above functions. In the learning, as the learning error (also referred to as the loss function) L, L = Latt + Lper is used. Here, Latt is the learning error regarding the classification result output by the attention unit 12, and Lper is the learning error regarding the classification result output by the recognition unit 14. Latt and Lper may be calculated by applying a combination of the Softmax function and cross entropy to their respective classification results. The feature extraction unit 11 is trained by passing through the gradients of the attention unit 12 and the recognition unit 14 in the error backpropagation method.
[0031] Hereinafter, the image classification process of the processing unit 5 using the thus configured trained neural network 10 will be described.
[0032] When a predetermined condition such as an execution start operation on the human operation device 2 is satisfied, the processing unit 5 starts the process shown in FIG. 4 defined in a predetermined program recorded in the memory 4. In this process, the processing unit 5 first reads the neural network 10 from the memory 4 in step 110.
[0033] Subsequently, in step 120, the input image 51 is acquired and the input image 51 is input to the neural network 10. The input image 51 may be an image selected by an operation on the human operation device 2 from among a plurality of images previously recorded in the memory 4, or may be an image received from another device via a communication network (not shown).
[0034] When the input image 51 is input to the neural network 10, as described above, the neural network 10 generates the feature map 52 and the classification result from the input image 51 by the feature extraction unit 11, and the attention unit 12 generates the attention map 53 from the feature map 52.
[0035] In step 130 following step 120, the processing unit 5 acquires the attention map 53 thus generated. That is, the attention map 53 generated in the memory 4 by the neural network 10 is copied or moved to another area in the memory 4.
[0036] Subsequently, in step 140, the processing unit 5 corrects the acquired (i.e., the destination of copying or moving) attention map 53 based on the correction operation on the human operation device 2. Thereby, the attention map 53 is corrected by human knowledge.
[0037] Specifically, the processing unit 5 causes the display device 3 to display the attention map 53 before correction and a pointer. The pointer is an image that moves within the display range of the attention map 53 displayed on the display device 3 according to the operation of a person on the operation device 2. The person corrects the value in the position range overlapping the pointer in the displayed attention map 53 by performing a predetermined correction operation (for example, deletion operation, addition operation, etc.) on the operation device 2.
[0038] At this time, as shown in FIG. 5, the processing unit 5 may cause the input image 51 to be transparently aligned and superimposed on the attention map 53, and reflect the correction according to the above correction operation on the attention map 53 in the state of being displayed on the display device 3. At this time, when the resolutions of the input image 51 and the attention map 53 are different, the processing unit 5 reduces the resolution of the input image 51 to match the attention map 53 and then transparently superimposes it on the attention map 53.
[0039] In FIG. 5, the input image 51 in which the dachshund is holding a soccer ball is transparently superimposed on the attention map 53.
[0040] In this attention map 53, the fixation region is in the region of the soccer ball. If this attention map 53 is input into the synthesis unit 13 as it is, and when the synthesized map 54, which is the synthesis result of the attention map 53 and the feature map 52, is input into the first layer of the recognition unit 14, the classification result generated by the recognition unit 14 will have the highest likelihood of the soccer ball. That is, the neural network 10 recognizes the input image 51 as an image of a soccer ball.
[0041] However, if the person using the image recognition device 1 wants to recognize the input image 51 as an image of a dalmatian, the fixation region in this attention map 53 should be in the region where the dalmatian is located.
[0042] Therefore, in such a case, the person uses the operation device 2 to modify the fixation region in the attention map 53. Specifically, first, the person uses the operation device 2 to erase the fixation region in the attention map 53. For example, while the person presses a predetermined erase button of the operation device 2, the pointer is moved to scan the entire fixation region in the attention map 53. As a result, the processing unit 5 lowers the pixel value of the attention map 53 in the region scanned by the pointer while pressing the erase button, so that it becomes a pixel value that does not form a fixation region, as shown in FIG. 6.
[0043] Then, afterwards, the person uses the operation device 2 to set the region in the attention map 53 that is desired to be the new fixation region. For example, while the person presses a predetermined add button of the operation device 2, the pointer is moved to scan the entire region in the attention map 53 that is desired to be the fixation region. As a result, the processing unit 5 raises the pixel value of the attention map 53 in the region scanned by the pointer while pressing the add button, so that it becomes a pixel value that forms a fixation region, as shown in FIG. 7. In the example of FIG. 7, the new fixation region designated by the person is the face part of the dalmatian.
[0044] In this way, by superimposing the input image 51 on the attention map 53 and displaying it on the display device 3, when a person can determine which part of the input image 51 should be the attention area, they can efficiently utilize that knowledge to easily specify the attention area in the attention map 53. Through the processing of step 140 like this, the attention map 53 obtained in step 130 is modified in the memory 4.
[0045] Subsequently, in step 150, the processing unit 5 inputs the attention map 53 modified in the immediately preceding step 140 into the synthesis unit 13. Then, the synthesis unit 13 synthesizes the feature map 52 and the attention map 53 as described above to generate a synthesis map 54 and inputs it into the first layer of the recognition unit 14. The recognition unit 14 into which the synthesis map 54 is input generates a classification result based on the synthesis map 54 as described above. In this classification result, the likelihood of Dalmation is the highest. That is, the neural network 10 recognizes the input image 51 as an image of a Dalmation.
[0046] In step 160 following step 150, the processing unit 5 acquires and outputs the classification result thus generated by the recognition unit 14. The output destination may be another device via a communication network (not shown), the memory 4, or the display device 3.
[0047] In this way, by modifying the attention map 53 using human knowledge, the recognition unit 14 is weighted more highly in the area intended by the person. As a result, image recognition can be performed in accordance with the person's intention. That is, by using the attention map manually corrected based on human knowledge, the recognition result can be adjusted.
[0048] Also at this time, the parameters of the neural network 10 are not changed. That is, it is not necessary to retrain the neural network 10, and the recognition result intended by the person can be obtained.
[0049] For example, when a fundus image is input into the neural network 10 as the input image 51, by modifying the attention area of the attention map 53 using the knowledge based on the doctor's experience, the classification of the grade of eye diseases can be more accurately identified. Thus, for example, in medical image diagnosis, the function of this embodiment is useful.
[0050] As described above, as shown in FIG. 8, the processing unit 5 of the image recognition apparatus 1 performs a correction according to a human correction operation on the attention map 53 generated by the neural network 10 into which the input image 51 is input (step 140). Then, the processing unit 5 outputs the recognition result of the input image 51 generated by the neural network 10 based on the corrected attention map 53 and the input image 51 (step 160).
[0051] In this way, by using human knowledge to correct the attention map 53, the recognition unit 14 emphasizes the area intended by the human. As a result, image recognition can be performed in accordance with the human intention. Also at this time, the parameters of the neural network 10 are not changed. That is, without the need for retraining of the neural network 10, a recognition result intended by the human can be obtained.
[0052] Also, the path through which image information propagates to generate the attention map 53 and the path through which image information propagates to generate the recognition result are shared in part (i.e., the feature extraction unit 11), and separated in other parts (i.e., the attention unit 12 and the recognition unit 14). Then, the synthesis unit 13 inputs a synthesis map 54 in which the corrected attention map 53 is reflected to the side of the recognition unit 14 of the separated part. In this way, since the input location of the synthesis map 54 based on the corrected attention map 53 is suitable for the structure of the neural network 10, the degree of improvement of the recognition result by the corrected attention map 53 is improved.
[0053] Further, in a state where the processing unit 5 superimposes the input image 51 transparently on the attention map 53 and causes the display device 3 to display it, the processing unit 5 reflects a correction according to a human correction operation on the attention map 53. A person can relatively easily determine which part of the input image 51 should be the attention area by looking at the input image 51. Therefore, by superimposing the input image 51 on the attention map 53 and displaying it on the display device 3, a person can efficiently utilize their knowledge visually and easily specify the attention area in the attention map 53.
[0054] In addition, in the present embodiment, the processing unit 5 functions as an input unit by executing step 120, functions as a map correction unit by executing step 140, and functions as an output unit by executing step 160.
[0055] (Second Embodiment) Next, the second embodiment will be described. In this embodiment, learning parameters such as the weights and biases of the neural network 10 are corrected based on the corrected attention map according to a human correction operation. That is, the neural network 10 is relearned based on the corrected attention map.
[0056] The hardware configuration of this embodiment is the same as that shown in FIG. 1 in the first embodiment. Also, the configuration of the learned neural network 10 stored in the memory 4 is the same as that in the first embodiment. Note that the image recognition device 1 of this embodiment corresponds to a training device.
[0057] One of the differences between this embodiment and the first embodiment is that instead of the processing unit 5 executing the processing of FIG. 4, the classification result corresponding to the input image 51 is generated by the neural network 10 without correcting the attention map 53.
[0058] That is, first, similar to the first embodiment, the processing unit 5 reads the neural network 10 from the memory 4, then acquires the input image 51, and inputs this input image 51 to the neural network 10.
[0059] Then, in the neural network 10, similar to the first embodiment, the feature extraction unit 11 and the attention unit 12 function, and the attention map 53 and the classification result are generated by the attention unit 12. This attention map 53 is input to the synthesis unit 13 without receiving a human correction operation, that is, without being corrected. The synthesis unit 13 generates a synthesis map 54 by synthesizing the feature map 52 and the attention map 53 that has not received a human correction operation. The recognition unit 14 generates a classification result based on this synthesis map 54 in the same manner as in the first embodiment. The processing unit 5 acquires and outputs this classification result in the same manner as in the first embodiment.
[0060] In addition to the process of acquiring the classification result of the input image 51 from the input image 51 using the neural network 10 as described above, the processing unit 5 executes the process shown in FIG. 9 in order to retrain the neural network 10. By this retraining, the neural network 10 is fine-tuned.
[0061] The processing unit 5 starts the process of FIG. 9 based on the fact that a predetermined retraining start operation by a person has been performed on the operation device 2. In this process, the processing unit 5 uses a dataset for retraining. The dataset for retraining has a plurality of groups (which may be 10, 100, or 100,000) consisting of learning images and teacher labels.
[0062] The learning image is data that is input to the feature extraction unit 11 like the input image 51. The teacher label is data that is the correct value of the classification result output from the attention unit 12 and the recognition unit 14 when the learning image of the same group is input to the feature extraction unit 11.
[0063] The dataset for re-learning may be generated in advance and recorded in the non-volatile storage medium of the memory 4, or may be acquired from a data server via a communication network (not shown). Also, as the learning images and teacher labels of the dataset for re-learning, the same ones as the dataset for learning used at the initial learning of the neural network 10 may be reused, or they may be different from the dataset for learning.
[0064] In the process of FIG. 9, the processing unit 5 first executes the loop processing of steps 210 and 220 for each group included in the dataset for re-learning. In each iteration of the loop processing, the processing unit 5 first inputs, in step 210, the learning image in the target group to the feature extraction unit 11. Subsequently, in step 220, the attention map 53 and classification result generated by the attention unit 12 based on the input learning image, and the classification result generated by the recognition unit 14 based on the learning image are acquired and recorded in the memory 4.
[0065] Note that the method by which the neural network 10 outputs the attention map 53 and two types of classification results based on the input learning image is equivalent to the above-described method in which the learning image is replaced with the input image 51. After step 220, one iteration of the loop processing ends.
[0066] When the loop processing ends for the number of groups, the processing of the processing unit 5 proceeds to step 230. At this point, for each group in all the datasets for re-learning, the attention map 53 and classification result generated by the attention unit 12, and the classification result generated by the recognition unit 14 are associated and recorded in the memory 4.
[0067] In step 230, the processing unit 5 extracts the group in which misrecognition has occurred from among the plurality of groups. The group extracted as having misrecognition is a group in which the class with the highest likelihood in the classification result output by the recognition unit 14 does not match the class indicated by the teacher label (that is, the class with the highest likelihood in the teacher label). Alternatively, a group in which the class with the highest likelihood in the classification result output by the attention unit 12 does not match the class indicated by the teacher label may be extracted as having misrecognition. Or alternatively, both of them may be extracted. The groups to be extracted are usually plural.
[0068] Subsequently, in step 240, the attention map 53 recorded in the memory 4 corresponding to each of the groups extracted in the immediately preceding step 230 is corrected based on human knowledge. Specifically, the attention map 53 is corrected based on the human correction operation on the operation device 2 by the same process as in step 140 of FIG. 4. Then, the processing unit 5 stores the corrected attention map 53 in the memory 4 as the teacher attention map belonging to the group.
[0069] The teacher attention map created in this way is data that is the correct value of the attention map 53 output from the attention unit 12 when the learning images of the same group are input to the feature extraction unit 11. By this process, the teacher attention map is added to the dataset for relearning.
[0070] Subsequently, in step 250, the processing unit 5 retrains the neural network 10 based on the two types of classification results, the attention map 53, and the dataset for relearning obtained in the processing of FIG. 9 this time. As described above, the dataset for relearning includes the teacher attention map and the teacher label.
[0071] Specifically, as shown in FIG. 10, an amount L = Latt + Lper + Lmap composed of the sum of the three learning errors Latt, Lper, and Lmap is used as the learning error, and the learning parameters such as the weights and biases of the attention unit 12 and the recognition unit 14 are updated by the error backpropagation method. In FIG. 10, the output layer 14b of the recognition unit 14 and a portion 14a in front of the output layer 14b of the recognition unit 14 are shown. Note that in the present embodiment, the learning parameters such as the weights and biases of the feature extraction unit 11 are not updated.
[0072] Here, Latt is an amount indicating the error between the classification result output by the attention unit 12 when the learning image 61 is input to the neural network 10 and the teacher label 60 belonging to the same group as the learning image 61.
[0073] Also, Lper is an amount indicating the error between the classification result output by the feature extraction unit 11 when the learning image 61 is input to the neural network 10 and the teacher label 60 belonging to the same group as the learning image 61.
[0074] Also, Lmap is an amount indicating the error between the attention map 53 output by the attention unit 12 when the learning image 61 is input to the neural network 10 and the teacher label 60 belonging to the same group as the learning image 61.
[0075] As the learning error Lmap, an L2 norm error may be adopted as in the following formula, or other forms of errors may be adopted. Lmap = γ × ||M’ - M|| 2 Here, M indicates the value of the attention map 53 output by the attention unit 12 when the learning image 61 is input to the neural network 10. M’ indicates the value of the corrected attention map corresponding to the same group as the learning image 61. By calculating the error for each element of these two attention maps, the attention unit 12 is trained to output an attention map closer to human knowledge.
[0076] Here, γ is a coefficient for adjusting the learning error Lmap. The value of Lmap is larger than those of Latt and Lper. Therefore, by multiplying γ by Lmap, the magnitudes of the three learning errors Lmap, Latt, and Lper can be adjusted. After step 250, the process of FIG. 9 is completed, and the re-learned neural network 10 is recorded in the memory 4.
[0077] In this way, by fine-tuning the neural network 10 based on the attention map corrected based on human knowledge, the image recognition function by the neural network 10 is improved. That is, when the processing unit 5 inputs various input images 51 to the neural network 10 after fine-tuning, the accuracy rate of the recognition results generated by the recognition unit 14 is improved.
[0078] As described above, the processing unit 5 re-learns the neural network 10 using the dataset for re-learning (step 250). And the dataset for re-learning includes a plurality of teacher attention maps.
[0079] In this way, when re-learning the neural network 10 that generates the attention map, the teacher attention map, which is the correct value of the attention map, is used. Since the teacher attention map is created based on human knowledge, it becomes possible to incorporate human knowledge into the learning of the neural network 10 in this way.
[0080] Also, the processing unit 5 obtains a plurality of attention maps corresponding to the plurality of learning images respectively by inputting the plurality of learning images to the neural network 10 (steps 210 and 220). And the processing unit 5 corrects those plurality of attention maps according to the human correction operation to obtain the teacher attention map (step 240).
[0081] In this way, based on the correction operation performed by a human on the attention map generated by the neural network 10, a teacher attention map can be generated. Therefore, more directly, it becomes possible to incorporate human knowledge into the learning of the neural network 10. Moreover, the correction operation is simpler compared to the case of creating a teacher attention map from scratch.
[0082] Also, the plurality of teacher attention maps included in the dataset for re-learning are only the training images misrecognized by the neural network 10 before re-learning. In this way, by using many teacher attention maps corresponding to the misrecognized training images for re-learning, re-learning can be performed with higher efficiency. This is because the attention map generated with the misrecognized training image as input is likely to be itself full of errors.
[0083] Also, the neural network 10 generates a recognition result of the input image 51 based on the input image 51 and the attention map 53. In the neural network 10 that feeds back not only the input image 51 but also the attention map 53 as information for image recognition in this way, the relevance between the recognition result of the input image 51 and the attention map 53 is strong. Therefore, in such a neural network 10, the degree to which the effect of re-learning using the teacher attention map contributes to the improvement of the recognition result of the input image 51 is high.
[0084] Also, the processing unit 5 re-learns the attention unit 12 without re-learning the feature extraction unit 11. In this way, by re-learning the part that is strongly related to the generation of the attention map 53 among the neural network 10, efficient fine-tuning of the neural network 10 is realized.
[0085] In the present embodiment, the processing unit 5 functions as a reading unit by executing step 205, functions as a training unit by executing step 250, functions as an acquisition unit by executing steps 210 and 220, and functions as a map correction unit by executing step 240.
[0086] (Other embodiments) Note that the present invention is not limited to the above-described embodiments, and can be appropriately modified. Also, the above-described embodiments are not independent of each other, and can be appropriately combined except in cases where the combination is clearly impossible. Further, in the above-described embodiments, the elements constituting the embodiments are not necessarily essential except in cases where it is clearly stated that they are essential and cases where they are considered to be clearly essential in principle. Also, in the above-described embodiments, when numerical values such as the number, numerical value, quantity, range, etc. of the components of the embodiments are mentioned, they are not limited to the specific number except in cases where it is clearly stated that they are essential and cases where they are clearly limited to a specific number in principle. Also, when a plurality of values are exemplified for a certain quantity, it is also possible to adopt a value between those plurality of values except in cases where it is specifically noted and cases where it is clearly impossible in principle.
[0087] Also, the present invention also allows the following modification examples and modification examples within the equivalent range for the above-described embodiments. Note that each of the following modification examples can independently select whether to apply or not apply to the above-described embodiments. That is, any combination of the following modification examples can be applied to the above-described embodiments.
[0088] (Modification example 1) The image recognition device 1 may have both functions of the first embodiment (i.e., image recognition using an attention map corrected based on human knowledge) and the second embodiment (i.e., relearning using an attention map corrected based on human knowledge).
[0089] (Modification example 2) In the above-described embodiment, as an example of the recognition result output by the attention unit 12 and the recognition unit 14, a classification result is given. However, the recognition result output by the attention unit 12 and the recognition unit 14 is not limited to the classification result, and may also be a result by regression. That is, the image recognition performed by the neural network 10 may be classification or regression.
[0090] (Modification Example 3) In the above-described first embodiment, the neural network 10 includes a feature extraction unit 11, an attention unit 12, a synthesis unit 13, and a recognition unit 14. However, the neural network for realizing image recognition using an attention map corrected based on human knowledge is not limited to such a configuration. That is, any neural network that generates an attention map based on the input image and generates a recognition result of the image based on the image and the attention map can improve the image recognition function by correcting the attention map.
[0091] (Modification Example 4) In the above-described second embodiment, the neural network 10 includes a feature extraction unit 11, an attention unit 12, a synthesis unit 13, and a recognition unit 14. However, it may be corrected based on human knowledge. However, the neural network for realizing re-learning using an attention map corrected based on human knowledge is not limited to such a configuration. That is, any neural network that generates an attention map and a recognition result of the image based on the input image can improve the image recognition function by re-learning using the corrected attention map. For example, a neural network such as CAM (Class Activation Mapping) described in Non-Patent Document 2 may be re-learned using an attention map corrected based on human knowledge.
[0092] (Modification Example 5) In the above-described first embodiment, the processing unit 5 transparently superimposes the input image 51 on the attention map 53 and, while it is being displayed on the display device 3, reflects the corrections according to the human correction operations in the attention map. However, it is not necessarily required to do so in this way. For example, the processing unit 5 may reflect the corrections according to the human correction operations in the attention map while arranging the input image 51 and the attention map 53 side by side without overlapping them and displaying them on the display device 3. Also, for example, the processing unit 5 may reflect the corrections according to the human correction operations in the attention map while displaying the attention map 53 on the display device 3 and not displaying the input image 51 on the display device 3.
[0093] (Modification Example 6) In the above-described first and second embodiments, as a method for correcting the attention map, a method is shown in which only the values of some of the pixels in the attention map generated by the attention unit 12 are changed and the values of the remaining pixels are not changed. That is, a method of making changes to the attention map generated by the attention unit 12 is shown.
[0094] However, the method for correcting the attention map is not necessarily limited to such a method. For example, a new attention map may be created from scratch separately from the attention map output by the attention unit 12 when the image is input to the neural network 10. In this case, in the first embodiment, this new attention map is input to the synthesis unit 13, and in the second embodiment, this new attention map becomes the teacher attention map.
[0095] As a method for creating a new attention map, for example, there are the following methods. First, a person determines the position range of the fixation area by looking at the image input to the neural network 10. Then, the person may create a new attention map that reflects the determined position range of the fixation area by operating a computer. This computer may be the image recognition device 1 or another device.
[0096] (Modification Example 7) In the above-described second embodiment, the teacher attention map used for re-learning is only the teacher attention map corresponding to the training image misrecognized by the neural network 10 before re-learning. However, the teacher attention map used for re-learning may include the teacher attention map corresponding to the training image correctly recognized by the neural network 10 before re-learning.
[0097] Even in this case, if the number of teacher attention maps corresponding to the misrecognized training images is larger than the number of teacher attention maps corresponding to the correctly recognized training images, the efficiency of re-learning can be improved.
[0098] Alternatively, the number of teacher attention maps corresponding to the misrecognized training images may be smaller than the number of teacher attention maps corresponding to the correctly recognized training images.
[0099] (Modification Example 8) In the above embodiment, in the re-learning of the neural network 10, the feature extraction unit 11 is not re-learned, and only the attention unit 12 and the recognition unit 14 are re-learned. The re-learning of the neural network 10 is not limited to this form. For example, the feature extraction unit 11 and the recognition unit 14 may not be re-learned, and only the attention unit 12 may be re-learned. Also, for example, only the feature extraction unit 11 may be re-learned, and the attention unit 12 and the recognition unit 14 may not be re-learned. Also, for example, the feature extraction unit 11 and the recognition unit 14 may be re-learned, and the attention unit 12 may not be re-learned.
[0100] Also, a form in which the feature extraction unit 11, the attention unit 12, and the recognition unit 14 are re-learned is also allowed. In this case, the re-learning of the neural network 10 is not fine-tuning.
[0101] (Modification Example 9) In the above-described embodiment, the relearning is performed using the error backpropagation method with three learning errors of Lmap, Latt, and Lper. However, it is not necessary to use all of Lmap, Latt, and Lper. For example, only Lmap may be used.
Explanation of Signs
[0102] 1…Image recognition device, 2…Operating device, 3…Display device, 4…Memory, 5…Processing unit, 10…Neural network, 11…Feature extraction unit, 12…Attention unit, 13…Synthesis unit, 14…Cognition unit, 51…Input image, 52…Feature map, 53…Attention map, 54…Synthesis map, 60…Teacher label, 61…Learning image
Claims
1. a readout unit for reading out an attention map (53) representing a gaze area of an input image (51) and a neural network (10) trained to generate a recognition result of the input image; A training unit (250) for retraining the neural network using a retraining data set; The re-learning dataset includes a plurality of teacher attention maps that are regarded as correct values of attention maps when a plurality of images are input to the neural network, A training device, wherein the plurality of teacher attention maps are attention maps generated by the neural network when the plurality of images are input to the neural network, and which have been modified by a human correction operation.
2. a readout unit for reading out an attention map (53) representing a gaze area of an input image (51) and a neural network (10) trained to generate a recognition result of the input image; A training unit (250) for retraining the neural network using a retraining data set; The re-learning dataset includes a plurality of teacher attention maps that are regarded as correct values of attention maps when a plurality of images are input to the neural network, The neural network generates a recognition result for the image based on the image and the attention map; and A training device, wherein the multiple teacher attention maps are either created as the correct answer by a person who looks at the multiple images input to the neural network, or are attention maps generated by the neural network when the multiple images are input to the neural network that have been corrected by a human correction operation.
3. The training device of claim 1 , wherein the neural network generates a recognition result for the image based on the image and the attention map.
4. an acquisition unit (210, 220) that acquires a plurality of attention maps corresponding to the plurality of learning images by inputting the plurality of learning images into the neural network; 4. The training device according to claim 1, further comprising a map correction unit (240) for correcting the plurality of attention maps in response to a correction operation by a person to generate a teacher attention map.
5. A training device as described in any one of claims 1 to 4, wherein, in the multiple teacher attention maps included in the re-learning dataset, there are more teacher attention maps corresponding to training images that were incorrectly recognized by the neural network before re-learning than teacher attention maps corresponding to training images that were correctly recognized by the neural network before re-learning.
6. The neural network includes a feature extraction unit (11), an attention unit (12), a synthesis unit (13), and a recognition unit (14), The feature extraction unit includes a plurality of convolution layers, and generates a feature map (52) including features of the image by propagating information of the image through the plurality of convolution layers; The attention unit generates the attention map based on the feature map; The synthesis unit synthesizes the feature map and the attention map to generate a synthesis map (54); The recognition unit generates the recognition result based on the composite map; 6. The training device according to claim 1, wherein the training section retrains the attention section without retraining the feature extraction section.
7. A program used in a training device for relearning a neural network, comprising: a reading unit for reading an attention map (53) representing a gaze area of an input image (51) and a neural network (10) trained to generate a recognition result of said image; and causing the training device to function as a training unit (250) that retrains the neural network using a retraining data set; The re-learning dataset includes a plurality of teacher attention maps that are regarded as correct values of attention maps when a plurality of images are input to the neural network, The program, wherein the multiple teacher attention maps are attention maps generated by the neural network when the multiple images are input to the neural network, and which have been modified by human correction operations.
8. A program used in a training device for relearning a neural network, comprising: a reading unit for reading an attention map (53) representing a gaze area of an input image (51) and a neural network (10) trained to generate a recognition result of said image; and causing the training device to function as a training unit (250) that retrains the neural network using a retraining data set; The re-learning dataset includes a plurality of teacher attention maps that are regarded as correct values of attention maps when a plurality of images are input to the neural network, The neural network generates a recognition result for the image based on the image and the attention map; and A program in which the multiple teacher attention maps are either created as the correct answer by a person who looks at the multiple images input to the neural network, or are attention maps generated by the neural network when the multiple images are input to the neural network that have been corrected by a human correction operation.
9. A method for creating a retrained neural network using an attention map (53) representing a fixation area of an input image (51) and a training device for retraining a pre-retrained neural network trained to generate a recognition result of the image, comprising: The training device reads the neural network (10) before retraining; The training device inputs training images to the pre-relearning neural network to cause the pre-relearning neural network to generate the attention map (220); correcting the attention map by a manual correction operation to create a teacher attention map that is a correct answer value of the attention map when the learning image is input to the neural network after the re-learning (240); and generating (250) the re-trained neural network by re-training the pre-re-trained neural network using a re-training dataset including the teacher attention map.
10. A method for creating a retrained neural network, comprising the steps of: generating an attention map (53) representing a gaze area of an input image (51) and retraining a pre-retrained neural network that has been trained to generate a recognition result of the image based on the image and the attention map, the method comprising the steps of: The training device reads the neural network (10) before retraining; The training device inputs training images to the pre-relearning neural network to cause the pre-relearning neural network to generate the attention map (220); A teacher attention map is created as a correct answer value of an attention map when the learning image is input to the neural network after the re-learning by a person who has viewed the learning image; and generating (250) the re-trained neural network by re-training the pre-re-trained neural network using a re-training dataset including the teacher attention map.
Citation Information
Patent Citations
Fine-grained object recognition in robotic systems
US20180250826A1
Medical image processing device, processor device, endoscope system, medical image processing method, and program
WO2020174747A1