Information processing device, information processing method and program

By dynamically selecting image conversion methods based on the current machine learning model, the apparatus enhances the performance of machine learning models in computer vision tasks, reducing the reliance on extensive annotated data and lowering annotation costs.

JP2025080634APending Publication Date: 2025-05-26CANON KK
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023193918
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-14
Publication Date
2025-05-26

AI Technical Summary

Technical Problem

Existing active learning techniques for machine learning models, such as those used in computer vision tasks, rely on fixed data conversion methods to calculate uncertainty, which may not dynamically adapt to improve model performance.

Method used

An information processing apparatus that performs active learning by dynamically selecting an appropriate image conversion method using a learned model, allowing for the selection of images that contribute most to improving the model's performance.

Benefits of technology

This approach enables the selection of images that effectively enhance the performance of the learning model, reducing the need for extensive annotated data and lowering the work cost associated with data annotation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025080634000001_ABST
    Figure 2025080634000001_ABST
Patent Text Reader

Abstract

To select an image contributing to the performance improvement of a learning model using an appropriate image conversion scheme.SOLUTION: An information processing device that performs active learning by repeating a selection of an image and re-learning by a learning model using the selected image includes: obtaining means for obtaining the learning model that has already learnt; first selecting means for selecting, using the learning model obtained by the obtaining means, an image conversion scheme to be applied to an image; and second selecting means for selecting an image applied for re-learning by the learning model using the image conversion scheme selected by the first selecting means and the learning model obtained by the obtaining means.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technology of machine learning.

Background Art

[0002] In recent years, CV (computer vision) tasks using machine learning methods have been utilized in various scenarios. In particular, methods using neural networks including deep learning can extract knowledge and features from a vast amount of data and achieve higher performance by learning. However, a machine learning model using deep learning needs to be trained using a large amount of annotated data, and creating such annotated data requires a huge amount of work. For example, in a region segmentation task, for each target object imaged in an image, a label indicating the classification of the object and the region of the object are shown in pixel units to create annotated data. Performing such work on a large amount of data requires a huge amount of manpower and time, and there is a problem that the work cost is high. Therefore, it is desirable to be able to generate a machine learning model with higher performance by learning with less annotated data. As one method for this, active learning is known. Active learning is a method of generating a high-performance machine learning model at a low work cost by using the current machine learning model to select data considered to contribute more to performance improvement from unannotated data, and performing learning by annotating only the selected data. Non-Patent Document 1 describes using the uncertainty about the classification class of an object and the uncertainty about the position of the object as indicators for data selection in an object detector that detects a specific object included in an image.

Prior Art Documents

Non-Patent Documents

[0003]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the technique described in Non-Patent Document 1, when selecting data, data conversion such as geometric transformation and noise addition is performed to calculate the uncertainty. However, among various data conversion methods, the data conversion method to be performed is fixed without being changed from the initially set one. For example, assuming that vertical flipping is set as the data conversion of image data and data selection and re-learning are repeatedly performed. In such a case, for the current machine learning model, the inference result of the image data with noise addition may be less certain than the inference result of the image data with vertical flipping. In such a case, it may be desirable to select the image data with noise addition from the viewpoint of improving the performance of the machine learning model.

[0005] Therefore, an object of the present invention is to select an image that contributes to improving the performance of a learning model by using an appropriate image conversion method.

Means for Solving the Problems

[0006] The present invention is an information processing apparatus for performing active learning by repeating image selection and re-learning of a learning model using the selected image, and includes an acquisition unit that acquires a learned learning model, a first selection unit that selects an image conversion method to be applied to an image using the learning model acquired by the acquisition unit, a second selection unit that selects an image to be used for re-learning of the learning model using the image conversion method selected by the first selection unit and the learning model acquired by the acquisition unit.

Effects of the Invention

[0007] According to the present invention, an image that contributes to improving the performance of a learning model can be selected by using an appropriate image conversion method.

Brief Description of the Drawings

[0008]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Mode for Carrying Out the Invention

[0009] Hereinafter, preferred embodiments of the present invention will be described with reference to the accompanying drawings. Note that the configurations shown in the following embodiments are merely examples, and the present invention is not limited to those configurations.

[0010] <Embodiment 1> The active learning system according to this embodiment performs active learning by repeating the selection of an image using the current learning model and the relearning using the selected image in order to execute a CV (computer vision) task. The CV task includes image classification, object detection, region segmentation, etc. In this embodiment, the case of applying it to a learning model for executing an image classification task will be described.

[0011] FIG. 1 shows an example of the hardware configuration of the active learning system 1. The active learning system 1 includes a CPU 100, a ROM 110, a RAM 120, an HDD 130, an input unit 140, a display unit 150, and a communication unit 160. These components are interconnected via a bus 170. The CPU (Central Processing Unit) 100 is a central processing unit that performs operations for various processes. The CPU 100 controls the overall operation of the active learning system 1. By executing programs stored in the ROM 110, HDD 130, etc., the processes of the flowchart described later are realized. The ROM (Read Only Memory) 110 stores a control program. The RAM (Random Access Memory) 120 is used as a temporary storage area such as the main memory and work area of the CPU 100.

[0012] The HDD (Hard Disk Drive) 130 stores an image data set, model parameters constituting a learned model, various programs, etc. An external storage device may be used to perform a similar role. Here, the external storage device can be realized, for example, by a medium (recording medium) and an external storage drive for accessing the medium. Such media include flexible disks (FDs), CD-ROMs, DVDs, USB memories, MOs, flash memories, etc. are known. Also, the external storage device may be a server device connected via a network.

[0013] The input unit 140 is composed of a keyboard, a touch panel, etc., and receives input from the user. The display unit 150 is composed of a liquid crystal display, etc., and can display various data and processing results to the user. The communication unit 160 is a network interface for communicating with an external device. The CPU 100 may receive an instruction from the user via the communication unit 160 and may also transmit the processing result to an external device. The active learning system 1 can be configured using a general-purpose information processing device having the above configuration.

[0014] Figure 2 shows an example of the functional configuration of the active learning system 1 according to the present embodiment. The active learning system 1 functions as a data acquisition unit 210, an image conversion method selection unit 220, an image selection unit 230, an annotation addition unit 240, and a model learning unit 250 by the CPU 100 executing programs stored in the HDD 130 or the like. Further, the HDD 130 stores an annotated image dataset 10, an unannotated image dataset 20, and a learned model 200.

[0015] The data acquisition unit 210 acquires the annotated image dataset 10, the unannotated image dataset 20, and the learned model 200 from the HDD 130. Here, when the learned model 200 cannot be acquired, the data acquisition unit 210 may generate the learned model 200 in the model learning unit 250 using the annotated image dataset 10. Further, the data acquisition unit 210 may acquire the annotated image dataset 10 and the unannotated image dataset 20 via the input unit 140 or the communication unit 160.

[0016] The image conversion method selection unit 220 selects an image conversion method to be used at the time of image selection from among candidates for the image conversion method using the learned model 200 and the annotated image dataset 10. Examples of the image conversion method here include geometric conversion, tone conversion, noise addition, blurring, mosaic, or a combination thereof. The image conversion method selection unit 220 is an example of the first selection means. The image selection unit 230 selects an image that is considered to contribute to improving the performance of the learned model 200 from the unannotated image dataset 20 using the image conversion method selected by the image conversion method selection unit 220 and the learned model 200 as the current learning model. The image conversion method selection unit 220 is an example of the second selection means. The annotation addition unit 240 adds an annotation to the image selected by the image selection unit 230. The annotation includes correct answer information for the image. The model learning unit 250 retrains the trained model 200 using the images selected by the image selection unit 230 and annotated by the annotation adding unit 240.

[0017] FIG. 3 is a flowchart showing the active learning process according to the present embodiment. In the present embodiment, the case of applying it to an image classification task will be described as an example. Image classification is to identify the classification class of an image or a specific object included in the image. Here, an image classifier that classifies an image into "car" and "motorcycle" will be described as an example. In the following description, the notation of each step (step) will be omitted by prefixing S to the beginning of each step (step).

[0018] FIG. 4 is a diagram for explaining an example of an annotated image. The image 400 shown in FIG. 4(a) is an original image. In FIG. 4(b), the classification class "car" is given as the correct answer information for image classification for the image 400. As shown in FIG. 4(b), an image to which an annotation including correct answer information is given for an image is called an annotated image.

[0019] Hereinafter, the active learning process flow (S1 to S6) in FIG. 3 will be described. The active learning process flow is repeatedly executed until it is determined that the end condition is satisfied at S7. In S1, the data acquisition unit 210 acquires the trained model 200. In the present embodiment, a trained image classifier is acquired as the trained model 200. As a method for acquiring the trained image classifier, in the initial state, the image classifier may be learned and acquired using a small amount of annotated image dataset, or a generally distributed trained image classifier may be acquired. Also, after the start of this flowchart, the current trained image classifier is acquired. That is, the data acquisition unit 210 acquires the trained image classifier that was retrained / updated in the previous S6.

[0020] In S2, the data acquisition unit 210 determines whether it is possible to acquire an image without annotation. Specifically, the data acquisition unit 210 determines whether there is an image as the image dataset 20 without annotation, or whether there is an image newly added to the image dataset 20 without annotation. If the data acquisition unit 210 determines that it cannot acquire an image without annotation, assuming there is no new image available for learning, the series of processes in this flowchart is terminated. If the data acquisition unit 210 determines that it can acquire an image without annotation, it acquires the image dataset 20 without annotation and the image dataset 10 with annotation, and proceeds to S3.

[0021] In S3, the image conversion method selection unit 220 calculates scores for each candidate of the image conversion method using the image dataset 10 with annotation and the learned model 200 acquired in S1, and selects an image conversion method based on the calculated scores.

[0022] FIG. 5 schematically shows the processing content of the selection process of the image conversion method in S3 of the present embodiment. The selection process of the image conversion method in S3 includes a score calculation process 54. The score calculation process 54 is a process in which the image conversion method selection unit 220 scores each candidate of the predefined image conversion method candidates 50 using the learned image classifier 51. The image conversion method candidates 50 include geometric conversions such as inversion and rotation, tone conversions such as saturation and brightness, noise addition, blurring, mosaic, and the like. The learned image classifier 51 is an example of the learned model 200.

[0023] The score list 56 is the processing result of the score calculation process 54, and represents the scores for each candidate (vertical inversion, saturation modulation, etc.) of the image conversion method candidates 50. The image conversion method selection unit 220 refers to the score list 56 and selects the top K image conversion methods in descending order of scores, or image conversion methods with an arbitrary threshold or higher as important image conversion methods.

[0024] The score calculation process 54 includes an uncertainty calculation process 52. The uncertainty calculation process 52 is a process of calculating the uncertainty of the classification result obtained by inputting the converted image obtained by subjecting the annotated image included in the annotated image dataset 10 to image conversion by the image conversion method selection unit 220 into the learned image classifier 51. In the score calculation process 54, the image conversion method selection unit 220 calculates the score of each candidate of the image conversion method candidates 50 using the classification uncertainty index 528 calculated in the uncertainty calculation process 52. Hereinafter, the calculation of the score for vertical flipping, which is one of the image conversion method candidates 50, for the annotated image 520 will be described as an example. The annotated image 520 includes the correct answer information 522 included in the annotation and the image 524. Here, the correct answer information 522 represents "car".

[0025] First, the image conversion method selection unit 220 subjects the image 524 to vertical flipping to generate a converted image 525. Next, the image conversion method selection unit 220 classifies the converted image 525 using the learned image classifier 51 to obtain a classification result 526. The classification result 526 is an example of the output result obtained by inputting it into the current learned model. Then, the classification result 526 is converted into a probability distribution. Here, for the conversion into a probability distribution, a softmax function or the like may be used.

[0026] Here, since an image classifier that classifies an image into "car" and "bike" is taken as an example, the image conversion method selection unit 220 uses the correct answer information 522 to generate a correct answer probability distribution 523 in which the probability of "car" is 1 and the probability of "bike" is 0. In the present embodiment, the correct answer probability distribution is generated using the correct answer information given to the annotated image, but the classification result obtained by classifying the image before conversion (here, the image 524) using the learned image classifier 51 may be used as the correct answer information. That is, the classification uncertainty index may be calculated using a small amount of unannotated image dataset.

[0027] Subsequently, the image conversion method selection unit 220 calculates an uncertainty index 528 of the classification by the learned image classifier 51 for the annotated image 520 with vertical flipping, using the true probability distribution 523 and the classification result 526 converted into a probability distribution. In the present embodiment, as the uncertainty index 528 of the classification, the distance between probability distributions is used, and the distance is calculated by the Jensen-Shannon distance. Note that the method used for calculating the distance is not particularly limited as long as it can measure the distance between probability distributions. The Kullback-Leibler distance, Pearson distance, relative Pearson distance, L 2 distance, etc. may be used. The distance calculated here reflects the result of comparing the true information 522 and the classification result 526.

[0028] Subsequently, the image conversion method selection unit 220 similarly calculates the uncertainty index 528 of the classification for each annotated image included in the annotated image dataset 10, and integrates them for all the annotated images, thereby performing score calculation 540 for the image conversion method (here, vertical flipping). The score is calculated, for example, by the following formula (1).

[0029]

Equation

[0030] The image conversion method selection unit 220 calculates a representative value (typical value) of the uncertainty (Uim) of each annotated image using the above formula (1), and obtains the L1 distance from an arbitrarily set base point β. For calculating the representative value, a statistical representative value including the arithmetic mean and geometric mean, or a positional representative value including the median and the mode is used. Here, when using the Jensen-Shannon distance for uncertainty calculation, the closer the representative value is to 1, the higher the uncertainty. Therefore, the closer the base point β is to 0, the higher the score of the image conversion method with high uncertainty. Conversely, the closer the base point β is to 1, the lower the score of the image conversion method with high uncertainty. An image conversion method with high uncertainty can be considered difficult for the learned image classifier 51. The value of the base point β is determined empirically according to the task.

[0031] For each candidate of the image conversion method candidates 50, the image conversion method selection unit 220 similarly performs score calculation 540 to generate a score list 56, and selects the top K image conversion methods in the order of scores, or image conversion methods with a score threshold or higher. Here, K and the score threshold may be arbitrary values and are determined empirically.

[0032] Returning to the description of FIG. 3. In S4, the image selection unit 230 uses the image conversion method selected in S3 and the learned model 200 acquired in S1 to select unannotated images that are considered to contribute to performance improvement from the unannotated image dataset 20 as the objects for annotation.

[0033] FIG. 6 schematically shows the content of the process of selecting an image to be annotated in S4 of the present embodiment. The process of selecting an image to be annotated in S4 includes a priority calculation process 62. The priority calculation process 62 is a process in which the image selection unit 230 assigns a priority to each unannotated image in the unannotated image dataset 20 using the learned image classifier 51. The priority list 64 is the result of the priority calculation process 62 and represents the priority for each unannotated image (image 1, image 2,...) in the unannotated image dataset 20. The image selection unit 230 refers to the priority list 64 and selects the top N images with the highest priority or images with a priority equal to or higher than an arbitrary threshold as important images. Hereinafter, the calculation of the priority will be described by taking as an example the case where vertical flipping is selected as the image conversion method.

[0034] First, the image selection unit 230 classifies the unannotated image 620 using the learned image classifier 51 and obtains a classification result 624. Then, the classification result 624 is converted into a probability distribution. Next, the image selection unit 230 performs vertical flipping on the unannotated image 620 to generate a transformed image 622. Then, the image selection unit 230 classifies the transformed image 622 using the learned image classifier 51 and obtains a classification result 626. Then, the classification result 626 is converted into a probability distribution. Subsequently, the image selection unit 230 calculates a classification uncertainty index 627 of the learned image classifier 51 for the unannotated image 620 subjected to vertical flipping using the classification result 624 converted into a probability distribution and the classification result 626 converted into a probability distribution. Here, the method of calculating the classification uncertainty index 627 is the same as the method of calculating the classification uncertainty index 528 in FIG. 5.

[0035] Then, the image selection unit 230 performs image priority calculation 628 using the calculated classification uncertainty index 627. The priority is represented by, for example, the following formula (2).

[0036]

Equation

[0037] The image selection unit 230 obtains the L1 distance from an arbitrarily set base point β using the uncertainty (Uim) of the image without annotation and the above formula (2). The closer the base point β is to 0, the higher the priority of the image with poor certainty of the learned image classifier 51 becomes. An image with high uncertainty can be considered a difficult image for the learned image classifier 51.

[0038] For each image without annotation included in the image dataset without annotation 20, the image selection unit 230 performs priority calculation 628 to generate a priority list 64, and selects the top N images in order of priority or images with a priority threshold or higher as the objects for annotation. Here, N and the priority threshold may be arbitrary values.

[0039] Returning to the description of FIG. 3. In S5, the annotation unit 240 annotates the image without annotation selected in S4. Here, annotation may be performed using the display unit 150 and the input unit 140. The annotation unit 240 adds the image annotated in S5 to the annotated image dataset 10 and updates the annotated image dataset 10.

[0040] In S6, the model learning unit 250 performs relearning of the learned model 200 obtained in S1 using the annotated image dataset 10 updated in S5. The model learning unit 250 updates the learned model 200 that has undergone relearning as a new learned model 200.

[0041] In S7, the CPU 100 determines whether or not the learning end condition is satisfied. Here, for the learning end condition, for example, the model learning unit 250 may determine whether or not a desired accuracy has been obtained for the learned model 200 for which relearning has been performed, and use the result of the determination. Alternatively, the CPU 100 may determine whether or not there is an image with a priority threshold or higher in the priority list 64 and use the result of the determination.

[0042] When the CPU 100 determines that the learning end condition is not satisfied, for example, when it indicates that the desired accuracy has not yet been obtained, it transitions to S1 and acquires the current learned model 200 (the learned model 200 for which relearning has been performed). That is, the CPU 100 repeats the active learning processing flow (S1 to S6) until the learning end condition is satisfied. When the CPU 100 determines that the learning end condition is satisfied, it ends the series of processes in this flowchart.

[0043] According to the present embodiment, in the process of performing active learning of a learning model for executing an image classification task, the learning transformation method used when selecting an image that contributes to improving the performance of the current learned model can be dynamically selected using the current learned model. That is, an image that contributes to improving the performance of the learned model can be selected using an appropriate image transformation method.

[0044] <Embodiment 2> In the present embodiment, the case of applying it to an object detection task will be described as an example. Object detection is to identify the position of a specific object included in an image and the classification class. Here, an object detector that detects "car" and "motorcycle" from an image will be described as an example. Hereinafter, the same parts as those in Embodiment 1 will be omitted from the description, and the description will focus on the parts different from those in Embodiment 1.

[0045] In FIG. 4(c), for the image 400, a bounding box indicating the position of the car is given as the correct answer information for object detection, and the classification class "car" is assigned.

[0046] Also in this embodiment, the flowchart shown in FIG. 3 is applied. FIG. 7 schematically shows the uncertainty calculation process 52 in the selection process of the image conversion method in S3 of this embodiment. The uncertainty calculation process 52 is a process in which the image conversion method selection unit 220 calculates the uncertainty of the detection result obtained by inputting the converted image obtained by performing image conversion on the annotated image included in the annotated image dataset 10 to the learned object detector 71. The learned object detector 71 is an example of the learned model 200. Hereinafter, the calculation of the uncertainty with respect to vertical flipping, which is one of the image conversion method candidates 50, for the annotated image 720 will be described as an example. The annotated image 720 includes the correct answer information 522 included in the annotation, the correct answer information 722 of the position, and the image 524. It is assumed that the correct answer information 522 and the image 524 are the same as those in FIG. 5.

[0047] First, the image conversion method selection unit 220 performs vertical flipping on the image 524 to generate a converted image 525. Next, the image conversion method selection unit 220 detects the objects included in the converted image 525 with the learned object detector 71 to obtain a detection result 726. Next, the image conversion method selection unit 220 performs vertical flipping on the correct answer information 722 of the position to generate the correct answer information 723 of the converted position.

[0048] Subsequently, the image conversion method selection unit 220 calculates the classification uncertainty index 528 of the learned object detector 71 for the annotated image 720 subjected to vertical flipping using the correct probability distribution 523 and the classification probability distribution of the detection result 726. The classification uncertainty index 528 is the same as that in FIG. 5. Next, the image conversion method selection unit 220 calculates the position uncertainty index 728 of the learned object detector 71 for the annotated image 720 subjected to vertical flipping using the correct answer information 723 of the converted position and the position information of the detection result 726.

[0049] In this embodiment, IoU (Intersection over Union) is used as the position uncertainty index 728. IoU is calculated, for example, by the following formula (3) for the regions A and B representing the positions of the objects in the image.

[0050]

Number

[0051] As shown in the above formula (3), IoU is an index representing the degree of overlap between two regions (region A and region B) by dividing the common part of the two regions by the union set. The value of IoU indicates that the closer it is to 0, the less the overlap and the higher the uncertainty of the position. Note that as the position uncertainty index 728, any index other than IoU may be used as long as it represents the degree of overlap of the position information.

[0052] Subsequently, the image conversion method selection unit 220 performs uncertainty integration 729 by combining the classification uncertainty index 528 and the position uncertainty index 728. In the uncertainty integration 729, both the classification uncertainty index 528 and the position uncertainty index 728 may be used, or either one of the classification uncertainty index 528 and the position uncertainty index 728 may be used. For the classification uncertainty index 528 (U im class ), when using the Jensen-Shannon distance and using IoU for the position uncertainty index 728 (U im local ), the uncertainty of the image (U im ) is represented by the following formula (4).

[0053]

Number

[0054] In the above formula (4), the closer the classification uncertainty index 528 (U im class ) is to 1, the closer the position uncertainty index 728 (U im local) Since the uncertainty increases as it approaches 0, the positive and negative correlations of the uncertainties are unified. The image conversion method selection unit 220 similarly performs uncertainty integration 729 on each annotated image included in the annotated image dataset 10 and integrates it across all annotated images to calculate the score of the image conversion method (here, vertical flipping). Further, the image conversion method selection unit 220 similarly calculates the score for each candidate of the image conversion method candidates 50, and selects the top K image conversion methods in order of score or image conversion methods with a score threshold or higher.

[0055] FIG. 8 schematically shows the content of the selection process of the image to be annotated in S4 of the present embodiment. The selection process of the image to be annotated in S4 includes a priority calculation process 62. The priority calculation process 62 is a process in which the image selection unit 230 assigns priorities to each unannotated image in the unannotated image dataset 20 using the learned object detector 71. Hereinafter, the calculation of the priority will be described by taking as an example the case where vertical flipping is selected as the image conversion method.

[0056] First, the image selection unit 230 detects the objects included in the unannotated image 620 using the learned object detector 71 and obtains the detection result 824. Then, the converted detection result 825 obtained by applying vertical flipping to the detection result 824 is obtained. Next, the image selection unit 230 applies vertical flipping to the unannotated image 620 to generate the converted image 622. Then, the image selection unit 230 detects the objects included in the converted image 622 using the learned object detector 71 and obtains the detection result 826 of the converted image.

[0057] Subsequently, the image selection unit 230 calculates the classification uncertainty index 627 of the learned object detector 71 for the vertically flipped annotation-free image 620 using the classification probability distribution of the detection result 825 after transformation and the classification probability distribution of the detection result 826 of the transformed image. Similarly, the image selection unit 230 calculates the position uncertainty index 828 of the learned object detector 71 for the vertically flipped annotation-free image 620 using the position information of the detection result 825 after transformation and the position information of the detection result 826 of the transformed image. Here, the method for calculating the position uncertainty index 828 is the same as the method for calculating the position uncertainty index 728 in FIG. 7.

[0058] The image selection unit 230 performs uncertainty integration 729 by combining the classification uncertainty index 627 and the position uncertainty index 828. Subsequently, the image selection unit 230 performs priority calculation 628 for each annotation-free image included in the annotation-free image dataset 20 using the value obtained by performing the uncertainty integration 729.

[0059] According to this embodiment, in the process of performing active learning of the learning model for executing the object detection task, the learning transformation method used when selecting an image that contributes to the performance improvement of the current learned model can be dynamically selected using the current learned model. That is, an appropriate image transformation method can be used to select an image that contributes to the performance improvement of the learned model.

[0060] <Embodiment 3> In this embodiment, the case of applying it to the region segmentation task will be described as an example. Region segmentation is to identify the region of a specific object included in an image and the classification class. Here, a region segmenter that segments the regions of "car" and "bike" from an image will be described as an example. Hereinafter, the parts that are the same as those in Embodiment 1 will be omitted from the description, and the description will focus on the parts that are different from Embodiment 1.

[0061] In FIG. 4(d), for the image 400, the region of the car and the classification class "car" are given as the correct information for region segmentation.

[0062] Also in this embodiment, the flowchart shown in FIG. 3 is applied. FIG. 9 schematically shows the uncertainty calculation process 52 in the selection process of the image conversion method in S3 of this embodiment. The uncertainty calculation process 52 is a process in which the image conversion method selection unit 220 calculates the uncertainty of the segmentation result obtained by inputting the converted image obtained by performing image conversion on the annotated image included in the annotated image dataset 10 into the learned region segmenter 91. The learned region segmenter 91 is an example of the learned model 200. Hereinafter, the calculation of the uncertainty with respect to the vertical inversion, which is one of the image conversion method candidates 50, for the annotated image 920 will be described as an example. The annotated image 920 includes the correct classification information 522 included in the annotation, the correct region information 922, and the image 524. The correct information 522 and the image 524 are the same as those in FIG. 5.

[0063] First, the image conversion method selection unit 220 performs vertical inversion on the image 524 to generate a converted image 525. Next, the image conversion method selection unit 220 uses the learned region segmenter 91 to segment the object regions included in the converted image 525 and obtains a segmentation result 926. Next, the image conversion method selection unit 220 performs vertical inversion on the correct region information 922 to generate the correct region information 923 after conversion.

[0064] Subsequently, the image conversion method selection unit 220 calculates the classification uncertainty index 528 of the learned region segmenter 91 for the annotated image 920 after vertical inversion using the correct probability distribution 523 and the classification probability distribution of the segmentation result 926. The classification uncertainty index 528 is the same as that in FIG. 5.

[0065] Next, the image conversion method selection unit 220 calculates an uncertainty index 928 of the regions of the learned region segmenter 91 for the annotated image 920 with vertical flipping, using the ground truth information 923 of the regions after conversion and the information of the regions of the segmentation result 926. Here, the IoU is used as the uncertainty index 928 of the regions. That is, the image conversion method selection unit 220 may calculate the uncertainty index 928 of the regions in the same manner as the uncertainty index 728 at the position in FIG. 7. The image conversion method selection unit 220 performs the same uncertainty integration 729 as in FIG. 7 for each annotated image included in the annotated image dataset 10, and integrates them over all annotated images to calculate the score of the image conversion method (here, vertical flipping). Further, the image conversion method selection unit 220 calculates the score in the same manner for each candidate of the image conversion method candidates 50, and selects the top K image conversion methods in order of score, or the image conversion methods with a score threshold or higher.

[0066] FIG. 10 schematically shows the content of the selection process of the image to be annotated in S4 of the present embodiment. The selection process of the image to be annotated in S4 includes a priority calculation process 62. The priority calculation process 62 is a process in which the image selection unit 230 assigns priorities to each unannotated image in the unannotated image dataset 20 using the learned region segmenter 91. Hereinafter, the calculation of the priority will be described by taking as an example the case where vertical flipping is selected as the image conversion method.

[0067] First, the image selection unit 230 divides the regions of the objects included in the unannotated image 620 using the learned region segmenter 91 to obtain a segmentation result 1024. Then, a transformed segmentation result 1025 is obtained by applying vertical flipping to the information of the regions of the segmentation result 1024. Next, the image selection unit 230 applies vertical flipping to the unannotated image 620 to generate a transformed image 622. Then, the image selection unit 230 divides the regions of the objects included in the transformed image 622 using the learned region segmenter 91 to obtain a segmentation result 1026 of the transformed image.

[0068] Subsequently, the image selection unit 230 calculates the classification uncertainty index 627 of the learned region divider 91 for the vertically flipped image without annotation 620 using the probability distribution of the classification included in the converted split result 1025 and the classification probability distribution of the split result 1026 of the converted image. Similarly, using the region information of the converted split result 1025 and the region information of the split result 1026 of the converted image, the region uncertainty index 1028 of the learned region divider 91 for the vertically flipped image without annotation 620 is calculated. Here, the method for calculating the region uncertainty index 1028 is the same as the method for calculating the region uncertainty index 928 in FIG. 9.

[0069] The image selection unit 230 performs uncertainty integration 729 by combining the classification uncertainty index 627 and the region uncertainty index 1028. Subsequently, using the value obtained by performing the uncertainty integration 729, priority calculation 628 is performed for each image without annotation included in the image dataset without annotation 20.

[0070] According to this embodiment, in the process of performing active learning of a learning model for executing a region segmentation task, the learning transformation method used when selecting an image that contributes to improving the performance of the current learned model can be dynamically selected using the current learned model. That is, an appropriate image transformation method can be used to select an image that contributes to improving the performance of the learned model.

[0071] <Embodiment 4> In this embodiment, in the process of performing active learning, a method for changing the currently set image transformation method when the change condition of the image transformation method is satisfied will be described. Hereinafter, the same parts as in Embodiment 1 will be omitted from the description, and the description will focus on the parts different from Embodiment 1.

[0072] FIG. 11 is a flowchart showing the active learning process according to this embodiment. S12 to S16 in FIG. 11 are repeatedly executed until it is determined in S17 that the end condition is satisfied. In S11, the data acquisition unit 210 acquires the image conversion method being set. As a method for acquiring the image conversion method, in the initial state, the default-set image conversion method may be acquired, or the image conversion method set by the user using the input unit 140 may be acquired. The image conversion method being set is stored in the ROM 120 or the like.

[0073] In S12, the data acquisition unit 210 acquires the learned model 200. In the present embodiment, a learned model corresponding to the task is acquired. As a method for acquiring the learned model, in the initial state, the learning model may be learned and acquired using a small amount of annotated image dataset, or a generally distributed learning model may be acquired. Also, after the start of this flowchart, the current learned model is acquired. That is, the data acquisition unit 210 acquires the learned model re-learned and updated in the previous S16.

[0074] In S13, the data acquisition unit 210 determines whether it is possible to acquire an image without annotation. The process of this step is the same as that of S2. If the data acquisition unit 210 determines that an image without annotation cannot be acquired, it is considered that there is no new image available for learning, and a series of processes of this flowchart are terminated. If the data acquisition unit 210 determines that an image without annotation can be acquired, the image dataset 20 without annotation and the image dataset 10 with annotation are acquired, and the process proceeds to S14.

[0075] In S14, the image selection unit 230 selects, as an object for annotation, an image without annotation that is considered to contribute to performance improvement from the image dataset 20 without annotation using the image conversion method being set and the learned model 200 acquired in S12. Here, the image selection unit 230 reads the image conversion method being set from the ROM 120 or the like. The process of selecting the image for annotation in this step is the same as the process described in FIGS. 6, 8, and 10.

[0076] In S15, the annotation adding unit 240 adds an annotation to the annotation-free image selected in S14. The processing of this step is the same as that of S5. The annotation adding unit 240 adds the image annotated in S15 to the annotated image dataset 10 and updates the annotated image dataset 10. In S16, the model learning unit 250 uses the annotated image dataset 10 updated in S15 to retrain the pre-trained model 200 obtained in S12. The model learning unit 250 updates the pre-trained model 200 after retraining as a new pre-trained model 200.

[0077] In S17, the CPU 100 determines whether the learning end condition is satisfied. The processing of this step is the same as that of S7. If the CPU 100 determines that the learning end condition is not satisfied, it transitions to S18. If the CPU 100 determines that the learning end condition is satisfied, the series of processes in this flowchart ends.

[0078] In S18, the CPU 100 determines whether the change condition of the image conversion method is satisfied. Here, for the change condition of the image conversion method, for example, the model learning unit 250 may determine whether the cumulative number of images used for retraining exceeds a predetermined number, and use the result of the determination. Alternatively, the model learning unit 250 may determine whether an index representing the proficiency of the current pre-trained model 200 (the pre-trained model 200 after retraining) exceeds a proficiency threshold, and use the result of the determination. The change condition of the image conversion method is not particularly limited as long as it is a condition related to the progress of learning of the pre-trained model 200 in the process of active learning. Also, a plurality of change conditions of the image conversion method are provided according to the progress stage of learning of the pre-trained model 200. For example, a plurality of proficiency thresholds are provided. If the CPU 100 determines that the change condition of the image conversion method is not satisfied, it transitions to S12 and obtains the current pre-trained model 200 (the pre-trained model 200 after retraining). If the CPU 100 determines that the change condition of the image conversion method is satisfied, it transitions to S19.

[0079] In S19, the image conversion method selection unit 220 calculates scores for each candidate of the image conversion method using the annotated image dataset 10 and the current learned model 200 (the learned model 200 for which relearning has been performed). Then, based on the calculated scores, the image conversion method is selected. The selection process of the image conversion method in this step is the same as the process described in FIGS. 5, 7, and 9. The image conversion method selection unit 220 updates the currently set image conversion method stored in the ROM 120 or the like with the selected image conversion processing method. The subsequent process then transitions to S12.

[0080] According to the present embodiment, in the process of performing active learning of the learning model, the image conversion method can be changed step by step according to the progress stage of the learning model. That is, using an appropriate image conversion method, an image that contributes to improving the performance of the learned model can be selected.

[0081] As described in detail above for each embodiment, the present invention can be implemented in embodiments such as a system, device, method, program, or recording medium (storage medium), for example. Specifically, it may be applied to a system composed of a plurality of devices (for example, a host computer, an interface device, an imaging device, a web application, etc.), or may also be applied to a device consisting of a single device.

[0082] Also, each embodiment is merely an example of the implementation of the present invention, and the technical scope of the present invention should not be construed in a limited manner by these. That is, the present invention can be implemented in various forms without departing from its technical idea or its main features.

[0083] The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (for example, an ASIC) that realizes one or more functions.

[0084] The disclosure of each of the above embodiments includes the following configurations, methods, and programs. (Configuration 1) An information processing apparatus for performing active learning by repeating image selection and re-learning of a learning model using the selected image, an acquisition means for acquiring a learned learning model, a first selection means for selecting an image conversion method to be applied to an image using the learning model acquired by the acquisition means, a second selection means for selecting an image for re-learning the learning model using the image conversion method selected by the first selection means and the learning model acquired by the acquisition means, An information processing apparatus characterized by comprising the above. (Configuration 2) The first selection means calculates a score for each candidate of the image conversion method using the annotated image with annotations including correct answer information and the learning model acquired by the acquisition means, and selects the image conversion method based on the score. The information processing apparatus according to Configuration 1. (Configuration 3) The first selection means a first calculation means for calculating the uncertainty of the output result obtained by inputting the converted image obtained by performing image conversion corresponding to each candidate of the image conversion method on the annotated image into the learning model acquired by the acquisition means, a second calculation means for calculating the score for each candidate of the image conversion method based on the uncertainty calculated by the first calculation means, An information processing apparatus according to Configuration 2, characterized by comprising the above. (Configuration 4) The first calculation means calculates the uncertainty by comparing the output result with the correct answer information. The information processing apparatus according to Configuration 3. (Configuration 5) The first selection means selects an image conversion method whose score is higher or equal to a threshold value. The information processing apparatus according to any one of Configurations 2 to 4. (Configuration 6) The second selection means selects an image from among the annotationless images without an annotation with correct answer information, The image selected by the second selection means is the information processing apparatus according to any one of Configurations 1 to 5, characterized in that it is a target for annotation. (Configuration 7) The second selection means For the annotationless image, the uncertainty of the output result obtained by inputting the converted image obtained by performing image conversion according to the image conversion method selected by the first selection means into the learning model acquired by the acquisition means is calculated. A third calculation means for calculating; A fourth calculation means for calculating the priority of the annotationless image based on the uncertainty calculated by the third calculation means; The information processing apparatus according to Configuration 6, characterized by having. (Configuration 8) The learning model is for performing at least one task including at least one of image classification, object detection, and region segmentation, and is the information processing apparatus according to any one of Configurations 1 to 7. (Configuration 9) When the classification result is included in the output result, the first calculation means calculates the uncertainty based on the distance of the probability distribution between the probability distribution obtained from the correct answer information and the classification result converted into the probability distribution. The information processing apparatus according to Configuration 3, characterized by the above. (Configuration 10) When the output result includes position or region information, the first calculation means calculates the uncertainty based on the degree of overlap between the position or region obtained from the correct answer information and the position or region included in the output result. The information processing apparatus according to Configuration 3, characterized by the above. (Configuration 11) When both the classification result and the position or region information are included in the output result, the second calculation means calculates the combination of the uncertainty calculated based on the classification result and the uncertainty calculated based on the position or region, or The information processing apparatus according to Configuration 3, 9, or 10, characterized in that the score is calculated based on one of the uncertainties. (Configuration 12) The first selection means selects an image conversion method from candidates including at least any one of geometric transformation, color tone transformation, noise addition, blurring, and mosaic, in the information processing apparatus according to any one of Configurations 1 to 11. (Configuration 13) The information processing apparatus further includes an update means for updating the learning model by performing re-learning of the learning model acquired by the acquisition means using the image selected by the second selection means. The acquisition means acquires the learning model updated by the update means, in the information processing apparatus according to any one of Configurations 1 to 12. (Configuration 14) An information processing apparatus for performing active learning by repeating image selection and re-learning of a learning model using the selected image, A setting means for setting an image conversion method to be applied to an image, An acquisition means for acquiring a learned learning model, A selection means for selecting an image to be used for re-learning of the learning model using the image conversion method set by the setting means and the learning model acquired by the acquisition means, and having The setting means changes the learning conversion method being set according to the progress of learning of the learning model acquired by the acquisition means, in the information processing apparatus. (Method) An information processing method for performing active learning by repeating image selection and re-learning of a learning model using the selected image, An acquisition step of acquiring a learned learning model, A first selection step of selecting an image conversion method to be applied to an image using the learning model acquired in the acquisition step, A second selection step of selecting an image to be used for re-learning of the learning model using the image conversion method selected in the first selection step and the learning model acquired in the acquisition step, and including (Program) A computer of an information processing apparatus for performing active learning by repeating selection of an image and re-learning of a learning model using the selected image, an acquisition means for acquiring a learned learning model, a first selection means for selecting an image conversion method to be applied to an image using the learning model acquired by the acquisition means, a second selection means for selecting an image to be used for re-learning of the learning model using the image conversion method selected by the first selection means and the learning model acquired by the acquisition means, a program for causing the above to function.

Claims

1. An information processing apparatus for performing active learning by repeating selection of an image and re-training of a learning model using the selected image, comprising: an acquisition means for acquiring a trained learning model; a first selection means for selecting an image conversion method to be applied to an image using the learning model acquired by the acquisition means; a second selection means for selecting an image for re-training the learning model using the image conversion method selected by the first selection means and the learning model acquired by the acquisition means; An information processing apparatus characterized by comprising the above.

2. The information processing apparatus according to claim 1, wherein the first selection means calculates a score for each candidate of the image conversion method using an annotated image with an annotation including correct answer information and the learning model acquired by the acquisition means, and selects the image conversion method based on the score.

3. The first selection means includes: a first calculation means for calculating the uncertainty of an output result obtained by inputting a converted image obtained by performing image conversion corresponding to each candidate of the image conversion method on the annotated image into the learning model acquired by the acquisition means; a second calculation means for calculating the score for each candidate of the image conversion method based on the uncertainty calculated by the first calculation means; The information processing apparatus according to claim 2, characterized by comprising the above.

4. The information processing apparatus according to claim 3, wherein the first calculation means calculates the uncertainty by comparing the output result and the correct answer information.

5. The information processing apparatus according to claim 2, wherein the first selection means selects an image conversion method having a score in the upper rank or equal to or higher than a threshold value.

6. The second selection means selects an image from among unannotated images without an annotation including correct answer information, The information processing apparatus according to claim 1, wherein the image selected by the second selection means is an object for annotation.

7. The second selection means includes: a third calculation means for calculating the uncertainty of an output result obtained by inputting a converted image obtained by performing image conversion corresponding to the image conversion method selected by the first selection means on the unannotated image into the learning model acquired by the acquisition means; a fourth calculation means for calculating the priority of the unannotated image based on the uncertainty calculated by the third calculation means; The information processing apparatus according to claim 6, characterized by comprising the above.

8. The information processing apparatus according to claim 1, wherein the learning model is for performing a task including at least any one of image classification, object detection, and region segmentation.

9. The information processing apparatus according to claim 3, wherein when the classification result is included in the output result, the first calculation means calculates the uncertainty based on the distance between the probability distribution obtained from the correct information and the classification result converted into the probability distribution.

10. The information processing apparatus according to claim 3, wherein when the output result includes position or region information, the first calculation means calculates the uncertainty based on the degree of overlap between the position or region obtained from the correct information and the position or region included in the output result.

11. The information processing apparatus according to claim 3, wherein when the output result includes both a classification result and position or region information, the second calculation means calculates the score based on a combination of the uncertainty calculated based on the classification result and the uncertainty calculated based on the position or region, or based on one of the uncertainties.

12. The information processing apparatus according to claim 1, wherein the first selection means selects an image conversion method from candidates including at least any one of geometric transformation, color tone transformation, noise addition, blurring, and mosaic.

13. The information processing apparatus further includes an update means for re-learning the learning model obtained by the acquisition means using the image selected by the second selection means to update the learning model, The information processing apparatus according to claim 1, wherein the acquisition means acquires the learning model updated by the update means.

14. An information processing apparatus for performing active learning by repeating image selection and re-learning of a learning model using the selected image, A setting means for setting an image conversion method to be applied to the image, An acquisition means for acquiring a learned learning model, A selection means for selecting an image to be used for re-learning the learning model using the image conversion method set by the setting means and the learning model acquired by the acquisition means, and The setting means is characterized in that it changes the learning conversion method being set according to the progress of learning of the learning model acquired by the acquisition means.

15. An information processing method for performing active learning by repeating selection of an image and re-learning of a learning model using the selected image, comprising: an acquisition step of acquiring a learned learning model; a first selection step of selecting an image conversion method to be applied to an image using the learning model acquired in the acquisition step; a second selection step of selecting an image to be used for re-learning the learning model using the image conversion method selected in the first selection step and the learning model acquired in the acquisition step; An information processing method characterized by including the above steps.

16. A computer of an information processing apparatus for performing active learning by repeating selection of an image and re-learning of a learning model using the selected image, causing the computer to function as: an acquisition means for acquiring a learned learning model; a first selection means for selecting an image conversion method to be applied to an image using the learning model acquired by the acquisition means; a second selection means for selecting an image to be used for re-learning the learning model using the image conversion method selected by the first selection means and the learning model acquired by the acquisition means; A program for causing the computer to function as described above.