Image selection method, learning method, image selection device, and program
Patent Information
- Application Number
- JP2025557751
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Priority Date
- 2024-05-14
- Filing Date
- 2024-10-30
- Publication Date
- 2025-05-22
AI Technical Summary
The annotation process for training machine learning models is costly, and existing methods do not effectively address the issue of undetected objects in image recognition, leading to decreased accuracy under domain shift.
An image selection method that estimates the degree of undetection in unlabeled images using an undetected object prediction model and selects images for labeling based on this estimation, thereby reducing annotation costs while improving detection performance.
This approach effectively reduces the annotation cost for training machine learning models while improving the detection performance for undetected objects, thereby enhancing the robustness of the model against domain shift.
Abstract
Description
Image selection method, learning method, image selection device, and program
[0001] The present disclosure relates to an image selection method, a learning method, an image selection device, and a program.
[0002] Various methods have been studied as a method for learning a machine learning model using a neural network, etc. For example, Patent Literature 1 discloses a learning device for a machine learning model that can accurately perform image recognition such as object detection.
[0003] JP 2024-006730 A
[0004] Incidentally, when training a machine learning model, images for training are collected and annotated, but there is a problem in that the annotation work is costly.
[0005] Therefore, the present disclosure provides an image selection method, a learning method, an image selection device, and a program that enable effective learning of a machine learning model while reducing the cost of annotation for learning the machine learning model.
[0006] An image selection method according to one aspect of the present disclosure is an image selection method executed by a processor using a memory, which inputs a plurality of unlabeled images into an undetected object prediction model to estimate a degree of undetection based on the number of undetected objects in each of the plurality of unlabeled images, and selects one or more unlabeled images to be labeled from the plurality of unlabeled images based on the degree of undetection in each of the plurality of unlabeled images.
[0007] A learning method according to one aspect of the present disclosure assigns a label to each of the one or more unlabeled images selected using the image selection method described above, and trains a target model to be learned using one or more labeled images that are the one or more unlabeled images to which labels have been assigned.
[0008] An image selection device according to one aspect of the present disclosure includes an estimation unit that estimates a degree of undetection based on the number of undetected objects in each of a plurality of unlabeled images by inputting the plurality of unlabeled images into an undetected object prediction model, and a selection unit that selects one or more unlabeled images to be labeled from among the plurality of unlabeled images based on the degree of undetection in each of the plurality of unlabeled images.
[0009] A program according to one aspect of the present disclosure is a program for causing a computer to execute the image selection method described above.
[0010] According to one aspect of the present disclosure, it is possible to realize an image selection method, etc., that enables effective training of a machine learning model while reducing the cost of annotation for training the machine learning model.
[0011] Fig. 1 is a block diagram showing the functional configuration of an information processing device according to an embodiment. Fig. 2 is a flowchart showing the operation of the information processing device according to an embodiment. Fig. 3 is a diagram for explaining the learning process of a target model according to an embodiment. Fig. 4 is a diagram for explaining the learning process of an FNPM according to an embodiment.
[0012] (Background to the Invention of the Present Disclosure) Before describing the present disclosure, the background to the invention of the present disclosure will be described.
[0013] If the domain, which is the concept of a collection of data in a certain environment, is different, the appearance in an image will differ, and there is a risk that the recognition performance of a machine learning model will deteriorate. Therefore, it is desirable to adapt the machine learning model from an existing environment (Source Domain) to a new environment (Target Domain). Note that adapting a machine learning model from an existing environment to a new environment is also referred to as domain adaptation (Active Domain Adaptation (ADA)). Furthermore, an example of a machine learning model is a machine learning model that performs image recognition using deep learning, particularly object detection. Object detection means outputting the position and class of an object in an image.
[0014] Conventionally, domain adaptation of a machine learning model involves collecting images in a new environment, annotating them with annotation information (teacher labels), and retraining the model. However, this annotation process is costly and requires a lot of man-hours. Furthermore, machine learning models are expected to be deployed in a variety of fields and perform accurate recognition. For example, it is desirable to prevent the recognition performance of a machine learning model from deteriorating due to differences in the application environment. Because the performance of a machine learning model depends on the images used during training, generating such a machine learning model is likely to require even greater annotation costs.
[0015] Here, domain adaptation is being considered by assigning only a small number (e.g., about a few percent) of teacher labels to images in a new environment (unlabeled images) and training (e.g., relearning) a machine learning model using the images in the new environment to which the teacher labels have been assigned. For example, ADA technology is being considered, which effectively samples and uses a small number of images in the target domain. Since only a small number of images need to be annotated, it is possible to reduce the annotation cost for training a machine learning model.
[0016] In this case, it is desirable to further improve (e.g., maximize) the performance of the machine learning model by selecting and labeling images that are effective for learning and using them for learning. For example, it is thought that selecting images that the machine learning model is likely to have difficulty with as images to label can contribute to improving learning efficiency. Images that the machine learning model is likely to have difficulty with can be identified based on the uncertainty of the detection results estimated by the machine learning model.
[0017] However, object detection involves a unique error known as missed detection. Conventionally, uncertainty is measured after a model detects an object, meaning that if the object is overlooked, the uncertainty cannot be measured at all. In particular, under circumstances where the environment changes and the appearance of the image changes (domain shift), an increase in undetected objects can lead to a decrease in accuracy, making it important to take measures against undetected objects. Note that a change in the environment refers to, for example, differences between the image captured during training of a machine learning model and the image captured in an environment where the machine learning model is used. This includes, for example, differences in at least one of the lighting environment (visible light, infrared), weather (sunny, cloudy, rain, snow, fog), time of day (morning, noon, night), and image type (CG, live action), etc.
[0018] Therefore, the inventors of the present application believed that by selecting images that are likely to be overlooked in a target model, which is an object detection model to be trained, (i.e., images containing undetected objects) as images to be labeled, it would be possible to effectively reduce undetected objects in the target model, and devised an image selection method etc. that can estimate the number of undetected objects in an image and select images to use for training based on the estimated number of undetected objects, i.e., an image selection method etc. that can effectively train a machine learning model while reducing the cost of annotation for training the machine learning model. The image selection method etc. is a method of selecting samples by quantifying the degree of undetection of objects in object detection.
[0019] It is believed that the performance degradation due to domain shift can be improved by selecting images captured in a new environment using the image selection method of the present application. For example, by using the image selection method of the present application, it is possible to achieve both performance degradation due to domain shift and reduced annotation costs.
[0020] An image selection method according to a first aspect of the present disclosure is an image selection method executed by a processor using a memory, in which a plurality of unlabeled images are input into an undetected object prediction model to estimate a degree of undetection based on the number of undetected objects in each of the plurality of unlabeled images, and one or more unlabeled images to be labeled are selected from the plurality of unlabeled images based on the degree of undetection in each of the plurality of unlabeled images.
[0021] As a result, images containing undetected objects are selected as images to be labeled, and when such images are used to train a machine learning model, the machine learning model's detection performance for undetected objects can be effectively improved. Furthermore, since it is not necessary to label each of multiple unlabeled images, the cost of annotation for training the machine learning model can be reduced. Therefore, according to the image selection method according to one aspect of the present disclosure, it is possible to effectively train the machine learning model while reducing the cost of annotation for training the machine learning model.
[0022] Furthermore, for example, the image selection method according to the second aspect may be the image selection method according to the first aspect, in which an input image is input into the undetected object prediction model to predict a first degree of undetection for the input image, a second degree of undetection for the input image is calculated based on correct answer information for an object region in the input image, and the undetected object prediction model is trained based on the first degree of undetection and the second degree of undetection.
[0023] This allows the undetected object prediction model to be trained so that the number of undetected objects can be accurately predicted. In other words, unlabeled images that can more effectively train the machine learning model can be selected as images to be labeled. This allows the machine learning model to be trained more effectively.
[0024] Also, for example, the image selection method according to the third aspect may be the image selection method according to the first or second aspect, and the one or more unlabeled images may be images selected from the plurality of unlabeled images based on the degree of non-detection.
[0025] This allows us to select images to be labeled based on the degree of undetection, and by using such images, we can more effectively train the machine learning model.
[0026] Also, for example, the image selection method according to the fourth aspect may be the image selection method according to the third aspect, wherein the one or more unlabeled images may be, among the plurality of unlabeled images, images whose degree of undetection is equal to or greater than a first predetermined degree, or the second predetermined number of images that are the highest.
[0027] This allows us to prioritize the selection of unlabeled images that contain many undetected objects and assign labels to them. By training a machine learning model using such images, we can realize a machine learning model that is robust against undetected objects.
[0028] Furthermore, for example, the image selection method according to the fifth aspect may be the image selection method according to any one of the first to fourth aspects, in which the one or more unlabeled images are further selected based on an index based on the reliability of the detected object in each of the plurality of unlabeled images.
[0029] This allows for the selection of images to be labeled based on the number of undetected objects and an index based on the reliability of the detected objects. Using such images can effectively improve the machine learning model's detection performance for both undetected objects and objects corresponding to the index. This allows for more effective learning of the machine learning model.
[0030] Also, for example, an image selection method according to a sixth aspect may be the image selection method according to the fifth aspect, wherein the index includes uncertainty indicating the degree of variability in predictions in the undetected object prediction model.
[0031] This allows us to select images to be labeled that can reduce uncertainty, i.e., false positives, and can effectively improve the machine learning model's detection performance for both undetected and false positive objects.
[0032] Furthermore, for example, the image selection method according to the seventh aspect is an image selection method according to any one of the first to sixth aspects, and the number of undetected objects may be the number of objects, out of multiple objects contained in the image, that cannot be detected by the target model to be learned when the image is input into the target model to be learned.
[0033] This allows labels to be added to images that contain objects that the target model cannot detect, thereby enabling more effective target model training.
[0034] Also, for example, the image selection method according to the eighth aspect may be the image selection method according to the seventh aspect, in which the target model is trained using images taken in a first environment, and the plurality of unlabeled images may include images taken in a second environment different from the first environment.
[0035] This effectively improves the detection performance for images captured in the second environment, i.e., images that look different from images captured in the first environment, thereby effectively suppressing performance degradation due to domain shift.
[0036] In addition, a learning method according to a ninth aspect assigns a label to each of the one or more unlabeled images selected using the image selection method according to any one of the first to eighth aspects, and trains a target model to be learned using one or more labeled images that are the one or more unlabeled images to which labels have been assigned.
[0037] This provides the same effect as the image selection method described above.
[0038] An image selection device according to a tenth aspect of the present disclosure includes an estimation unit that estimates a degree of undetection based on the number of undetected objects in each of a plurality of unlabeled images by inputting the plurality of unlabeled images into an undetected object prediction model, and a selection unit that selects one or more unlabeled images to be labeled from the plurality of unlabeled images based on the degree of undetection in each of the plurality of unlabeled images.
[0039] This provides the same effect as the image selection method described above.
[0040] A program according to an eleventh aspect of the present disclosure is a program for causing a computer to execute the image selection method according to any one of the first to eighth aspects.
[0041] This provides the same effect as the image selection method described above.
[0042] These general or specific aspects may be realized as a system, a method, an integrated circuit, a computer program, or a non-transitory recording medium such as a computer-readable CD-ROM, or as any combination of the system, method, integrated circuit, computer program, or recording medium. The program may be pre-stored in the recording medium, or may be supplied to the recording medium via a wide area communication network including the Internet.
[0043] Hereinafter, the embodiments will be specifically described with reference to the drawings.
[0044] The embodiments described below are all comprehensive or specific examples. The numerical values, shapes, components, component placement and connection configurations, steps, and step order shown in the following embodiments are merely examples and are not intended to limit the present disclosure. Furthermore, among the components in the following embodiments, components not described in independent claims are described as optional components.
[0045] Furthermore, each figure is a schematic diagram and is not necessarily an exact illustration. Therefore, for example, the scales of the figures do not necessarily match. Furthermore, in each figure, substantially the same components are given the same reference numerals, and redundant explanations are omitted or simplified.
[0046] Furthermore, in this specification, terms indicating relationships between elements such as "same," as well as numerical values and numerical ranges, are not expressions that only express a strict meaning, but are expressions that also include a substantially equivalent range, for example, a difference of about several percent (or about 10%).
[0047] Furthermore, in this specification, ordinal numbers such as "first" and "second" do not refer to the number or order of components unless otherwise specified, but are used for the purpose of avoiding confusion and distinguishing between components of the same type.
[0048] (Embodiment) Hereinafter, an information processing device according to the present embodiment will be described with reference to FIGS.
[0049] 1. Configuration of Information Processing Apparatus First, the configuration of an information processing apparatus according to this embodiment will be described with reference to Fig. 1. Fig. 1 is a block diagram showing the functional configuration of an information processing apparatus 100 according to this embodiment.
[0050] As shown in FIG. 1 , the information processing device 100 includes, as its functional configuration, a first learning processing unit 10, an image selection unit 20, a label assignment unit 30, a second learning processing unit 40, and a storage unit 50. The information processing device 100 also includes, as its hardware configuration, a non-volatile memory storing a program, a volatile memory serving as a temporary storage area for executing the program, an input / output port, a communication interface, a processor for executing the program, and the like. The memory may be a read-only memory (ROM) or a random access memory (RAM), and can store a program to be executed by the processor. The first learning processing unit 10, the image selection unit 20, the label assignment unit 30, and the second learning processing unit 40 are realized by a processor that executes a program stored in a memory (e.g., the storage unit 50), and the like.
[0051] To achieve active learning that takes undetected objects into consideration, the first learning processing unit 10 executes a process of learning a False Negative Prediction Module (FNPM13 shown in FIG. 3 , which will be described later) that is a machine learning model that estimates the number of undetected objects (the number of false negatives) for an input image. The learning process method and the detailed configuration of the first learning processing unit 10 will be described later with reference to FIG. 4 .
[0052] The image selection unit 20 uses the FNPM 13 trained by the first learning processing unit 10 to execute a process of selecting a small number of target images to be labeled from a plurality of images (unlabeled images). The image selection unit 20 actively samples images containing undetected objects. The plurality of images may include, for example, images captured in a new environment different from images captured in the existing environment used to train the target model. The number of target images to be selected is set in advance and may be, for example, approximately several percent of the total number of unlabeled images. The image selection unit 20 functions as a selection unit.
[0053] The target model is a machine learning model using a neural network such as deep learning (for example, a convolutional neural network (CNN)). For example, Fast R-CNN is exemplified, but other machine learning models such as R-CNN, Faster R-CNN, SSD (Single Shot MultiBox Detector), and YOLO (You Only Look Once) may also be used.
[0054] The image selection unit 20 inputs a plurality of unlabeled images (specifically, feature quantities of a plurality of unlabeled images) into the FNPM 13, estimates a degree of undetection based on the number of undetected objects for each of the plurality of unlabeled images, and selects a target image from the plurality of unlabeled images based on the estimation result. The image selection unit 20 may select, as the target image, from the plurality of unlabeled images, images whose degree of undetection is equal to or greater than a first predetermined degree, or that is among the top two predetermined numbers of images. Note that the number of target images may be one or more. The first predetermined degree and the second predetermined number are, for example, set in advance and stored in the storage unit 50. In this way, the image selection unit 20 functions as an estimation unit that estimates the degree of undetection using the FNPM 13.
[0055] The image selection unit 20 may select a target image based on an index based on the reliability of the detected object in each of the plurality of unlabeled images, in addition to the degree of non-detection. The index based on the reliability includes, for example, uncertainty. For example, the image selection unit 20 may select a target image from the plurality of unlabeled images based on the degree of non-detection and the uncertainty. In other words, the image selection unit 20 may select a target image from the plurality of unlabeled images taking into account undetected objects and falsely detected objects. It is expected that uncertainty will be high in images of a novel environment even if an object is detected. By selecting images with high uncertainty, images of a novel environment can be effectively selected.
[0056] Note that "non-detection" includes not detecting a target object in a machine learning model (missed detection). False positives include misclassifying an object, misaligning the detected position of an object in an image, and detecting the background (i.e., an area where the target object does not exist) as the target. Furthermore, "uncertainty" refers to the degree of variability in predictions made in deep learning. For example, the smaller the difference in probability distribution between classes, the greater the variability in predictions, and the greater the uncertainty. Note that reliability-based metrics are not limited to uncertainty.
[0057] The labeling unit 30 performs a process of assigning a label to a target image selected by the image selecting unit 20 from among the plurality of unlabeled images. The labeling unit 30 assigns a label only to the target image from among the plurality of unlabeled images, for example. In other words, the labeling unit 30 prohibits the assignment of labels to unlabeled images other than the target image from among the plurality of unlabeled images.
[0058] The labeling unit 30 may receive label information input from a user and assign a label to a target image. In this case, the labeling unit 30 is connected to a reception unit that receives input from the user and is configured to be able to acquire the input received by the reception unit. The labeling unit 30 may be configured to include, for example, a communication interface that communicates with the reception unit. The reception unit may be, for example, but is not limited to, a button, a keyboard, a touch panel, a microphone, or the like. Furthermore, the labeling unit 30 may be configured to automatically assign labels using an image segmentation model that classifies each pixel into a class.
[0059] The second learning processing unit 40 executes a process of learning an object model using at least the target images (labeled images) to which labels have been assigned by the label assignment unit 30. In the present embodiment, the second learning processing unit 40 executes a process of learning an object model using the target images (labeled images) and images (unlabeled images) that have not been selected by the image selection unit 20 from among the multiple images. Note that learning by the second learning processing unit 40 also includes relearning.
[0060] The storage unit 50 is a storage device that stores various information for training the target model. The storage unit 50 stores the FNPM 13, a plurality of images, etc. The storage unit 50 is realized by, for example, a semiconductor memory or a hard disk drive (HDD), but is not limited to these.
[0061] In the information processing device 100, for example, the first learning processing unit 10, the image selection unit 20, and the second learning processing unit 40 may each be realized as a separate device. For example, the first learning processing unit 10 may be realized as a learning device that executes a learning process for learning the FNPM 13. The image selection unit 20 may be realized as an image selection device that selects target images to be labeled from multiple images using the FNPM 13. The second learning processing unit 40 may be realized as a learning device that executes a process for learning (e.g., relearning) a target model using ADA technology corresponding to the object detection model.
[0062] 2. Operation of Information Processing Device Next, the operation of the information processing device 100 configured as described above will be described with reference to Figs. 2 to 4. Fig. 2 is a flowchart showing the operation (image selection method and learning method) of the information processing device 100 according to this embodiment. Note that at the time of step S10 shown in Fig. 2, the feature extraction unit 11 has completed learning, and the learning process step of the feature extraction unit 11 is omitted in Fig. 2.
[0063] The information processing device 100 mainly processes (i) an existing environment image D S (See FIG. 3) and new environment image D T (Unlabeled image D UT and labeled image D LT ) (see FIG. 3 ), an initial model is trained by an unsupervised domain adaptation technique, and (ii) an unlabeled image D UT The budgeted amount is sampled from the image D, and labels are added to the image D. LT and (iii) the labeled existing environment image D S and labeled image D LT and unlabeled image D UT (ii) and (iii) are repeated multiple times. (i) corresponds to step S10, (ii) corresponds to steps S20 to S50, and (iii) corresponds to step S60.
[0064] As shown in FIG. 2, the second learning processing unit 40 generates an existing environmental image D S and new environment image D T The second learning processing unit 40 performs a learning process on the teacher model 210 and the student model 220 (see FIG. 3).
[0065] FIG. 3 is a diagram for explaining the learning process of the target model according to this embodiment. FIG. 3 shows the overall configuration for executing the learning process of the target model, which has a Teacher-Student structure. The teacher model 210 ("Teacher" shown in FIG. 3) and the student model 220 ("Student" shown in FIG. 3) have the same model configuration but different internal parameters, with the teacher model 210 having higher performance than the student model 220. The student model 220 is an example of a target model.
[0066] The teacher model 210 has a feature extraction unit 211, a domain identification unit 213, and a detection and estimation unit 214, and the student model 220 has a feature extraction unit 221, a domain identification unit 223, and a detection and estimation unit 224. In addition, the student model 220 has GRLs (Gradient Reversal Layers) between the feature extraction unit 221 and the domain identification unit 223.
[0067] In this embodiment, data augmentation is performed to improve the generalization performance of the machine learning model. The first learning processing unit 10 includes, for example, a first data augmentation unit 230 ("Week Aug." shown in FIG. 3 ) and a first data augmentation unit 230 augments the input unlabeled image D UT The first degree of data augmentation (weak data augmentation) is performed on the unlabeled image D (shown in FIG. 3 as “Unlabeled Target Data”). UT to the feature extraction unit 211. The first degree of data augmentation processing includes, for example, performing at least one of inversion, horizontal movement, and vertical movement on the object shown in the image. The second learning processing unit 40 also has, for example, a second data augmentation unit 240 ("Strong Aug." shown in FIG. 3 ), which performs a second degree of data augmentation processing (strong data augmentation processing) on the input image, which is higher than the first degree, and outputs the image on which the data augmentation processing has been performed to the feature extraction unit 221. The second degree of data augmentation processing is, for example, data augmentation processing using reinforcement learning.
[0068] The bold-line frame shown in FIG. 3 shows a plurality of unlabeled images D UT A target image to which a label is to be assigned is selected from the list, and a labeled image D is generated by assigning a label to the selected target image.LT 3A and 3B show a schematic diagram of a process for generating the Labeled Target Data (shown in FIG. 3).
[0069] The feature extraction unit 211 of the teacher model 210 extracts the unlabeled image D UT When the label is input, the unlabeled image D UT The feature extraction unit 221 has the same function as the feature extraction unit 211, and outputs a feature map 212 of the unlabeled image D UT When the label is input, the unlabeled image D UT The feature extraction units 211 and 221 have the same functions as the feature extraction unit 11.
[0070] The domain identification unit 213 outputs an identification result that identifies an existing environment or a new environment for each element of the feature map 212 based on the feature map 212. The identification result has a size of one channel and includes, for example, a real value for each pixel. The real value is, for example, a real value between "0" and "1". The domain identification unit 223 has the same function as the domain identification unit 213.
[0071] The detection and estimation unit 214 performs object detection on the image and outputs the detection result. The detection and estimation unit 214 outputs the class and position of the target object in the image. Note that the detection and estimation unit 224 has the same function as the detection and estimation unit 214.
[0072] As shown in FIG. 3, the second learning processing unit 40 first acquires an existing environmental image D S and new environment image D T Using unsupervised domain adaptation, the initial parameters common to the teacher model 210 and the student model 220 are initialized. Active sampling is performed based on values obtained by evaluating data in a new environment using an acquisition function. If the model used as the acquisition function does not have knowledge of the data in the new environment, the acquisition function cannot perform appropriate evaluation. Therefore, the second learning processing unit 40 first learns a model adapted to the new environment using unsupervised domain adaptation. Specifically, domain alignment at the feature level is performed using adversarial learning by a gradient inversion layer and a domain identification unit 223.
[0073] When the parameter of the student model 220 is θs and the parameter of the domain identification unit 223 is φs, the target loss function in model initialization ("Adversarial Loss" shown in FIG. 3) is expressed by the following equation 1.
[0074]
[0075] where λ is L adv It is a hyperparameter that controls the weighting of the adversarial loss.
[0076]
[0077] is the supervised learning loss in the existing environment, and is expressed by the following Equation 2.
[0078]
[0079] Here, (x', y') is the image expanded by the second data expansion unit 240. Here, the superscript s indicates the existing environment, and the subscript i indicates the ith image. det is the loss of the student model 220 ("Detection Loss" shown in FIG. 3), and is expressed by the following Equation 3:
[0080]
[0081] where:
[0082]
[0083] is the true value y i is the number of bounding boxes contained in
[0084]
[0085] indicates the classification loss in the Region Proposal Network (RPN). The RPN is located between the feature extraction unit 211 and the detection and estimation unit 214 shown in FIG. 3, and receives the feature map 212 as an input.
[0086]
[0087] denotes the bounding box regression loss.
[0088]
[0089] denotes the classification loss in the detection and estimation unit 224 (ROI Head).
[0090]
[0091] is the bounding box regression loss in the detection and estimation unit 224. i denotes the i-th image, and b i,j denotes the j-th bounding box coordinate in the i-th image, and c i,j denotes the class index of the j-th bounding box in the i-th image. In this case, the adversarial loss is expressed as Equation 4 below.
[0092]
[0093] Here, F enc represents the feature extraction unit 221 of the student model 220, and D represents the domain identification unit 223 of the student model 220. The second learning processing unit 40 trains the domain identification unit 223 so as to identify the existing environment as 1 and the new environment as 0. Then, the second learning processing unit 40 extracts the existing environment image D S and new environment image D T After completing the learning process for unsupervised domain adaptation for all data, the parameters of the student model 220 are copied (θt←θs, φt←φs) to the parameters (θt, φt) of the teacher model 210. This causes the information processing device 100 to proceed to the ADA step.
[0094] 2 again, the first learning processing unit 10 executes a learning process for the undetected object prediction model (S20). The first learning processing unit 10 executes the learning process for the FNPM 13 as the undetected object prediction model. The images used in the learning process in step S10 are annotated images, that is, images including labels (annotation information).
[0095] FIG. 4 is a diagram for explaining the learning process of the FNPM 13 according to this embodiment.
[0096] As shown in Fig. 4, the first learning processing unit 10 has a feature extraction unit 11, a detection / estimation unit 14, an undetected degree calculation unit 15, and a loss calculation unit 16. In Fig. 4, the feature extraction unit 11 is referred to as "Backbone," the detection / estimation unit 14 is referred to as "ROI (Region of Interest) Head," the undetected degree calculation unit 15 is referred to as "False Negative Calculation," and the loss calculation unit 16 is referred to as "False Negative Prediction Loss."
[0097] The feature extraction unit 11 constitutes a pre-stage of the FNPM 13 and is configured to receive an image and output a feature map 12 (intermediate feature values) of the image. The feature extraction unit 11 is configured, for example, by a convolutional neural network and has a function of extracting feature values from the input image. The feature extraction unit 11 may be, for example, a visual geometry group (VGG) 16 pre-trained using an image database such as Image.NET, but is not limited to this. The feature extraction unit 11 may also be trained to extract domain-invariant feature values. The feature map 12 is output to both the FNPM 13 and the detection / estimation unit 14.
[0098] The feature extraction unit 11 may be shared with a feature extraction unit included in the target model (for example, the feature extraction unit 221 shown in FIG. 3 ), which has the advantage of enabling undetected objects to be estimated with an extremely small number of parameters.
[0099] Next, the first learning processing unit 10 inputs the feature map 12 output from the feature extraction unit 11 into the FNPM 13 and obtains the degree of undetection of the image corresponding to the feature map 12 as the output (estimated result) of the FNPM 13. The FNPM 13 is a machine learning model that receives the feature values of an input image and outputs a degree of undetection based on the number of undetected objects contained in the input image, for example, the number of undetected objects when the input image is input to a predetermined machine learning model or target model. The FNPM 13 is an example of an undetected object prediction model, and the output of the FNPM 13 here is an example of a first degree of undetection. While such an FNPM 13 can be realized using any network, the configuration shown in FIG. 4 will be described as an example.
[0100] The FNPM 13 is configured to be able to output a predicted value of the degree of non-detection (e.g., a predicted value of the number of undetected objects) through a global average pooling layer ("GAP" shown in FIG. 4) and a fully connected layer ("Sigmoid" shown in FIG. 4). The global average pooling layer is a layer for downsampling the input feature map 12. The fully connected layer is a layer for scaling the output of the FNPM 13 to a value between 0 and 1 using a Sigmoid function or the like. The value obtained by scaling the number of undetected objects to a value between 0 and 1 is an example of the degree of non-detection.
[0101] The FNPM 13 alternately includes a linear layer ("Linear" in FIG. 4) that multiplies an input value by a weight and outputs a value obtained by adding a bias, and an activation layer ("ReLU" in FIG. 4) that converts the output value using an activation function. Note that the activation function used in the activation layer is exemplified by ReLU, but is not limited to this, and may be a step function, a sigmoid function, or the like.
[0102] Note that the FNPM 13 is not limited to the above configuration, and can be configured using any network that receives image features indicated by the feature map 12 as input and outputs a scalar real value (an example of the degree of undetection) that normalizes the number of undetected objects.
[0103] The detection and estimation unit 14 receives the feature map 12 output from the feature extraction unit 11 and estimates the object class and object position for a region determined to be likely to be an object using RoI pooling (region of interest pooling). For example, when the feature map 12 is input to a predetermined machine learning model or target model, the detection and estimation unit 14 estimates the object class and object position that the model can detect.
[0104] The undetection degree calculation unit 15 calculates the number of undetected objects in the image based on the output of the detection / estimation unit 14 and the image label (correct answer information) corresponding to the feature map 12. The undetection degree calculation unit 15 identifies undetected objects and falsely detected objects based on the output of the detection / estimation unit 14 and the image correct answer information, and outputs only the number of undetected objects among the undetected and falsely detected objects. The undetection degree calculation unit 15 then calculates a degree of undetection between 0 and 1 by dividing the calculated number of undetected objects by a predetermined value, and outputs the calculated degree of undetection as correct answer information. The degree of undetection is a value based on the number of objects that would be undetected if the image were input to the target model, and can be used as correct answer information for the output of the FNPM 13 during training of the FNPM 13. The output of the undetection degree calculation unit 15 is an example of a second degree of undetection.
[0105] The loss calculation unit 16 calculates an error, which is the magnitude of deviation between the predicted value ("Prediction" shown in FIG. 4) that is the output of the FNPM 13 and the true value ("Ground Truth" shown in FIG. 4) that is the output of the undetection degree calculation unit 15, and adjusts the parameters of the FNPM 13 using the error as a loss. For example, a loss function that calculates the loss is expressed by the following equation 5.
[0106]
[0107] where G and ψ represent FNPM13 and its parameters, and F head represents the head (ROI Head) of the detection model (e.g., target model). FN(...,...) is a function that calculates the number of undetected objects for the detection result. head(x; θ) is compared with the true value y, and bounding boxes of true values to which no detection result having an IoU (Intersection over Union) of the same class equal to or greater than a threshold is assigned are counted as undetected, and the number of such undetected bounding boxes is calculated.
[0108] Because the true value of the number of undetected objects is calculated based on the detection model, the true value also changes when the detection model is updated, resulting in a dual optimization problem, making it difficult to stably converge FNPM13. In response to this, inspired by reinforcement learning methods, the detection model and FNPM13 are alternately optimized. Specifically, FNPM13 is used only for active sampling, and is not updated during detection model training (see Figure 3). Before active sampling, the parameters of the detection model are fixed, and only FNPM13 is updated. This makes it possible to simply and stably optimize both the detection model and FNPM13.
[0109] The method for adjusting the parameters of the FNPM 13 is not particularly limited, and known methods such as back-propagation may be used.
[0110] 2 again, next, the image selection unit 20 predicts the degree of non-detection of the unlabeled image using the FNPM 13 generated by the first learning processing unit 10 (S30). The image selection unit 20 inputs the feature map 212 output by the feature extraction unit 211 to the FNPM 13, thereby detecting the degree of non-detection of the unlabeled image D. UT Obtain a predicted result of the degree of non-detection in
[0111] The acquisition unit 250 included in the image selection unit 20 acquires output from the FNPM 13 using an acquisition function ("Acquisition Function" shown in FIG. 3). The acquisition unit 250 may further acquire outputs from the domain identification unit 213 and the uncertainty estimation unit 260 ("Uncertainty Estimation" shown in FIG. 3). The acquisition unit 250 may be configured to include a communication interface.
[0112] Next, the image selection unit 20 selects the plurality of unlabeled images D acquired by the acquisition unit 250. UTBased on the degree of non-detection, a plurality of unlabeled images D UT The image selection unit 20 selects images to be labeled (target images) from the images D10 and D20, and selects images having a degree of undetection equal to or higher than a first predetermined degree, or a second predetermined number of images having a high degree of undetection, as images to be labeled. UT Among these, images with a degree of undetection equal to or greater than a first predetermined degree, or a second predetermined number of images with a high degree of undetection may be selected as images to be labeled.
[0113] In the example of FIG. 3, the acquisition unit 250 acquires the unlabeled image D UT The image selection unit 20 acquires information indicating the uncertainty for the unlabeled image D from the uncertainty estimation unit 260. UT The image selection unit 20 may select images to be labeled based on information indicating uncertainty in addition to the degree of non-detection corresponding to the image. UT For each of the unlabeled images, a score is calculated based on a first score based on the degree of non-detection and a second score based on the uncertainty. UT The image to be labeled may be selected based on each score. When the degree of non-detection is high, the first score has a large value, and when the uncertainty is high, the second score has a large value, the image selection unit 20 selects a plurality of unlabeled images D UT , images having a score equal to or higher than the first score, or a third predetermined number of images having a higher score may be selected as images to be labeled.
[0114] The method for selecting images to be labeled is not limited to the above, and examples using other indices will be described below.
[0115] In object detection, the uncertainty of bounding box localization is also important, but unlike the entropy of class probability, it may be difficult to calculate the uncertainty from normal estimated coordinates. Therefore, in this embodiment, the parameters of the student model 220 are captured as a probability distribution by variational inference using Monte Carlo Dropout (MCDropout), and the fluctuations in estimated coordinates due to model fluctuations may be quantified and used as an index of uncertainty.
[0116] The detection and estimation unit 214 of the teacher model 210 and the detection and estimation unit 224 of the student model 220 each have an MCDropout layer, and the predicted coordinates and predicted class probabilities of the student model 220 are respectively expressed as
[0117]
[0118]
[0119] Then, the predicted coordinates and predicted class probabilities of the teacher model 210 are calculated by the following equation 6.
[0120]
[0121] where Ber(η) is the Bernoulli distribution of the dropout rate η.
[0122] By performing variational inference multiple times, multiple prediction results can be obtained. In this embodiment, the second learning processing unit 40 calculates the average and variance from the prediction results performed multiple times, and sets these averages as prediction results that take into account fluctuations in the model.
[0123]
[0124] and the variance of the estimated coordinates as the uncertainty of localization.
[0125]
[0126] and quantify the uncertainty.
[0127]
[0128]
[0129] M indicates the number of variational inferences.
[0130] Here, the image selection unit 20 may calculate a score by combining three indices in addition to the degree of non-detection (Active Sampling Strategy). The diversity, entropy, and position uncertainty shown below are examples of indices based on reliability.
[0131] The degree of non-detection (False Negative) is a score that estimates the degree of detection failure of the machine learning model for the input image, and is calculated using the following Equation 9.
[0132]
[0133] Diversity is a score based on the idea that a high density of distribution of a new environment is more important. The image selection unit 20 calculates the diversity of each image using the following equation 10.
[0134]
[0135] Entropy is a score that estimates the uncertainty in class probability. The higher the entropy, the more difficult the image is for the model to predict, and therefore the more useful it is for learning. The image selection unit 20 calculates the entropy of each image using the following equation 11. Equation 11 calculates the uncertainty in class prediction.
[0136]
[0137] Localization uncertainty is a score that estimates the uncertainty of the position of a bounding box in object detection. The fluctuation of the coordinates of the bounding box estimated by variational inference is quantified and defined as the localization uncertainty. The image selection unit 20 calculates the localization uncertainty of each image using the following equation 12.
[0138]
[0139] The final metrics are calculated for each image using these four metrics (undetected, diversity, entropy, and position uncertainty), but since each metric has a different range of possible values, there is a possibility that an metric with a larger value will dominate. Therefore, the image selection unit 20 normalizes each metric based on the following equation 13.
[0140]
[0141] Here, m∈{fn, div, ent, loc}, and μ and σ represent the mean and standard deviation, respectively. Although preliminary evaluation has revealed that each index follows a unimodal normal distribution, the presence of outliers makes it difficult to accurately normalize using the maximum and minimum values. Therefore, the image selection unit 20 reduces the influence of outliers by calculating the mean and standard deviation and scaling by 6σ based on the normal distribution. The image selection unit 20 calculates the product of the scores, as shown in the following equation 14, to calculate the final score for the image.
[0142]
[0143] The image selection unit 20 may select images to which labels are assigned based on the final score for each image. Note that the image selection unit 20 may calculate the score of each image based on the degree of undetection and at least one index of diversity, entropy, and positional uncertainty.
[0144] Next, the labeling unit 30 performs labeling processing on the selected image (S50). The labeling unit 30, for example, presents the selected image to the user U, acquires annotation information input from the user U accepted by the accepting unit, and generates an unlabeled image D UT By adding the acquired annotation information (adding a label) to the image, the labeled image D LT The labeling unit 30 generates an unlabeled image D that was not selected by the image selection unit 20. UT In this way, the image selection unit 20 and the label assignment unit 30 perform active sampling.
[0145] Next, the second learning processing unit 40 generates the labeled image D LT In this embodiment, the second learning processing unit 40 further performs a learning process on the target learning model (the target model, which is the student model 220 shown in FIG. 3) using the unlabeled image D UT and existing environment image D S and execute a learning process for the target model. S For example, the image set includes images captured in an existing environment that was used in a past learning process for the target model.
[0146] Here, in semi-supervised domain adaptation, learning is performed by adding images of a new environment to which labels are assigned by a budget (for example, a preset number) through active sampling. LT Not only unlabeled image D UT The second learning processing unit 40 executes the learning process in a semi-supervised learning framework that also utilizes the existing environmental image D S and the labeled image D of the new environment LT Supervised learning is performed using
[0147] The second learning processing unit 40 also generates an unlabeled image D UT The second learning processing unit 40 performs unsupervised learning by assigning pseudo labels 270 (see FIG. 3 ) to the unlabeled images D , which are assigned pseudo labels 270 by using the teacher model 210, taking into account the uncertainty of the model. The pseudo labels 270 are labels based on the classification results of the teacher model 210. UT Specifically, the second learning processing unit 40 utilizes the uncertainty of the model to exclude images with highly uncertain pseudo labels 270. The second learning processing unit 40 uses the unlabeled images D UTAmong these, images with uncertainty equal to or greater than a predetermined value are excluded from the images used to train student model 220. This makes student model 220 less susceptible to errors in pseudo label 270 during training. Second learning processing unit 40 calculates various values using the following equations 15 to 20.
[0148] The target loss function is expressed as Equation 15 below.
[0149]
[0150]
[0151] is the labeled image D in the new environment. LT is the supervised learning loss in , and is calculated in the same way as in Equation 3.
[0152]
[0153] is the unlabeled image D in the new environment. UT is the unsupervised detection loss in the image ("Unsupervised Loss" shown in FIG. 3), and is expressed by the following Equation 16:
[0154]
[0155] where:
[0156]
[0157] is an exponential function and is expressed by the following equation 17.
[0158]
[0159] Here, γ is a threshold for using pseudo labels 270 whose variance is equal to or less than a certain value.
[0160]
[0161] is an exponential function and is expressed by the following equation 18.
[0162]
[0163] Here, τ is a threshold for using the pseudo label 270 when the maximum value of the class probability is equal to or greater than a certain value. The pseudo label 270 is defined by the following equation 19.
[0164]
[0165] The second learning processing unit 40 learns the student model 220 and updates the teacher model 210 using the Exponential Moving Average (EMA). The update of the teacher model 210 is expressed by the following equation 20.
[0166] θt←αθt+(1-α)θs, φt←αφt+(1-α)φt...Formula 20
[0167] Here, α represents the update rate.
[0168] As described above, the second learning processing unit 40 repeatedly executes the learning of the student model 220 and the updating of the teacher model 210 .
[0169] (Other Embodiments) While the image selection device and the like according to one or more aspects have been described above based on the embodiments, the present disclosure is not limited to these embodiments. As long as they do not deviate from the spirit of the present disclosure, various modifications conceivable by a person skilled in the art to the present embodiments and embodiments constructed by combining components of different embodiments may also be included in the present disclosure.
[0170] For example, the degree of undetection in the above embodiment may be the number of undetected objects itself.
[0171] Each component may be configured with dedicated hardware, or may be realized by executing a software program suitable for that component. Each component may be realized by a program execution unit such as a CPU or processor reading and executing a software program recorded on a recording medium such as a hard disk or semiconductor memory.
[0172] The order in which the steps in the flowchart are executed is merely an example for specifically explaining the present disclosure, and other orders may be used. Some of the steps may be executed simultaneously (in parallel) with other steps, or some of the steps may not be executed.
[0173] The division of functional blocks in the block diagram is an example, and multiple functional blocks may be realized as a single functional block, one functional block may be divided into multiple blocks, or some functions may be moved to another functional block.Furthermore, the functions of multiple functional blocks having similar functions may be processed in parallel or in time-sharing by a single piece of hardware or software.
[0174] Furthermore, the information processing device (e.g., image selection device) according to the above-described embodiments may be realized as a single device or may be realized by multiple devices. When the information processing device is realized by multiple devices, the components of the information processing device may be distributed in any manner among the multiple devices. When the information processing device is realized by multiple devices, the communication method between the multiple devices is not particularly limited, and may be wireless communication or wired communication. Furthermore, wireless communication and wired communication may be combined between the devices. The same applies to the image selection device and each learning device.
[0175] Furthermore, each component described in the above embodiments may be implemented as software or, typically, as an LSI, which is an integrated circuit. These components may be individually integrated into a single chip, or some or all of them may be integrated into a single chip. While LSI is used here, it may also be referred to as an IC, system LSI, super LSI, or ultra LSI depending on the level of integration. Furthermore, the integrated circuit implementation is not limited to LSI, and may be implemented using a dedicated circuit (a general-purpose circuit that executes a dedicated program) or a general-purpose processor. A field programmable gate array (FPGA), which can be programmed after LSI fabrication, or a reconfigurable processor, which allows the connection or settings of circuit cells within an LSI to be reconfigured, may also be used. Furthermore, if an integrated circuit implementation technology that replaces LSI emerges due to advances in semiconductor technology or a derivative technology, that technology may naturally be used to integrate the components.
[0176] A system LSI is an ultra-multifunctional LSI manufactured by integrating multiple processing units on a single chip, and is specifically a computer system comprising a microprocessor, ROM, RAM, etc. The ROM stores computer programs. The system LSI achieves its functions when the microprocessor operates in accordance with the computer programs.
[0177] Furthermore, one aspect of the present disclosure may be a computer program that causes a computer to execute each of the characteristic steps included in the image selection method and learning method shown in FIG. 2 .
[0178] Furthermore, for example, the program may be a program to be executed by a computer. Another aspect of the present disclosure may be a computer-readable non-transitory recording medium on which such a program is recorded. For example, such a program may be recorded on a recording medium and distributed or circulated. For example, the distributed program may be installed in a device having another processor, and the program may be executed by the processor, thereby causing the device to perform each of the above processes.
[0179] The present disclosure is useful for an information processing device or the like that performs learning of an object detection model.
[0180] DESCRIPTION OF SYMBOLS 10 First learning processing unit 11, 211, 221 Feature extraction unit 12, 212, 222 Feature map 13 FNPM (undetected object prediction model) 14, 214, 224 Detection estimation unit 15 Undetection degree calculation unit 16 Loss calculation unit 20 Image selection unit (estimation unit, selection unit) 30 Label assignment unit 40 Second learning processing unit 50 Storage unit 100 Information processing device (image selection device) 210 Teacher model 213, 223 Domain identification unit 220 Student model 230 First data extension unit 240 Second data extension unit 250 Acquisition unit 260 Uncertainty estimation unit 270 Pseudo label D LT Labeled image D S Existing environment image D UTUnlabeled image U User
Claims
1. An image selection method executed by a processor using a memory, comprising: inputting a plurality of unlabeled images into an undetected object prediction model to estimate a degree of undetection based on the number of undetected objects in each of the plurality of unlabeled images; and selecting one or more unlabeled images to which a label is to be assigned from among the plurality of unlabeled images based on the degree of undetection for each of the plurality of unlabeled images.
2. The image selection method according to claim 1, further comprising the steps of: predicting a first degree of undetection for an input image by inputting the input image into the undetected object prediction model; calculating a second degree of undetection for the input image based on correct answer information for an object region of the input image; and training the undetected object prediction model based on the first degree of undetection and the second degree of undetection.
3. The image selection method according to claim 1 or 2, wherein the one or more unlabeled images are images selected from the plurality of unlabeled images based on the degree of non-detection.
4. The image selection method according to claim 3, wherein the one or more unlabeled images are images among the plurality of unlabeled images whose degree of undetection is equal to or greater than a first predetermined degree, or which are among a second predetermined number of images that are in the top ranking.
5. The image selection method according to claim 1 or 2, wherein the one or more unlabeled images are further selected based on an index based on the reliability of the detected object in each of the plurality of unlabeled images.
6. The image selection method according to claim 5, wherein the index includes an uncertainty indicating a degree of variability in predictions in the undetected object prediction model.
7. The image selection method according to claim 1 or 2, wherein the number of undetected objects is the number of objects, out of a plurality of objects contained in the image, that cannot be detected by the target model to be learned when the image is input into the target model.
8. The image selection method of claim 7, wherein the target model is trained using images captured in a first environment, and the plurality of unlabeled images includes images captured in a second environment different from the first environment.
9. A learning method comprising: assigning a label to each of the one or more unlabeled images selected using the image selection method according to claim 1 or 2; and training a target model to be trained using one or more labeled images which are the one or more unlabeled images to which a label has been assigned.
10. An image selection device comprising: an estimation unit that estimates a degree of undetection based on the number of undetected objects in each of a plurality of unlabeled images by inputting the plurality of unlabeled images into an undetected object prediction model; and a selection unit that selects one or more unlabeled images to which a label is to be assigned from among the plurality of unlabeled images based on the degree of undetection in each of the plurality of unlabeled images.
11. A program for causing a computer to execute the image selection method according to claim 1 or 2.