Data collection program, data collection device, and data collection method
The data collection program addresses the high human cost of labeling in machine learning by using data augmentation and controlled data acquisition to automate the labeling process, thereby reducing the need for manual labeling.
Patent Information
- Application Number
- JP2023549254
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-09-24
- Publication Date
- 2025-06-18
- Estimated Expiration
- 2041-09-24
AI Technical Summary
Existing machine learning methods require significant human cost for labeling datasets, especially when using active learning, and self-learning methods are limited in reducing this cost due to the proximity of unlabeled data to labeled data.
A data collection program that executes data augmentation for unlabeled data, assigns a specific label indicating that all labels of the augmented data are the same, and controls the position and orientation of the data acquirer to maximize the likelihood of accurate labeling.
The proposed solution reduces the human cost of labeling for machine learning datasets by automating the labeling process through data augmentation and controlled data acquisition, effectively addressing the limitations of existing methods.
Smart Images

Figure 0007694678000001 
Figure 0007694678000002 
Figure 0007694678000003
Abstract
Description
Technical Field
[0001] The present invention relates to a data collection program, a data collection device, and a data collection method.
Background Art
[0002] In machine learning, supervised learning using labeled data for learning may be applied to product classification problems and the like.
[0003] FIG. 1 is a diagram for explaining an example of assigning correct labels to a data set.
[0004] For a data set representing an image of a car indicated by reference sign A1, a correct label is assigned as shown by reference sign A2. In the example shown in FIG. 1, as the correct labels, a taxi and an Electric Vehicle (EV) are assigned. Then, as shown by reference sign A3, training for a learning model is performed using the labeled data.
[0005] Since the assignment of correct labels to such a data set is usually performed manually, the collection cost of the labeled data is higher than that of the unlabeled data.
[0006] FIG. 2 is a diagram for explaining active learning.
[0007] Active learning may be performed in which unlabeled data is divided into known data (in other words, data that the model being learned can estimate labels with high confidence) and unknown data (in other words, data that the model being learned cannot classify), and labeling is requested for the unknown data.
[0008] As shown by reference sign B1, for unlabeled data of automobile images, prediction is performed using a learning model, and a confidence level is calculated. In the example of calculating the confidence level shown by reference sign B2, the confidence level that the automobile image is a taxi is higher than the confidence level that it is another type of automobile such as an EV vehicle. On the other hand, in the example of calculating the confidence level shown by reference sign B3, the confidence level that the automobile image is a taxi, the confidence level that it is an EV vehicle, and the confidence level that it is another type of automobile are all approximately the same value. Labeling by a person may be required only for such low-confidence data.
[0009] Figure 3 is a diagram for explaining self-learning.
[0010] Self-learning (in other words, label propagation) may be performed to automatically label unlabeled data using the assumption that data close to labeled data has the same label.
[0011] As shown by reference sign C1, for unlabeled data of automobile images, prediction is performed using a learning model, and a confidence level is calculated. In the example of calculating the confidence level shown by reference sign C2, since the confidence level that the automobile image is a taxi is higher than the confidence level that it is another type of automobile such as an EV vehicle, a taxi is assigned as a pseudo ground-truth label. On the other hand, in the example of calculating the confidence level shown by reference sign C3, since the confidence level that the automobile image is an EV vehicle is higher than the confidence level that it is another type of automobile such as a taxi, an EV vehicle is assigned as a pseudo ground-truth label.
Prior Art Documents
Patent Documents
[0012]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0013] Figure 4 is a diagram for explaining the problem of labeling for a dataset.
[0014] As described above, even when using active learning, human cost is required for labeling. Even if self-learning (in other words, label propagation) is used to reduce this human cost, most of the data close to the labeled data is known data for which the model can estimate the label with high confidence, and the effect of reducing the human cost is limited.
[0015] In the example shown by reference sign D1, since the unlabeled data U is close to the labeled data La, it can be automatically labeled by label propagation. On the other hand, in the example shown by reference sign D2, since the unlabeled data U is far from the labeled data Lb, even the path data that requires labeling cannot be automatically labeled.
[0016] In one aspect, it aims to reduce the human cost of labeling for the dataset of the machine learning model.
Means for Solving the Problem
[0017] In one aspect, the data collection program executes data augmentation for unlabeled data, assigns a specific label indicating that all the labels of the augmented data are the same to the group of augmented data generated by the data augmentation, and when the label for any one of the augmented data in the group of augmented data is determined, the same label as the determined label is assigned to the augmented data with the same specific label as the any one of the augmented data And control the position and orientation of the data acquirer that acquires training data so that the possibility of attaching the label or specific label becomes the highest to execute the to a computer processing.
Effect of the Invention
[0018] In one aspect, the human cost of labeling for the dataset of the machine learning model can be reduced.
Brief Description of the Drawings
[0019]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Mode for Carrying Out the Invention
[0020] 〔A〕Embodiment Hereinafter, an embodiment will be described with reference to the drawings. However, the embodiment shown below is merely an example, and there is no intention of excluding various modifications and applications of technologies not explicitly shown in the embodiment. That is, the present embodiment can be variously modified and implemented without departing from its gist. In addition, each figure is not intended to include only the components shown in the figure, and can include other functions and the like.
[0021] Hereinafter, in the drawings, since the same reference numerals denote the same parts, the description thereof will be omitted.
[0022] 〔A-1〕Configuration Example FIG. 5 is a diagram for explaining labeling by data augmentation.
[0023] By creating data close to the target data and reducing the distance between the data, labels are assigned to the dataset. Therefore, it is assumed that data without / with labels is expanded and label propagation is performed.
[0024] In the example shown by symbol E1, data augmentation is performed on a plurality of unlabeled data U close to the labeled data Lb, and they are grouped together for labeling. This enables labeling by label propagation. However, if the data augmentation is too strong, there is a risk of mislabeling by other labeled data La nearby. In particular, it may become unstable in the initial stage of learning.
[0025] FIG. 6 is a diagram for explaining labeling by the control of a data acquirer in the embodiment.
[0026] In the present embodiment, data (in other words, weakly labeled data) having the same label as the unlabeled data is continuously acquired by controlling the movement or focus of a data acquirer such as the camera 161 (described later with reference to FIG. 7).
[0027] In the example shown by symbol E2, a plurality of consecutive unlabeled data close to the labeled data La are acquired to distinguish them from other labeled data Lb.
[0028] FIG. 7 is a block diagram schematically showing an example of the hardware configuration of the data collection device 1 in the embodiment.
[0029] As shown in FIG. 7, the data collection device 1 includes a CPU 11, a memory unit 12, a display control unit 13, a storage device 14, an input Interface (IF) 15, an external recording medium processing unit 16, and a communication IF 17.
[0030] The memory unit 12 is an example of a storage unit, and is exemplified by a Read Only Memory (ROM) and a Random Access Memory (RAM), etc. Programs such as a Basic Input / Output System (BIOS) may be written in the ROM of the memory unit 12. The software program of the memory unit 12 may be appropriately read and executed by the CPU 11. Further, the RAM of the memory unit 12 may be used as a temporary recording memory or a working memory.
[0031] The display control unit 13 is connected to the display device 131 and controls the display device 131. The display device 131 is, for example, a liquid crystal display, an Organic Light-Emitting Diode (OLED) display, a Cathode Ray Tube (CRT), an electronic paper display, etc., and displays various information for an operator or the like. The display device 131 may be combined with an input device, for example, a touch panel.
[0032] The storage device 14 is a high IO performance storage device. For example, a Dynamic Random Access Memory (DRAM), an SSD, a Storage Class Memory (SCM), or an HDD may be used.
[0033] The input IF 15 is connected to input devices such as a mouse 151 and a keyboard 152, and may control the input devices such as the mouse 151 and the keyboard 152. The mouse 151 and the keyboard 152 are examples of input devices, and through these input devices, an operator performs various input operations.
[0034] The external recording medium processing unit 16 is configured such that a recording medium 160 can be attached. The external recording medium processing unit 16 is configured to be able to read information recorded on the recording medium 160 when the recording medium 160 is attached. In this example, the recording medium 160 has portability. For example, the recording medium 160 is a flexible disk, an optical disk, a magnetic disk, a magneto-optical disk, or a semiconductor memory, etc. Further, a camera 161 is connected to the external recording medium processing unit 16, and the external recording medium processing unit 16 may acquire an image captured by the camera 161 and control the position and orientation of the camera 161.
[0035] The communication IF 17 is an interface for enabling communication with an external device.
[0036] The CPU 11 is an example of a processor and is a processing device that performs various controls and calculations. The CPU 11 realizes various functions by executing the Operating System (OS) and programs read into the memory unit 12.
[0037] The device for controlling the operation of the entire data collection device 1 is not limited to the CPU 11. For example, it may be any one of an MPU, DSP, ASIC, PLD, or FPGA. Also, the device for controlling the operation of the entire data collection device 1 may be a combination of two or more types among the CPU, MPU, DSP, ASIC, PLD, and FPGA. Note that MPU is an abbreviation for Micro Processing Unit, DSP is an abbreviation for Digital Signal Processor, and ASIC is an abbreviation for Application Specific Integrated Circuit. Also, PLD is an abbreviation for Programmable Logic Device, and FPGA is an abbreviation for Field Programmable Gate Array.
[0038] FIG. 8 is a block diagram schematically showing an example of the software configuration of the data collection device 1 shown in FIG. 7.
[0039] The CPU 11 of the data collection device 1 shown in FIG. 7 functions as a parameter prediction unit 111, an unlabeled data processing unit 112, a label prediction unit 113, a label detection unit 114, a label learning unit 115, and a parameter learning unit 116.
[0040] When the unlabeled sensor information 141 is acquired from the camera 161, it may be transmitted to the parameter prediction unit 111 and stored in the HDD 140. Note that the HDD 140 is an example of the storage device 14.
[0041] Based on the sensor information without labels 141 from the camera 161 or the sensor information without labels 141 stored in the HDD 140, the parameter prediction unit 111 calculates parameters for controlling the camera 161 so as to increase the likelihood of label detection. The calculated parameters are transmitted to the label - free data processing unit 112 and stored in the HDD 140. The details of the processing in the parameter prediction unit 111 will be described later with reference to FIGS. 11 and 12, etc.
[0042] The parameter learning unit 116 performs learning of the first control parameter prediction model (to be described later with reference to FIG. 14, etc.).
[0043] The label - free data processing unit 112 acquires a plurality of pieces of label - free data Un. The label - free data processing unit 112 assigns a weak label indicating that all the plurality of pieces of label - free data Un match. The label - free data processing unit 112 labels Un∈U using a learning model (in other words, the label detection unit 114) or a model under learning (in other words, the label prediction unit 113). The label - free data processing unit 112 stores the label - free data Un in the HDD 140. The details of the processing in the label - free data processing unit 112 will be described later with reference to FIG. 9, etc.
[0044] During training, the label prediction unit 113 performs data augmentation processing and label propagation processing. During prediction, the label prediction unit 113 uses the first product classification model (to be described later with reference to FIG. 17, etc.). The label prediction unit 113 stores the predicted label in the HDD 140. During prediction, based on the acquired test data, the details of the processing in the label prediction unit 113 will be described later with reference to FIG. 10, etc.
[0045] The label detection unit 114 performs a labeling process on the label - free data Un. The label detection unit 114 stores the success or failure of the labeling in the HDD 140. The details of the processing in the label detection unit 114 will be described later with reference to FIG. 10, etc.
[0046] The label learning unit 115 reads the training dataset from the HDD 140, performs label learning, and stores the learning result in the HDD 140. The label learning unit 115 uses the first product classification model (described later with reference to FIG. 17 etc.) during learning.
[0047] In this way, the data collection device 1 executes data augmentation on the unlabeled data, and assigns a specific label indicating that all the labels of the augmented data match to the group of augmented data generated by the data augmentation. Further, when the label for any one of the augmented data in the group of augmented data is determined, the data collection device 1 assigns the same label as the determined label to the augmented data to which the same specific label (in other words, weak label) as any one of the augmented data is assigned.
[0048] FIG. 9 is a diagram for briefly explaining the labeling process in the embodiment.
[0049] The unlabeled data processing unit 112 acquires a plurality of unlabeled data related to the target data, and assigns a weak label indicating that the labeled data has the same label. The unlabeled data processing unit 112 acquires a plurality of weakly labeled data by performing continuous data augmentation on, for example, a video. As shown by the reference sign F1, it becomes possible to assign a label to the entire unlabeled data by labeling one of the plurality of weak labels.
[0050] FIG. 10 is a diagram for explaining the labeling process using similarity in the embodiment.
[0051] Labeling may be performed using the similarity between the measured data or the output labels. By using two types of prediction paths, a low-confidence learning model and a high-confidence learned model, it is possible to cover the labeling errors of label propagation. By using a high-confidence label detector based on image processing such as a barcode reader, it is possible to supplement the low-confidence labeling of label propagation. In combination with the process shown in FIG. 9, it is possible to address the problem that labeling cannot be performed by label propagation for data for which labeling is required in active learning.
[0052] In the initial stage of learning shown by reference sign G1, the error of the label prediction process that assigns a low-confidence label is corrected by the label detection process that assigns a high-confidence label. On the other hand, in the later stage of learning shown by reference sign G2, the escape from detection in the label detection process can be avoided by the label prediction process that assigns a high-confidence label.
[0053] FIG. 11 is a diagram for explaining the labeling process by the control of the data acquirer in the embodiment.
[0054] The parameter prediction unit 111 improves the efficiency of control by predicting the control result of a data acquirer such as the camera 161. In the initial stage of learning, the control was random, but as learning progresses, the control is changed so as to image the object surface with label information such as a barcode.
[0055] The camera 161 is installed on a robot 162 capable of controlling the position and orientation of the camera 161. The camera 161 can be preferentially changed from the initial orientation to an effective orientation by the parameter prediction process. In the example shown in FIG. 11, the pose #1 shown by reference sign H1 is preferentially selected over the pose #2 shown by reference sign H2, and a pseudo label is detected from the captured video.
[0056] The robot 162 for product classification is, for example, a robot that recognizes products on a line in a factory or the like. Machine learning, particularly deep learning, may be used as an apparatus for identifying products. Deep learning can easily construct a highly accurate identifier by preparing a large amount of training data that pairs an input with a required output and performing supervised learning.
[0057] However, in a factory, the products handled change over time, and it is costly to manually label the training data each time.
[0058] Therefore, in the present embodiment, for unlabeled data, weakly labeled data collection is performed, and for a part of the collected data, labeling using data augmentation and label propagation, or high-confidence label detection is performed to automatically perform labeling. This reduces the human cost related to labeling the training data used in machine learning.
[0059] FIG. 12 is a diagram for explaining a modification example of the labeling process by the control of the data acquirer shown in FIG. 11.
[0060] In the example shown in FIG. 11, as a typical data acquirer, a visual sensor (in other words, the camera 161) is used, but an example using another sensor is also conceivable. In the example shown in FIG. 12, a contact sensor 163 is installed as another sensor. By installing the contact sensor 163, a material classification problem can be considered. The label detection unit 114 adopts a model that has been pre-learned. What was randomly controlled in the initial stage of learning will be controlled at characteristic locations as learning progresses.
[0061] The contact sensor 163 is installed on a robot 162 that can control the position and orientation of the contact sensor 163. The contact sensor 163 can be preferentially changed from the initial orientation to an effective orientation by parameter prediction processing. In the example shown in FIG. 12, pose #2 shown by reference sign I2 is preferentially selected over pose #1 shown by reference sign I1, and a pseudo label is detected from the acquired contact data.
[0062] FIG. 13 is a diagram for explaining an installation example of an object to be acquired with data in an embodiment.
[0063] The training data set (refer to reference sign J2) collected by the robot 162 (refer to reference sign J1) for product classification is an image showing a product flowing on a conveyor. Here, it is assumed that an Augmented Reality (AR) marker (refer to reference sign J3) capable of identifying a product class is attached to the product.
[0064] The "AR marker" here may be one capable of identifying another product class. For example, there are a "logo" in manufacturer classification, a "barcode" for product reading, etc. The AR marker may be a one-dimensional code or a two-dimensional code.
[0065] Also, as hardware, a robot 162 with an RGB-format camera 161 attached to its hand part is used.
[0066] First, a target product a_n1 flows on the conveyor and is automatically installed in front of the robot 162. As an initial posture of the robot 162, a posture in which the camera 161 faces vertically downward from above the conveyor is taken. The conveyor is stopped at a position where the center of the object coincides with the center of the image of the camera 161, and the next process is moved to.
[0067] FIG. 14 is a diagram for explaining an example of use of a first control parameter prediction model in an embodiment.
[0068] The parameter prediction unit 111 determines a camera posture p = (x, y, z, roll, pitch, yaw) for taking a series of images. The parameter prediction unit 111 acquires an image i_n1 (refer to reference sign K2) taken from an initial posture (refer to reference sign K1). Transitable camera postures (for example, p_1, p_2 ··· p_N2) may be prepared in advance.
[0069] The parameter prediction unit 111 adjusts so that the object center is located at the center of the image that can be acquired with a movable camera pose. The number of camera poses to be prepared may be adjusted based on the number of objects and the like.
[0070] The parameter prediction unit 111 inputs the possible camera poses into the first control parameter prediction model (see reference symbol K3) and predicts the presence or absence of label information.
[0071] Then, as shown by reference symbol K4, the parameter prediction unit 111 performs a full search and identifies the pose p_n2 with the highest confidence of being able to acquire a label (c‘1).
[0072] The first control parameter prediction model is a learning device that predicts whether label information can be obtained by taking an image and the control parameters of the data acquisition device as inputs. Acquisition device parameter learning means: Among the above, those used during the learning of the first control parameter prediction model.
[0073] For the first control parameter prediction model, a deep learning device composed of 3 layers of Convolution + 3 layers of Multilayer perceptron (MLP) may be used, or other models may be used.
[0074] The image is input into the Convolution, the extracted feature amount is combined with the camera parameters, and then input into the 3 - layer MLP.
[0075] FIG. 15 is a table illustrating camera parameter candidates when using the first control parameter prediction model shown in FIG. 14.
[0076] In the camera parameter candidates shown in FIG. 15, as shown by reference symbol L1, the pose candidate p_n2 with the highest confidence of being able to acquire a label (c‘1) having a confidence of 0.9 is identified as the pose with the highest confidence.
[0077] FIG. 16 is a diagram for explaining the movement process of the camera pose in the embodiment.
[0078] By changing the posture of the robot 162 and performing imaging with the camera 161 during that time, a plurality of consecutive object images are acquired.
[0079] As shown by reference sign M1, the object center is aligned with the center of the imaging range of the camera 161 in the initial posture.
[0080] As shown by reference sign M2, the camera 161 is controlled to move toward the posture p_n2 estimated by the process described above with reference to FIGS. 14 and 15, and images are always acquired during the movement. During the movement, the camera 161 is always centered on the object.
[0081] As shown by reference sign M3, a plurality of acquired images U_n1 always show the same object a_n1, and weak labels indicating the same object class are assigned.
[0082] FIG. 17 is a diagram for explaining a usage example of the first product classification model in the embodiment.
[0083] For the data U_n1 with weak labels (see reference sign N1), data augmentation and label propagation as shown by reference sign N2 are performed, and labels are estimated. For u_n1 ∈ U_n1, random data augmentation (for example, Gaussian blur, crop, rotation, luminance and chroma conversion) is performed, and u'_n1 is created as shown by reference sign N3.
[0084] As shown by reference sign N4, u'_n1 is input to the first product classification model, and the class label l1_u'_n1 is estimated.
[0085] The same process is performed for all the images in U_n1. If the l1_u'_n1 with the highest confidence exceeds the threshold t1, a pseudo label L1 with low confidence assigned by the first product classification model is used. In the example shown in FIG. 17, the confidence of c3 is high as shown by reference sign N5.
[0086] The first product classification model is a learner that predicts class labels with an image as input. For the first product classification model, ResNet or other models may be used.
[0087] FIG. 18 is a table illustrating estimation results when using the first product classification model shown in FIG. 17.
[0088] As indicated by reference sign O1, in the image u3’_n1, c3 with a confidence level of 0.9, which is the highest, is set as the pseudo label L1 as the label with the highest confidence level.
[0089] FIG. 19 is a diagram for explaining the label detection process in the embodiment.
[0090] For the image u3_n1 (see reference sign P2) that satisfies u_n1∈U_n1 (see reference sign P1), a label detection process such as an AR marker (see reference sign P3) is performed.
[0091] Then, the same process is performed for all the images in U_n1, and the class label with the largest number of detections is set as the high-confidence pseudo label L2 attached by the label detection unit 114.
[0092] FIG. 20 is a table illustrating estimation results when the label detection process shown in FIG. 19 is executed.
[0093] As indicated by reference sign Q1, in the estimation results, the class level c2 with the largest number of detection results as the pseudo label L2 is specified.
[0094] FIG. 21 is a diagram for explaining a training example of the first control parameter prediction model in the embodiment.
[0095] Let F be the acquisition success or failure of whether the pseudo-label L2 is attached. When the pseudo-label L2 is not attached (F = 0), data collection and label prediction are repeatedly performed based on the following procedure. If N2 camera parameters have not been tried yet, the camera 161 is returned to its initial position, the posture of the camera 161 is controlled and shooting is performed, and pseudo-labeling is redone. If N2 camera parameters have been tried, the process proceeds to the branch process when the pseudo-label is attached.
[0096] Even when the pseudo-label is attached (F = 1), data collection and label prediction are repeatedly performed based on the following procedure. If data collection processing has been performed on N1 objects, the entire process ends. Otherwise, a new object a_(n1 + 1) is installed, and the search for camera parameters resumes.
[0097] Regardless of whether the pseudo-label is attached, the first data set (i_n1, p_n2, F) may be added as training data. The label learning unit 115 uses the first data set to train the model. The training of the model may be performed at any timing. It may be performed every time the data set reaches a specified number (for example, 100 collections). The accuracy gradually improves as data is collected.
[0098] The first data set is used during the learning of the first control parameter prediction model and when collecting the data set.
[0099] In the example shown in FIG. 21, as indicated by the reference numeral R1, the image data i_n1, the camera parameter p_n2, and the image acquisition success or failure F are stored in a storage device such as the HDD 140 and used for training.
[0100] When an image is input as shown by symbol R2, 3-layer Convolution is performed as shown by symbol R3. Then, based on the result of the 3-layer Convolution and the camera parameters, as shown by symbol R4, a 3-layer MLP is performed. Although an error occurs between the prediction confidence and the teaching signal, this error is fed back to the 3-layer Convolution and the 3-layer MLP. As shown by symbol R5, the higher the number of processed data, the better the system performance and the smaller the error.
[0101] FIG. 22 is a diagram for explaining a training example of the first product classification model in the embodiment.
[0102] When the L2 label is attached, set L = L2, and in other cases, set L = L1. Then, the data u_al estimated by active learning and the label L in the second data set are added as training data. The label learning unit 115 uses the second data set to train the model. The training of the model may be performed at any timing. It may be performed every time the data set reaches a specified number (for example, 100 collections). The accuracy gradually improves as the data is collected.
[0103] In the example shown in FIG. 22, as shown by symbol T1, u_al as data without a label U is stored in a storage device such as the HDD 140, and since the L2 label is given as c2, the pseudo label L = c2 is stored in the storage device such as the HDD 140 and used for training. When an image is input as shown by symbol T2, ResNet is performed as shown by symbol T3. Although an error occurs between the prediction confidence and the teaching signal, this error is fed back to ResNet. As shown by symbol T4, the higher the number of processed data, the better the system performance and the smaller the error.
[0104] 〔A-2〕Operation The training process of the machine learning model in the embodiment will be described according to the flowchart shown in FIG. 23 (steps S1 to S16, S21 to S27, S31 to S37).
[0105] The object for data acquisition is installed (step S1).
[0106] The object for data acquisition is photographed in the initial posture (step S2).
[0107] Camera parameters are selected (step S3).
[0108] Prediction candidates are calculated using the first control parameter prediction model (step S4).
[0109] It is determined whether a label can be obtained (step S5).
[0110] If a label cannot be obtained (refer to the NO route of step S5), the process returns to step S3.
[0111] On the other hand, if a label can be obtained (refer to the YES route of step S5), the posture of the camera 161 is moved and photographed (step S6).
[0112] Label information is detected (step S7).
[0113] In parallel with the processing of steps S6 and S7, the following processing of steps S8 and S9 is performed.
[0114] Data expansion processing of the label is performed (step S8).
[0115] Prediction candidates for the label are calculated (step S9) Whether the acquisition of label L2 is successful or not is added to the first dataset (step S10).
[0116] It is determined whether the label acquisition is successful or not (step S11).
[0117] If the label cannot be obtained (refer to the NO route of step S11), it is determined whether a certain number of camera parameters have been tried (step S12).
[0118] If a certain number of camera parameters have not been tried (refer to the NO route in step S12), the process returns to step S3.
[0119] On the other hand, if a certain number of camera parameters have been tried (refer to the YES route in step S12), the process proceeds to step S16.
[0120] If a label is obtained in step S11 (refer to the YES route in step S11), additional data is selected by active learning (AL) (step S13).
[0121] Labeling is performed (step S14).
[0122] The assigned label is added to the second dataset (step S15).
[0123] It is determined whether the processing has been completed for all data acquisition target objects (step S16).
[0124] If there is an object for which the processing has not been completed among the data acquisition target objects (refer to the NO route in step S16), the process returns to step S1.
[0125] On the other hand, if the processing has been completed for all data acquisition target objects (refer to the YES route in step S16), the training process of the machine learning model ends.
[0126] In parallel with the processing in steps S1 to S16, the processing in the following steps S21 to S27 is executed.
[0127] The first control parameter prediction model is initialized (step S21).
[0128] The first dataset is read (step S22).
[0129] Prediction candidates are calculated (step S23).
[0130] The error between the prediction confidence and the teaching signal is calculated (step S24).
[0131] The error is fed back (step S25).
[0132] It is determined whether the processing of the first data set is completed (step S26).
[0133] If the processing of the first data set is not completed (refer to the NO route of step S26), the processing returns to step S22.
[0134] On the other hand, if the processing of the first data set is completed (refer to the YES route of step S26), the parameters are saved (step S27). Then, the training process of the machine learning model ends.
[0135] In parallel with the processing in steps S1 to S16, the processing in the following steps S31 to S37 is executed.
[0136] The first product classification model is initialized (step S31).
[0137] The second data set is read (step S32).
[0138] Prediction candidates are calculated (step S33).
[0139] The error between the prediction confidence and the teaching signal is calculated (step S34).
[0140] The error is fed back (step S35).
[0141] It is determined whether the processing of the second data set is completed (step S36).
[0142] If the processing of the second data set is not completed (refer to the NO route of step S36), the processing returns to step S32.
[0143] On the other hand, when the processing of the second data set is completed (refer to the YES route in step S36), the parameters are saved (step S37). Then, the training process of the machine learning model ends.
[0144] Next, the prediction process of the test data in the embodiment will be described according to the flowchart shown in FIG. 24 (steps S41 to S43).
[0145] The learning result is read (step S41).
[0146] The test data is read (step S42).
[0147] Using the learning model, prediction candidates are calculated (step S43). Then, the prediction process of the test data ends.
[0148] 〔B〕Effect According to the data collection program, the data collection device 1, and the data collection method in the above-described embodiment, for example, the following operational effects can be achieved.
[0149] The data collection program executes data augmentation on the unlabeled data, and assigns a specific label indicating that all the labels of the augmented data generated by the data augmentation match to the group of augmented data. Then, when the label for any one of the augmented data in the group of augmented data is determined, the same label as the determined label is assigned to the augmented data with the same specific label.
[0150] Thereby, the manual cost of labeling the data set of the machine learning model can be reduced.
[0151] For data that could not be automatically labeled by conventional methods, correct labels are assigned. By using weak labels, which means that all the labels of the extended data match, the labels assigned to high-confidence data can be treated as the labels for the entire unlabeled data. During automation, the time required for labeling is made more efficient. By including data acquisition device control parameters as learning targets even when acquiring weakly labeled data, the time required for labeling (in other words, the time required for overall learning) can be shortened.
[0152] 〔C〕Others The disclosed technology is not limited to the above-described embodiments, and can be implemented with various modifications without departing from the spirit of the present embodiment. Each configuration and each process of the present embodiment can be selectively adopted as necessary, or can be appropriately combined.
Explanation of Signs
[0153] 1: Data collection device 11: CPU 12: Memory unit 13: Display control unit 14: Storage device 16: External recording medium processing unit 111: Parameter prediction unit 112: Unlabeled data processing unit 113: Label prediction unit 114: Label detection unit 115: Label learning unit 116: Parameter learning unit 131: Display device 140: HDD 141: Unlabeled sensor information 151: Mouse 152: Keyboard 160: Recording medium 161: Camera 162: Robot 163: Contact sensor 15: Input IF 17: Communication IF
Claims
1. Execute data augmentation on the unlabeled data, Assign a specific label indicating that all the labels of the augmented data match to the group of augmented data generated by the data augmentation, When the label for any of the augmented data in the group of augmented data is determined, assign the same label as the determined label to the augmented data with the same specific label assigned, Control the position and orientation of the data acquirer that acquires the training data so that the possibility of assigning the label or the specific label is maximized, A data collection program that causes a computer to execute the process.
2. The data acquirer is a camera, The data collection program according to claim 1.
3. The data acquirer is a contact sensor, The data collection program according to claim 1.
4. Execute data augmentation on the unlabeled data, Assign a specific label indicating that all the labels of the augmented data match to the group of augmented data generated by the data augmentation, When the label for any of the augmented data in the group of augmented data is determined, assign the same label as the determined label to the augmented data with the same specific label assigned, Control the position and orientation of the data acquirer that acquires the training data so that the possibility of assigning the label or the specific label is maximized, A data collection device including a processor.
5. Execute data augmentation on the unlabeled data, For the group of augmented data generated by the data augmentation, all the labels of the augmented data are all Assign a specific label indicating that it is to be done, When the label for any of the extended data in the extended data group is determined, for the extended data to which the same specific label as the any of the extended data is assigned, assign the same label as the determined label, Control the position and orientation of the data acquirer that acquires the training data so that the possibility of assigning the label or the specific label is the highest, A data collection method in which a computer executes the processing.
Citation Information
Patent Citations
State estimation apparatus
JP2019191644A
Information processing apparatus, inference model construction method, information processing method, inference model, program, and recording medium
JP2021140445A
Systems and methods for training data generation for object identification and self-checkout Anti-theft
US20200151692A1