Learning data generation device, robot system, learning data generation method, and learning data generation program
Patent Information
- Application Number
- JP2025509286
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Filing Date
- 2025-03-13
- Publication Date
- 2025-12-16
AI Technical Summary
Manual annotation of image data for supervised learning is labor-intensive and time-consuming, requiring high labor costs and significant personnel effort, especially when handling large datasets with varying judgments among annotators.
A learning data generation device and method that uses a robot system equipped with a camera and data processing unit to estimate area information of workpieces from images, generating teaching data without the need for extensive human intervention, by acquiring images of workpieces, processing them to extract area information, and storing the data for machine learning.
This approach reduces labor costs and time required for annotation, enabling efficient generation of learning data while maintaining accuracy through automated data processing and user correction, thus facilitating supervised learning without the need for extensive manual annotation.
Abstract
Description
Learning data generation device, robot system, learning data generation method, and learning data generation program
[0001] The present disclosure relates to a training data generation device, a robot system, a training data generation method, and a training data generation program.
[0002] In recent years, machine learning has been used and put to practical use in a variety of fields. One of the machine learning techniques is known as "supervised learning." In "supervised learning," a large number of pairs of features (variables that represent the characteristics of data: clues for prediction) and training data (labels, correct answer data) are prepared and given to a computer (machine learning device). One method of generating training data is annotation (annotation teaching).
[0003] Annotation requires the generation of training data (teaching data) by labeling large amounts of image data, audio data, video data, text data, etc. with relevant tags and metadata. Specifically, annotation involving images is typically done manually by a person clicking on the image to provide information about the area of the object (work). Thus, performing annotation manually requires high labor costs and a huge amount of work time.
[0004] Conventionally, various proposals have been made for technologies to easily generate training data to be used in "supervised learning."
[0005] JP 2022-118300 A JP 2014-059729 A
[0006] As mentioned above, annotation has traditionally been performed manually, resulting in high labor costs and a significant amount of work time. Furthermore, for example, when annotating the position, orientation, and shape regions of objects in a large amount of image data, it is necessary to minimize discrepancies in judgments made by each individual (annotator). This requires the creation of detailed annotation rules, as well as the training and selection of each individual, which further increases labor costs and lengthens the work time.
[0007] Therefore, there is a demand for a training data generation device, a robot system, a training data generation method, and a training data generation program that can perform annotation without incurring high labor costs and huge amounts of work time.
[0008] According to an embodiment of the present disclosure, there is provided a learning data generation device including a data acquisition unit, a data processing unit, and a data storage unit, and configured to generate learning data to be used in machine learning. The data acquisition unit acquires images of areas where multiple workpieces exist, the data processing unit estimates at least area information reflecting the area range of all or part of at least one workpiece based on the acquired images, and generates teaching data including the estimated area information, and the data storage unit stores the generated teaching data and images as learning data.
[0009] FIG. 1 is a diagram schematically illustrating an example of an entire robot system for explaining an example of a training data generation device according to the present embodiment. FIG. 2 is a functional block diagram for explaining an example of a training data generation device according to the present embodiment. FIG. 3 is a flowchart for explaining an example of processing in a first example of a training data generation program according to the present embodiment. FIG. 4 is a flowchart for explaining an example of processing in a second example of a training data generation program according to the present embodiment. FIG. 5 is a flowchart for explaining an example of processing in a third example of a training data generation program according to the present embodiment. FIG. 6 is a flowchart for explaining an example of processing in a fourth example of a training data generation program according to the present embodiment. FIG. 7 is a diagram for explaining an example of processing in which a user corrects teaching data in the training data generation method according to the present embodiment. FIG. 8 is a flowchart for explaining an example of processing in a fifth example of a training data generation program according to the present embodiment. FIG. 9 is a diagram illustrating an example of a workpiece in a robot system in which an example of a training data generation device according to the present embodiment is used. FIG. 10 is a diagram for explaining the shape region of a workpiece in the example of a training data generation device according to the present embodiment.
[0010] Hereinafter, examples of a training data generation device, a robot system, a training data generation method, and a training data generation program according to the present embodiments will be described in detail with reference to the accompanying drawings. In each drawing, identical or similar components are assigned identical or similar reference numerals. Furthermore, the embodiments described below do not limit the technical scope and meaning of the terms of the invention described in the claims.
[0011] 1 is a diagram schematically illustrating an example of an entire robot system for explaining an example of a training data generation device according to this embodiment. As shown in FIG. 1, the robot system 100 includes a robot 1, a robot control device 2, a training data generation device 3, and a camera 4. The robot 1 includes a robot mechanism 10, an arm 11, and an end effector (hand) 12.
[0012] In the robot system 100 shown in FIG. 1 , a machine learning device for performing machine learning (supervised learning) is not shown because it is built into the robot control device 2. However, if the amount of calculation or data is large and it is difficult to build the machine learning device into the robot control device 2, the machine learning device can be configured, for example, as a dedicated workstation installed near the robot control device 2, or a higher-level computer or general-purpose computer installed in a location remote from the robot system 100. Furthermore, when a large amount of data is input into a learning model for learning, a general-purpose computer or processor may be used, but using a GPGPU (General-Purpose Computing on Graphics Processing Units) or a large-scale PC cluster, etc., enables faster processing.
[0013] The robot 1 is configured as, for example, a multi-axis robot, and an end effector 12 is provided at the tip of an arm 11. In Fig. 1, the end effector 12 is a suction device (suction hand), but it goes without saying that this can be changed to various types depending on the workpiece (object) and work content used by the robot system 100. The robot mechanism unit 10 is for making the robot 1 perform predetermined operations based on control commands from the robot control device 2.
[0014] The robot control device 2 receives the output of the camera 4 and the output of the learning data generation device 3, and generates control commands for causing the robot 1 to perform a predetermined operation, for example, based on a program, control data, etc. stored in advance in an internal storage device, and outputs the control commands to the robot mechanism unit 10. As mentioned above, the machine learning device that performs supervised learning does not need to be built into the robot control device 2, but may be provided as a separate device near the robot control device 2 or at a location remote from the robot system 100, depending on the amount of calculation and data, for example.
[0015] The learning data generation device 3 receives multiple images of workpieces and generates teaching data (teacher data) including area information (information about the shape area of the workpieces) of the workpieces D1 to D9 in each image. Furthermore, the learning data generation device 3 outputs the generated teaching data and corresponding images as learning data to the robot control device 2 (machine learning device). Here, the learning data generation device 3 receives image data of the workpieces D1 to D9 captured by the camera 4. However, the image data provided to the learning data generation device 3 is not limited to images captured by the camera 4 of the robot system 100. For example, various image data such as images of the workpieces captured in advance or images obtained by another robot system may be provided. Furthermore, the images input to the learning data generation device 3 are not limited to two-dimensional images. For example, three-dimensional data (three-dimensional images, three-dimensional point cloud data, three-dimensional measurement data) may also be input, as described in detail below.
[0016] The camera 4 is configured to acquire two-dimensional images of the area where multiple workpieces (e.g., multiple cardboard boxes) D1-D9 are present, or two-dimensional images and three-dimensional data (three-dimensional point cloud data). It includes two cameras 4a and 4b and a projector 4c. The projector 4c projects a predetermined pattern onto the area where the multiple workpieces D1-D9 are present, and the two cameras 4a and 4b capture images of the area where the multiple workpieces are present onto which the predetermined pattern is projected by the projector 4c, thereby measuring the three-dimensional shapes of the workpieces D1-D9. In this way, the camera 4 may be configured to measure the three-dimensional shapes of the area where the multiple workpieces D1-D9 are present and acquire three-dimensional point cloud data. However, it can also be configured as a single two-dimensional camera. The camera 4 in FIG. 1 can also acquire two-dimensional images of the area where the multiple workpieces D1-D9 are present using images captured by one of the two cameras 4a and 4b. That is, the camera 4 only needs to be able to capture appropriate images according to, for example, the work or work content that the robot 1 is performing, or according to annotations in supervised learning.
[0017] FIG. 2 is a functional block diagram illustrating an example of a training data generation device according to this embodiment. As shown in FIG. 2, the training data generation device 3, which generates training data for use in machine learning, includes a data acquisition unit 31, a data processing unit (arithmetic processing device) 32, a data storage unit 33, a reception unit 34, and a display unit 35. The data acquisition unit 31 acquires images (two-dimensional images, or two-dimensional images and three-dimensional point cloud data) of the areas where multiple workpieces (D1 to D9) exist and a trained model. The data processing unit 32 estimates, based on the images acquired by the data acquisition unit 31, at least area information reflecting the area range of all or part of at least one workpiece (D0, D), and generates training data including the estimated area information. Here, the trained model may be, for example, a training model generated by another training data generation device 3, or a training model (previous training model) generated by the training data generation device 3 itself. Note that, as will be described in detail later with reference to FIGS. 3 to 8, if a trained model is not used depending on each example according to this embodiment, the data acquisition unit 31 does not need to acquire a trained model. Similarly, in each example according to this embodiment, if three-dimensional point cloud data is not used, the data acquisition unit 31 does not need to acquire three-dimensional point cloud data.
[0018] The area information of the workpiece (D0) includes, for example, at least one of outer shape area information reflecting the outer shape of the workpiece, pick-up area information for adsorbing, suctioning, or gripping the workpiece, and local area information on the workpiece. Note that the local area information of the workpiece (D0) includes, for example, at least one of flat surfaces, curved surfaces, or large-area areas on the workpiece, non-slip areas on the workpiece, and high-density areas on the workpiece. This local area information of the workpiece will be described in detail later with reference to FIG. 9.
[0019] The data storage unit 33 stores the teaching data and images (e.g., two-dimensional images, or two-dimensional images and three-dimensional point cloud data) generated by the data processing unit 32 as learning data. The reception unit 34 receives, for example, correction information based on area information of at least one workpiece input by a worker (user). At this time, the data processing unit 32 corrects and outputs the teaching data based on, for example, the correction information received by the reception unit 34. The display unit 35 displays the images and teaching data, and allows, for example, the user to further correct the teaching data (learning data).
[0020] Here, when the data acquisition unit 31 acquires a trained model, the data processing unit 32 estimates area information of the workpiece based on the image and the trained model and generates teaching data. Also, when the data acquisition unit 31 acquires two-dimensional images of the areas where multiple workpieces (D0 to D9, D) exist, the data processing unit 32 estimates area information based on the two-dimensional images acquired by the data acquisition unit 31 and generates teaching data. Furthermore, when the data acquisition unit 31 acquires two-dimensional images and three-dimensional data (three-dimensional point cloud data) of the areas where multiple workpieces (D0 to D9, D) exist, the data processing unit 32 estimates area information based on the two-dimensional images and three-dimensional point cloud data acquired by the data acquisition unit 31 and generates teaching data.
[0021] The data processing unit 32 can also estimate 3D region information of the workpiece based on the results of 3D point cloud data analysis, such as at least one of plane analysis, curved surface analysis, blob analysis, coordinate system analysis, posture analysis, depth analysis, scale analysis, feature analysis, neighboring point analysis, mesh analysis, and voxel analysis, and can compare the 3D point cloud data with a 2D image to estimate the region information of the workpiece in the 2D image and generate teaching data. Furthermore, the data processing unit 32 can correct and output the teaching data based on the workpiece region information estimated based on the analysis results of the 3D point cloud data. The data processing unit 32 can also perform image processing based on at least one workpiece region information and image received by the receiving unit 34 to extract features, and can match the extracted features to estimate the region information of the workpiece in the image and generate teaching data.
[0022] In the above, the data processing unit 32 can perform image processing based on the image, and can generate teaching data by estimating workpiece area information based on the image and the results of the image processing. The data processing unit 32 can also estimate workpiece area information based on the results of image processing performed using at least one of the following: pattern matching, plane matching, curved surface matching, blob analysis, feature analysis, gradient analysis, edge analysis, contrast analysis, histogram analysis, and color information analysis. Furthermore, the data processing unit 32 can correct and output teaching data based on the workpiece area information estimated based on the results of the image processing.
[0023] Here, before explaining first to fifth examples of the learning data generation program (learning data generation method) according to this embodiment with reference to Figures 3 to 8, an example of a workpiece (object) D0 to be worked on will be explained with reference to Figure 9, and the term "shape area of the workpiece" in this specification will be explained with reference to Figure 10.
[0024] FIG. 9 is a diagram showing an example of a workpiece in a robot system using an example of a training data generation device according to this embodiment. In the robot system according to this embodiment, the workpiece (object) D0 to be worked on is not limited to rectangular parallelepiped objects made of the same material, such as the cardboard boxes D1 to D9 shown in FIG. 1 , but various other objects are possible. Specifically, the workpiece D0 shown in FIG. 9 is an air joint, with regions Da, Db, and Dd formed from plastic (e.g., PTFE: polytetrafluoroethylene) and region Dc formed from metal (e.g., brass or stainless steel). Here, the metal forming region Dc has a higher specific gravity than the plastic forming regions Da, Db, and Dd. Furthermore, region Dc is designed to surround regions Db to Dd with a hexagonal plane to facilitate tightening operations using tools such as wrenches.
[0025] Incidentally, local regions on a workpiece can be various, such as flat surfaces, curved surfaces, large areas, non-slip areas, and high-density (heavy) areas. For example, when a workpiece is picked up using a suction hand (12), the success rate of picking up varies depending on the area of the workpiece being picked up. Therefore, it is preferable for the learning data (teaching data) to include area information surrounding the outline of the local region on the workpiece. Specifically, in the case of the air joint D0 shown in FIG. 9 , for example, picking up the flat area Dc using the suction hand 12 is considered to have a higher success rate of picking up than picking up the curved areas Da, Db, and Dd. Furthermore, picking up the area Dc made of a high-density (heavy) metal material is considered to have a higher success rate of picking up than picking up the areas Da, Db, and Dd made of a low-density (light) plastic material, because picking can be performed at a position closer to the center of gravity of the entire workpiece D0. That is, the suction hand 12 sucks the flat area Dc formed of metal of the air joint D0, which is the workpiece, and thereby the workpiece D0 can be stably picked up.
[0026] Here, as in the fourth embodiment described with reference to Figures 6 and 7, it is preferable for the user to refer to the image and teaching data displayed on the display unit 35 and correct the teaching data (learning data). Even in this case, the user only needs to perform some of the processing, which requires much less effort than the conventional method of manually clicking on an image to provide object area information. In this way, it is preferable for the area information of the workpiece D0 to include at least one of outer shape area information reflecting the outer shape of the workpiece D0, pick-up area information for adsorbing, suctioning, or gripping the workpiece D0, and local area information on the workpiece D0.
[0027] Next, FIG. 10 is a diagram illustrating the shape region of a workpiece in one example of the training data generation device according to this embodiment, where FIG. 10(a) shows the shape region of workpiece D, and FIG. 10(b) shows the region including the background of workpiece D. In this specification, the term "shape region of workpiece" refers to, for example, the region indicated by reference character A1 in the two-dimensional image P3 output from camera 4, which includes only workpiece D but not its background (the region of workpiece D itself), as shown in FIG. 10(a), and not the region including workpiece D and its background, as shown in FIG. 10(b). In other words, the "shape region of workpiece" in this specification is "region information reflecting the shape / outline of the workpiece," and this information can be used to calculate the shape / outline of the workpiece.
[0028] Below, first to fifth examples of the training data generation program (training data generation method) according to this embodiment will be described with reference to Figures 3 to 8. Note that the following description is based on the functional block diagram of the training data generation device 3 described with reference to Figure 2, but it goes without saying that the training data generation device 3 is not limited to that shown in Figure 2.
[0029] 3 is a flowchart illustrating an example of processing in a first example of a training data generation program (training data generation method) according to this embodiment, showing a case where the data acquisition unit 31 acquires a two-dimensional image and a trained model. As shown in FIG. 3, when an example of processing in the training data generation program of the first example starts (START), in step ST11, the camera 4 captures a two-dimensional image of the presence area of the workpiece D. That is, as described with reference to FIG. 1, the camera 4 captures two-dimensional images of the presence areas of the workpieces (e.g., multiple cardboard boxes) D1 to D9 and outputs the images to the training data generation device 3.
[0030] Next, the process proceeds to step ST12, where the data acquisition unit 31 acquires a two-dimensional image and a trained model and outputs them to the data processing unit 32. Here, the two-dimensional image received by the data acquisition unit 31 is a two-dimensional image of the area where the workpiece is located, captured by the camera 4, as described with reference to FIG. 2 . The trained model received by the data acquisition unit 31 may be, for example, a model prepared in advance by a provider that provides the robot system. Alternatively, as described above, the trained model may be, for example, a training model generated by another training data generation device 3 or a training model (previous training model) generated by the training data generation device 3 itself.
[0031] The process then proceeds to step ST13, where the data processing unit 32 estimates the work area based on the two-dimensional image from the data acquisition unit 31 and the trained model, and then the process proceeds to step ST14, where the data processing unit 32 generates teaching data (teaching data) based on the estimated information of the shape area of the workpiece D. That is, the data processing unit 32 estimates the shape area (area information) of the workpiece based on the two-dimensional image and the trained model, and generates teaching data. Here, the shape area of the workpiece D corresponds to the shape area A1 in FIG. 10(a), for example, as described with reference to FIG. 10.
[0032] Then, the process proceeds to step ST15, where the data storage unit 33 stores the teaching data and the 2D image as learning data in the data storage unit 33, and the process ends (END). Here, step ST16 is a process of displaying the 2D image data from step ST12 and the teaching data from step ST14 on the display unit 35. For example, the user can refer to the 2D image and teaching data displayed on the display unit 35 and modify the teaching data. Needless to say, it is also possible to simply display the 2D image and teaching data on the display unit 35 without performing any processing by the user. The process of step ST16 in this first embodiment is the same as steps ST26, ST36, ST47, and ST55 in each embodiment described below.
[0033] In this way, according to the learning data generation program (learning data generation method) of the first embodiment, learning data (teaching data) can be easily generated without incurring high labor costs and huge amounts of work time.
[0034] 4 is a flowchart illustrating an example of processing in a second example of the training data generation program according to this embodiment, showing a case in which the data acquisition unit 31 acquires a two-dimensional image but does not acquire a trained model. As shown in FIG. 4, when an example of processing in the training data generation program according to the second example begins (START), in step ST21, the camera 4 captures a two-dimensional image of the area where the workpiece D (D1 to D9) is present. The process then proceeds to step ST22, in which the data acquisition unit 31 acquires the two-dimensional image and outputs it to the data processing unit 32. Here, the two-dimensional image received by the data acquisition unit 31 is a two-dimensional image of the area where the workpiece is present, as described with reference to FIG. 2, captured by the camera 4.
[0035] Next, the process proceeds to step ST23, where the data processing unit 32 performs image processing on the two-dimensional image from the data acquisition unit 31 to estimate the shape region of the workpiece D. The process then proceeds to step ST24, where the data processing unit 32 generates teaching data based on the estimated information about the shape region of the workpiece D. The process then proceeds to step ST25, where the data storage unit 33 stores the teaching data and the two-dimensional image as learning data, and the process ends (END). Note that step ST26 corresponds to step ST16 in the first embodiment described above, and a description thereof will be omitted.
[0036] As described above, the training data generation program of the second embodiment does not use the trained model of the first embodiment, which simplifies the processing. However, since the trained model is not used, the accuracy of the teaching data may be somewhat lower than that of the first embodiment. Therefore, it is preferable to determine whether to implement the second embodiment based on the shape of the target workpiece, the content of the work, or the time until the robot system actually starts operating.
[0037] Here, as a modified example of the learning data generation program (learning data generation method) according to this embodiment, for multiple workpieces (D0 to D9, D), for example, a user (operator) teaches region information (workpiece shape region) for only one workpiece (D0, D) for one type of workpiece. Furthermore, image processing is performed based on the taught region information and the image (2D image) to extract features. Then, a matching process of the extracted features is performed to estimate the regions of all (multiple) workpieces on the 2D image, and teaching data can be generated.
[0038] 5 is a flowchart for explaining an example of processing in a third example of the training data generation program according to this embodiment, showing a case where the data acquisition unit 31 acquires only a two-dimensional image without acquiring a trained model, the reception unit 34 receives area information (shape area) of the workpiece, and the data processing unit 32 performs image processing. As shown in FIG. 5, when an example of processing of the training data generation program according to the third example starts (START), in step ST31, the camera 4 captures a two-dimensional image of the area where the workpiece D exists. Then, the process proceeds to step ST32, where the data acquisition unit 31 acquires the two-dimensional image and outputs it to the data processing unit 32.
[0039] Next, the process proceeds to step ST33, where the data processing unit 32 performs image processing based on the two-dimensional image from the data acquisition unit 31 and the area information of the workpiece D, and estimates the shape areas of all (multiple) workpieces D1 to D9 (D0, D) on the image. Here, the image processing performed by the data processing unit 32 is, for example, at least one of pattern matching, plane matching, curved surface matching, blob analysis, feature analysis, gradient analysis, edge analysis, contrast analysis, histogram analysis, and color information analysis.
[0040] Next, the process proceeds to step ST34, where the data processing unit 32 generates teaching data based on the estimated information about the shape region of the workpiece D. Then, the process proceeds to step ST35, where the data storage unit 33 stores the teaching data and the two-dimensional image as learning data, and the process ends (END). Note that the process of step ST36 corresponds to steps ST16 and ST26 described above. For example, as in steps ST45 to ST47 of the fourth embodiment described with reference to FIG. 6, the user can also modify the teaching data by referring to the two-dimensional image and teaching data displayed on the display unit 35. Here, the teaching data can be modified by the user based on, for example, the surface shape and condition of the workpiece D or the material (material) forming the region, as described with reference to FIG. 9. However, it is also possible to automate the determination of various regions on the workpiece D, such as flat surfaces, curved surfaces, large-area regions, non-slip regions, and high-density (heavy) regions, using a machine learning device or the like.
[0041] 6 is a flowchart for explaining an example of processing in a fourth example of the training data generation program according to this embodiment, showing a case where the data acquisition unit 31 acquires a two-dimensional image and a trained model, and the user modifies the teaching data. As shown in FIG. 6, when an example of processing of the training data generation program according to the fourth example starts (START), in step ST41, the camera 4 captures a two-dimensional image of the area where the workpiece D exists. Then, the process proceeds to step ST42, where the data acquisition unit 31 acquires the two-dimensional image and the trained model and outputs them to the data processing unit 32.
[0042] Next, the process proceeds to step ST43, where the data processing unit 32 estimates the shape region of the workpiece D based on the two-dimensional image from the data acquisition unit 31 and the trained model, and then the process proceeds to step ST44. In step ST44, the data processing unit 32 generates teaching data based on the estimated information of the shape region of the workpiece D.
[0043] The process then proceeds to step ST45, where the user (operator) modifies the teaching data while viewing (referring to) the display unit 35. The process then proceeds to step ST46, where the data storage unit 33 saves the teaching data and the 2D image as learning data, and the process ends (END). The process of step ST47 corresponds to steps ST16, ST26, and ST36 described above. In the fourth embodiment, as shown in steps ST45 to ST47, the user can modify the teaching data by referring to the 2D image and teaching data displayed on the display unit 35. The modification of the teaching data can be performed by the user based on, for example, the surface shape and condition of the workpiece D or the material (materials) forming the region, as described with reference to FIG. 9 . However, it is also possible to automate the determination of various regions on the workpiece D, such as flat surfaces, curved surfaces, large-area regions, non-slip regions, and high-density (heavy) regions, using a machine learning device or the like.
[0044] 7A and 7B are diagrams illustrating an example of a user-initiated correction process for teaching data in the learning data generation method according to this embodiment. The upper figures (7A, 7B, and 7C) illustrate cases in which the teaching data is correct and no correction by the user is required, while the lower figures (7D, 7E, and 7F) illustrate cases in which the teaching data is incorrect and correction by the user is required. Furthermore, Figures 7A and 7D are two-dimensional images of multiple workpieces, Figures 7B and 7E are images showing the results of contour extraction by image processing superimposed on the two-dimensional images, and Figures 7C and 7F are images showing the final teaching results (final learning data).
[0045] First, a camera 4 is used to capture an image of an area where multiple workpieces (e.g., multiple cardboard boxes) are lined up. Then, while changing the arrangement of the workpieces, multiple 2D images are captured of the area where the multiple workpieces are lined up in different arrangements (FIGS. 7(a) and 7(d)). Next, image processing is performed on each image to extract the contour line characteristics of each workpiece in each image (FIGS. 7(b) and 7(e)). Here, the contour line can be extracted, for example, by binarizing the 2D image to generate a black-and-white image and then searching for the black-and-white boundary line in the image. For example, if the workpiece is a cardboard box, the color within the workpiece area is nearly uniform, and the greater the color difference between the workpiece area and the background area, the more accurately the contour line of the workpiece can be determined.
[0046] First, let us consider the case where the contour lines extracted by image processing are correct. For example, the case where image processing is performed on the captured images of workpieces a11 to a14 shown in FIG. 7(a), resulting in the extraction of contour line features of workpieces b11 to b14 shown in FIG. 7(b). In this case, the user views (references) the 2D image of the workpiece and the image of the teaching data shown in FIG. 7(b) displayed on the display unit 35. Here, the user determines whether the learning data (2D image and teaching data) shown in FIG. 7(b) displayed on the display unit 35 is correct. If the user determines that the learning data is correct, the final learning data (teaching data c11 to c14) shown in FIG. 7(c) is obtained without modifying the learning data (teaching data), and this learning data is stored in the data storage unit 33.
[0047] On the other hand, when the contours extracted by image processing are incorrect, for example, when image processing is performed on a captured image of workpieces d11 to d14 as shown in FIG. 7(d), the contour features of workpieces e11 to e14 as shown in FIG. 7(e) are extracted. That is, as shown in FIG. 7(d), for example, if the color within the area of workpiece d13 in the image is very close to the background color, and workpieces d12 and d14 are joined together and the boundary between them is unclear, as shown in FIG. 7(e), the image processing fails to recognize the left and bottom contours of workpiece e13, and the contour between workpieces e12 and e14 is not recognized, resulting in the incorrect identification of them as a single workpiece. In this case, the user sees a two-dimensional image of the workpiece and an image of the teaching data as shown in FIG. 7(e) displayed on the display unit 35. The user then determines whether the learning data (two-dimensional image and teaching data) displayed on the display unit 35 as shown in FIG. 7(e) is correct.
[0048] If the user determines that the training data shown in Figure 7(e) is incorrect, for example, if the user determines that no workpiece is recognized in e13 in Figure 7(e), the missing contour line in e13 can be added with a mouse or other device to correct the area where the workpiece is present. Furthermore, if the user determines that e12 and e14 in Figure 7(e) are not a single workpiece but two, the missing boundary line between e12 and e14 can be added with a mouse or other device to correct the area to two works. Furthermore, if the user determines that the contour lines of the label in the area e11 in Figure 7(e) were extracted and the label enclosed by these contour lines was erroneously recognized as a single cardboard box, the contour line enclosing the label can be deleted. Furthermore, if the user determines that a contour line was erroneously extracted in a background area where no cardboard box is present, the excess contour line can be deleted. In this way, the final learning data (teaching data f11 to f14) corrected by the user as shown in Figure 7(f) is obtained, and this learning data is stored in the data storage unit 33. Note that the correction of the teaching data by the user is preferably used, for example, to determine areas of metal and plastic materials, which can be determined immediately by the user (person) referring to an image (for example, a color image or a grayscale image), as described with reference to Figure 9.
[0049] Thus, various problems can arise, such as, for example, if the color within the work area is not uniform, a contour line will be extracted within the work area; if the color difference between the work area and the background area is not large, part of the work's contour line will be missing; or, if multiple workpieces are closely spaced, adjacent contour lines cannot be extracted. Furthermore, if the resolution of the captured image is low or if the lighting conditions at the time of capture are poor and noise is likely to occur, a contour line may be erroneously extracted in a background area where no workpiece is present. Therefore, it is preferable to superimpose such extraction results on the image and display them on the display unit 35 for the user, so that the user can visually delete / correct incorrect contour lines and add missing contour lines to modify the teaching data to reflect the correct work area, thereby generating learning data.
[0050] In other words, if training data including teaching data with an incorrect work area is used for machine learning, correct learning will be impossible, making it difficult to generate a trained model that can correctly infer and calculate the work area. In response to this, the accuracy (reliability) of the teaching data can be improved by presenting the generated teaching data and images (training data) to the user and allowing the user to modify the teaching data based on their own judgment. While the above-described processing is performed by the user, the user does not need to perform the teaching process from scratch; instead, the user only needs to modify the incorrect parts, which obviously does not require high labor costs or extensive work time. Note that the above-described user-initiated correction of teaching data is not limited to the fourth example and can be widely used in the examples of the training data generation device, robot system, training data generation method, and training data generation program according to this embodiment.
[0051] 8 is a flowchart for explaining an example of processing in a fifth example of the learning data generation program according to this embodiment. As shown in FIG. 8, when an example of processing in the learning data generation program according to the fifth example starts (START), in step ST51, the camera 4 captures a two-dimensional image of the area where the workpiece D exists, and the three-dimensional sensor acquires three-dimensional point cloud data (three-dimensional data) of the area where the workpiece D exists. Then, the process proceeds to step ST52, where the data acquisition unit 31 performs image processing on the two-dimensional image to estimate the shape area.
[0052] Next, the process proceeds to step ST53, where the area information of the workpiece D is corrected based on the three-dimensional point cloud data. Here, correcting the area information of the workpiece D based on the three-dimensional point cloud data means, for example, comparing the three-dimensional point cloud data with the two-dimensional images of the shape areas calculated in the two-dimensional images of the multiple workpieces. Furthermore, by utilizing the difference in three-dimensional position contained in the three-dimensional point cloud data, background areas, obstacle areas, areas of adjacent workpieces, etc. that are erroneously included in the calculation results are excluded, and the corrected teaching data for the workpiece is obtained.
[0053] The process then proceeds to step ST54, where the data storage unit 33 stores the teaching data and the two-dimensional image as learning data, and the process ends (END). When correcting the region information of the workpiece D from the two-dimensional image based on the three-dimensional point cloud data, it is preferable to extract features such as grooves, gaps, steps, three-dimensional planes, and three-dimensional curved surfaces from the two-dimensional image. This makes it possible to improve the reliability (accuracy) of the teaching data by comparing it with the three-dimensional point cloud data, even if the two-dimensional image cannot distinguish the workpiece and its surrounding background, or the boundaries between workpieces. In the fifth embodiment, step ST55 corresponds to step ST47 of the fourth embodiment described with reference to FIG. 6, and, as described above, the processes of steps ST45 and ST46 in the fourth embodiment can be added.
[0054] The training data generation program according to the present embodiment described above may be provided by recording it on a computer-readable non-transitory recording medium or a non-volatile semiconductor memory, or may be provided via a wired or wireless connection. Examples of the computer-readable non-transitory recording medium include optical disks such as CD-ROMs (Compact Disc Read Only Memory) and DVD-ROMs, or hard disk drives. Examples of the non-volatile semiconductor memory include PROMs (Programmable Read Only Memory) and flash memory. Furthermore, distribution from a server device may be via a wired or wireless local area network (LAN) or a wide area network (WAN) such as the Internet.
[0055] As described above in detail, the training data generation device, robot system, training data generation method, and training data generation program according to this embodiment make it possible to perform annotation without incurring high labor costs and huge amounts of work time.
[0056] Although the present disclosure has been described in detail, the present disclosure is not limited to the individual embodiments described above. Various additions, substitutions, modifications, partial deletions, etc. are possible in these embodiments without departing from the gist of the present disclosure or the spirit of the present disclosure derived from the content of the claims and their equivalents. These embodiments can also be implemented in combination. For example, in the above-described embodiments, the order of each operation and the order of each process are shown as examples and are not limited to these. The same applies when numerical values or mathematical expressions are used in the description of the above-described embodiments.
[0057] The following supplementary notes are further disclosed regarding the above-described embodiments and variations. [Supplementary Note 1] A learning data generation device (3) for generating learning data to be used in machine learning, comprising: a data acquisition unit (31) for acquiring images of regions where multiple workpieces (D0 to D9, D) exist; a data processing unit (32) for estimating, based on the acquired images, region information reflecting the region range of all or part of at least one of the workpieces (D0, D), and generating teaching data including the estimated region information; and a data storage unit (33) for saving the generated teaching data and the image as the learning data. [Supplementary Note 2] The learning data generation device according to Supplementary Note 1, wherein the region information of the workpiece (D0, D) includes at least one of outer shape region information reflecting the outer shape of the workpiece (D0, D), removal region information for adsorbing, suctioning, or gripping the workpiece (D0, D), and local region information on the workpiece (D0, D). [Supplementary Note 3] The learning data generation device according to Supplementary Note 2, wherein the local region information of the workpiece (D0, D) includes at least one of a flat surface, a curved surface, or a region with a large area on the workpiece (D0, D), a region with a low slipperiness on the workpiece (D0, D), and a region with a high density on the workpiece (D0, D). [Supplementary Note 4] The learning data generation device according to any one of Supplementary Notes 1 to 3, further comprising a reception unit (34), wherein the reception unit (34) receives correction information based on region information of at least one of the workpieces (D0, D), and the data processing unit (32) corrects and outputs the teaching data based on the teaching information received by the reception unit (34). [Supplementary Note 5] The training data generation device according to any one of Supplementary Note 1 to Supplementary Note 4, wherein the data acquisition unit (31) further acquires a trained model, and the data processing unit (32) estimates area information of the work (D0, D) based on the image and the trained model, and generates the teaching data.[Supplementary Note 6] The learning data generation device according to any one of Supplements 1 to 5, wherein the data acquisition unit (31) acquires two-dimensional images of areas where the plurality of workpieces (D0 to D9, D) exist, and the data processing unit (32) estimates the area information based on the two-dimensional images, and generates the teaching data. [Supplementary Note 7] The learning data generation device according to any one of Supplements 1 to 5, wherein the data acquisition unit (31) acquires two-dimensional images and three-dimensional data of areas where the plurality of workpieces (D0 to D9, D) exist, and the data processing unit (32) estimates the area information based on the two-dimensional images and the three-dimensional data, and generates the teaching data. [Supplementary Note 8] The learning data generation device according to Supplementary Note 7, wherein the data processing unit (32) analyzes the three-dimensional data, estimates three-dimensional area information of the workpieces based on the analysis results, compares the three-dimensional data with the two-dimensional image, estimates area information of the workpieces in the two-dimensional image, and generates the teaching data. [Supplementary Note 9] The training data generation device according to Supplementary Note 8, wherein the data processing unit (32) estimates three-dimensional region information of the workpiece (D0, D) based on a processing result of performing at least one of plane analysis, curved surface analysis, blob analysis, coordinate system analysis, posture analysis, depth analysis, scale analysis, feature analysis, neighboring point analysis, mesh analysis, and voxel analysis as an analysis of the three-dimensional data. [Supplementary Note 10] The training data generation device according to Supplementary Note 8 or Supplementary Note 9, wherein the data processing unit (32) corrects and outputs the teaching data based on the region information of the workpiece (D0, D) estimated based on the analysis result of the three-dimensional data. [Supplementary Note 11] The learning data generation device according to any one of Supplementary Note 4 to Supplementary Note 6, wherein the data processing unit (32) performs image processing based on area information of at least one of the works (D0, D) received by the receiving unit (34) and the image to extract features, performs matching processing of the extracted features to estimate area information of the plurality of works (D0, D) on the image, and generates the teaching data.[Supplementary Note 12] The training data generation device according to any one of Supplements 1 to 11, wherein the data processing unit (32) performs image processing based on the image, and estimates region information of the workpiece (D0, D) based on the image and a result of the image processing, thereby generating the teaching data. [Supplementary Note 13] The training data generation device according to Supplementary Note 11 or 12, wherein the data processing unit (32) estimates region information of the workpiece (D0, D) based on a processing result of at least one of pattern matching, plane matching, curved surface matching, blob analysis, feature analysis, gradient analysis, edge analysis, contrast analysis, histogram analysis, and color information analysis performed as the image processing. [Supplementary Note 14] The training data generation device according to any one of Supplements 11 to 13, wherein the data processing unit (32) corrects and outputs the teaching data based on the region information of the workpiece (D0, D) estimated based on the result of the image processing. [Supplementary Note 15] The training data generation device according to any one of Supplements 1 to 14, further comprising a display unit (35), wherein the display unit (35) displays the image and the teaching data. [Supplementary Note 16] The training data generation device according to Supplementary Note 15, wherein the data processing unit (32) includes a process of correcting the teaching data by referring to the image and the teaching data displayed on the display unit (35). [Supplementary Note 17] A robot system (100) comprising: a robot (1) performing a predetermined process on a plurality of the workpieces (D0 to D9, D); a camera (4) capturing an image of an area where the plurality of workpieces (D0 to D9, D) are present and outputting the image; the training data generation device (3) according to any one of Supplements 1 to 16, receiving the image from the camera (4) and outputting the training data; a machine learning device receiving the training data from the training data generation device (3) and performing machine learning; and a robot control device (2) receiving an output from the machine learning device and controlling the robot (1).[Supplementary Note 18] A learning data generation method for generating learning data to be used in machine learning, comprising: a data acquisition step of acquiring images of areas where a plurality of workpieces (D0 to D9, D) exist, a data processing step of estimating at least area information reflecting an area range of all or part of at least one of the workpieces (D0, D) based on the acquired images, and generating teaching data including the estimated area information, and a data storage step of storing the generated teaching data and the images as the learning data. [Supplementary Note 19] The learning data generation method according to Supplementary Note 18, wherein the data acquisition step further acquires a trained model generated by prior training, and the data processing step estimates area information of the workpiece (D0, D) based on the images and the trained model, and generates the teaching data. [Supplementary Note 20] The learning data generation method according to Supplementary Note 18 or Supplementary Note 19, further comprising a receiving step, wherein the receiving step receives area information of at least one of the workpieces (D0, D), and the data processing step performs image processing based on the received area information and the image for a plurality of the workpieces (D0 to D9, D) to extract features, and performs matching processing of the extracted features to estimate area information of the plurality of workpieces (D0 to D9, D) on the image, thereby generating the teaching data. [Supplementary Note 21] The learning data generation method according to any one of Supplementary Note 18 to Supplementary Note 20, wherein the data processing step performs image processing based on the image, and generates the teaching data by estimating area information of the workpieces (D0 to D9, D) based on a processing result of the image processing and the image. [Supplementary Note 22] The learning data generation method according to any one of Supplementary Note 18 to Supplementary Note 21, wherein the data acquisition step further acquires two-dimensional images and three-dimensional data of areas where the plurality of workpieces (D0 to D9, D) exist, and the data processing step estimates the area information based on the two-dimensional images and the three-dimensional data, and generates the teaching data.[Supplementary Note 23] The training data generation method according to any one of Supplementary Note 18 to Supplementary Note 22, further comprising a receiving step of receiving correction information based on area information of at least one of the workpieces (D0, D), and the data processing step of correcting the teaching data based on the received correction information. [Supplementary Note 24] A training data generation program for generating training data to be used in machine learning, comprising: a data acquiring step of acquiring images of areas where a plurality of workpieces (D0 to D9, D) exist; a data processing step of estimating, based on the acquired images, at least area information reflecting an area range of all or part of at least one of the workpieces (D0, D), and generating teaching data including the estimated area information; and a data saving step of saving the generated teaching data and the images as the training data.
[0058] REFERENCE SIGNS LIST 1 Robot 2 Robot control device 3 Learning data generation device 4 Camera 10 Robot mechanism 11 Arm 12 End effector (hand) 31 Data acquisition unit 32 Data processing unit 33 Data storage unit 34 Reception unit 35 Display unit 100 Robot system A1 Shape area A2 Areas D, D1 to D9 Workpiece (object)
Claims
1. A learning data generation device that generates learning data to be used in machine learning, a data acquisition unit that acquires images of areas where multiple workpieces exist; a data processing unit that estimates at least area information reflecting an area range of all or part of at least one of the workpieces based on the acquired image, and generates teaching data including the estimated area information; a data storage unit that stores the generated teaching data and the image as the learning data; A training data generation device comprising:
2. 2. The learning data generation device according to claim 1, wherein the area information of the workpiece includes at least one of outer shape area information reflecting the outer shape of the workpiece, removal area information for adsorbing, suctioning or grasping the workpiece, and local area information on the workpiece.
3. 3. The learning data generation device according to claim 2, wherein the local area information of the workpiece includes at least one of a flat surface, a curved surface, or a large area on the workpiece, a non-slip area on the workpiece, and a high-density area on the workpiece.
4. Furthermore, a reception unit is provided, the receiving unit receives correction information based on area information of at least one of the works; The training data generating device according to claim 1 , wherein the data processing unit corrects and outputs the teaching data based on the correction information received by the receiving unit.
5. The data acquisition unit further acquires a trained model, The learning data generation device according to claim 1 , wherein the data processing unit estimates area information of the workpiece based on the image and the trained model, and generates the teaching data.
6. the data acquisition unit acquires two-dimensional images of areas where the plurality of workpieces are present, The training data generating device according to claim 1 , wherein the data processing unit estimates the region information based on the two-dimensional image and generates the teaching data.
7. the data acquisition unit acquires two-dimensional images and three-dimensional data of the areas where the plurality of workpieces are present, The training data generation device according to claim 1 , wherein the data processing unit estimates the region information based on the two-dimensional image and the three-dimensional data, and generates the teaching data.
8. The learning data generation device according to claim 7, wherein the data processing unit analyzes the three-dimensional data, estimates three-dimensional area information of the workpiece based on the analysis results, compares the three-dimensional data with the two-dimensional image to estimate area information of the workpiece in the two-dimensional image, and generates the teaching data.
9. 9. The learning data generation device according to claim 8, wherein the data processing unit estimates three-dimensional region information of the workpiece based on processing results obtained by analyzing the three-dimensional data using at least one of plane analysis, curved surface analysis, blob analysis, coordinate system analysis, posture analysis, depth analysis, scale analysis, feature analysis, neighboring point analysis, mesh analysis, and voxel analysis.
10. The learning data generating device according to claim 8 , wherein the data processing unit corrects and outputs the teaching data based on area information of the workpiece estimated based on an analysis result of the three-dimensional data.
11. The learning data generation device according to claim 4, wherein the data processing unit performs image processing based on the area information of at least one of the workpieces received by the receiving unit and the image to extract features, performs a matching process of the extracted features to estimate area information of the multiple workpieces on the image, and generates the teaching data.
12. The learning data generation device according to any one of claims 1 to 3, wherein the data processing unit performs image processing based on the image, and estimates area information of the work based on the processing result of the image processing and the image, thereby generating the teaching data.
13. 12. The training data generation device according to claim 11, wherein the data processing unit estimates the region information of the workpiece based on a processing result obtained by performing, as the image processing, at least one of pattern matching, plane matching, curved surface matching, blob analysis, feature analysis, gradient analysis, edge analysis, contrast analysis, histogram analysis, and color information analysis.
14. The learning data generating device according to claim 11 , wherein the data processing unit corrects and outputs the teaching data based on area information of the workpiece estimated based on a processing result of the image processing.
15. Further, the device is provided with a display unit, The training data generation device according to claim 1 , wherein the display unit displays the image and the teaching data.
16. The training data generation device according to claim 15 , wherein the data processing unit includes a process for correcting the teaching data by referring to the image and the teaching data displayed on the display unit.
17. a robot that performs a predetermined process on a plurality of the workpieces; a camera that captures an image of an area where the plurality of workpieces are present and outputs the image; a training data generation device according to claim 1 , which receives the image from the camera and outputs the training data; a machine learning device that receives the training data from the training data generation device and performs machine learning; a robot control device that receives the output of the machine learning device and controls the robot.
18. A learning data generation method for generating learning data to be used in machine learning, comprising: a data acquisition step of acquiring images of areas where a plurality of workpieces are present; a data processing step of estimating at least area information reflecting an area range of all or part of at least one of the workpieces based on the acquired image, and generating teaching data including the estimated area information; a data storage step of storing the generated teaching data and the image as the learning data; A training data generation method comprising:
19. The data acquisition step further includes acquiring a trained model that has been trained and generated in advance, The training data generation method according to claim 18 , wherein the data processing step estimates area information of the workpiece based on the image and the trained model, and generates the teaching data.
20. The method further includes a receiving step, wherein the receiving step receives area information of at least one of the works, The learning data generation method according to claim 18 or 19, wherein the data processing step performs image processing based on the received area information and the image for a plurality of the works to extract features, performs a matching process for the extracted features to estimate area information for the plurality of works on the image, and generates the teaching data.
21. The learning data generation method according to claim 18 or claim 19, wherein the data processing step performs image processing based on the image, and estimates area information of the work based on the processing result of the image processing and the image to generate the teaching data.
22. The data acquisition step further acquires two-dimensional images and three-dimensional data of the areas where the plurality of workpieces are present, 20. The learning data generating method according to claim 18, wherein the data processing step estimates the region information based on the two-dimensional image and the three-dimensional data, and generates the teaching data.
23. The method further includes a receiving step, wherein the receiving step receives correction information based on area information of at least one of the works; 20. The learning data generating method according to claim 18, wherein the data processing step corrects the teaching data based on the received correction information.
24. A learning data generation program for generating learning data to be used in machine learning, The processing unit a data acquisition step of acquiring images of areas where a plurality of workpieces are present; a data processing step of estimating at least area information reflecting an area range of all or part of at least one of the workpieces based on the acquired image, and generating teaching data including the estimated area information; a data saving step of saving the generated teaching data and the image as the learning data.