Information processing device, robot system, and program
Patent Information
- Application Number
- PCT/JP2025/012487
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2026-10-01
Smart Images

Figure JP2025012487_01102026_PF_FP_ABST
Abstract
Description
Information processing apparatus, robot system, and program
[0001] The present disclosure relates to an information processing apparatus, a robot system, and a program.
[0002] Robot systems that detect the position of an object based on image information acquired by a visual sensor and enable the robot to perform a work of picking up the object based on the detection result are widely used (for example, Patent Document 1). Further, in such a robot system, in order to improve the recognition rate of the object by the visual sensor, a system having a function of detecting the object from a captured image using a learning model constructed by machine learning using training images has also been proposed (for example, Patent Document 2 and Patent Document 3).
[0003] Japanese Unexamined Patent Application Publication No. 2018-144167 International Publication No. 2022 / 138545 Japanese Unexamined Patent Application Publication No. 2021-3782
[0004] In the above-described robot system that detects an object from a captured image using a learning model constructed by machine learning, it is desired to evaluate whether the learning model constructed by learning is appropriate. As an evaluation index for a learning model, for example, in the field of object detection, Average Precision (AP), which represents the degree of matching between the inference result by the learning model and the ground truth data, is known. However, it is difficult for an ordinary user to grasp whether a learning model is suitable for application to a robot system from such an index. There is a demand for a technology that can provide an evaluation index that allows even an ordinary user to easily grasp whether or not a learning model is suitable for application to a robot system.
[0005] One aspect of the present disclosure is an information processing device comprising: a data acquisition unit that acquires a set of image data and annotation data representing an object in the image data for a plurality of images; a learning model acquisition unit that acquires a learning model for detecting an object from image information; and a model evaluation unit that evaluates the learning model by determining the occurrence of undetected and falsely detected objects based on an inference result of an object obtained by inputting at least a portion of the acquired image data for the plurality of images into the learning model and the annotation data of the input image data.
[0006] These and other objects, features, and advantages of the present invention will become even clearer from the detailed description of typical embodiments of the present invention shown in the accompanying drawings.
[0007] This figure shows the configuration of a robot system including an image processing device according to one embodiment. This is a functional block diagram showing the functions of the image processing device and the robot control device. This is a flowchart showing the overall flow of processing for evaluating a learning model. This figure shows a first example of how collected image data and annotation data are used for learning and evaluating a learning model. This figure shows a second example of how collected image data and annotation data are used for learning and evaluating a learning model. This figure shows an example of a settings screen for setting evaluation conditions for a learning model. This figure shows an example of an output screen for evaluation results. This figure shows a first example of detection result images where undetected and falsely detected items occurred. This figure shows a second example of detection result images where undetected and falsely detected items occurred.
[0008] Next, embodiments of the present disclosure will be described with reference to the drawings. In the drawings, similar components or functional parts are given the same reference numerals. For ease of understanding, the scale of these drawings has been appropriately changed. Furthermore, the embodiments shown in the drawings are just one example of how to carry out the present invention, and the present invention is not limited to the illustrated embodiments.
[0009] Figure 1 is a diagram showing the configuration of a robot system 100 including an image processing device 20 according to one embodiment. As shown in Figure 1, the robot system 100 comprises a robot 10, a robot control device 50 that controls the robot 10, a teaching control panel 40 connected to the robot control device 50, a vision sensor 70 connected to the image processing device 20, and an image processing device 20 connected to the robot control device 50.
[0010] In this embodiment, the robot system 100 can be configured as a system that detects an object based on inference results obtained by inputting images captured by the vision sensor 70 into a learning model, and then uses a hand 15 mounted on the robot 10 to pick up and transport the object (box 90). In the learning stage of such a robot system, learning data is created by adding annotation data indicating the position of the object as a correct label to each of a large number of images of the object (box 90 in this embodiment), and then learning is performed. The robot system 100 according to this embodiment has a function to provide evaluation indicators that allow even ordinary users to easily understand the quality of the learning model thus constructed.
[0011] The image processing device 20 is responsible for controlling and learning the visual sensor 70, detecting objects using the learned model, and performing various image processing operations. In this embodiment, the image processing device 20 may be composed of various information processing devices such as a PC (personal computer), tablet terminal, or smartphone. The image processing device 20 may have a hardware configuration as a general computer, including a processor 21, memory (ROM, RAM, non-volatile memory, etc.), storage unit 22, input / output interface, network interface, display unit 23, operation unit 24, etc. (Figure 2).
[0012] Alternatively, the image processing device 20 may be configured as a dedicated image processing device. Alternatively, the functions of the image processing device 20 may be incorporated into the robot control device 50.
[0013] In this embodiment, robot 10 is a vertical articulated robot, but various types of robots may be used as robot 10 depending on the work to be performed, such as a horizontal articulated robot, a parallel link type robot, or a dual-arm robot. Robot 10 can perform desired tasks by an end effector attached to its wrist. The end effector is an external device that can be replaced depending on the application, such as a hand, a welding gun, or a tool. Figure 1 shows an example in which a hand is used as the end effector. Various types of hands can be used as the hand, such as a multi-fingered gripping hand or a suction-type hand, but in this embodiment, a suction-type hand is used.
[0014] The visual sensor 70 may be a two-dimensional camera that acquires two-dimensional images, or it may be a three-dimensional camera that acquires three-dimensional position information of an object. In this embodiment, since a three-dimensional object such as a box 90 is assumed to be the object of image recognition using the visual sensor 70, the visual sensor 70 is assumed to have the function of a three-dimensional camera that can acquire distance images. Even if the visual sensor 70 is a two-dimensional camera, the functions according to this embodiment described below can be basically realized in the same way. A distance image is an image that represents the distance to the object using shades of gray. As the three-dimensional camera, a stereo camera or a TOF (Time of Flight) camera that acquires three-dimensional position information of an object using the optical time-of-flight method can be used. The visual sensor 70 is also assumed to have the function of a two-dimensional camera that acquires normal two-dimensional images.
[0015] In this specification, unless otherwise specified, two-dimensional images and depth images may be simply referred to as images, image information, image data, etc.
[0016] In this embodiment, it is assumed that the vision sensor 70 is attached to the tip of the arm of the robot 10, as shown in Figure 1. The vision sensor 70 is calibrated, and its position relative to the robot 10's reference coordinate system (e.g., flange coordinate system) is known to the robot control device 50. That is, the position and orientation of the vision sensor 70 relative to the robot 10 is known information. The image processing device 20 can use the calibration data to determine the position on the image captured by the vision sensor 70 as the position on a coordinate system (such as the robot coordinate system) set in the workspace. The calibration data (internal parameters and external parameters) is stored, for example, in the storage unit 22 of the image processing device 20 (see Figure 2).
[0017] Although Figure 1 shows an example configuration in which the vision sensor 70 is mounted on the robot 10, the vision sensor 70 may also be fixed in a position in the workspace where it can capture images of the target object.
[0018] The robot control device 50 controls the operation of the robot 10 according to an operation program or commands from the teaching control panel 40. The robot control device 50 may have a hardware configuration as a general computer, including a processor, memory (ROM, RAM, non-volatile memory, etc.), storage unit, operation unit, input / output interface, network interface, etc.
[0019] The teaching control panel 40 is connected to the robot control device 50 and also to the image processing device 20 via the robot control device 50. The teaching control panel 40 is used as an operating terminal for teaching (program creation) the operation program of the robot 10 and for making various settings related to teaching. The teaching control panel 40 may be composed of a tablet terminal or the like. The teaching control panel 40 may have a hardware configuration as a general computer, including a processor, memory (ROM, RAM, non-volatile memory, etc.), storage unit, operation unit, display unit 43, input / output interface, network interface, etc. (see Figure 2). The display unit 43 is composed of, for example, a liquid crystal display.
[0020] Figure 2 is a functional block diagram showing the functions of the image processing device 20 and the robot control device 50. As shown in Figure 2, the robot control device 50 includes an motion control unit 151. The motion control unit 151 controls the movement of the robot 10 according to the motion program or according to commands from the teaching control panel 40. The motion program may include a vision program (a program related to imaging by a visual sensor and image processing of the captured images). The robot control device 50 includes a servo control unit (not shown) that performs servo control to the servo motors of each axis according to commands for each axis generated by the motion control unit 151.
[0021] As shown in Figure 2, the image processing device 20 includes a visual sensor control unit 121, a detection unit 122, a data acquisition unit 123, a learning model acquisition unit 124, a model evaluation unit 125, a setting unit 126, and a learning unit 127. These functional elements may also be realized by the processor 21 of the image processing device 20 executing software.
[0022] Figure 2 shows the storage unit 22 as a hardware component. The storage unit 22 is a storage device consisting of, for example, non-volatile memory or a hard disk drive. The storage unit 22 stores data 80 in which annotation data has been added to the image data (hereinafter also referred to as annotated image data), a learning model 128 constructed by the learning unit 127, and various other information necessary for learning and image processing.
[0023] The visual sensor control unit 121 has the function of performing various controls related to the imaging operation on the visual sensor 70.
[0024] The learning unit 127 can construct a learning model 128 by learning from a large number of annotated images to which information indicating the position of the box as an object (outer shape position of the box) is attached as a correct label to images captured by the visual sensor 70. Supervised learning, a type of machine learning technique, can be used for learning. Deep learning techniques using a multilayer neural network (DNN) may also be incorporated into the learning process.
[0025] As a multilayer neural network, a convolutional neural network (CNN), suitable for the field of image recognition, may be used. In the field of image recognition, CNNs are used for various tasks such as image classification, object detection, and semantic segmentation. In the task of image classification, the output layer of the trained CNN outputs a probability score for each class as an inference result. A CNN-based object detector also outputs the bounding box position and class score as an inference result. In CNN-based semantic segmentation, a configuration called an encoder-decoder model is used. The final feature map of the decoder is converted into the probability of the object class for each pixel via a softmax layer and output as an inference result.
[0026] In the training phase of various image recognition models, an error function (loss function) is calculated using training data, and the optimal weight values are obtained by repeatedly performing gradient descent calculations to minimize the error function.
[0027] In this embodiment, the learning unit 127 is not limited to any particular learning framework. The learning unit 127 only needs to be configured to construct a learning model that infers the positional information of an object (for example, the external position of a box) from an image captured by a visual sensor. In this embodiment, the learning image used by the learning unit 127 for learning is assumed to be a distance image captured by the visual sensor 70 so that the box as the object is visible. The learning unit 127 may also be configured to construct a learning model using a learning image as a two-dimensional image.
[0028] The detection unit 122 detects an object from image information acquired by the visual sensor 70 by inference using a learning model 128. For example, the detection unit 122 may be configured to output a final detection result based on the inference result obtained by applying the image captured by the visual sensor 70 to the learning model 128 as an estimator, and predetermined detection conditions.
[0029] The data acquisition unit 123, the learning model acquisition unit 124, the model evaluation unit 125, and the setting unit 126 are components related to the function of providing evaluation indicators for users to understand the quality of the learning model 128.
[0030] The data acquisition unit 123 acquires annotated image data 80 of a large number of images. The data acquisition unit 123 may be configured to acquire annotated image data 80 from the storage unit 22, or it may be configured to acquire annotated image data from an external device. Alternatively, the data acquisition unit 123 may have a function to display a graphical user interface on the display unit 23 of the image processing device 20, accept user operations to add annotations to images on this graphical user interface, and generate annotated image data as a result.
[0031] The learning model acquisition unit 124 acquires the learning model 128. The learning model acquisition unit 124 may be configured to acquire the learning model 128 from the storage unit 22, or it may be configured to acquire the learning model from an external device.
[0032] The model evaluation unit 125 has the function of evaluating the learning model 128 by determining the occurrence of at least one of undetected or falsely detected events in the inference result, based on the inference result obtained by inputting at least a portion of the annotated image data 80 into the learning model 128 and the annotation data of the input image data.
[0033] The setting unit 126 provides a function for setting evaluation conditions when the model evaluation unit 125 performs an evaluation. The setting unit 126 may also have a function for displaying a user interface screen for setting evaluation conditions. Alternatively, the setting unit 126 may have a function for receiving evaluation conditions from an external device.
[0034] Figure 3 is a flowchart showing the overall flow of the process for evaluating the learning model executed on the image processing device 20 (hereinafter also referred to as the learning model evaluation process). The learning model evaluation process in Figure 3 is mainly executed under the control of the processor 21 of the image processing device 20.
[0035] First, a process is performed to collect images to be used for learning (step S1). Here, as an example, the user may collect a large number of images by arranging the box 90 as the object in various positions and taking images of the box 90 with the visual sensor 70.
[0036] Next, the process of annotating the image, that is, the process of adding annotation data to the image, is performed (step S2). Here, as an example, the image processing device 20 (processor 21) may accept an operation in which the user inputs images of the outlines of each box 90 as annotation data. Assume that annotated image data 80 containing a large number of images has been created by the processing from steps S1 to S2.
[0037] Next, the setting unit 126 accepts user input to set parameters (evaluation parameters) as evaluation conditions (step S3).
[0038] Figure 5 shows an example of a setting screen 200 displayed on the display unit 23 by the setting unit 126. The setting screen 200 includes a setting button 201 for setting that training images be used as evaluation images as well, and a specification field 202 for specifying the percentage of images included in the annotated image data 80 to be used for evaluation. When the setting button 201 is set to use training images as evaluation images as well, as shown in Figure 4A, the entire annotated image data 80 is used for training, and the percentage of data specified in the specification field 202 from the entire annotated image data 80 is also used for evaluation by the model evaluation unit 125.
[0039] The usage configuration of the annotated image data 80 shown in Figure 4A has the advantage of saving the effort of collecting new images for evaluation, as some of the training images are also used as evaluation images, thus keeping the total number of images in the annotated image data 80 relatively low.
[0040] If the setting button 201 does not configure the system to use training images as evaluation images, the entire annotated image data 80 is divided into training data 80a for training by the training unit 127 and evaluation data 80b for evaluation by the model evaluation unit 125, according to the proportion specified in the specification field 202, as shown in Figure 4B. In this case, as shown in Figure 5, if 50% is specified in the specification field, 50% of the entire annotated image data 80 is used for evaluation, and the remaining annotated image data (the remaining 50%) is used for training.
[0041] The usage of the annotated image data 80 shown in Figure 4B has the advantage of allowing the learning model to be evaluated using images that are unknown to the learning model.
[0042] In step S4, the learning unit 127 performs training using all or part of the annotated image data 80 configured as described above. This constructs the learning model 128.
[0043] Next, in step S5, the model evaluation unit 125 performs processing to evaluate the learning model 128. As shown in Figure 5, the setting screen 200 includes a specification field 202 for specifying judgment conditions for determining whether the detection results (inference results) by the learning model are correct. Based on the judgment conditions specified in the specification field 202, the model evaluation unit 125 can determine whether or not there are undetected and false detections, and can also calculate evaluation indicators such as the undetected rate and the false detection rate.
[0044] The judgment conditions that can be specified in the specification field 202 and the parameters for the judgment may include, for example, one or more of the following. Note that the judgment condition refers to a condition for judging that an inference result obtained by the learning model 128 is correct. (r1) The maximum allowable value of the position error of the vertices of the circumscribed rectangle for each of an inference result obtained by the learning model (e.g., the outer shape of an object) and annotation data (the outer shape of the object). In this case, if the position error is within the maximum allowable value, it is determined that "the inference result is correct". (r2) A threshold for the overlapping area ratio between an inference result obtained by the learning model (e.g., the outer shape of an object) and annotation data (the outer shape of the object). The ratio in this case may be calculated as a ratio of an overlapping area of the object as the inference result to the area of the object according to the annotation data. In this case, if the overlapping area ratio is equal to or higher than the threshold, it is determined that "the inference result is correct". (r3) The maximum allowable value of at least one error among the position (e.g., center of gravity position) and angle of the circumscribed rectangle for each of an inference result obtained by the learning model (e.g., the outer shape of an object) and annotation (the outer shape of the object). In this case, if the position error or the angle error is within the maximum allowable value, it is determined that "the inference result is correct". Note that, in the above description, the circumscribed rectangle is, for example, a rectangle circumscribed to the inference result or the annotation data, and is selected to have the minimum area.
[0045] Note that, although the above judgment conditions (r1) to (r3) are used herein based on the fact that the object appears as a rectangular shape in an image, even when the object appears in any shape in the image, from a similar viewpoint, conditions relating to position error, vertex position error, angle error, or the degree of overlapping area between the arbitrary shape of the object as an inference result and the predetermined shape of the object as annotation data may be set.
[0046] FIG. 5 shows a situation where 5 pixels are specified as the above condition (r1) as a determination condition. In this case, the model evaluation unit 125 may determine that the estimation result obtained by the learning model 128 is correct when the average value of positional errors between the vertices of the circumscribed rectangles for the inference result (e.g., the outer shape of the object) obtained by the learning model 128 and the annotation (the outer shape of the object) is within 5 pixels. Note that conditions (r1) to (r3) may be applied based on an OR condition or an AND condition.
[0047] The model evaluation unit 125 determines occurrence of non-detection and false detection using the above determination conditions. Here, false detection refers to a case where, although a detection result of an object (outer shape of a box) is output as the inference result obtained by the learning model 128, no true object (annotation data) corresponding to the inference result exists. The model evaluation unit 125 may perform brute-force comparison and determination on one inference result on an image and annotation data of all boxes 90 on the image according to the above determination conditions, and determine that the inference result is a false detection when all comparison results indicate "the inference result is incorrect". Non-detection refers to a case where no inference result corresponding to a true object (annotation data) exists. The model evaluation unit 125 may perform brute-force comparison and determination on an annotation of one object (true outer shape of a box) on an image and all inference results on the image according to the above determination conditions, and determine that the annotation (true outer shape of the box) is not detected when all comparison results indicate "the inference result is incorrect".
[0048] The model evaluation unit 125 determines occurrence of non-detection and false detection for each of the images used for evaluation. Based on these determination results, the model evaluation unit 125 can obtain evaluation indices for evaluating the learning model 128, such as a false detection rate and a non-detection rate.
[0049] Next, in step S6, the model evaluation unit 125 outputs the results of the undetection and false detection determination (evaluation results). Figure 6 shows an example of the output screen 300 of the results of the undetection and false detection determination by the model evaluation unit 125 (evaluation results). The model evaluation unit 125 outputs the undetection rate and false detection rate as evaluation results, that is, as indicators representing the quality of the learning model 128. Specifically, as shown in Figure 6, the model evaluation unit 125 has the function of outputting the undetection rate and false detection rate (indicated by symbols 311, 321) based on the total number of images used for evaluation, and the undetection rate and false detection rate (indicated by symbols 312, 322) based on the total number of boxes. The undetection rate and false detection rate based on the total number of images used for evaluation represent the ratio of the number of images in which undetection occurred and the ratio of the number of images in which false detection occurred, relative to the total number of images used for evaluation. The undetected rate and false positive rate, based on the total number of boxes, represent the percentage of boxes that were not detected and the percentage of boxes that were false positives, relative to the total number of boxes in all images used for evaluation.
[0050] The output screen 300 in Figure 6 displays the following evaluation indicators: the total number of images used for evaluation was 100, with 5 images being undetected (undetection rate 5.0%) and 1 image being falsely detected (false detection rate 1.0%). The output screen 300 also displays the total number of boxes in all images used for evaluation being 1000, with 12 boxes being undetected (undetection rate 1.2%) and 1 box being falsely detected (false detection rate 0.1%).
[0051] The non-detection rate and false positive rate are numerical indicators that even an average user can easily understand as to whether the learning model is suitable for the user's system application. Therefore, by looking at the non-detection rate and false positive rate displayed on the output screen 300, the user can easily grasp the quality of the learning model 128. Consequently, the user does not need to examine each detection result image individually to understand the overall situation of the detection results by the learning model.
[0052] Next, in step S7, the model evaluation unit 125 displays images in which undetected and falsely detected items have occurred, for example, on the display unit 23. Figure 7A shows a first example of a detection result image including undetected and falsely detected items (detection result image G1). Detection result image G1 is a two-dimensional image obtained by the visual sensor 70 with the detection results superimposed. Detection result image G1 shows a total of 10 boxes 90 (only some are labeled), and 9 of these boxes have solid line borders F1 (only some are labeled) indicating correct estimation results. One of the boxes (labeled 90a) has a dashed line border (labeled E1) indicating undetected because there is no corresponding estimation result. Also, since there is no corresponding annotation data for the estimation result labeled E2, this estimation result is displayed with a dashed line border E2 indicating a false detection.
[0053] Figure 7B shows a second example of a detection result image (detection result image G2) that includes undetected and falsely detected items. Detection result image G2 is a two-dimensional image captured by the visual sensor 70 with the detection results superimposed. Detection result image G2 shows a total of 10 boxes 90 (only some are labeled), and 8 of these boxes are shown with solid line borders F1 (only some are labeled) as correct detection results. In the example in Figure 7B, the entire area of two adjacent boxes 90b and 90c in the upper left of the image is detected as a single box and shown with a dashed line border E2. That is, border E2 indicates a false detection for which there is no corresponding annotation. As a result, since there are no estimated results corresponding to the annotations of boxes 90b and 90c, the outlines (annotations) of boxes 90b and 90c are shown with dashed line borders E1 indicating undetected items.
[0054] The model evaluation unit 125 may simultaneously display the evaluation result output screen 300 and the detection result images where no detection or false detection occurred on the display unit 23.
[0055] These detection result images allow users to easily understand the quality of the learned model. In particular, the detection result images allow users to identify the circumstances under which undetected or falsely detected objects occur, that is, to pinpoint the causes of undetected or falsely detected objects. Therefore, based on the causes of undetected or falsely detected objects obtained from the detection result images, users can take necessary measures to improve the recognition rate of the learned model, such as retraining with new training images or improving the operating environment of the robot system (lighting, camera placement, etc.). By configuring the system to display only images where undetected or falsely detected objects occur, users can check only the problematic images without having to review all of the detection result images.
[0056] Note that the learning model evaluation process shown in Figure 3 is an example of processing starting from the stage where the user collects images to be used for training. When annotated image data and a learning model are prepared in advance, the process for evaluating the learning model executed by the image processing device 20 can also be as follows: (Procedure 1) The data acquisition unit 123 acquires sets of image data and annotation data representing the objects in the image data for multiple images. (Procedure 2) The learning model acquisition unit 124 acquires a learning model for detecting objects from image information. (Procedure 3) The model evaluation unit 125 evaluates the learning model by determining the occurrence of undetected and falsely detected objects based on the object inference results obtained by inputting at least a portion of the image data for the multiple acquired images into the learning model and the annotation data of the input image data.
[0057] As described above, this embodiment provides an evaluation index that allows even ordinary users to easily understand whether or not a learning model used in a robot system application is appropriate, and thereby users can easily understand whether or not the learning model is appropriate.
[0058] Since the images used for training can also be used as evaluation images to evaluate the trained model, it becomes unnecessary to actually operate the robot system and perform object detection tests to confirm the performance of the trained model.
[0059] Here, we will explain the flexibility of the system configuration in realizing the functions of the image processing device 20 in the above-described embodiment. In the above-described embodiment, we described an example configuration in which the functions related to the evaluation of the learning model are consolidated in the image processing device 20, but this example configuration can be modified in various ways. For example, there may be an example configuration in which some or all of the functions in the image processing device 20 are placed in the teaching operation panel 40 or the robot control device 50. In other words, the functions of the image processing device 20 in the above-described embodiment can be mounted on various types of information processing devices. The function blocks shown in the function block of Figure 2 may be integrated as appropriate.
[0060] The model evaluation unit 125 may display the numerical information of the evaluation results or the detection result images shown in Figures 6, 7A, and 7B on the display unit 43 of the teaching operation panel 40.
[0061] If the learning model is provided to the robot system 100 from an external device, the image processing device 20 does not need to have the function of a learning unit 127.
[0062] The model evaluation unit 125 may further include a function to determine whether the evaluation value (at least one of the non-detection rate and false detection rate) is suitable for the robot system according to a preset threshold, and to automatically apply the learning model to the robot system if the determination result is OK.
[0063] In the above-described embodiment, a configuration was described in which the model evaluation unit 125 performs an evaluation on one learning model. However, the model evaluation unit 125 may be configured to calculate the above-described evaluation values (at least one of the non-detection rate and false detection rate) for each of multiple learning models. In this case, the user can select the learning model that is most suitable for the robot system from among the multiple learning models based on the evaluation values. Alternatively, the model evaluation unit 125 may operate to automatically apply the most suitable learning model to the robot system based on the evaluation values.
[0064] The configuration of the above-described embodiment can be applied to various industrial machine systems.
[0065] In the functional block diagram of Figure 2, each functional block described as a function of an image processing device, a robot control device, or a teaching control panel may be realized by one or more processors of these devices executing various software stored in a memory device, or in this case, part of the function may be made up of hardware such as discrete circuits (i.e., the functional block may be realized by a combination of a processor and discrete circuits), or the functions shown in the functional block diagram may be realized by a hardware-based configuration such as an ASIC (Application Specific Integrated Circuit).
[0066] The computer programs for executing various processes such as the learning model evaluation process in the above-described embodiments, or the computer programs for executing the processes of each part of the processor of the robot control device, may be provided in the form of program products recorded on various computer-readable recording media (for example, semiconductor memories such as ROM, EEPROM, and flash memory, magnetic recording media, or optical recording media such as CD-ROM and DVD-ROM).
[0067] While this disclosure has been described in detail, it is not limited to the individual embodiments described above. These embodiments can be added, replaced, modified, partially deleted, etc., in any way that does not depart from the gist of this disclosure or from the spirit of this disclosure derived from the claims and their equivalents. Furthermore, these embodiments can be implemented in combination. For example, the order of operations and processes in the embodiments described above are given as examples only and are not limited thereto. The same applies when numerical values or mathematical formulas are used in the description of the embodiments described above.
[0068] The following additional notes are provided with respect to the above embodiments and modified examples. (Note 1) An information processing device (20) comprising: a data acquisition unit (123) that acquires a set of image data and annotation data representing an object in the image data for a plurality of images; a learning model acquisition unit (124) that acquires a learning model for detecting an object from image information; and a model evaluation unit (125) that evaluates the learning model by determining the occurrence of undetected and falsely detected objects based on the inference result of an object obtained by inputting at least a portion of the acquired image data for the plurality of images into the learning model and the annotation data of the input image data. (Note 2) The information processing device (20) according to Note 1, further comprising: a learning unit (127) that generates the learning model using at least a portion of the acquired image data and annotation data for the plurality of images as learning data, wherein the learning model acquisition unit (124) acquires the learning model generated by the learning unit. (Note 3) The information processing device (20) according to Note 2, wherein the learning unit (127) uses all of the acquired image data and annotation data for the plurality of images for learning, and the model evaluation unit (125) uses a portion of the acquired image data and annotation data for the plurality of images for evaluation. (Note 4) The information processing device (20) according to Note 2, wherein the learning unit (127) uses a portion of the acquired image data and annotation data for the plurality of images for learning, and the model evaluation unit (125) uses a portion of the remaining acquired image data and annotation data for the plurality of images for evaluation. (Note 5) The information processing device (20) according to Note 2, further comprising a setting unit (126) for setting conditions for evaluation by the model evaluation unit (125).(Note 6) The information processing device (20) according to Note 5, wherein the conditions include at least one of the following: (1) whether the acquired image data and annotation data for the plurality of images are used for both learning by the learning unit and evaluation by the model evaluation unit; (2) the ratio of the image data and annotation data used by the model evaluation unit to the acquired image data and annotation data for the plurality of images, or the image data and annotation data used by the learning unit for learning; and (3) a determination condition for determining whether the inference result is correct. (Note 7) The information processing device (20) according to Note 6, wherein the determination condition includes a condition relating to the degree of overlap of position error, vertex position error, angle error, or area between an arbitrary shape of the object as the inference result and a predetermined shape of the object as annotation data. (Note 8) The object is depicted as a rectangular object on the image corresponding to the image data, and the determination condition includes conditions relating to the degree of overlap in position, vertex position, angle, or area between the bounding rectangle of an arbitrary shape of the object as the inference result and the rectangular shape of the object as annotation data, as described in the information processing device (20) according to any one of the items in Notes 1 to 8, wherein the model evaluation unit (125) determines at least one of the undetected rate and the false detection rate based on the number of images used for evaluation or the total number of objects in the images used for evaluation, and displays it on the display screen, as described in the information processing device (20) according to any one of the items in Notes 1 to 9, wherein the model evaluation unit (125) displays on the display screen an image corresponding to the image data in which at least one of the undetected or false detected objects has occurred, as described in the information processing device (20) according to any one of the items in Notes 1 to 9.(Note 11) A robot system (100) comprising: a visual sensor (70); a robot (10); a robot control device (50) for controlling the robot; a data acquisition unit (123) for acquiring sets of image data and annotation data representing objects in the image data for multiple images; a learning model acquisition unit (124) for acquiring a learning model for detecting objects from image information provided by the visual sensor (70); and a model evaluation unit (125) for evaluating the learning model by determining the occurrence of undetected and falsely detected objects based on the object inference result obtained by inputting at least a portion of the acquired image data for the multiple images into the learning model and the annotation data of the input image data. (Note 12) A computer program that causes a computer processor to execute the following steps: a procedure for acquiring a set of image data and annotation data representing an object in the image data for multiple images; a procedure for acquiring a learning model for detecting an object from image information; and a procedure for evaluating the learning model by determining the occurrence of undetected and falsely detected objects based on the object inference result obtained by inputting at least a portion of the acquired image data for the multiple images into the learning model and the annotation data of the input image data.
[0069] 10 Robot 15 Hand 20 Image processing device 21 Processor 22 Memory unit 23 Display unit 24 Operation unit 40 Teaching control panel 43 Display unit 50 Robot control device 70 Vision sensor 100 Robot system 121 Vision sensor control unit 122 Detection unit 123 Data acquisition unit 124 Learning model acquisition unit 125 Model evaluation unit 126 Setting unit 127 Learning unit 151 Motion control unit
Claims
1. An information processing device comprising: a data acquisition unit that acquires sets of image data and annotation data representing objects in the image data for multiple images; a learning model acquisition unit that acquires a learning model for detecting objects from image information; and a model evaluation unit that evaluates the learning model by determining the occurrence of undetected and falsely detected objects based on the object inference result obtained by inputting at least a portion of the acquired image data for the multiple images into the learning model and the annotation data of the input image data.
2. The information processing apparatus according to claim 1, comprising a learning unit that generates the learning model using at least a portion of the image data and annotation data of the plurality of images acquired as learning data, and the learning model acquisition unit that acquires the learning model generated by the learning unit.
3. The information processing apparatus according to claim 2, wherein the learning unit uses all of the acquired image data and annotation data for the plurality of images for learning, and the model evaluation unit uses a portion of the acquired image data and annotation data for the plurality of images for evaluation.
4. The information processing apparatus according to claim 2, wherein the learning unit uses a portion of the acquired image data and annotation data for the plurality of images for learning, and the model evaluation unit uses the remaining portion of the acquired image data and annotation data for the plurality of images for evaluation.
5. The information processing apparatus according to claim 2, further comprising a setting unit for setting conditions for evaluation by the model evaluation unit.
6. The information processing apparatus according to claim 5, wherein the conditions include at least one of the following: (1) whether the acquired image data and annotation data for the plurality of images are used for both learning by the learning unit and evaluation by the model evaluation unit; (2) the ratio of the image data and annotation data used by the model evaluation unit to the acquired image data and annotation data for the plurality of images, or the image data and annotation data used by the learning unit for learning; and (3) a determination condition for determining whether the inference result is correct.
7. The information processing apparatus according to claim 6, wherein the determination conditions include conditions relating to positional errors, vertex position errors, angular errors, or the degree of area overlap between an arbitrary shape of the object as an inference result and a predetermined shape of the object as annotation data.
8. The information processing apparatus according to claim 6, wherein the object is depicted as a rectangular object on the image corresponding to the image data, and the determination condition includes conditions relating to positional errors, vertex position errors, angular errors, or the degree of overlap in area between the bounding rectangle of an arbitrary shape of the object as the inference result and the rectangular shape of the object as annotation data.
9. The information processing apparatus according to any one of claims 1 to 8, wherein the model evaluation unit determines at least one of the non-detection rate and the false detection rate based on the number of images used for evaluation or the total number of objects in the images used for evaluation, and displays them on a display screen.
10. The information processing apparatus according to any one of claims 1 to 9, wherein the model evaluation unit displays an image on a display screen that corresponds to image data in which at least one of the cases of undetected or falsely detected has occurred.
11. A robot system comprising: a visual sensor; a robot; a robot control device for controlling the robot; a data acquisition unit for acquiring sets of image data and annotation data representing objects in the image data for multiple images; a learning model acquisition unit for acquiring a learning model for detecting objects from image information provided by the visual sensor; and a model evaluation unit for evaluating the learning model by determining the occurrence of undetected and falsely detected objects based on the object inference result obtained by inputting at least a portion of the acquired image data for the multiple images into the learning model and the annotation data of the input image data.
12. A computer program that causes a computer processor to execute the following steps: a procedure for acquiring a set of image data and annotation data representing an object in the image data for multiple images; a procedure for acquiring a learning model for detecting an object from image information; and a procedure for evaluating the learning model by determining the occurrence of undetected and falsely detected objects based on the object inference result obtained by inputting at least a portion of the acquired image data for the multiple images into the learning model and the annotation data of the input image data.