Processing device, program, and robot system
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- KYOCERA CORP
- Filing Date
- 2026-01-21
- Publication Date
- 2026-07-30
Smart Images

Figure JP2026001747_30072026_PF_FP_ABST
Abstract
Description
Processing Device, Program, and Robot System
[0001] The present disclosure relates to a learning model.
[0002] Patent Document 1 describes a technology related to a learning model.
[0003] Japanese Patent Application Laid-Open No. 2024-20743
[0004] A processing device, a program, and a robot system are disclosed. In one embodiment, the processing device includes a processing unit. The processing unit compares the inference accuracies of a first learning model and a second learning model using a CG (Computer Graphics) image.
[0005] Also, in one embodiment, the program is a program for causing a computer device to function as the processing unit of the above-described processing device.
[0006] Also, in one embodiment, the robot system includes the above-described processing device and a robot controlled by the processing device.
[0007] FIG. 1 is a schematic diagram showing an example of the configuration of the processing device. FIG. 2 is a schematic diagram showing an example of the configuration of the processing unit. FIG. 3 is a schematic diagram showing an example of a state where a camera photographs a plurality of objects in a tray. FIG. 4 is a flowchart showing an example of the operation of the processing unit. FIG. 5 is a schematic diagram showing an example of a photographed image and a mask. FIG. 6 is a schematic diagram showing an example of a photographed image and a mask. FIG. 7 is a schematic diagram showing an example of the configuration of the processing unit. FIG. 8 is a schematic diagram showing an example of a CG object image and a CG reference mask. FIG. 9 is a schematic diagram showing an example of a first partial image and a first mask. FIG. 10 is a schematic diagram showing an example of a second partial image and a second mask. FIG. 11 is a schematic diagram showing an example of a robot system. FIG. 12 is a schematic diagram for explaining an example of the operation of the processing unit. FIG. 13 is a schematic diagram for explaining an example of the operation of the processing unit. FIG. 14 is a schematic diagram for explaining an example of the operation of the processing unit.
[0008] Figure 1 is a schematic diagram showing an example of the processing unit 1. The processing unit 1 can compare the estimation accuracy of a first learning model and a second learning model using, for example, computer graphics (CG) images. CG images can also be called, for example, composite images. The learning models can also be called, for example, machine learning models.
[0009] As shown in Figure 1, the processing unit 1 comprises, for example, a processing unit 2, a storage unit 3, and an interface 4. The processing unit 1 can also be described as, for example, a computer device. Alternatively, the processing unit 1 can also be described as, for example, a processing circuit.
[0010] Interface 4 can communicate with devices outside of the processing unit 1 (also simply called external devices). Processing unit 2 can communicate with external devices through interface 4. Interface 4 can also be called, for example, an interface circuit, a communication unit, or a communication circuit. Interface 4 may communicate with external devices via wired communication or wireless communication.
[0011] The processing unit 2 can comprehensively manage the operation of the processing unit 1 by controlling other components of the processing unit 1. The processing unit 2 can also be called, for example, a control unit. Alternatively, the processing unit 2 can also be called, for example, a processing circuit or a control circuit. The processing unit 2 includes at least one processor to provide control and processing capabilities for performing various functions, as will be described in more detail below.
[0012] According to various embodiments, at least one processor may be implemented as a single integrated circuit (IC) or as a plurality of communicably connected integrated circuits IC and / or discrete circuits. At least one processor can be implemented according to various known techniques.
[0013] In one embodiment, the processor includes one or more circuits or units configured to perform one or more data computation procedures or processes by, for example, executing instructions stored in associated memory. In other embodiments, the processor may be firmware (e.g., discrete logic components) configured to perform one or more data computation procedures or processes.
[0014] According to various embodiments, the processor may include one or more processors, controllers, microprocessors, microcontrollers, application-specific integrated circuits (ASICs), digital signal processing devices, programmable logic devices, field-programmable gate arrays, or any combination of these devices or configurations, or other known combinations of devices and configurations, and may perform the functions described below.
[0015] The processing unit 2 may, for example, include a CPU (Central Processing Unit) as a processor. The processing unit 2 can compare the estimation accuracy of the first and second learning models using computer graphics (CG) images.
[0016] The storage unit 3 may include non-temporary recording media that can be read by the CPU of the processing unit 2, such as ROM (Read Only Memory) and RAM (Random Access Memory). The storage unit 3 stores, for example, a program 3a for controlling the processing unit 1. Various functions of the processing unit 2 are realized, for example, by the CPU of the processing unit 2 executing the program 3a in the storage unit 3.
[0017] The configuration of the processing unit 1 is not limited to the above example. For example, the processing unit 2 may include multiple CPUs. The processing unit 2 may also include at least one DSP (Digital Signal Processor). Furthermore, all or some of the functions of the processing unit 2 may be implemented by hardware circuits that do not require software to implement those functions. The storage unit 3 may also include a computer-readable, non-temporary recording medium other than ROM and RAM. The storage unit 3 may include, for example, a small hard disk drive and an SSD (Solid State Drive).
[0018] Furthermore, the processing unit 1 may include a display unit controlled by the processing unit 2. The display unit may be, for example, a liquid crystal display, an organic electroluminescent (EL) display, or a plasma display. The display unit of the processing unit 1 may, for example, display an image generated by the processing unit 2.
[0019] Furthermore, the processing unit 1 may include an input unit for receiving input from the user. The input unit may include, for example, a mouse and a keyboard. The input unit may also include a touch sensor for receiving user touch operations. The input unit may also include a microphone for receiving user voice input.
[0020] Furthermore, the processing unit 1 may be composed of multiple computer devices. Alternatively, the processing unit 1 may be a cloud server. In this case, the interface 4 of the processing unit 1 may communicate with external devices via a network, including the Internet.
[0021] <Example of Processing Unit Configuration> Figure 2 is a schematic diagram showing an example of the configuration of the processing unit 2. As shown in Figure 2, the processing unit 2 has, for example, a determination unit 20, a first learning model 21a, and a second learning model 21b as functional blocks. The CPU of the processing unit 2 executes the program 3a in the storage unit 3, thereby forming the determination unit 20, the first learning model 21a, and the second learning model 21b in the processing unit 2.
[0022] The determination unit 20 compares the estimation accuracy of the first learning model 21a and the second learning model 21b. Each of the first learning model 21a and the second learning model 21b is capable of recognizing objects (also called real objects) in a captured image, for example, based on an image acquired by a camera. Each of the first learning model 21a and the second learning model 21b can also be called a trained model, as they are both trained learning models. The captured image can also be called a camera image. The captured image may be a color image such as an RGB image, or a grayscale image. Hereafter, when describing the common content of the first learning model 21a and the second learning model 21b, they will each be referred to as learning model 21.
[0023] The learning model 21 is composed of, for example, a neural network that performs instance segmentation. The learning model 21 may be, for example, a Mask Scoring R-CNN or a Mask R-CNN. R-CNN is an abbreviation for Region Based Convolutional Neural Networks. The learning model 21 may also be any other learning model that performs instance segmentation.
[0024] The learning model 21 can, for example, generate a mask of an object in a captured image. It can also be said that the learning model 21 infers the mask of an object in a captured image. A mask is a type of image, and can also be called a mask image. The learning model 21 sets a mask for an object in a captured image. The position and shape of the mask set for an object in a captured image, generated by the learning model 21, correspond to the position and shape of the object in the captured image. The recognition result of object 50 by the learning model 21 includes the mask of object 50.
[0025] The accuracy of the mask generated by the learning model 21 can be said to be either the recognition accuracy of the learning model 21 for the object 50, or the inference accuracy of the learning model 21. The accuracy of the mask can be said to be either the inference accuracy of the mask, or the generation accuracy of the mask. The determination unit 20 compares, for example, the recognition accuracy of the first learning model 21a and the second learning model 21b. For example, the determination unit 20 compares the accuracy of the masks generated by the first learning model 21a and the second learning model 21b.
[0026] The learning model 21 may generate bounding boxes for objects in the captured image. The learning model 21 may also generate a classification score representing the accuracy of the classification of the regions within the bounding boxes (in other words, the confidence level, likelihood, or certainty of the classification). The classification score indicates the likelihood that the region within the bounding box is the target object. The classification score is, for example, a numerical value between 0 and 1. A higher classification score indicates a higher likelihood that the region within the bounding box is the target object. The learning model 21 may also generate a mask score representing the accuracy of the mask (in other words, the inference accuracy of the mask). The mask score is, for example, a numerical value between 0 and 1. A higher mask score indicates higher mask accuracy. Hereafter, "captured image" simply refers to the captured image input to the learning model 21.
[0027] The camera 30 that generates the captured image acquires the image by photographing the multiple objects 50 scattered inside the tray 40 from directly above the tray 40, as shown in Figure 3. The captured image shows the multiple objects 50 scattered inside the tray 40. The captured image shows the multiple objects 50 inside the tray 40 as seen from directly above the tray 40. In the example in Figure 3, the objects 50 are cylindrical, but the shape of the objects 50 is not limited to this.
[0028] The learning model 21 can individually recognize each of the multiple objects 50 that appear in the captured image. The objects 50 can also be called the objects to be recognized. The learning model 21 generates a mask for each of the multiple objects 50 that appear in the captured image. Hereafter, the mask generated by the first learning model 21a may be referred to as the first mask. The mask generated by the second learning model 21b may be referred to as the second mask.
[0029] The second learning model 21b is, for example, a further trained version of the first learning model 21a. Therefore, the original learning models for the first learning model 21a and the second learning model 21b are the same. The first learning model 21a may be a further trained version of a pre-trained model. The second learning model 21b may also be further trained. Further training can also be called retraining.
[0030] The learning model 21 may be further trained, for example, when the shooting environment of the camera 30 changes. The shooting environment can also be said to be, for example, the environment surrounding the camera 30 and the object 50. Since the recognition accuracy (in other words, estimation accuracy) of the learning model 21 may decrease when the shooting environment changes, the recognition accuracy of the learning model 21 can be improved by further training the learning model 21 when the shooting environment changes.
[0031] For example, the learning model 21 may be further trained when the brightness of the shooting environment changes. In this case, the second learning model 21b is the first learning model 21a that has been further trained when the brightness of the shooting environment changes. The learning model 21 may be further trained when the shooting environment becomes brighter, or when the shooting environment becomes darker. When the learning model 21 is further trained, multiple images taken in shooting environments with different brightness levels may be used as training images.
[0032] Furthermore, the learning model 21 may be further trained when the location of the camera 30 and tray 40 changes and the shooting environment changes. Also, the learning model 21 may be further trained when the environment of the room in which the camera 30 and tray 40 are located changes and the shooting environment changes. For example, the learning model 21 may be further trained when the door to the room in which the camera 30 and tray 40 are located is left open and light from outside the room enters the room. Also, the learning model 21 may be further trained when the way light enters the room through the window of the room in which the camera 30 and tray 40 are located changes due to a change in season.
[0033] Furthermore, the learning model 21 may be further trained when camera 30 is replaced with another camera.
[0034] The additional training of the learning model 21 may be performed by the processing unit 1, or by a device other than the processing unit 1. The device that performs the additional training of the learning model 21 can be said to be a learning device.
[0035] The recognition results (in other words, inference results) of the learning model 21 may be used in any way. For example, the robot may be controlled based on the recognition results of the learning model 21. For example, the robot may be controlled based on the recognition results of the learning model 21 so that it holds the object 50 in the tray 40. In this case, the object 50 can be said to be the work object that the robot will work on. However, the methods of using the recognition results of the learning model 21 are not limited to this.
[0036] <Example of operation of the determination unit> The determination unit 20 performs a decision process to determine which of the first learning model 21a and the second learning model 21b to use, based on a predetermined number of captured images. Hereafter, each of the predetermined number of captured images used in the decision process may be referred to as a determination image. The predetermined number of determination images may include multiple determination images taken in different environments.
[0037] Figure 4 is a flowchart showing an example of the decision process. In the decision process, steps s1 and s2 are executed individually for each of a predetermined number of judgment images. The first learning model 21a sets multiple first masks for each of the predetermined number of judgment images, based on the judgment image. Similarly, the second learning model 21b sets multiple second masks for each of the predetermined number of judgment images, based on the judgment image, for each of the multiple objects 50 depicted in the judgment image. Hereafter, one judgment image being explained may be referred to as the target judgment image. Note that the judgment image input to the first learning model 21a and the judgment image input to the second learning model 21b may be the same image.
[0038] In step s1, the determination unit 20 performs a degree of agreement acquisition process to acquire the degree of agreement between the first mask and the second mask of each of the multiple objects 50 captured in the target determination image. In step s2, the determination unit 20 performs a mask accuracy comparison process to compare the accuracy of the first mask and the second mask based on the degree of agreement between the first mask and the second mask of each of the multiple objects 50 captured in the target determination image, based on the degree of agreement between the first mask and the second mask acquired in the degree of agreement acquisition process. In the mask accuracy comparison process, if the degree of agreement between the first mask and the second mask acquired in the degree of agreement acquisition process is greater than a threshold, the determination unit 20 determines that the accuracy of the first mask and the second mask is about the same. On the other hand, if the degree of agreement between the first mask and the second mask acquired in the degree of agreement acquisition process is less than or equal to a threshold, the determination unit 20 determines which of the first mask and the second mask has higher accuracy.
[0039] When the degree of agreement determination process in step s1 and the mask accuracy comparison process in step s2 are performed for each of a predetermined number of judgment images (determined as YES in step s3), step s4 is executed. In step s4, the determination unit 20 performs a usage target determination process to determine which of the first learning model 21a and the second learning model 21b to use, based on the results of the mask accuracy comparison process in step s2 for the predetermined number of judgment images. For example, if the robot is controlled based on the recognition results of the learning model 21, the determination unit 20 can be said to determine which of the first learning model 21a and the second learning model 21b to use for robot control, based on the results of the mask accuracy comparison process in step s2 for a predetermined number of judgment images. For example, the recognition results of the learning model 21 determined by the determination unit 20 to be used are used for controlling the robot 10.
[0040] The following provides a detailed explanation of examples of the matching degree determination process, mask accuracy comparison process, and target determination process.
[0041] <Example of matching degree acquisition process> In the matching degree acquisition process, the determination unit 20 acquires the IoU of the first mask and the second mask for all combinations of the first mask and the second mask for the target determination image, using the multiple first masks generated by the first learning model 21a and the multiple second masks generated by the second learning model 21b. IoU is an abbreviation for Intersection Over Union. The IoU of the first mask and the second mask indicates the degree of overlap between the first mask and the second mask. The IoU of the first mask and the second mask is a value obtained by dividing the area of the logical AND of the first mask and the second mask on the target determination image by the area of the logical OR of the first mask and the second mask on the target determination image. The IoU is a value between 0 and 1. If the first mask and the second mask are a perfect match on the target determination image, the IoU is 1.
[0042] Next, for each of the plurality of second masks for the target determination image, the determination unit 20 sets the first mask with the largest IoU with the second mask among the plurality of first masks for the target determination image as the first mask having the same position and shape as the second mask. Hereinafter, the combination of the second mask and the first mask having the same position and shape as the second mask may be referred to as a similar mask pair. In the consistency determination process, the determination unit 20 identifies a plurality of similar mask pairs. Note that even for the first mask and the second mask (i.e., the similar mask pair) having the largest IoU, if the IoU value is small, there is a possibility that the first mask and the second mask do not represent the same object 50, or that there is a defect. Therefore, first masks and second masks (i.e., similar mask pairs) having an IoU smaller than a predetermined value may be excluded from subsequent processing.
[0043] The determination unit 20 determines that, for each of the plurality of similar mask pairs, the first mask and the second mask constituting the similar mask pair are masks for the same object 50 shown in the target determination image. It can also be said that the determination unit 20 determines that the second mask and the first mask having the same position and shape as the second mask are masks for the same object 50. As a result, for each of the plurality of objects 50 shown in the target determination image, a pair of the first mask and the second mask of the object 50 is identified. Hereinafter, the pair of the first mask and the second mask of the same object 50 may be referred to as the same object mask pair. In this example, the similar mask pair becomes the same object mask pair.
[0044] In the consistency acquisition process, the determination unit 20 identifies a plurality of same object mask pairs corresponding to the plurality of objects 50 shown in the target determination image. The same object mask pair corresponding to a certain object 50 shown in the target determination image is composed of the first mask and the second mask of the certain object 50.
[0045] In the consistency acquisition process, for each of a plurality of identical object mask pairs, the determination unit 20 uses the IoU of the first mask and the second mask of the identical object 50 that constitute the identical object mask pair as the consistency between the first mask and the second mask. As a result, for each of the plurality of objects 50 shown in the target determination image, the consistency between the first mask and the second mask of the object 50 is acquired. It can be said that the consistency between the first mask and the second mask of the identical object 50 is the consistency of the shapes of the first mask and the second mask of the identical object 50.
[0046] Hereinafter, the consistency of an identical object mask pair means the consistency between the first mask and the second mask that constitute the identical object mask pair. The consistency of an identical object mask pair is a value between 0 and 1.
[0047] FIG. 5 is a schematic diagram showing an example of the first mask 200a and the second mask 200b that constitute an identical object mask pair with a large consistency (for example, about 0.97). FIG. 6 is a schematic diagram showing an example of the first mask 200a and the second mask 200b that constitute an identical object mask pair with a not-so-large consistency (for example, about 0.90). A part of an example of the target determination image 300 is shown above FIG. 5. In the center and below FIG. 5, examples of the first mask 200a and the second mask 200b of the identical object 50 shown in the upper target determination image 300 are shown respectively. Also, in the center and below FIG. 5, the bounding box 210a generated by the first learning model 21a and the bounding box 210b generated by the second learning model 21b are shown respectively. The same applies to FIG. 6.
[0048] In the example of FIG. 5, the consistency between the first mask 200a and the second mask 200b of the identical object 50 is large, and the shapes of the first mask 200a and the second mask are substantially the same. On the other hand, in the example of FIG. 6, the consistency between the first mask 200a and the second mask 200b of the identical object 50 is not so large, and the shapes of the first mask 200a and the second mask are slightly different.
[0049] <Example of mask accuracy comparison processing> In the mask accuracy comparison processing, the determination unit 20 compares the accuracy of the first mask and the second mask that constitute the identical object mask pair for each of the multiple identical object pairs identified in the matching degree acquisition processing.
[0050] The determination unit 20 determines that the accuracy of the first mask and the second mask constituting the identical mask pair is about the same if the degree of agreement of the identical object mask pair acquired by the agreement acquisition unit is greater than a threshold. For example, suppose the threshold is 0.95. In this case, the determination unit 20 determines that the accuracy of the first mask and the accuracy of the second mask are about the same if the degree of agreement of the first mask and the second mask constituting the identical object mask pair is greater than 0.95. In the example of Figure 5 above, it is determined that the accuracy of the first mask 200a and the accuracy of the second mask 200b are about the same. As in this example, if the second learning model 21b is an additionally trained version of the first learning model 21a, and the degree of agreement of the first and second masks of a certain object 50 is greater than a threshold, then with respect to that certain object 50, the recognition accuracy of the second learning model 21b is not degraded compared to the recognition accuracy of the first learning model 21a. In other words, if the shapes of the first mask and the second mask of an object 50 are similar, then the quality of the mask generated by the second learning model 21b for that object 50 is not inferior to the quality of the mask generated by the first learning model 21a.
[0051] On the other hand, if the degree of agreement of a pair of identical object masks acquired by the degree of agreement acquisition unit is below a threshold, the determination unit 20 compares the accuracy of the first mask and the second mask constituting the pair of identical object masks based on the CG image. For example, suppose the threshold is 0.95. In this case, if the degree of agreement of the first mask and the second mask constituting the pair of identical object masks is 0.95 or less, the determination unit 20 determines that the degree of agreement of the first mask and the second mask is low, or that the first mask and the second mask do not match, and compares the accuracy of the first mask and the accuracy of the second mask based on the CG image. In the example of Figure 6 above, since the degree of agreement of the first mask 200a and the second mask 200b is 0.95 or less, the accuracy of the first mask 200a and the accuracy of the second mask 200b are compared based on the CG image. The determination unit 20 also compares the accuracy of the first mask and the accuracy of the second mask based on the CG image when the shapes of the first and second masks constituting the same object mask pair are different from each other.
[0052] In the mask accuracy comparison process, the determination unit 20 performs a process (also called CG usage comparison process) that compares the accuracy of the first mask and the second mask constituting the identical object mask pair based on the CG image for each identical object mask pair whose degree of agreement is below a threshold.
[0053] Hereafter, an object 50 corresponding to an identical object mask pair with a degree of matching greater than a threshold may be referred to as the first object 50. Similarly, an object 50 corresponding to an identical object mask pair with a degree of matching less than or equal to a threshold may be referred to as the second object 50. An identical object mask pair composed of the first and second masks of a given object 50 can be said to be an identical object mask pair corresponding to that object 50.
[0054] Furthermore, in the explanation of the CG usage comparison process, the first mask refers to the first mask that constitutes a pair of identical object masks with a degree of matching below a threshold, that is, the first mask of the second object 50. Furthermore, in the explanation of the CG usage comparison process, the second mask refers to the second mask that constitutes a pair of identical object masks with a degree of matching below a threshold, that is, the second mask of the second object 50.
[0055] In the CG usage comparison process, a reference mask is used for each second object 50, which serves as the basis for the first and second masks of the second object 50. The reference mask is, for example, a CG image or a mask based on a CG image. Hereafter, the reference mask will be referred to as the CG reference mask. Furthermore, when we say the CG reference mask of the second object 50, we mean the CG reference mask that serves as the basis for the first and second masks of the second object 50. The CG reference mask of the second object 50 can be said to be the ideal mask of the second object 50, the correct mask of the second object 50, or it can be said to represent the desired appearance or state of the mask of the second object 50.
[0056] In the CG usage comparison process, the determination unit 20 obtains a first similarity between the CG reference mask of the second object 50 and the first mask for each second object 50, as the accuracy of the first mask of the second object 50. In addition, in the CG usage comparison process, the determination unit 20 obtains a second similarity between the CG reference mask of the second object 50 and the second mask for each second object 50, as the accuracy of the second mask of the second object 50. Hereafter, the first similarity of the second object 50 refers to the first similarity between the CG reference mask of the second object 50 and the first mask of the second object 50. Similarly, the second similarity of the second object 50 refers to the second similarity between the CG reference mask of the second object 50 and the second mask of the second object 50. Furthermore, the second object 50 being explained may be referred to as the target second object 50.
[0057] Based on the results of the CG usage comparison process, the determination unit 20 obtains, for a plurality of second objects 50, the number of second objects 50 in which the accuracy of the first mask is higher than the accuracy of the second mask (also called the first number) and the number of second objects 50 in which the accuracy of the second mask is higher than the accuracy of the first mask (also called the second number). The first number can also be said to be the number of second objects 50 in which the first similarity is greater than the second similarity among the plurality of second objects 50. The second number can also be said to be the number of second objects 50 in which the second similarity is greater than the first similarity among the plurality of second objects 50. Once the first number and the second number have been obtained for the target determination image, the mask accuracy comparison process is completed and step s3 is executed.
[0058] In the CG usage comparison process, the determination unit 20 determines a CG reference mask for each second object 50 from a plurality of CG masks. A CG mask is a CG image. One CG mask is, for example, the mask of one object 50 that appears in a CG image (also called a randomly stacked CG image) that shows a plurality of randomly stacked objects 50. The CG mask can be said to be the ideal mask of the object 50 that appears in the randomly stacked CG image, the correct mask of the object 50, and the ideal appearance or state of the mask of the object 50. The object 50 that appears in the randomly stacked CG image is a CG image.
[0059] For example, 3DCG technology is used to generate a randomly stacked CG image. 3DCG is an abbreviation for 3-Dimensional Computer Graphics. When a randomly stacked CG image is generated, 3DCG technology is used to model object 50 based on its 3D CAD data, and a 3DCG model of object 50 (also called an object model) is generated. The object model can also be said to be a CG image. Then, multiple object models are randomly stacked in a virtual 3D space. For example, using simulation software, multiple object models are made to fall one by one in a virtual 3D space, and multiple object models are stacked on top of each other. After that, the randomly stacked multiple object models are photographed from above by a virtual camera, and the resulting virtual image is considered a randomly stacked CG image. However, the method of generating a randomly stacked CG image is not limited to this.
[0060] Multiple CG masks are generated based on multiple randomly stacked CG images, each showing a different arrangement of multiple objects 50 (in other words, multiple object models). In each randomly stacked CG image, the mask of each object 50 (in other words, each object model) depicted in that randomly stacked CG image becomes the CG mask. Each CG mask is a CG image generated on a computer.
[0061] Multiple CG masks are stored in the storage unit 3. Multiple CG masks may be generated by the processing unit 2, or they may be generated by a device other than the processing unit 1. The determination unit 20 determines the third similarity between the CG mask and the first mask for all combinations of the multiple CG masks in the storage unit 3 and the first masks of the multiple second objects 50. Similarly, the determination unit 20 determines the fourth similarity between the CG mask and the second mask for all combinations of the multiple CG masks in the storage unit 3 and the second masks of the multiple second objects 50.
[0062] As shown in Figure 7, the determination unit 20 includes a feature acquisition unit 20a that acquires feature quantities from an image. The feature acquisition unit 20a can also be called a feature extraction unit that extracts feature quantities from an image. The determination unit 20 uses the feature acquisition unit 20a in the CG usage comparison process.
[0063] The feature acquisition unit 20a may be, for example, an encoder provided by an autoencoder. An autoencoder is a type of learning model. An autoencoder may be composed of a neural network, such as a convolutional neural network. An autoencoder comprises an encoder and a decoder. The encoder converts the input image it receives into feature quantities. The encoder outputs the feature quantities of the input image. The feature quantities output by the encoder are represented as vectors. Feature quantities can also be called feature vectors. The decoder reconstructs the input image based on the feature vectors (in other words, feature quantities) of the input image output from the encoder. The learning of the autoencoder may be performed in the processing unit 2, or it may be performed in a device other than the processing unit 1.
[0064] The feature acquisition unit 20a acquires the feature quantities of each CG mask. The feature acquisition unit 20a acquires the feature quantities of the first mask of each second object 50. The feature acquisition unit 20a acquires the feature quantities of the second mask of each second object 50.
[0065] The determination unit 20 calculates the cosine similarity between the feature quantities (in other words, feature vectors) of the CG mask and the feature quantities (in other words, feature vectors) of the first mask for all combinations of the CG mask and the first mask, and uses this as the third similarity between the CG mask and the first mask. The determination unit 20 also calculates the cosine similarity between the feature quantities (in other words, feature vectors) of the CG mask and the feature quantities (in other words, feature vectors) of the second mask for all combinations of the CG mask and the second mask, and uses this as the fourth similarity between the CG mask and the second mask.
[0066] The third similarity may be calculated from the L2 distance (in other words, the Euclidean distance) between the position of the CG mask feature and the position of the first mask feature in the feature space. Similarly, the fourth similarity may be calculated from the L2 distance between the position of the CG mask feature and the position of the second mask feature in the feature space.
[0067] Next, the determination unit 20 identifies the CG mask with the highest third similarity to the first mask of each of the multiple second objects 50 (also called the first CG mask) from among the multiple CG masks. The determination unit 20 also identifies the CG mask with the highest fourth similarity to the second mask of each of the multiple second objects 50 (also called the second CG mask) from among the multiple CG masks.
[0068] Hereafter, the first CG mask of a certain second object 50 will be called the first most similar CG mask, and the second CG mask of a certain second object 50 will be called the second most similar CG mask. As described above, in the matching score acquisition process, if a pair of first and second masks whose IoU is smaller than a predetermined value is excluded in the matching score acquisition process, the first CG mask and the second CG mask for a certain second object 50 may be considered the same CG mask. In this case, the first most similar CG mask and the second most similar CG mask may simply be called the most similar CG mask.
[0069] The determination unit 20 determines, for each second object 50, that the first most similar CG mask of the second object 50 is the CG reference mask of the first mask of the second object 50, and that the second most similar CG mask of the second object 50 is the CG reference mask of the second mask of the second object 50. The determination unit 20 then determines, for each second object 50, that the third similarity between the first mask of the second object 50 and the first CG mask (in other words, the first most similar CG mask) is the first similarity between the first mask and the CG reference mask. Furthermore, the determination unit 20 determines, for each second object 50, that the fourth similarity between the second mask of the second object 50 and the second CG mask (in other words, the second most similar CG mask) is the second similarity between the second mask and the CG reference mask.
[0070] <Example of the process for determining which model to use> In the process for determining which model to use in step s4, the determination unit 20 determines which of the first learning model 21a and the second learning model 21b to use based on the first and second numbers obtained from the mask accuracy comparison process for a predetermined number of judgment images. If we collectively refer to the mask accuracy comparison process for a predetermined number of judgment images as the comparison process, then it can be said that the determination unit 20 determines which of the first learning model 21a and the second learning model 21b to use based on the results of the comparison process.
[0071] In the process of determining the objects to be used, for example, the determination unit 20 obtains the total number of objects 50 (also called the total number of objects) that appear in a predetermined number of determination images. If there are 50 objects 50 in one determination image, and the predetermined number is 100 images, then the total number of objects will be 5000.
[0072] Furthermore, the determination unit 20 obtains the sum of a first number (first sum) obtained from the mask accuracy comparison process for a predetermined number of determination images. The first sum indicates the number of second objects 50 in all second objects 50 in the predetermined number of determination images for which the accuracy of the first mask is higher than the accuracy of the second mask. Furthermore, the determination unit 20 obtains the sum of a second number (second sum) obtained from the mask accuracy comparison process for a predetermined number of determination images. The second sum indicates the number of second objects 50 in all second objects 50 in the predetermined number of determination images for which the accuracy of the second mask is higher than the accuracy of the first mask.
[0073] Next, the determination unit 20 obtains the ratio of the second sum to the total number of objects (also called the second ratio). The second ratio represents the ratio of the total number of objects 50 in a predetermined number of determination images to the total number of objects 50 in a predetermined number of determination images in which the accuracy of the second mask is higher than the accuracy of the first mask. If the second learning model 21b is an additionally trained version of the first learning model 21a, the second ratio can also be said to represent the improvement rate of the recognition accuracy of objects 50 in the second learning model 21b compared to the recognition accuracy of objects 50 in the first learning model 21a.
[0074] Furthermore, the determination unit 20 obtains a first sum ratio (also called the first ratio) to the total number of objects. The first ratio represents the ratio of the total number of objects 50 in a predetermined number of determination images in which the accuracy of the first mask is higher than the accuracy of the second mask, relative to the total number of objects 50 in a predetermined number of determination images. If the second learning model 21b is an additionally trained version of the first learning model 21a, the first ratio can also be said to represent the rate of deterioration of the object recognition accuracy in the second learning model 21b relative to the object recognition accuracy in the first learning model 21a.
[0075] The determination unit 20 determines, for example, that the second learning model 21b is the target model if the second ratio is greater than or equal to the first predetermined value and the first ratio is less than the second predetermined value. The first predetermined value is, for example, 0.3. The second predetermined value is a value smaller than the first predetermined value, for example, 0.03. If the second learning model 21b is an additionally trained version of the first learning model 21a, the determination unit 20 can also be said to determine that the second learning model 21b is the target model if, for example, the improvement rate is greater than or equal to the first predetermined value and the deterioration rate is less than the second predetermined value. On the other hand, the determination unit 20 determines that the first learning model 21a is the target model if the second ratio is less than the first predetermined value, and determines that the first learning model 21a is the target model if the first ratio is greater than or equal to the second predetermined value.
[0076] However, the method for determining which learning model 21 to use is not limited to this. For example, the determination unit 20 may determine the second learning model 21b to use when the second ratio is greater than the first ratio, and determine the first learning model 21a to use when the first ratio is greater than the second ratio. If the second ratio and the first ratio are equal, the determination unit 20 may determine the first learning model 21a to use. Also, if the second ratio and the first ratio are equal, the determination unit 20 may determine the second learning model 21b to use.
[0077] Here, if the second proportion is greater than the first proportion, it can be said that the recognition accuracy of the object 50 in the entire captured image in the second learning model 21b (also called the second overall recognition accuracy) is higher than the recognition accuracy of the object 50 in the entire captured image in the first learning model 21a (also called the first overall recognition accuracy). On the other hand, if the first proportion is greater than the second proportion, it can be said that the first overall recognition accuracy is higher than the second overall recognition accuracy. The recognition accuracy of the object 50 in the entire captured image can also be said to be, for example, the overall recognition accuracy of multiple objects 50 that appear in the captured image.
[0078] In step s4, if the second ratio is greater than the first ratio, the determination unit 20 may determine that the second overall recognition accuracy is higher than the first overall recognition accuracy and may use the second learning model 21b. Alternatively, if the first ratio is greater than the second ratio, the determination unit 20 may determine that the first overall recognition accuracy is higher than the second overall recognition accuracy and may use the first learning model 21a. In such cases, it can also be said that in step s4, the determination unit 20 compares the first overall recognition accuracy and the second overall recognition accuracy based on the results of the mask accuracy comparison process (in other words, the results of the comparison process) for a predetermined number of images used for determination.
[0079] In the CG usage comparison process, a CG reference mask and a CG image (also called a CG object image) showing the original object 50 (in other words, the object model) of the CG reference mask may be used to obtain the accuracy of the first and second masks. The CG object image is composed of a part of the randomly stacked CG image. Hereafter, the CG usage comparison process that uses a CG reference mask and a CG object image showing the original object 50 (also called object 50a) of the CG reference mask may be referred to as the second CG usage comparison process. The above CG usage comparison process may also be referred to as the first CG usage comparison process. The second CG usage comparison process will be explained based on Figures 8 and 9.
[0080] Figure 8 is a schematic diagram showing an example of a CG reference mask 500 and a CG object image 550 in which the original object 50 (also called object 50a) of the CG reference mask 500 is depicted. An example of the CG object image 550 is shown in the upper part of Figure 8. In the lower part of Figure 8, an example of the CG reference mask 500 is shown superimposed on the original object 50a depicted in the CG object image 550.
[0081] As shown in Figure 8, the CG object image 550 captures not only the original object 50a but also the objects 50 surrounding the original object 50a. In the CG object image 550, the original object 50a is captured such that its center coincides with the center of the CG object image 550. The CG object image 550 is generated by cropping a partial image from a randomly stacked CG image that includes the original object 50a and the objects 50 surrounding it.
[0082] Hereafter, the CG reference mask and the CG object image in which the original object 50a of the CG reference mask is depicted may be collectively referred to as the CG reference image. In the second CG usage comparison process, multiple CG reference images, each containing a CG reference mask for multiple second objects 50, are used. When referring to the CG reference image of a second object 50, it means the CG reference image containing the CG reference mask of that second object 50.
[0083] Figure 9 is a schematic diagram showing an example of a first mask 200a and a first partial image 320a corresponding to the first mask 200a in the determination image 300. Figure 10 is a schematic diagram showing an example of a second mask 200b and a second partial image 320b corresponding to the second mask 200b. Figures 9 and 10 show examples of the first mask 200a and the second mask 200b for the same second object 50, respectively.
[0084] In the second CG usage comparison process, to obtain the accuracy of the first mask 200a, a first partial image 320a that includes the first region 310a at the same position as the first mask 200a in the judgment image 300 is used. Similarly, to obtain the accuracy of the second mask 200b, a second partial image 320b that includes the second region 310b at the same position as the second mask 200b in the judgment image 300 is used. When referring to the first partial image 320a corresponding to the first mask 200a, it means the first partial image 320a that includes the first region 310a at the same position as the first mask 200a in the judgment image 300. Likewise, when referring to the second partial image 320b corresponding to the second mask 200b, it means the second partial image 320b that includes the second region 310b at the same position as the second mask 200b in the judgment image 300.
[0085] An example of the first partial image 320a is shown in the upper part of Figure 9. In the lower part of Figure 9, the first mask 200a is shown superimposed on the first region 310a included in the first partial image 320a. In the first partial image 320a shown in the lower part of Figure 9, the region over which the first mask 200a is superimposed is the first region 310a.
[0086] An example of the second partial image 320b is shown at the top of Figure 10. At the bottom of Figure 10, the second mask 200b is shown superimposed on the second region 310b included in the second partial image 320b. In the second partial image 320b shown at the bottom of Figure 10, the region over which the second mask 200b is superimposed is the second region 310b. In the example of Figure 10, the accuracy of the second mask 200b is not very high, and the second mask 200b is located across two objects 50.
[0087] As shown in Figure 9, the first partial image 320a includes not only the first region 310a but also the region surrounding the first region 310a in the determination image 300. The first partial image 320a is generated such that the position in the determination image 300 that coincides with the center of the first mask 200a coincides with the center of the first partial image 320a. The first partial image 320a is generated by cropping a partial image from the determination image 300 that includes the first region 310a and the region surrounding it.
[0088] As shown in Figure 10, the second partial image 320b includes not only the second region 310b in the determination image 300, but also the region surrounding the second region 310b. The second partial image 320b is generated such that the position in the determination image 300 that coincides with the center of the second mask 200b coincides with the center of the second partial image 320b. The second partial image 320b is generated by cropping a partial image from the determination image 300 that includes the second region 310b and its surrounding region.
[0089] Hereafter, the first mask and the first partial image corresponding to the first mask may be collectively referred to as the first image. Similarly, the second mask and the second partial image corresponding to the second mask may be collectively referred to as the second image. Furthermore, as mentioned above, the CG reference mask and the CG object image corresponding to the CG reference mask may be collectively referred to as the CG reference image. In the second CG usage comparison process, for each of the multiple second objects 50, the first mask of the second object 50, the first partial image corresponding to the first mask, the second mask of the second object 50, and the second partial image corresponding to the second mask are used. When referring to the first image of the second object 50, it means the first image including the first mask of the second object 50 and the first partial image corresponding to the first mask. Similarly, when referring to the second image of the second object 50, it means the second image including the second mask of the second object 50 and the second partial image corresponding to the second mask.
[0090] In the second CG usage comparison process, the determination unit 20 obtains the fifth similarity between the first image and the CG reference image of each second object 50 as the accuracy of the first mask of the second object 50. In addition, in the second CG usage comparison process, the determination unit 20 obtains the sixth similarity between the second image and the CG reference image of each second object 50 as the accuracy of the second mask of the second object 50.
[0091] In the second CG usage comparison process, the feature acquisition unit 20a of the determination unit 20 acquires the feature quantities of each first image, each second image, and each CG reference image.
[0092] The feature acquisition unit 20a in this example can individually acquire the feature quantities (in other words, feature vectors) of each of the two input images. The feature acquisition unit 20a individually acquires the feature quantities of the first mask and the first partial image that constitute the first image. The feature quantities of the first image consist of the feature quantities of the first mask included in the first image and the feature quantities of the first partial image included in the first image. A single feature vector obtained by arranging multiple elements that constitute the feature vector of the first mask included in the first image and multiple elements that constitute the feature vector of the first partial image included in the first image in a line becomes the feature vector (in other words, feature quantity) of the first image.
[0093] The feature acquisition unit 20a individually acquires the feature quantities of the second mask and the second partial image that constitute the second image. The feature quantities of the second image consist of the feature quantities of the second mask included in the second image and the feature quantities of the second partial image included in the second image.
[0094] The feature acquisition unit 20a individually acquires the feature quantities of the CG reference mask and the CG object image that constitute the CG reference image. The feature quantities of the CG reference image consist of the feature quantities of the CG reference mask included in the CG reference image and the feature quantities of the CG object image included in the CG reference image.
[0095] The determination unit 20, for example, uses the cosine similarity between the feature quantities of the first image of the second object 50 and the feature quantities of the CG reference image of the second object 50 as the fifth similarity between the first image of the second object 50 and the CG reference image. In this example, the cosine similarity between the feature quantities of the first image of the second object 50 and the feature quantities of the CG reference image of the second object 50 is used as the precision of the first mask of the second object 50.
[0096] Furthermore, the determination unit 20 uses, for example, the cosine similarity between the feature quantities of the second image of the target second object 50 and the feature quantities of the CG reference image of the target second object 50 as the sixth similarity between the second image of the target second object 50 and the CG reference image. In this example, the cosine similarity between the feature quantities of the second image of the target second object 50 and the feature quantities of the CG reference image of the target second object 50 is used as the accuracy of the second mask of the target second object 50.
[0097] Thus, in the second CG usage comparison process, the fifth similarity between the first image and the CG reference image is obtained as the accuracy of the first mask, and the sixth similarity between the second image and the CG reference image is obtained as the accuracy of the second mask. As a result, the accuracy of the first and second masks can be obtained more accurately.
[0098] For example, consider a case where the position of the first mask of the second object 50 is not shifted from the position of the second object 50, but the position of the second mask of the second object 50 is shifted from the position of the second object 50. In this case, the first mask can be said to have higher accuracy than the second mask. In the first CG usage comparison process, the first similarity between the first mask and the CG reference mask is obtained as the accuracy of the first mask, and the second similarity between the second mask and the CG reference mask is obtained as the accuracy of the second mask. Therefore, the relative positions of the first and second masks and the second object 50 are not taken into consideration. Consequently, the first CG usage comparison process may not determine that the accuracy of the first mask is higher than that of the second mask.
[0099] In contrast, in the second CG usage comparison process, the fifth similarity between the first image and the CG reference image is obtained as the accuracy of the first mask, and the sixth similarity between the second image and the CG reference image is obtained as the accuracy of the second mask. This means that the relative positions of the first and second masks and the target second object 50 are taken into consideration. Therefore, in the second CG usage comparison process, there is a high probability that the accuracy of the first mask will not be determined to be higher than the accuracy of the second mask.
[0100] The CG object image 550 of the target second object 50 does not need to contain information about areas other than the area in which the original object 50a of the CG reference mask of the target second object 50 is captured. Specifically, in the CG object image 550 of the target second object 50, the pixel values of areas other than the area in which the original object 50a is captured may be set to zero. Similarly, the first partial image 320a does not need to contain information about areas other than the first region 310a. Specifically, in the first partial image 320a, the pixel values of areas other than the first region 310a may be set to zero. Similarly, the second partial image 320b does not need to contain information about areas other than the second region 310b. Specifically, in the second partial image 320b, the pixel values of areas other than the second region 310b may be set to zero. The determination unit 20, in the same manner as described above, uses a CG reference image including a CG object image 550 that does not contain information about areas other than the area in which the original object 50a is depicted, a first image including a first partial image 320a that does not contain information about areas other than the first area 310a, and a second image including a second partial image 320b that does not contain information about areas other than the second area 310b to obtain the accuracy of the first and second masks. This makes it possible to obtain the accuracy of the first and second masks more accurately.
[0101] In the example above, the second learning model was a model that had been further trained on the first learning model, but the application of this disclosure is not limited to that. The second learning model may be a learning model generated using the same process as the first learning model, or it may be a learning model generated using a different process. Furthermore, the second learning model may be a learning model trained from a different pre-training model than the first learning model, or it may be a learning model trained from the same pre-training model as the first learning model. Furthermore, the second learning model may be a learning model trained using different training data than the first learning model, or it may be a learning model trained using the same training data as the first learning model.
[0102] <An Example of Robot Control Based on the Recognition Results of a Learning Model> Below, we will describe an example of how the recognition results of the learning model 21 are used to control a robot. Below, we will describe an example in which the processing unit 1 controls the robot based on the recognition results of the learning model 21.
[0103] Figure 11 is a schematic diagram showing an example of a robot system 100. As shown in Figure 11, the robot system 100 comprises a robot 10, a processing unit 1 for controlling the robot 10, and a camera 30 for photographing multiple objects 50 that are loosely packed in a tray 40.
[0104] The interface 4 of the processing unit 1 can communicate with the robot 10. The processing unit 2 can control the robot 10 through the interface 4.
[0105] The robot 10, under the control of the processing unit 2, can, for example, hold an object 50 in a tray 40 and move the held object 50 to a location other than the tray 40. The robot 10 is, for example, an arm-type robot and comprises an arm 11 and an end effector 12 connected to the arm 11. The processing unit 2 can control the arm 11 and the end effector 12.
[0106] The end effector 12 is capable of holding the object 50 in the tray 40 under the control of the processing unit 2. In this case, the end effector 12 can also be called a holding mechanism. Specifically, the end effector 12 only needs to be capable of holding the object 50 by suction. In this case, the end effector 12 can also be called a suction mechanism. The end effector 12 includes, for example, a suction part 12a that suctions the object 50. The suction part 12a is also called a suction pad. More specifically, the end effector 12 may be a gripping mechanism that grips the object 50 with multiple finger-like parts.
[0107] The arm 11, for example, has multiple joints. The posture of the arm 11 changes as the amount of rotation of at least one of the multiple joints changes. The processing device 1 can change the posture of the arm 11. As the posture of the arm 11 changes, the position and posture of the end effector 12 changes. Also, as the posture of the arm 11 changes, the position and posture of the object 50 held by the end effector 12 changes.
[0108] The end effector 12, for example, uses its suction unit 12a to hold each object 50 in the tray 40 one by one. The robot 10 then changes the posture of its arm 11 to move the held object 50 from the tray 40 to another location.
[0109] The camera 30 is fixed to, for example, the end effector 12. The processing unit 2 controls the posture of the arm 11 so that the camera 30 fixed to the end effector 12 can photograph multiple objects 50 in the tray 40.
[0110] The processing unit 2 controls the robot 10 based on the recognition result of the learning model 21 selected for use in the usage target determination process in step s4 above. For example, when the second learning model 21b is selected for use in the usage target determination process, the processing unit 2 controls the robot 10 based on the recognition result of the second learning model 21b. On the other hand, when the first learning model 21a is selected for use in the usage target determination process, the processing unit 2 controls the robot 10 based on the recognition result of the second learning model 21b.
[0111] The processing unit 2 determines, for example, a holding position in which the end effector 12 holds the object 50 based on the mask of the object 50 generated by the learning model 21. Then, the processing unit 2 controls the arm 11 and the end effector 12 so that the end effector 12 holds the object 50 at the determined holding position.
[0112] The robot 10 may be controlled by a device other than the processing unit 1. In this case, the control device that controls the robot 10 is equipped with the learning model 21 that was selected for use in the usage target determination process. The control device controls the robot 10 based on the authentication result of the learning model 21 that was selected for use.
[0113] <An Example of a Method for Deciding to Execute Additional Learning of a Learning Model> Next, an example of a method for deciding to execute additional learning of the learning model 21 will be explained. Here, using the robot system 100 described above as an example, an example of a method for deciding to execute additional learning of the learning model 21 used in the control of the robot 10 will be explained. In this example, the processing unit 2 of the processing unit 1 performs additional learning of the learning model 21.
[0114] In the robot system 100, the camera 30 repeatedly generates captured images 300. The learning model 21 in the processing unit 1 repeatedly performs recognition processing to recognize objects 50 that appear in the captured images 300 acquired by the camera 30.
[0115] The processing unit 2 of the processing unit 1 repeatedly performs an execution decision process to determine, for example, whether to perform additional training on the learning model 21.
[0116] Here, the period from the start time of the execution decision process (in other words, the present) to the first timing one hour before the start time of the execution decision process is called the first period. The first time may be set to, for example, one week or two weeks. The period from the first timing to the second timing two hours before the first timing is called the second period. The second time may be longer than the first time, for example, three weeks or one month. The period from the start time of the execution to the third timing three hours before the start time of the execution is called the third period. The third time may be longer than the first and second times, for example, several months or one year. Figure 12 is a schematic diagram showing an example of the first period, second period, and third period.
[0117] In each execution decision process, the processing unit 2 executes a process (also called a short-term decision process) to determine whether to perform additional learning of the learning model 21 based on a plurality of captured images 300 used in controlling the robot 10 during a relatively short second period and a plurality of captured images 300 used in controlling the robot 10 during the most recent first period. In addition, in each execution decision process, the processing unit 2 executes a long-term decision process to determine whether to perform additional learning of the learning model 21 based on a plurality of captured images 300 used in controlling the robot 10 during a relatively long third period and a plurality of captured images 300 used in controlling the robot 10 during the most recent first period. The processing unit 2 executes both the short-term decision process and the long-term decision process in each execution decision process.
[0118] Hereafter, each of the multiple captured images 300 used to control the robot 10 during the first period may be referred to as the first captured image 300. Similarly, each of the multiple captured images 300 used to control the robot 10 during the second period may be referred to as the second captured image 300. And each of the multiple captured images 300 used to control the robot 10 during the third period may be referred to as the third captured image 300.
[0119] In the short-term decision process, the processing unit 2, for example, acquires the feature quantities of each of the multiple first captured images 300. Then, the processing unit 2 calculates the average value of the feature quantities of the multiple first captured images 300. Hereafter, the average value (statistical value) of the feature quantities of the multiple first captured images 300 will be called the first statistical feature. The first statistical feature can be said to be the average vector of the feature vectors of the multiple first captured images 300.
[0120] The acquisition of features from the captured image 300 may be performed, for example, by a Backbone network included in a learning model 21 composed of a Mask Scoring R-CNN. A Backbone network is a type of feature acquisition unit that can acquire features from an image based on that image.
[0121] Furthermore, in the short-term decision process, the processing unit 2 acquires the feature quantities of each of the multiple second captured images 300. Then, the processing unit 2 calculates, for example, the average value of the feature quantities of the multiple second captured images 300. Hereafter, the average value (statistical value) of the feature quantities of the multiple second captured images 300 will be called the second statistical feature. The second statistical feature can be said to be the average vector of the feature vectors of the multiple second captured images 300.
[0122] In the short-term decision process, the processing unit 2 obtains the cosine similarity between the first statistical feature and the second mean feature. Then, if the obtained cosine similarity is below a threshold, the processing unit 2 decides to perform additional training on the learning model 21. That is, if the first statistical feature and the second statistical feature are not similar, the processing unit 2 may determine that the first image 300 taken in the first period has undergone a data shift relative to the first image 300 taken in the second period, and then perform additional training on the learning model 21.
[0123] In the long-term decision processing, the processing unit 2 divides the third period into multiple subperiods, as shown in Figure 13. The length of each subperiod is, for example, the same as the length of the first period. The last of the multiple subperiods coincides with the first period.
[0124] The processing unit 2 obtains, for each subperiod, the average value (statistical value) (also called the third statistical feature) of the feature quantities of multiple third captured images 300 used to control the robot 10 during that subperiod. Hereafter, the first subperiod among the multiple subperiods, that is, the subperiod with the oldest time, will be called the first subperiod. The subperiods other than the first subperiod among the multiple subperiods will be called second subperiods. The number of second subperiods will be represented by N.
[0125] In this example, the N second subperiods are numbered from 1 to N, starting with the oldest. For example, of the N second subperiods, the oldest is second subperiod 1, and the newest second subperiod (i.e., period 1) is second subperiod N.
[0126] In the long-term decision process, the processing unit 2 obtains the cosine similarity between the third statistical feature of each of the N second subperiods and the third statistical feature of the first subperiod (in other words, the oldest subperiod). Then, the processing unit 2 creates a scatter plot 600 of the cosine similarity obtained for each of the N second subperiods.
[0127] Figure 14 is a schematic diagram showing an example of scatter plot 600 for the case N=23. In scatter plot 600, the horizontal axis represents the number of the second partial period, and the vertical axis represents the cosine similarity. In scatter plot 600, the cosine similarity for N second partial periods is arranged in chronological order from the oldest second partial period.
[0128] In the long-term decision process, the processing unit 2 creates a scatter plot 600 and then calculates a regression line 610 (see Figure 14) for multiple points shown in the scatter plot 600. When the absolute value of the slope of the regression line 610 is greater than or equal to a threshold, that is, when the slope of the regression line 610 is large, the processing unit 2 determines that the third captured image 300 is data-shifted and decides to perform additional training on the learning model 21.
[0129] In each execution decision process, if the processing unit 2 decides to perform additional training of the learning model 21 in at least one of the short-term decision process and the long-term decision process, it performs additional training of the learning model 21. The processing unit 2 may, for example, perform additional training of the learning model 21 using each of the multiple first captured images 300 used in the first period as training images. As in this example, the continuous performance of additional training of the learning model 21 is sometimes called continuous training of the learning model 21.
[0130] Thus, in this example, since the execution of additional training of the learning model 21 is determined based on multiple first captured images 300 used in a relatively short second period, the processing unit 2 can automatically perform additional training of the learning model 21 if the shooting environment of the captured images 300 (in other words, the shooting environment of the camera 30) changes in a short period of time. For example, when the brightness of the shooting environment changes, the processing unit 2 can automatically perform additional training of the learning model 21. Also, when the shooting environment changes due to a change in the placement of the camera 30 and tray 40, the processing unit 2 can automatically perform additional training of the learning model 21. Furthermore, when the shooting environment changes due to a change in the environment of the room in which the camera 30 and tray 40 are placed, the processing unit 2 can automatically perform additional training of the learning model 21.
[0131] Furthermore, in this example, since the execution of additional training of the learning model 21 is determined based on multiple third captured images 300 used over a relatively long third period, the processing unit 2 can automatically perform additional training of the learning model 21 if the shooting environment of the captured images 300 changes slowly over a long period. For example, by setting the length of the third period to about three months, the processing unit 2 can automatically perform additional training of the learning model 21 in response to seasonal changes.
[0132] As described above, the processing apparatus and robotic system have been described in detail, but the above description is illustrative in all respects, and this disclosure is not limited thereto. Furthermore, the various examples described above can be combined and applied insofar as they do not contradict each other. And it is understood that countless examples not illustrated can be conceived without falling outside the scope of this disclosure.
[0133] This disclosure includes the following:
[0134] In one embodiment, (1) the processing unit includes a processing unit that compares the inference accuracy of a first learning model and a second learning model using computer graphics (CG) images.
[0135] (2) The processing device described in (1) above, wherein the first learning model and the second learning model each recognize real objects in the captured image based on the captured image, and the processing device compares the recognition accuracy of the first learning model and the second learning model using the CG image.
[0136] (3) The processing device of (2) above, wherein the first learning model generates a first mask of the real object, the second learning model generates a second mask of the real object, and the processing device compares the accuracy of the first mask and the second mask using the CG image.
[0137] (4) The processing apparatus of (3) above, wherein the CG image includes a CG reference mask which serves as a reference for the first mask and the second mask, and the processing apparatus obtains a first similarity between the CG reference mask and the first mask as the accuracy of the first mask, and obtains a second similarity between the CG reference mask and the second mask as the accuracy of the second mask.
[0138] (5) The processing apparatus of (4) above, wherein the processing apparatus obtains the first similarity based on the feature quantities of the CG reference mask and the feature quantities of the first mask, and obtains the second similarity based on the feature quantities of the CG reference mask and the feature quantities of the second mask.
[0139] (6) The processing apparatus of (3) above, wherein the CG image includes a CG reference mask that serves as a reference for the first mask and the second mask, and a CG object image in which the original CG object of the CG reference mask is depicted, and the processing apparatus obtains a first similarity between a first image having a first partial image including a first region at the same position as the first mask in the captured image and the first mask, and a CG reference image including the CG object image and the CG reference mask, as the accuracy of the first mask, and obtains a second similarity between a second image having a second partial image including a second region at the same position as the second mask in the captured image and the second mask, and the CG reference image, as the accuracy of the second mask.
[0140] (7) The processing apparatus of (6) above, wherein the processing apparatus obtains the first similarity based on the feature quantities of the first image and the feature quantities of the CG reference image, and obtains the second similarity based on the feature quantities of the second image and the feature quantities of the CG reference image.
[0141] (8) The processing apparatus according to (6) or (7) above, wherein the CG object image does not contain information about areas other than the area in which the CG object is depicted, the first partial image does not contain information about areas other than the first area, and the second partial image does not contain information about areas other than the second area.
[0142] (9) Any one of the processing devices described in (4) to (8) above, wherein the processing device determines the CG reference mask from a plurality of CG masks.
[0143] (10) Any one of the processing devices described in (3) to (9) above, wherein the processing device obtains the degree of agreement between the first mask and the second mask, and compares the accuracy of the first mask and the second mask based on the degree of agreement.
[0144] (11) Any one of the processing devices described in (3) to (10) above, wherein the first learning model generates a plurality of first masks for a plurality of real objects captured in the captured image, the second learning model generates a plurality of second masks for the plurality of real objects, and the processing device uses the CG image to compare the accuracy of the first masks and second masks for the same real object in the plurality of first masks and the plurality of second masks.
[0145] (12) The processing device of (11) above, wherein the first learning model generates a plurality of first masks for each of the plurality of captured images for each of the plurality of real objects captured in the captured image, the second learning model generates a plurality of second masks for each of the plurality of captured images for each of the plurality of real objects captured in the captured image, and the processing device performs a comparison process for each of the plurality of captured images, using the CG image to compare the accuracy of the first mask and the second mask for the same real object in the plurality of first masks and the plurality of second masks.
[0146] (13) The processing device of (12) above, wherein the processing device determines which of the first learning model and the second learning model to use based on the result of the comparison process.
[0147] (14) The processing device of (13) above, wherein the physical object is a work object on which the robot performs work, and the processing device controls the robot based on the recognition result of the learning model that has been determined to be the target of use from among the first learning model and the second learning model.
[0148] (15) Any one of the processing devices described in (1) to (14) above, wherein the second learning model is an additional learning version of the first learning model.
[0149] In one embodiment, the program (16) is used to cause the computer device to function as the processing unit of any one of the processing units (1) to (15) described above.
[0150] In one embodiment, the (17) robot system comprises the processing device described in (14) and a robot controlled by the processing device.
[0151] 1 Processing unit 2 Processing unit 3a Program 10 Robot 21 Learning model 21a First learning model 21b Second learning model 50 Object (real object) 100 Robot system 300 Captured image
Claims
1. A processing unit comprising a processing unit that compares the inference accuracy of a first learning model and a second learning model using computer graphics (CG) images.
2. A processing apparatus according to claim 1, wherein each of the first learning model and the second learning model recognizes a real object depicted in a captured image based on the captured image, and the processing apparatus compares the recognition accuracy of the first learning model and the second learning model using the CG image.
3. The processing apparatus according to claim 2, wherein the first learning model generates a first mask of the real object, the second learning model generates a second mask of the real object, and the processing unit compares the accuracy of the first mask and the second mask using the CG image.
4. The processing apparatus according to claim 3, wherein the CG image includes a CG reference mask that serves as a reference for the first mask and the second mask, and the processing apparatus obtains a first similarity between the CG reference mask and the first mask as the accuracy of the first mask, and obtains a second similarity between the CG reference mask and the second mask as the accuracy of the second mask.
5. The processing apparatus according to claim 4, wherein the processing apparatus obtains a first similarity based on the feature quantities of the CG reference mask and the feature quantities of the first mask, and obtains a second similarity based on the feature quantities of the CG reference mask and the feature quantities of the second mask.
6. The processing apparatus according to claim 3, wherein the CG image includes a CG reference mask that serves as a reference for the first mask and the second mask, and a CG object image in which the original CG object of the CG reference mask is depicted, and the processing apparatus obtains a first similarity between a first image having a first partial image including a first region at the same position as the first mask in the captured image and the first mask, and a CG reference image including the CG object image and the CG reference mask, as the accuracy of the first mask, and obtains a second similarity between a second image having a second partial image including a second region at the same position as the second mask in the captured image and the second mask, and the CG reference image, as the accuracy of the second mask.
7. The processing apparatus according to claim 6, wherein the processing apparatus obtains a first similarity based on the feature quantities of the first image and the feature quantities of the CG reference image, and obtains a second similarity based on the feature quantities of the second image and the feature quantities of the CG reference image.
8. The processing apparatus according to claim 6 or claim 7, wherein the CG object image does not contain information about areas other than the area in which the CG object is depicted, the first partial image does not contain information about areas other than the first area, and the second partial image does not contain information about areas other than the second area.
9. A processing apparatus according to any one of claims 4 to 8, wherein the processing apparatus determines a CG reference mask from a plurality of CG masks.
10. A processing apparatus according to any one of claims 3 to 9, wherein the processing apparatus obtains the degree of agreement between the first mask and the second mask, and compares the accuracy of the first mask and the second mask based on the degree of agreement.
11. A processing apparatus according to any one of claims 3 to 10, wherein the first learning model generates a plurality of first masks for a plurality of real objects captured in the captured image, the second learning model generates a plurality of second masks for the plurality of real objects, and the processing apparatus compares the accuracy of the first masks and second masks for the same real object using the CG image.
12. The processing apparatus according to claim 11, wherein the first learning model generates a plurality of first masks for each of a plurality of captured images for a plurality of real objects captured in the captured image, the second learning model generates a plurality of second masks for each of the plurality of captured images for a plurality of real objects captured in the captured image, and the processing apparatus performs a comparison process for each of the plurality of captured images, using the CG image, to compare the accuracy of the first masks and second masks for the same real object in the plurality of first masks and the plurality of second masks.
13. The processing apparatus according to claim 12, wherein the processing apparatus determines, based on the result of the comparison process, which of the first learning model and the second learning model to use.
14. The processing apparatus according to claim 13, wherein the physical object is a work object on which the robot performs work, and the processing apparatus controls the robot based on the recognition result of the learning model selected as the work object from among the first learning model and the second learning model.
15. A processing apparatus according to any one of claims 1 to 14, wherein the second learning model is a first learning model that has been further trained.
16. A program for causing a computer device to function as the processing unit of the processing unit described in any one of claims 1 to 15.
17. A robot system comprising the processing device described in claim 14 and a robot controlled by the processing device.