Control device, control system, robot system, control method, and computer program

The robot system uses image data and three-dimensional position data to enhance object recognition and control, addressing inefficiencies in processing multiple objects, especially those arranged irregularly, by improving identification and manipulation accuracy.

WO2026069666A1PCT designated stage Publication Date: 2026-04-02NIKON CORP
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-30
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

Existing control systems for robots struggle to accurately identify and process multiple objects, particularly when they are randomly or irregularly arranged, leading to inefficiencies in tasks such as picking, placing, and processing.

Method used

A robot system equipped with an imaging system and a control device that generates control signals based on image data and three-dimensional position data to precisely identify and process objects, using techniques like machine learning and interpolation to enhance object recognition and control accuracy.

Benefits of technology

The system enables precise identification and processing of objects, improving task efficiency by accurately selecting and manipulating objects even in complex arrangements, enhancing the robot's ability to handle multiple objects with varying shapes and positions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024034956_02042026_PF_FP_ABST
    Figure JP2024034956_02042026_PF_FP_ABST
Patent Text Reader

Abstract

A control device according to the present invention generates a control signal for controlling a processing device that subjects an object to processing and a robot provided with the processing device. The control device uses image data generated by an imaging system that captures images of the object or first three-dimensional position data to identify an object region in which a target object is present, uses the identification result of the object region and second three-dimensional position data to generate first position / orientation information indicating the position and orientation of the target object, uses the first position / orientation information and third three-dimensional position data to generate second position / orientation information indicating the position and orientation of the target object, and generates a control signal for executing a process on the target object on the basis of the second position / orientation information.
Need to check novelty before this filing date? Find Prior Art

Description

Control device, control system, robot system, control method, and computer program

[0001] The present invention relates, for example, to the technical fields of control devices, control systems, robot systems, control methods, and computer programs capable of generating control signals for controlling a robot.

[0002] Patent Document 1 describes an example of a control device for controlling a robot equipped with an end effector capable of holding parts. Such a control device is required to properly control the robot.

[0003] U.S. Patent Application Publication No. 2013 / 0230235

[0004] According to the first aspect, there is provided a robot provided with an imaging system capable of generating image data by imaging at least one of a plurality of objects, and a processing device that performs processing on at least one of the plurality of objects, and a control device that generates a control signal for controlling at least one of the processing devices. The control device includes an arithmetic device that generates the control signal, and an output device that outputs the control signal generated by the arithmetic device. The arithmetic device uses at least one of the first three-dimensional position data generated from the first image data, which is the image data generated by the imaging system imaging at least one of the plurality of objects, and the second image data, which is the image data generated by the imaging system imaging at least one of the plurality of objects, to identify an object region where a target object to be processed by the processing device exists among the plurality of objects. Using the identification result of the object region and the second three-dimensional position data indicating the three-dimensional position generated from the third image data, which is the image data generated by the imaging system imaging at least one of the plurality of objects, the first position and orientation information indicating at least one of the position and orientation of the target object is generated. Using the first position and orientation information and the third three-dimensional position data indicating the three-dimensional position generated from the fourth image data, which is the image data generated by the imaging system imaging at least one of the plurality of objects, the second position and orientation information indicating at least one of the position and orientation of the target object is generated. A control device is provided that generates the control signal for performing the processing on the target object based on the second position and orientation information.

[0005] According to a second embodiment, a robot is provided which is equipped with a processing device that performs processing on at least one object, and a control device that generates a control signal for controlling at least one of the processing devices, wherein the control device comprises a calculation device that generates the control signal and an output device that outputs the control signal generated by the calculation device, the calculation device identifies a first data portion in which the object exists from first three-dimensional position data generated by an imaging system capable of generating image data by imaging the object, generates position and orientation information of the object based on a second data portion that interpolates the first data portion calculated using a three-dimensional model of the object and the first data portion, and generates the control device that generates the control signal for performing the processing on the object based on the position and orientation information.

[0006] According to a third aspect, a robot is provided which is equipped with a processing device that performs processing on at least one of a plurality of objects, and a control device that generates a control signal for controlling at least one of the processing devices, wherein the control device comprises a calculation device that generates the control signal and an output device that outputs the control signal generated by the calculation device, and the calculation device uses at least one of first three-dimensional position data generated from first image data which is image data generated when an imaging system capable of generating image data by imaging the plurality of objects images at least one of the plurality of objects, and a learning model generated by machine learning to select a target object from the plurality of objects to be processed by the processing device, and generates the control signal for performing the processing on the selected target object.

[0007] According to a fourth aspect, a robot is provided which is equipped with a processing device that performs processing on at least one object, and a control device that generates a control signal for controlling at least one of the processing devices, wherein the control device comprises a calculation device that generates the control signal and an output device that outputs the control signal generated by the calculation device, the calculation device generates first position and orientation information indicating the position and orientation of the object using first three-dimensional position data generated from first image data which is image data generated by an imaging system capable of generating image data by imaging the object, second position and orientation information indicating the position and orientation of the object based on second three-dimensional position data generated from second image data which is image data generated by the imaging system that imaging the object, and at least one of the position determined based on the first position and orientation information and the orientation determined based on the first position and orientation information, and the control device generates the control signal based on the second position and orientation information.

[0008] According to the fifth aspect, a control system is provided comprising a control device provided by any one of the first to fourth aspects and an imaging system.

[0009] According to the sixth aspect, a robot system is provided comprising a control device provided by any one of the first to fourth aspects, an imaging system, and the robot.

[0010] According to the seventh aspect, a robot is provided with a processing device that processes at least one of a plurality of objects and an imaging system capable of generating image data by imaging at least one of the plurality of objects, and a control method for generating a control signal for controlling at least one of the processing device, wherein the method uses at least one of first image data, which is the image data generated by the imaging system imaging at least one of the plurality of objects, and first three-dimensional position data generated from the second image data, which is the image data generated by the imaging system imaging at least one of the plurality of objects, to identify an object region in which an object to be processed by the processing device exists among the plurality of objects, and the object region A control method is provided which includes generating first position and orientation information indicating at least one of the position and orientation of the target object using the results of identifying the area and second three-dimensional position data indicating a three-dimensional position generated from third image data which is image data generated when the imaging system images at least one of the plurality of objects; generating second position and orientation information indicating at least one of the position and orientation of the target object using the first position and orientation information and third three-dimensional position data indicating a three-dimensional position generated from fourth image data which is image data generated when the imaging system images at least one of the plurality of objects; and generating the control signal for executing the processing on the target object based on the second position and orientation information.

[0011] According to the eighth aspect, a robot is provided which is equipped with a processing device that performs processing on at least one object, and a control method for generating a control signal for controlling at least one of the processing devices, the control method being provided which includes: identifying a first data portion in which the object exists from first three-dimensional position data generated by an imaging system capable of generating image data by imaging the object; generating position and orientation information of the object based on a second data portion that interpolates the first data portion calculated using a three-dimensional model of the object and the first data portion; and generating the control signal for the object to perform the processing based on the position and orientation information.

[0012] According to the ninth aspect, a robot is provided which is equipped with a processing device that performs processing on at least one of a plurality of objects, and a control method for generating a control signal for controlling at least one of the processing devices, the control method being provided which includes selecting a target object from the plurality of objects to be processed by the processing device, using at least one of first image data which is image data generated when an imaging system capable of generating image data by imaging at least one of the plurality of objects, and first three-dimensional position data generated from second image data which is image data generated when the imaging system captures at least one of the plurality of objects, and a learning model generated by machine learning, and generating the control signal for performing the processing on the selected target object.

[0013] According to the tenth aspect, a robot is provided which is equipped with a processing device that performs processing on at least one object, and a control method for generating a control signal for controlling at least one of the processing devices, the control method being provided which includes generating first position and orientation information indicating the position and orientation of an object using first three-dimensional position data generated from first image data which is image data generated by an imaging system capable of generating image data by imaging the object; generating second position and orientation information indicating the position and orientation of an object based on second three-dimensional position data generated from second image data which is image data generated by the imaging system by imaging the object, and at least one of a position determined based on the first position and orientation information and an orientation determined based on the first position and orientation information; and generating the control signal based on the second position and orientation information.

[0014] According to the eleventh aspect, a computer program is provided that causes a computer to execute a control method provided by any one of the seventh to tenth aspects.

[0015] Figure 1 is a block diagram showing the configuration of the robot system of this embodiment. Figure 2 is a side view showing the external appearance of the robot of this embodiment. Figure 3 is a block diagram showing the configuration of the control device of this embodiment. Figure 4 is a block diagram showing the configuration of the robot control device of this embodiment. Figure 5 is a block diagram showing the configuration of the model generation device of this embodiment. Figures 6A to 6C are side views showing the positional relationship between the robot and an object at a certain point in time during the period in which a holding process is being performed to hold an object placed on a mounting device. Figure 7 is a flowchart showing the flow of the model generation process. Figure 8A shows the data structure of the training dataset, and Figure 8B schematically shows the object region in which an object or sample object is captured in the image shown by the training image data. Figure 9 is a cross-sectional view showing the simulation space in which the object model is placed. Figure 10 schematically shows the object region information and the correct label output by the object recognition model that receives the training image data as input. Figure 11 is a flowchart showing the overall flow of the robot control process. Figure 12 is a flowchart showing the flow of the process to calculate at least one of the position and orientation of an object based on the image data and the object recognition model in step S22 of Figure 11. Figure 13 schematically shows the object region information output by an object recognition model that has been input with training image data. Figure 14 shows the object region where an object exists in a 2D image, the object region where an object exists in a 2D imaging coordinate system, and the object region where an object exists in a 3D imaging coordinate system. Figure 15 shows the data structure of the training dataset. Figure 16 shows the initial position and initial orientation of the three-dimensional template model. Figure 17 shows the data structure of the training dataset. Figure 18 schematically shows the interpolation model, the point cloud data before interpolation, and the point cloud data after interpolation. Figure 19 shows the second object and the first object overlapping the second object. Figure 20 shows the data structure of the training dataset. Figure 21 shows the data structure of the training dataset. Figure 22 is a block diagram showing the configuration of a robot system equipped with a measurement system. Figure 23 is a side view showing the appearance of a modified example of the robot.

[0016] Next, embodiments of the control device, control system, robot system, control method, and computer program will be described with reference to the drawings. In the following, embodiments of the control device, control system, robot system, control method, and computer program will be described using the robot system SYS.

[0017] (1) Configuration of the SYS robot system First, we will explain the configuration of the SYS robot system.

[0018] (1-1) Overall Configuration of the SYS Robot System First, the overall configuration of the SYS robot system will be explained with reference to Figure 1. Figure 1 is a block diagram showing the overall configuration of the SYS robot system.

[0019] As shown in Figure 1, the robot system SYS comprises a robot 1, an imaging system 2, a control device 3, an end effector 4, and a model generation device 5.

[0020] Robot 1 is a device capable of performing predetermined processing on an object OBJ. An example of robot 1 is shown in Figure 2. Figure 2 is a side view showing the external appearance of robot 1. As shown in Figure 2, robot 1 includes, for example, a base 11, a robot arm 12, and a robot control device 13.

[0021] The base 11 is a fundamental component of the robot 1. The base 11 is placed on a support surface S, such as a floor. The base 11 may be fixed to the support surface S. Alternatively, the base 11 may be movable relative to the support surface S. Figure 2 shows an example where the base 11 is fixed to the support surface S.

[0022] The robot arm 12 is attached to the base 11. The robot arm 12 is a device in which a plurality of links 121 are connected via joints 122. The joints 122 have actuators built into them. The links 121 may be rotatable around an axis defined by the joints 122 by the actuators built into the joints 122. At least one link 121 may be extendable and retractable along the direction in which the link 121 extends. The device including the plurality of links 121 connected via joints 122 and the base 11 may also be referred to as the robot arm 12.

[0023] An end effector 4 is attached to (or rather, provided with) the robot arm 12. In other words, the robot 1 has an end effector 4 attached to it. In the example shown in Figure 2, the end effector 4 is attached to the tip of the robot arm 12. The end effector 4 is movable by the movement of the robot arm 12. In other words, the robot arm 12 moves the end effector 4. In other words, the robot 1 moves the end effector 4.

[0024] The end effector 4 is a device that performs a predetermined process (in other words, a predetermined operation) on the object OBJ. The end effector 4 that performs the predetermined process on the object OBJ may also be called a processing device.

[0025] For example, the end effector 4 may perform a holding process to hold an object OBJ as an example of a predetermined process. In this case, the end effector 4 may be considered to be performing a holding process on the object OBJ that the end effector 4 is to hold. The end effector 4 capable of performing a holding process may be called a holding device. Furthermore, holding an object OBJ may be considered equivalent to picking up an object OBJ. The holding process for holding an object OBJ may be considered equivalent to the pickup process for picking up an object OBJ.

[0026] For example, holding an object OBJ may include gripping the object OBJ. For example, holding an object OBJ may include gripping the object OBJ using a hand gripper, which is an example of an end effector 4. Holding an object OBJ may also include attracting the object OBJ. For example, holding an object OBJ may include attracting (vacuum attracting) the object OBJ using a vacuum gripper, which is an example of an end effector 4. For example, holding an object OBJ may include attracting the object OBJ using a magnetic gripper (e.g., a magnetic gripper), which is an example of an end effector 4.

[0027] For example, the end effector 4 may perform a release process (in other words, a release operation) to release (i.e., separate) the object OBJ it is holding, as an example of a predetermined process. In this case, the end effector 4 may be considered to be performing a release process on the object OBJ it is holding. The end effector 4 capable of performing a release process may be called a release device. The release process may also be called a placement process.

[0028] An example of an end effector 4 capable of holding and releasing is a hand gripper. A hand gripper is an end effector 4 capable of holding (e.g., gripping) an object OBJ by physically gripping the object OBJ with multiple (e.g., two, three, or four) finger or claw members. Another example of an end effector 4 capable of holding and releasing is a vacuum gripper (e.g., a vacuum suction type gripper). A vacuum gripper is an end effector 4 capable of holding (e.g., suctioning) an object OBJ by vacuum suction. Another example of an end effector 4 capable of holding and releasing is a magnetic gripper (e.g., a magnetic suction type gripper). Figure 2 shows an example where the end effector 4 is a magnetic gripper. However, the end effector 4 capable of holding and releasing is not limited to the above examples, and may be other existing end effectors capable of holding and releasing an object OBJ. For example, the end effector 4 may be a Bernoulli chuck capable of holding an object OBJ without contact. Furthermore, the end effector 4 capable of performing at least one of the holding and releasing processes is not limited to the end effector described above, but may be any other existing end effector.

[0029] Robot 1 may use an end effector 4 capable of holding and releasing to perform a placement process (in other words, a placement operation) to position an object OBJ at a desired location as an example of a predetermined process. For example, Robot 1 may use the end effector 4 to hold a first object OBJ, which is a first example of an object OBJ, and then perform a placement process to position the first object OBJ held by the end effector 4 at a desired location on a second object OBJ, which is a second example of an object OBJ and different from the first object OBJ. In this case, the end effector 4 may be considered to be performing a release process on the second object OBJ to which the end effector 4 is to place the first object OBJ. Similarly, the end effector 4 may be considered to be performing a release process on the first object OBJ to which the end effector 4 is to release.

[0030] As an example of a second object OBJ, a jig that supports (and in some cases holds) the first object OBJ may be provided. In this case, robot 1 may perform a placement process to position (in other words, set) the first object OBJ held by end effector 4 onto the jig. As an example of a second object OBJ, a machining device (for example, the stage of the machining device) that processes the first object OBJ may be provided. In this case, robot 1 may perform a placement process to position (in other words, set) the first object OBJ held by end effector 4 onto the machining device (for example, the stage of the machining device). As an example of a second object OBJ, a measuring device (for example, the stage of the measuring device) that measures the first object OBJ may be provided. In this case, robot 1 may perform a placement process to position (in other words, set) the first object OBJ held by end effector 4 onto the measuring device (for example, the stage of the measuring device).

[0031] Robot 1 may use an end effector 4 capable of holding and releasing to perform a fitting process (in other words, a fitting operation) to fit a first object OBJ into a second object OBJ that is different from the first object OBJ, as one specific example of a placement process (in other words, a placement operation). The fitting process may include, for example, a process to fit (i.e., insert) the first object OBJ (for example, a protrusion of the first object OBJ) into a recess (for example, a hole) formed in the second object OBJ. The fitting process may also include, for example, a process to fit the first object OBJ (for example, a recess of the first object OBJ) into a protrusion (for example, a rod) formed in the second object OBJ. In this case, the second object OBJ may be the object (workpiece) into which the first object OBJ is to be fitted.

[0032] Robot 1 may use an end effector 4 capable of holding and releasing to perform a placement process (in other words, a placement operation) to place a first object OBJ onto a second object OBJ that is different from the first object OBJ, as one specific example of a placement process (in other words, a placement operation). In this case, the second object OBJ may be the object (workpiece) on which the first object OBJ is placed. Robot 1 may also use an end effector 4 capable of holding and releasing to perform an attachment process (in other words, an attachment operation) to attach the first object OBJ to a second object OBJ that is different from the first object OBJ, as one specific example of a placement process (in other words, a placement operation). In this case, the second object OBJ may be the object (workpiece) on which the first object OBJ is attached. Robot 1 may use an end effector 4 capable of holding and releasing to perform an adhesive process (in other words, an adhesive operation) to bond a first object OBJ to a second object OBJ different from the first object OBJ, as one specific example of a placement process (in other words, a placement operation). In this case, the second object OBJ may be the object (workpiece) to which the first object OBJ is to be bonded. Furthermore, another end effector capable of dispensing adhesive for the adhesive process may be provided on robot 1 or another robot. Robot 1 may also use an end effector 4 capable of holding and releasing to perform a welding process (in other words, a welding operation) to weld a first object OBJ to a second object OBJ different from the first object OBJ, as one specific example of a placement process (in other words, a placement operation). In this case, the second object OBJ may be the object (workpiece) to which the first object OBJ is to be welded. Furthermore, other end effectors (for example, processing devices for welding with an energy beam) for welding the first object OBJ and the second object OBJ may be provided on robot 1 or another robot.Robot 1 may use an end effector 4 capable of holding and releasing to perform a screw fastening process (in other words, a screw fastening operation) as a specific example of a placement process (in other words, a placement operation) to fasten a first object OBJ, which is capable of functioning as a screw, into a screw hole formed in a second object OBJ, which is different from the first object OBJ. In this case, the first object OBJ may be a screw member capable of functioning as a screw, such as a bolt or nut, and the second object OBJ may be an object (workpiece) to which the first object OBJ is to be screwed. The end effector 4 may also be a tool such as a screwdriver capable of screw fastening. At least one of the bonding process, adhesive process, welding process and screw fastening process may be referred to as a processing process.

[0033] Robot 1 may use an end effector 4 capable of holding and releasing to perform a movement process (in other words, a movement operation) to move an object OBJ as an example of a predetermined process. For example, Robot 1 may use the end effector 4 to hold a first object OBJ, which is a first example of an object OBJ, and then move the end effector 4 to perform a movement process to move the first object OBJ held by the end effector 4. In this case, the end effector 4 may be considered to be performing a movement process on the first object OBJ that the end effector 4 is to move.

[0034] Robot 1 may use an end effector 4 capable of holding and releasing to perform a disposal process (in other words, a disposal operation) for the object OBJ as a specific example of a placement process (in other words, a placement operation).

[0035] The end effector 4 may perform a predetermined process on each of the multiple object OBJs. In other words, the end effector 4 may sequentially perform a predetermined process on multiple object OBJs. In this case, the robot 1 moves the end effector 4 to a first position where it can perform a predetermined process on one object OBJ, and after the end effector 4 moves to the first position, the end effector 4 may perform a predetermined process on one object OBJ. Subsequently, the robot 1 moves the end effector 4 to a second position where it can perform a predetermined process on other object OBJs different from the first object OBJ, and after the end effector 4 moves to the second position, the end effector 4 may perform a predetermined process on the other object OBJs.

[0036] The shapes of multiple objects OBJ may be identical to each other. Alternatively, multiple objects OBJ may include at least two objects OBJ with different shapes. For example, if an object OBJ includes a workpiece W as described later, multiple objects OBJ (i.e., multiple workpieces W) may include at least two objects OBJ (i.e., at least two workpieces W) with different shapes.

[0037] For example, the end effector 4 may perform a predetermined process on each of the multiple parts of a single object OBJ. That is, the end effector 4 may sequentially perform a predetermined process on multiple parts of a single object OBJ. In this case, the robot 1 moves the end effector 4 to a third position where it can perform a predetermined process on a first part of the object OBJ, and after the end effector 4 moves to the third position, the end effector 4 may perform a predetermined process on the first part of the object OBJ. Subsequently, the robot 1 moves the end effector 4 to a fourth position where it can perform a predetermined process on a second part of the object OBJ, and after the end effector 4 moves to the fourth position, the end effector 4 may perform a predetermined process on the second part of the object OBJ.

[0038] The object OBJ that the end effector 4 performs a predetermined process on may include a workpiece W, as shown in Figure 2. The workpiece W may include, for example, a part or component used to manufacture a desired product. The workpiece W may include, for example, a part or component that is processed to manufacture a desired product. The workpiece W may include, for example, a part or component that is transported to manufacture a desired product. The workpiece W may include, for example, a part or component that is being moved by transport to manufacture a desired product. The workpiece W may include, for example, a part or component that is being moved to manufacture a desired product.

[0039] The object OBJ on which the end effector 4 performs a predetermined process may include a mounting device T on which the workpiece W is placed, as shown in Figure 2. An example of a mounting device T is at least one of a container (storage box) CB and a jig J. Figure 2 shows an example where the mounting device T is a container CB. The container CB may be a mounting device T on which the workpiece W is placed in a storage space SP enclosed by the bottom wall BS and the side wall SS, with the container CB comprising a bottom wall BS and a side wall SS protruding upward from the bottom wall BS. The container CB may also be a mounting device T capable of accommodating the workpiece W in a storage space SP enclosed by the bottom wall BS and the side wall SS. However, the container CB does not have to have a side wall SS. A container CB without a side wall SS may be called a pallet. Furthermore, the mounting device T is not limited to a container CB or a pallet, but may be any existing object on which the workpiece W can be placed. The mounting device T may also be referred to as a mounting member.

[0040] The mounting device T may be positioned on a support surface S. The mounting device T may be fixed to the support surface S. Alternatively, at least a part of the mounting device T may be movable relative to the support surface S. A first example in which at least a part of the mounting device T is movable relative to the support surface S is an example in which the mounting device T is supported by a movable (in other words, transportable) transport device. In this case, the transport device may be, for example, a belt conveyor. A second example in which at least a part of the mounting device T is movable relative to the support surface S is an example in which the mounting device T is supported by a movable device (mounting movable device). Examples of mounting movable devices include at least one of an automated guided vehicle (AGV), an autonomous mobile robot, an unmanned aerial vehicle (e.g., a drone), and a submersible. A third example in which at least a part of the mounting device T is movable relative to the support surface S is an example in which the mounting device T functions as a movable mounting device. In other words, a third example in which at least a part of the mounting device T is movable relative to the support surface S is an example in which a movable mounting device is used as the mounting device T. Figure 2 shows an example in which the mounting device T is self-propelled on the support surface S.

[0041] If object OBJ includes a workpiece W and a mounting device T, the above-described holding process may include a process of holding the workpiece W placed on a stationary or moving mounting device T. The above-described holding process may also include a process of holding the workpiece W placed on a support surface S. The above-described release process may include a process of releasing the workpiece W held by the end effector 4 in order to position the workpiece W held by the end effector 4 at a desired position on the stationary or moving mounting device T. The above-described release process may include a process of releasing the workpiece W held by the end effector 4 in order to position the workpiece W held by the end effector 4 at a desired position on the support surface S. The above-described release process may include a process of releasing the first workpiece W held by the end effector 4 in order to fit the first workpiece W held by the end effector 4 into a second workpiece W placed on the stationary or moving mounting device T. The above-described release process may include a process of releasing the first workpiece W held by the end effector 4 in order to fit the first workpiece W held by the end effector 4 into the second workpiece W placed on the support surface S. The above-described release process may include a process of releasing the first workpiece W held by the end effector 4 in order to fit (i.e., insert) the first workpiece W held by the end effector 4 into a hole formed in the second workpiece W placed on a stationary or moving mounting device T. The above-described release process may include a process of releasing the first workpiece W held by the end effector 4 in order to fit (i.e., insert) the first workpiece W held by the end effector 4 into a hole formed in the second workpiece W placed on the support surface S.

[0042] The robot control device 13 controls the movement of the robot 1.

[0043] Specifically, the robot control device 13 may control the operation of the robot arm 12. For example, the robot control device 13 may control the operation of the robot arm 12 so that the desired link 121 rotates around the axis defined by the desired joint 122. For example, the robot control device 13 may control the operation of the robot arm 12 so that the end effector 4 attached to the robot arm 12 is located at a desired position. For example, the robot control device 13 may control the operation of the robot arm 12 so that the end effector 4 attached to the robot arm 12 moves to a desired position.

[0044] In addition to controlling the operation of the robot 1, the robot control device 13 may control the operation of the end effector 4 (processing by the end effector 4) attached to the robot 1. For example, the robot control device 13 may control the operation of the end effector 4 so that the end effector 4 holds the object OBJ at a desired timing. That is, the robot control device 13 may control the operation of the end effector 4 so that the end effector 4 performs a holding process at a desired timing. For example, the robot control device 13 may control the operation of the end effector 4 so that the end effector 4 releases the object OBJ held at a desired timing. That is, the robot control device 13 may control the operation of the end effector 4 so that the end effector 4 performs a release process at a desired timing. In order for the end effector 4 to hold or release the object OBJ, when the end effector 4 is a hand gripper, the robot control device 13 may control the timing at which the hand gripper opens and closes. When the end effector 4 is a vacuum gripper, the robot control device 13 may control the timing (or vacuum adsorption force) at which the vacuum device of the vacuum gripper is turned on and off. When the end effector 4 is a magnet gripper, the robot control device 13 may control the timing (or magnetic force) at which the magnetic adsorption device of the magnet gripper is turned on and off.

[0045] Figure 2 shows an example where robot 1 is a robot arm 12 (i.e., a vertical articulated robot). However, robot 1 may be a robot other than a vertical articulated robot. For example, robot 1 may be a SCARA robot (i.e., a horizontal articulated robot). For example, robot 1 may be a parallel kinematic machine (PKM) type (in other words, a parallel link type) robot. For example, robot 1 may be a dual-arm robot equipped with two robot arms 12. For example, robot 1 may be a Cartesian coordinate robot. For example, robot 1 may be a cylindrical coordinate robot.

[0046] Robot 1 may be installed on a movable device (robot movable device) that allows it to move. Examples of robot movable devices include at least one of an automated guided vehicle (AGV), an autonomous mobile robot, an unmanned aerial vehicle (e.g., a drone), and a submersible. If robot 1 is installed on a robot movable device different from robot 1, the device including robot 1 and the robot movable device on which robot 1 is installed may be referred to as the movable device (robot movable device). When robot 1 is installed on a robot movable device, the robot control device 13 may control the operation of the robot movable device on which robot 1 is installed, in addition to controlling the operation of robot 1. Furthermore, since robot 1 moves its robot arm 12, robot 1 itself may be considered a movable device.

[0047] Again in Figure 1, the imaging system 2 images the object OBJ. To image the object OBJ, the imaging system 2 includes an imaging device 21, an imaging device 22, and an illumination device 23.

[0048] Each of imaging devices 21 and 22 is a camera capable of imaging object OBJ. For example, each of imaging devices 21 and 22 may image object OBJ under the control of control device 3. Imaging device 21 generates image data IMD by imaging object OBJ. That is, each of imaging devices 21 and 22 generates image data IMD which is the imaging result of object OBJ. In the following description, as necessary, the image data IMD generated by imaging device 21 is referred to as image data IMD_2D, and the image data IMD generated by imaging device 22 is referred to as image data IMD_3D to distinguish the two. Also, in the following description, image data IMD is assumed to mean at least one of image data IMD_2D and IMD_3D. The image data IMD_2D and IMD_3D respectively generated by imaging devices 21 and 22 are output from imaging devices 21 and 22 to control device 3 respectively. As a result, control device 3 acquires image data IMD_2D and IMD_3D.

[0049] Imaging device 21 may be a monocular camera. In this case, imaging device 21 may be capable of imaging object OBJ using a monocular camera (in other words, an image sensor). Specifically, because imaging device 21 is a monocular camera, imaging device 21 generates one image data generated by the monocular camera as image data IMD_2D. In this case, the image data IMD_2D corresponding to one image data (that is, indicating one image) may be referred to as monocular image data.

[0050] On the other hand, imaging device 22 may be a stereo camera. Specifically, imaging device 22 may be a stereo camera capable of imaging object OBJ using two monocular cameras (in other words, two image sensors). In this case, imaging device 22 generates image data IMD_3D including two image data respectively generated by the two monocular cameras. In this case, the image data IMD_3D including two image data (that is, indicating two images) may be referred to as stereo image data.

[0051] At least one of the imaging devices 21 and 22 may image the entire object OBJ. Alternatively, at least one of the imaging devices 21 and 22 may image a portion of the object OBJ. In other words, at least one of the imaging devices 21 and 22 may image a portion of the object OBJ while not imaging other portions of the object OBJ.

[0052] At least one of the imaging devices 21 and 22 may image a single object OBJ. That is, the image shown by the image data IMD may contain a single object OBJ. Alternatively, at least one of the imaging devices 21 and 22 may image multiple object OBJs. That is, the image shown by the image data IMD may contain multiple object OBJs. In this case, as will be described in detail later, the control device 3 may determine (in other words, select) one of the multiple object OBJs imaged by the imaging device 21 as the target object OBJ_tgt that the end effector 4 will actually perform a predetermined processing on.

[0053] When at least one of the imaging devices 21 and 22 images multiple objects OBJ, the multiple objects OBJ imaged by at least one of the imaging devices 21 and 22 may be arranged such that at least two of the multiple objects OBJ overlap at least partially. For example, if the objects OBJ are workpieces W corresponding to parts used to manufacture a desired product, the multiple workpieces W (i.e., multiple parts) may be arranged such that at least two of the multiple workpieces W overlap at least partially. In this case, the end effector 4 may perform the above-described holding process by holding at least one workpiece W (i.e., target object OBJ_tgt) from the multiple randomly arranged workpieces W. In other words, the robot 1 may perform random picking, picking workpieces W one by one from the multiple randomly arranged workpieces W. Note that the multiple randomly arranged workpieces W can also be described as multiple irregularly arranged workpieces W, multiple haphazardly arranged workpieces W, or multiple randomly arranged workpieces W. For example, if multiple workpieces W are housed in the container CB described above, the end effector 4 may perform the above-described holding process to hold at least one workpiece W (i.e., the target object OBJ_tgt) from among the multiple workpieces W that are haphazardly placed in the container CB.

[0054] In this embodiment, the state of "two objects overlapping" may include the state of "two objects overlapping while in contact." The state of "two objects overlapping" may include the state of "two objects overlapping while not in contact." The state of "two objects overlapping" may include the state of "two objects entangled." The state of "two objects overlapping" may include the state of "one of the two objects at least partially covering the other of the two objects." The state of "two objects overlapping" may include the state of "one of the two objects being located on top of at least a portion of the other of the two objects."

[0055] Alternatively, the multiple objects OBJ imaged by at least one of the imaging devices 21 and 22 may be arranged regularly. For example, if the object OBJ is a workpiece W corresponding to a part used to manufacture a desired product, then multiple workpieces W (i.e., multiple parts) may be arranged in a matrix. In this case, the end effector 4 may perform the above-described holding process to hold at least one workpiece W (i.e., target object OBJ_tgt) from the multiple regularly arranged workpieces W. In other words, the robot 1 may pick workpieces W one by one from the multiple regularly arranged workpieces W. Note that the multiple regularly arranged workpieces W can also be described as multiple workpieces W arranged in an orderly manner, or multiple workpieces W arranged according to a certain arrangement rule.

[0056] Furthermore, the placement process for arranging the object OBJ held by the end effector 4 may include a placement process that arranges the object OBJ randomly (in other words, haphazardly), or it may include a placement process that arranges the object OBJ regularly (in other words, in an orderly manner), similar to the holding process. For example, the robot 1 may use the end effector 4 to hold one workpiece W from among a plurality of workpieces W arranged regularly or randomly in the first mounting device T, and then arrange the workpiece W held by the end effector 4 regularly or randomly in the second mounting device T.

[0057] When each of the imaging devices 21 and 22 images multiple objects OBJ, the multiple objects OBJ imaged by imaging device 21 may be the same as the multiple objects OBJ imaged by imaging device 22. In other words, the multiple objects OBJ captured in the image shown by image data IMD_2D may be the same as the multiple objects OBJ captured in the image shown by image data IMD_3D. Alternatively, at least one of the multiple objects OBJ imaged by imaging device 21 may be different from at least one of the multiple objects OBJ imaged by imaging device 22. In other words, at least one of the multiple objects OBJ captured in the image shown by image data IMD_2D may be the same as at least one of the multiple objects OBJ captured in the image shown by image data IMD_3D.

[0058] The illumination device 23 is a device capable of irradiating an object OBJ with illumination light. If there are multiple objects OBJ, the illumination device 23 is a device capable of irradiating at least one of the multiple objects OBJ with illumination light. For example, the illumination device 23 may irradiate an object OBJ with illumination light under the control of the control device 3. In particular, the illumination device 23 is a device capable of illuminating an object OBJ with illumination light by irradiating the object OBJ with illumination light. In this case, at least one of the imaging devices 21 and 22 may image the object OBJ that is illuminated with illumination light. However, the illumination device 23 does not have to irradiate an object OBJ with illumination light. In this case, the imaging system 2 (robot system SYS) does not have to be equipped with the illumination device 23.

[0059] The illumination device 23 may be a device capable of projecting a desired projection pattern onto an object OBJ by irradiating the object OBJ with illumination light. The illumination device 23 capable of projecting a projection pattern onto an object OBJ may also be referred to as a projection device. The desired projection pattern may include, for example, a random pattern. The random pattern may include a random dot pattern. The desired projection pattern may include, for example, a one-dimensional or two-dimensional grid pattern. The desired projection pattern may include, for example, a linear pattern. The desired projection pattern may include, for example, a striped pattern. The desired projection pattern may include other projection patterns. The desired projection pattern can also be described as light with a desired intensity distribution.

[0060] At least one of the imaging devices 21 and 22 may image the object OBJ onto which the projection pattern from the illumination device 23 is projected. In this case, the image shown in the image data IMD will include the object OBJ onto which the projection pattern is projected. For example, imaging device 22 may image the object OBJ onto which the projection pattern is projected. On the other hand, imaging device 21 does not have to image the object OBJ onto which the projection pattern is projected.

[0061] The imaging system 2, like the end effector 4, is attached to (or rather, provided on) the robot arm 12. That is, the imaging devices 21 and 22 and the illumination device 23 are attached to the robot arm 12. For example, as shown in Figure 2, the imaging devices 21 and 22 and the illumination device 23 may be attached to the tip of the robot arm 12, like the end effector 4. In this case, the imaging devices 21 and 22 and the illumination device 23 are movable by the movement of the robot arm 12. That is, the robot arm 12 moves the imaging devices 21 and 22 and the illumination device 23.

[0062] The imaging device 21 may image object OBJ during a period when either the imaging device 21 or object OBJ is moving relative to the other. A state in which either the imaging device 21 or object OBJ is not moving relative to the other may mean that the relative positional relationship between the imaging device 21 and object OBJ is changing. In this case, the imaging device 21 does not need to be stationary in order to image object OBJ, so the robot system SYS can efficiently perform predetermined processing on object OBJ using the end effector 4. Similarly, the imaging device 22 may image object OBJ during a period when either the imaging device 22 or object OBJ is moving relative to the other. A state in which either the imaging device 22 or object OBJ is not moving relative to the other may mean that the relative positional relationship between the imaging device 22 and object OBJ is changing. In this case, the imaging device 22 does not need to remain stationary in order to image the object OBJ, so the robot system SYS can efficiently perform predetermined processing on the object OBJ using the end effector 4.

[0063] Alternatively, the imaging device 21 may image object OBJ during a period when neither the imaging device 21 nor object OBJ is moving relative to the other. Note that the state in which neither the imaging device 21 nor object OBJ is moving relative to the other may mean that the relative positional relationship between the imaging device 21 and object OBJ is fixed. Similarly, the imaging device 22 may image object OBJ during a period when neither the imaging device 22 nor object OBJ is moving relative to the other. Note that the state in which neither the imaging device 22 nor object OBJ is moving relative to the other may mean that the relative positional relationship between the imaging device 22 and object OBJ is fixed.

[0064] The imaging devices 21 and 22 may image the object OBJ while synchronizing with each other. For example, the imaging devices 21 and 22 may image the object OBJ simultaneously. However, the imaging devices 21 and 22 do not have to image the object OBJ simultaneously.

[0065] The control device 3 performs robot control processing. The robot control processing includes the process of generating robot control signals for controlling the robot 1. Specifically, the control device 3 generates robot control signals based on image data IMD output from the imaging system 2. In this embodiment, the control device 3 calculates at least one of the position and orientation of object OBJ within the reference coordinate system of the robot system SYS based on the image data IMD, and generates robot control signals based on at least one of the calculated position and orientation of object OBJ.

[0066] The reference coordinate system is the coordinate system that serves as the reference for the robot system SYS. For example, the reference coordinate system may be the coordinate system that serves as the reference for robot 1. Furthermore, the reference coordinate system can also be said to be the coordinate system used to control robot 1. A global coordinate system (in other words, a world coordinate system) may be used as the reference coordinate system. The global coordinate system may be a coordinate system determined with respect to the support surface S on which robot 1 is placed. That is, the global coordinate system may be a coordinate system fixed to the support surface S on which robot 1 is placed. A robot coordinate system may be used as the reference coordinate system. The robot coordinate system may be a coordinate system determined with respect to robot 1. That is, the robot coordinate system may be a coordinate system fixed to robot 1 (for example, fixed to the base 11 of robot 1). A 2D imaging coordinate system may be used as the reference coordinate system. The 2D imaging coordinate system may be a coordinate system determined with respect to the imaging device 21. That is, the 2D imaging coordinate system may be a coordinate system fixed to the imaging device 21. One example of a 2D imaging coordinate system is a coordinate system determined based on the optical axis AX21 (see Figure 2) of the optical system (particularly the terminal optical elements such as the objective lens) of the imaging device 21. Another example of a 2D imaging coordinate system is a coordinate system in which one of the three coordinate axes constituting the imaging coordinate system is an axis along the optical axis AX21 of the optical system of the imaging device 21. A 3D imaging coordinate system may be used as a reference coordinate system. The 3D imaging coordinate system may be a coordinate system determined with respect to the imaging device 22. In other words, the 3D imaging coordinate system may be a coordinate system fixed to the imaging device 22. One example of a 3D imaging coordinate system is a coordinate system determined based on the optical axis AX22 (see Figure 2) of the optical system (particularly the terminal optical elements such as the objective lens) of the imaging device 22. Another example of a 3D imaging coordinate system is a coordinate system in which one of the three coordinate axes constituting the imaging coordinate system is an axis along the optical axis AX22 of the optical system of the imaging device 22.

[0067] In the following explanation, unless otherwise specified, the X-axis, Y-axis, and Z-axis may refer to the X-axis, Y-axis, and Z-axis in the reference coordinate system, respectively.

[0068] In addition to or instead of performing robot control processing, the control device 3 may perform end-effector control processing. The end-effector control processing may include processing to generate an end-effector control signal for controlling the end-effector 4. For example, the control device 3 may generate an end-effector control signal based on image data IMD output from the imaging system 2. Specifically, the control device 3 may generate an end-effector control signal based on at least one of the position and orientation of the object OBJ calculated from the image data IMD.

[0069] Furthermore, the end effector control process may or may not be included in the robot control process. In other words, the end effector control signal generated by the control device 3 may or may not be included in the robot control signal. For the sake of clarity, the following explanation will describe an example in which the end effector control process is included in the robot control process (i.e., the end effector control signal is included in the robot control signal). For this reason, in the following explanation, the robot control process may mean the process of generating at least one of the robot control signal and the end effector control signal. Also, in the following explanation, the robot control signal may mean at least one of the signals for controlling the robot 1 and the signals for controlling the end effector 4. Furthermore, the robot control signal may simply be called a control signal. The robot control signal may also be called robot control information or control information.

[0070] As described above, if the robot 1 is installed in a robotic mobile device (for example, at least one of an automated guided vehicle, an autonomous mobile transport robot, an unmanned aerial vehicle, and a submarine), the control device 3 may perform mobile device control processing in addition to, or instead of, performing robot control processing. The mobile device control processing may include processing to generate a mobile device control signal for controlling the robotic mobile device. For example, the control device 3 may generate a mobile device control signal based on image data IMD output from the imaging system 2 (for example, the imaging device 21). Specifically, the control device 3 may generate a mobile device control signal based on at least one of the position and orientation of the object OBJ calculated from the image data IMD.

[0071] Furthermore, the movable device control process may or may not be included in the robot control process. In other words, the movable device control signal generated by the control device 3 may or may not be included in the robot control signal. For the sake of explanation, the following description will explain an example in which the movable device control process is included in the robot control process (i.e., the movable device control signal is included in the robot control signal). For this reason, in the following description, the robot control process may mean the process of generating at least one of the robot control signal, the end effector control signal, and the movable device control signal. Also, in the following description, the robot control signal may mean at least one of the signals for controlling the robot 1, the signals for controlling the end effector 4, and the signals for controlling the robot's movable device.

[0072] Thus, the control device 3 and the imaging system 2 are used to control the robot 1. For this reason, the system including the control device 3 and the imaging system 2 may be referred to as a robot control system or a control system.

[0073] The robot control signal generated by the control device 3 is output to the robot control device 13 of the robot 1. The robot control device 13 controls the operation of the robot 1 based on the robot control signal generated by the control device 3. For this reason, the robot control signal may include signals for controlling the operation of the robot 1.

[0074] As described above, if the robot control signal includes a signal for controlling the robot arm 12, the robot control device 13 may control the robot arm 12 based on the robot control signal. For example, the robot control device 13 may control the operation of the robot arm 12 by controlling the operation of an actuator built into the joint 122 based on the robot control signal.

[0075] For example, as described above, the robot arm 12 moves the end effector 4. In this case, the robot control signal may include signals for controlling the robot arm 12 so that the end effector 4 is positioned at a desired location. The robot control signal may include signals for controlling the robot arm 12 so that the end effector 4 moves to a desired location. The robot control signal may include signals for controlling the robot arm 12 so that the positional relationship between the end effector 4 and the object OBJ becomes a desired positional relationship. In this case, the robot control device 13 may control the robot arm 12 based on the robot control signal so that the end effector 4 is positioned at a desired location. The robot control device 13 may control the robot arm 12 based on the robot control signal so that the end effector 4 moves to a desired location. The robot control device 13 may control the robot arm 12 based on the robot control signal so that the positional relationship between the end effector 4 and the object OBJ becomes a desired positional relationship.

[0076] For example, when the end effector 4 performs a holding process to hold an object OBJ, the robot control signal may include a signal to control the robot arm 12 so that the end effector 4 moves toward (i.e., approaches) a first desired position in which the end effector 4 can hold the object OBJ. In other words, the robot control signal may include a signal to control the robot arm 12 so that the end effector 4 is positioned at the first desired position. In this case, the robot control device 13 may control the robot arm 12 based on the robot control signal so that the end effector 4 moves toward (i.e., approaches) the first desired position. In other words, the robot control signal may control the robot arm 12 so that the end effector 4 is positioned at the first desired position.

[0077] For example, when the end effector 4 performs a holding process to hold an object OBJ, the robot control signal may include a signal to control the robot arm 12 so that the posture of the end effector 4 is a first desired posture in which the end effector 4 can hold the object OBJ. In this case, the robot control device 13 may control the robot arm 12 based on the robot control signal so that the posture of the end effector 4 is the first desired posture.

[0078] As another example, when performing a release process to release an object OBJ held by the end effector 4, the robot control signal may include a signal to control the robot arm 12 so that the end effector 4 moves toward (i.e., approaches) a second desired position from which the object OBJ held by the end effector 4 should be released. In other words, the robot control signal may include a signal to control the robot arm 12 so that the end effector 4 is positioned at the second desired position. In this case, the robot control device 13 may control the robot arm 12 based on the robot control signal so that the end effector 4 moves toward (i.e., approaches) the second desired position. In other words, the robot control signal may control the robot arm 12 so that the end effector 4 is positioned at the second desired position.

[0079] As another example, when performing a release process to release an object OBJ held by the end effector 4, the robot control signal may include a signal to control the robot arm 12 so that the posture of the end effector 4 is a second desired posture in which the end effector 4 can release the object OBJ. In this case, the robot control device 13 may control the robot arm 12 based on the robot control signal so that the posture of the end effector 4 is the second desired posture.

[0080] As described above, if the robot control signal includes a signal for controlling the end effector 4, the robot control device 13 may control the end effector 4 based on the robot control signal. For example, the robot control device 13 may control the operation of the end effector 4 by controlling the operation of an actuator that moves a hand gripper constituting the end effector 4 based on the robot control signal. For example, the robot control device 13 may control the operation of the end effector 4 by controlling the operation of a vacuum device of a vacuum gripper constituting the end effector 4 based on the robot control signal. For example, the robot control device 13 may control the operation of the end effector 4 by controlling the operation of a magnetic adsorption device of a magnet gripper constituting the end effector 4 based on the robot control signal.

[0081] As an example, when the end effector 4 performs a holding process to hold the object OBJ, the robot control signal may include a signal to control the end effector 4 so that the end effector 4, when positioned at the first desired position and / or in the first desired posture described above, holds the object OBJ. In this case, the robot control device 13 may control the end effector 4 based on the robot control signal so that the end effector 4, when positioned at the first desired position and / or in the first desired posture, holds the object OBJ.

[0082] As another example, when performing a release process to release an object OBJ held by the end effector 4, the robot control signal may include a signal to control the end effector 4 to release the object OBJ held by the end effector 4 when it is located at the second desired position and / or in the second desired posture described above. In this case, the robot control device 13 may control the end effector 4 based on the robot control signal to release the object OBJ held by the end effector 4 when it is located at the second desired position and / or in the second desired posture.

[0083] As described above, if the robot control signal includes signals for controlling the robotic mobile device (e.g., at least one of an automated guided vehicle, an autonomous mobile transport robot, an unmanned aerial vehicle, and a submarine) on which the robot 1 is installed, the robot control device 13 may control the robotic mobile device based on the robot control signal. For example, the robot control device 13 may control the power source (e.g., a motor or engine) of the robotic mobile device based on the robot control signal so that the robot 1 moves to the target position indicated directly or indirectly by the robot control signal.

[0084] In the above description, an example is given in which a robot 1 equipped with a robot arm 12 is installed on a robotic movable device. In this case, the end effector 4 attached to the robot arm 12 may be considered to be installed on the robotic movable device via the robot 1 (for example, via the robot arm 12). On the other hand, the end effector 4 may be installed on the robotic movable device without going through the robot 1 (for example, without going through the robot arm 12). In this case, the robot 1 (for example, the robot arm 12) may not be installed on the robotic movable device. In this case, the control device 3 may generate at least one of a movable device control signal for controlling the robotic movable device on which the end effector 4 is installed and an end effector control signal for controlling the end effector 4 installed on the robotic movable device, without generating a robot control signal for controlling the robot 1.

[0085] Furthermore, the imaging system 2 may be installed on the robotic movable device. In this case, the control device 3 may generate at least one of the movable device control signals for controlling the robotic movable device on which the end effector 4 is installed, and an end effector control signal for controlling the end effector 4 installed on the robotic movable device, based on the image data IMD generated by the imaging system 2 installed on the robotic movable device. Also, if the end effector 4 (and, in some cases, the imaging system 2) is installed on the robotic movable device, the control device 3 may also be located inside the robotic movable device.

[0086] The robot control signal may include signals that can be directly used by the robot control device 13 to control the operation of the robot 1. The robot control signal may include signals that can be directly used as robot drive signals that the robot control device 13 uses to control the operation of the robot 1. In this case, the robot control device 13 may use the robot control signal as is to control the operation of the robot 1. For example, the control device 3 may generate a drive signal for an actuator built into the joint 122 of the robot arm 12 as a robot control signal, and the robot control device 13 may use the robot control signal generated by the control device 3 as is to control the actuator built into the joint 122 of the robot arm 12.

[0087] The robot control signal may include signals that can be directly used by the robot control device 13 to control the operation of the end effector 4. The robot control signal may include signals that can be directly used as end effector drive signals used by the robot control device 13 to control the operation of the end effector 4. In this case, the robot control device 13 may use the robot control signal as is to control the operation of the end effector 4. For example, the control device 3 may generate a drive signal (end effector drive signal) for the actuator that moves the hand gripper constituting the end effector 4 as a robot control signal, and the robot control device 13 may use the robot control signal generated by the control device 3 as is to control the actuator of the end effector 4. For example, the control device 3 may generate a drive signal (end effector drive signal) for driving the vacuum device of the vacuum gripper constituting the end effector 4 as a robot control signal, and the robot control device 13 may use the robot control signal generated by the control device 3 as is to control the vacuum device of the end effector 4. For example, the control device 3 may generate a drive signal (end effector drive signal) as a robot control signal to drive the magnetic adsorption device of the magnet gripper constituting the end effector 4, and the robot control device 13 may use the robot control signal generated by the control device 3 as is to control the magnetic adsorption device of the end effector 4.

[0088] The robot control signal may include signals that can be directly used by the robot control device 13 to control the operation of the robot movable device on which the robot 1 is installed. The robot control signal may also include signals that can be directly used as movable device drive signals used by the robot control device 13 to control the operation of the robot movable device. In this case, the robot control device 13 may use the robot control signal as is to control the operation of the robot movable device. For example, the control device 3 may generate a power source drive signal (movable device drive signal) for driving the power source (e.g., a motor or engine) of the robot movable device as a robot control signal, and the robot control device 13 may use the robot control signal generated by the control device 3 as is to control the power source of the robot movable device.

[0089] As described above, if the robot control signal includes signals that can be directly used by the robot control device 13 to control the operation of at least one of the robot 1, the end effector 4, and the robot movable device, the robot 1 does not need to be equipped with the robot control device 13. In this case, the control device 3 may control an actuator built into the joint 122 of the robot arm 12 by the robot control signal. Alternatively, for example, the control device 3 may control an actuator that moves a hand gripper constituting the end effector 4 by the robot control signal (end effector drive signal). For example, the control device 3 may control the vacuum device of a vacuum gripper constituting the end effector 4 by the robot control signal (end effector drive signal). For example, the control device 3 may control the magnetic adsorption device of a magnet gripper constituting the end effector 4 by the robot control signal (end effector drive signal). For example, the control device 3 may control the robot movable device on which the robot 1 is installed by the robot control signal (movable device drive signal).

[0090] Alternatively, the robot control signal may include signals that the robot control device 13 can use to generate robot drive signals for controlling the movement of robot 1. In this case, the robot control device 13 may generate robot drive signals for controlling the movement of robot 1 based on the robot control signal, and control the movement of robot 1 based on the generated robot drive signals. For example, the robot control device 13 may generate robot drive signals for driving an actuator built into the joint 122 of the robot arm 12 based on the robot control signal, and control the actuator built into the joint 122 of the robot arm 12 based on the generated robot drive signals.

[0091] The robot control signal may include signals that the robot control device 13 can use to generate an end effector drive signal for controlling the operation of the end effector 4. In this case, the robot control device 13 may generate an end effector drive signal for controlling the operation of the end effector 4 based on the robot control signal, and control the operation of the end effector 4 based on the generated end effector drive signal. For example, if the end effector 4 is a hand gripper, the robot control device 13 may generate an end effector drive signal for driving the actuator of the hand gripper based on the robot control signal, and control the actuator of the hand gripper based on the generated end effector drive signal. For example, if the end effector 4 is a magnet gripper, the robot control device 13 may generate an end effector drive signal for driving the magnetic adsorption device of the magnet gripper based on the robot control signal, and control the magnetic adsorption device based on the generated end effector drive signal.

[0092] Furthermore, if the robot system SYS is equipped with a control device (end effector control device) for controlling the end effector 4 separately from the robot control device 13, the robot control signal may include signals that can be used by the end effector control device to generate end effector drive signals for controlling the operation of the end effector 4. In this case, the end effector control device may generate end effector drive signals for controlling the operation of the end effector 4 based on the robot control signal, and control the operation of the end effector 4 based on the generated end effector drive signals.

[0093] The robot control signal may include signals that can be used by the robot control device 13 to generate a movable device drive signal for controlling the operation of the robot movable device on which the robot 1 is installed. In this case, the robot control device 13 may generate a movable device drive signal for controlling the operation of the robot movable device based on the robot control signal, and control the operation of the robot movable device based on the generated movable device drive signal. For example, the robot control device 13 may generate a power source drive signal (movable device drive signal) for driving the power source (e.g., a motor or engine) of the robot movable device based on the robot control signal, and control the power source of the robot movable device based on the generated movable device drive signal.

[0094] Furthermore, if the robot system SYS is equipped with a control device (robot motion control device) for controlling the robot's movable device, separate from the robot control device 13, the robot control signal may include signals that can be used by the robot motion control device to generate movable device drive signals for controlling the operation of the robot's movable device. In this case, the robot motion control device may generate movable device drive signals for controlling the operation of the robot's movable device based on the robot control signal, and control the operation of the robot's movable device based on the generated movable device drive signals.

[0095] The signals available to the robot control device 13 for generating robot drive signals may include signals representing at least one of the position and orientation of the object OBJ in the reference coordinate system. The signals available to the robot control device 13 for generating robot drive signals may also include signals representing a desired positional relationship between the robot 1 and the object OBJ in the reference coordinate system.

[0096] The signals available to the robot control device 13 for generating robot drive signals may include signals representing a target value (target position) for the position of the end effector 4 in the reference coordinate system. An example of a target position is the processing position where the end effector 4 should process the object OBJ. For example, the target position may include the position where the end effector 4 should hold the object OBJ. For example, the target position may include the position where the end effector 4 should release the object OBJ. The signals available to the robot control device 13 for generating robot drive signals may include signals representing a target value for the position of the tip of the robot arm 12 (e.g., the tool center point) in the reference coordinate system. The signals available to the robot control device 13 for generating robot drive signals may include signals representing a target value for the position of the imaging system 2 in the reference coordinate system.

[0097] The signals available to the robot control device 13 for generating robot drive signals may include signals representing a target value for the pose of the end effector 4 in the reference coordinate system (target pose). An example of a target pose is the pose that the end effector 4 should take when it processes an object OBJ (processing pose). For example, the target pose may include the pose that the end effector 4 should take when it holds an object OBJ. For example, the target pose may include the pose that the end effector 4 should take when it releases an object OBJ. The signals available to the robot control device 13 for generating robot drive signals may include signals representing a target value for the pose of the tip of the robot arm 12 (e.g., the tool center point) in the reference coordinate system. The signals available to the robot control device 13 for generating robot drive signals may include signals representing a target value for the pose of the imaging system 2 in the reference coordinate system.

[0098] The signals available to the robot control device 13 for generating robot drive signals may include signals representing the amount and direction of movement from the current position of the end effector 4 to the target position of the end effector 4.

[0099] Next, the model generation device 5 is a device capable of generating (in other words, constructing) a learning model by machine learning. That is, the model generation device 5 is a device capable of performing a model generation process to generate a learning model by machine learning. The learning model may also be called a computational model or a machine learning model. An example of a learning model that can be generated by machine learning is a computational model including a neural network (so-called artificial intelligence (AI)).

[0100] In this embodiment, the model generation device 5 can generate an object recognition model 50, which is an example of a learned model, by machine learning. The object recognition model 50 is a learned model that, when an image is input, can output object region information OAI (see Figure 12) relating to the object region WA (see Figure 12) in which an object OBJ exists within the input image.

[0101] In this embodiment, the object recognition model 50 is a learning model that, when image data IMD_2D is input, outputs object region information OAI (see Figure 12) relating to the object region WA (see Figure 12) in which an object OBJ exists within the image shown by the input image data IMD_2D. For the purposes of the following explanation, the image shown by the image data IMD_2D will be referred to as the 2D image IMG_2D. For example, if a workpiece W, which is an example of an object OBJ, is captured in the 2D image IMG_2D, the object recognition model 50 may, when image data IMD_2D is input, output object region information OAI relating to the object region WA (which may be called the work region in this case) in which the workpiece W exists (i.e., is captured) within the 2D image IMG_2D shown by the input image data IMD_2D. For example, if a mounting device T, which is an example of an object OBJ, is captured in the 2D image IMG_2D, the object recognition model 50 may, upon receiving the image data IMD_2D, output object region information OAI relating to the object region WA (which may be called the mounting device region in this case) in the 2D image IMG_2D indicated by the input image data IMD_2D in which the mounting device T exists (i.e., is captured).

[0102] The model generation device 5 may perform machine learning to generate the object recognition model 50 such that, when image data IMD_2D is input, the object recognition model 50 generated by machine learning outputs object region information OAI related to the object region WA in which the object OBJ is captured within the 2D image IMG_2D shown by the input image data IMD_2D.

[0103] The object recognition model 50 generated by the model generation device 5 is output from the model generation device 5 to the control device 3. The control device 3 uses the object recognition model 50 obtained from the model generation device 5 to calculate at least one of the position and orientation of the object OBJ.

[0104] Details of the model generation device 5, the object recognition model 50, the operation for generating the object recognition model 50, and the operation for calculating at least one of the position and orientation of the object OBJ using the object recognition model 50 will be described in detail later, so their explanation here is omitted.

[0105] (1-2) Configuration of the control device 3 Next, the configuration of the control device 3 will be described with reference to Figure 3. Figure 3 is a block diagram showing the configuration of the control device 3.

[0106] As shown in Figure 3, the control device 3 comprises an arithmetic unit 31, a storage device 32, and a communication device 33. Furthermore, the control device 3 may also include an input device 34 and an output device 35. However, the control device 3 does not have to include at least one of the input device 34 and the output device 35. The arithmetic unit 31, the storage device 32, the communication device 33, the input device 34, and the output device 35 may be connected via a data bus 36. The arithmetic unit 31, the storage device 32, the communication device 33, the input device 34, and the output device 35 may also be referred to as the arithmetic unit, the storage unit, the communication unit, the input unit, and the output unit, respectively.

[0107] The arithmetic unit 31 is hardware that includes at least one circuit (for example, at least one of an electronic circuit and an electrical circuit). For this reason, the arithmetic unit 31 may be referred to as a circuit group.

[0108] The arithmetic unit 31 includes at least one processor (i.e., one or more processors) as hardware. The processor may include, for example, a processor conforming to a von Neumann computer architecture. A processor conforming to a von Neumann computer architecture may include at least one of a CPU (Central Processing Unit) and a GPU (Graphics Processing Unit). The processor may also include, for example, a processor conforming to a non-von Neumann computer architecture. A processor conforming to a non-von Neumann computer architecture may include at least one of an FPGA (Field Programmable Gate Array) and an ASIC (Application Specific Circuit). The processor may be implemented by a group of circuits (e.g., at least one of an electronic circuit and an electrical circuit).

[0109] The arithmetic unit 31 reads a computer program 321 which includes at least one of computer program code and computer program instructions. For example, the arithmetic unit 31 may read a computer program 321 stored in a storage device 32. For example, the arithmetic unit 31 may read a computer program 321 stored in a computer-readable and non-temporary recording medium using a recording medium reader (not shown) provided by the control device 3. The computer program 321 read from the recording medium may be stored in the storage device 32. The arithmetic unit 31 may obtain (i.e., download or read) a computer program 321 from a device (not shown) located outside the control device 3 via a communication device 33 (or other communication device). The downloaded computer program 321 may be stored in the storage device 32.

[0110] The arithmetic unit 31 executes the loaded computer program 321. As a result, a logical functional block for executing the processing that the control device 3 should perform (for example, the robot control processing described above) is realized within the arithmetic unit 31. In other words, the arithmetic unit 31, together with the storage device 32 on which the computer program 321 is recorded (in other words, together with the storage device 32 and the computer program 321 recorded in the storage device 32, etc.), can function as a controller or computer for realizing a logical functional block for executing the processing that the control device 3 should perform. That is, together with at least one processor in the arithmetic unit 31, the memory (recording medium) in the storage device 32, etc., and the computer program 321 are configured so that the control device 3 performs the processing that the control device 3 should perform (for example, the robot control processing described above).

[0111] The arithmetic unit 31 may include a single processor. In this case, the arithmetic unit 31 may use a single processor to perform the processing that the control device 3 should perform (for example, the robot control processing described above). For example, if the arithmetic unit 31 performs a first operation (for example, a first process that is part of the robot control processing) and a second operation (for example, a second process that is another part of the robot control processing), the arithmetic unit 31 may use a single processor to perform both the first and second operations. Alternatively, the arithmetic unit 31 may include multiple processors. In this case, the arithmetic unit 31 may use any one of the multiple processors to perform the processing that the control device 3 should perform (for example, the robot control processing described above). For example, if the arithmetic unit 31 includes a first and a second processor and performs the first and second operations, the arithmetic unit 31 may use any one of the first and second processors to perform the first and second operations, respectively. For example, the arithmetic unit 31 may perform a first operation using the first processor, or a second operation using the first processor, or a first operation using the second processor, or a second operation using the second processor.

[0112] The computing device 31 may implement a computational model that can be constructed by machine learning by executing a computer program 321. An example of a computational model that can be constructed by machine learning is a computational model that includes a neural network (so-called artificial intelligence (AI)). In this case, the learning of the computational model may include learning the parameters of the neural network (for example, at least one of the weights and biases). The computing device 31 may use the computational model to perform robot control processing. In other words, the operation of performing robot control processing may include the operation of performing robot control processing using the computational model. Furthermore, the computing device 31 may implement a computational model that has been constructed by offline machine learning using training data. In addition, the computational model implemented in the computing device 31 may be updated by online machine learning on the computing device 31. Alternatively, the arithmetic unit 31 may perform robot control processing using an arithmetic model implemented in an external device (i.e., a device provided outside the control device 3) in addition to or instead of the arithmetic model implemented in the arithmetic unit 31.

[0113] Furthermore, as the recording medium for recording the computer program 321 executed by the arithmetic unit 31, at least one of the following may be used: optical discs such as CD-ROM, CD-R, CD-RW, flexible disk, MO, DVD-ROM, DVD-RAM, DVD-R, DVD+R, DVD-RW, DVD+RW, and Blu-ray (registered trademark); magnetic media such as magnetic tape; magneto-optical disks; semiconductor memory such as USB memory; and any other medium capable of storing a program. The recording medium may also include equipment capable of recording the computer program 321 (for example, general-purpose or dedicated equipment on which the computer program 321 is implemented in an executable state in at least one form such as software and firmware). Furthermore, each process and function included in the computer program 321 may be realized by logical processing blocks implemented within the arithmetic unit 31 (i.e., the processor) when the arithmetic unit 31 executes the computer program 321, or by hardware such as a predetermined gate array (FPGA (Field Programmable Gate Array), ASIC (Application Specific Integrated Circuit)) provided by the arithmetic unit 31, or in a form in which logical processing blocks and partial hardware modules that realize some elements of the hardware are mixed.

[0114] Figure 3 shows an example of a logical functional block implemented within the arithmetic unit 31 for executing robot control processing. As shown in Figure 3, the arithmetic unit 31 implements a position and orientation calculation unit 311 and a signal generation unit 312. The processing performed by the position and orientation calculation unit 311 and the signal generation unit 312 will be described in detail later, so their explanation is omitted here.

[0115] The storage device 32 includes at least one memory capable of storing desired data. In other words, the storage device 32 includes at least one memory containing desired data. The memory may be implemented by a group of circuits (for example, at least one of an electronic circuit and an electrical circuit). For example, the storage device 32 may store a computer program 321 executed by the arithmetic unit 31. In this case, the storage device 32 (memory) may be used as the recording medium described above for recording the computer program 321 executed by the arithmetic unit 31. The storage device 32 may temporarily store data that the arithmetic unit 31 temporarily uses when the arithmetic unit 31 is executing the computer program 321. The storage device 32 may store data that the control device 3 stores long-term. The storage device 32 may include at least one of RAM (Random Access Memory), ROM (Read Only Memory), hard disk drive, magneto-optical disk drive, SSD (Solid State Drive), and disk array drive. In other words, the storage device 32 may include a non-temporary recording medium.

[0116] The communication device 33 can communicate with the robot 1 and the imaging system 2 via a communication network (not shown). Alternatively, the communication device 33 may communicate with other devices different from the robot 1 and the imaging system 2, in addition to or instead of at least one of the robot 1 and the imaging system 2, via a communication network (not shown). In this embodiment, the communication device 33 may receive (i.e., acquire) image data IMD from the imaging system 2. Furthermore, the communication device 33 may receive (i.e., acquire) an object recognition model 50 from the model generation device 5. Furthermore, the communication device 33 may transmit (i.e., output) robot control signals to the robot 1. The communication device 33 that outputs robot control signals to the robot 1 may be referred to as an output unit or output device.

[0117] The input device 34 is a device that receives information input to the control device 3 from outside the control device 3. For example, the input device 34 may include an operating device that can be operated by the user of the control device 3 (for example, at least one of a keyboard, mouse, and touch panel). For example, the input device 34 may include a recording medium reader that can read information recorded as data on a recording medium that can be attached externally to the control device 3.

[0118] Furthermore, information can be input as data to the control device 3 from an external device via the communication device 33. In this case, the communication device 33 may function as an input device that receives information input to the control device 3 from an external source.

[0119] The output device 35 is a device that outputs information to the outside of the control device 3. For example, the output device 35 may output information as an image. That is, the output device 35 may include a display device (so-called display) 37 capable of displaying an image. In this case, the arithmetic unit 31 may generate a display signal for displaying the image on the display device 37. Specifically, the arithmetic unit 31 may generate a display control signal as a display signal for controlling the display device 37 to display an image. The arithmetic unit 31 may output the generated display signal (display control signal) to the display device 37 via the data bus 36. The display device 37 may display the image based on the display signal (display control signal) generated by the arithmetic unit 31.

[0120] The output device 35 may be an output device different from the display device 37. For example, the output device 35 may output information as sound. That is, the output device 35 may include an audio device (so-called speaker) capable of outputting sound. For example, the output device 35 may output information onto paper. That is, the output device 35 may include a printing device (so-called printer) capable of printing desired information onto paper. For example, the output device 35 may output information as data to a recording medium that can be attached externally to the control device 3.

[0121] Furthermore, the control device 3 can output information as data to an external device via the communication device 33. In this case, the communication device 33 may function as an output device that outputs information to an external device of the control device 3.

[0122] Furthermore, the robot control device 13 provided by the robot 1 described above may also have the same configuration as the control device 3. That is, the robot control device 13, as shown in block diagram 4, comprises a computing device 131, a storage device 132, and a communication device 133. In addition, the robot control device 13 may also comprise an input device 134 and an output device 135. However, the robot control device 13 does not have to comprise at least one of the input device 134 and the output device 135. The computing device 131, the storage device 132, the communication device 133, the input device 134, and the output device 135 may be connected via a data bus 136.

[0123] Furthermore, the characteristics of the arithmetic unit 131, memory device 132, communication device 133, input device 134, and output device 135 may be the same as those of the arithmetic unit 31, memory device 32, communication device 33, input device 34, and output device 35, respectively. The above-mentioned description of the control device 3 can be reused as a description of the robot control device 13 by replacing the terms "control device 3," "arithmetic unit 31," "memory device 32," "communication device 33," "input device 34, and output device 35" with the terms "robot control device 13," "arithmetic unit 131," "memory device 132," "communication device 133," "input device 134, and output device 135," respectively. For this reason, in order to avoid redundant explanations, the description of the robot control device 13 will be omitted.

[0124] Furthermore, the model generation device 5 described above may also have the same configuration as the control device 3. That is, the model generation device 5 comprises a arithmetic unit 51, a storage device 52, and a communication device 53. In addition, the model generation device 5 may also include an input device 54 and an output device 55. However, the model generation device 5 does not have to include at least one of the input device 54 and the output device 55. The arithmetic unit 51, the storage device 52, the communication device 53, the input device 54, and the output device 55 may be connected via a data bus 56.

[0125] Furthermore, the characteristics of the arithmetic unit 51, memory device 52, communication device 53, input device 54, and output device 55 may be the same as those of the arithmetic unit 31, memory device 32, communication device 33, input device 34, and output device 35, respectively. The above-mentioned description of the control device 3 can be reused as a description of the model generation device 5 by replacing the terms "control device 3," "arithmetic unit 31," "memory device 32," "communication device 33," "input device 34, and output device 35," respectively, with the terms "model generation device 5," "arithmetic unit 51," "memory device 52," "communication device 53," "input device 54, and output device 55." For this reason, in order to avoid redundant explanations, the description of the model generation device 5 will be omitted.

[0126] (2) Operation of the Robot System SYS Next, the operation of the robot system SYS will be described. In this embodiment, as described above, the model generation device 5 performs model generation processing to generate the object recognition model 50. Furthermore, the control device 3 performs robot control processing to control at least one of the robot 1, the end effector 4, and the robot movable device. For this reason, the following description will explain the model generation processing and the robot control processing in order.

[0127] For the sake of clarity, the following description will explain the model generation process performed when the holding process shown in Figures 6A to 6C is performed. In other words, for the sake of clarity, the following description will explain the model generation process for generating an object recognition model 50 used by the control device 3 to perform robot control processing for the holding process shown in Figures 6A to 6C. Furthermore, for the sake of clarity, the following description will explain the robot control process in which the control device 3 controls the robot 1, the end effector 4, and at least one of the robot movable devices so that the holding process shown in Figures 6A to 6C is performed. However, the model generation process and robot control process described below may also be performed even if a process different from the holding process shown in Figures 6A to 6C is performed. In other words, the situations in which the model generation process and robot control process described below are performed are not limited to the situations in which the holding process shown in Figures 6A to 6C is performed.

[0128] In the example shown in Figures 6A to 6C, the control device 3 may control the robot 1, the end effector 4, and at least one of the robot's movable devices so that the end effector 4 performs a holding process to hold one object OBJ selected as the target object OBJ_tgt from among a plurality of objects OBJ (for example, a plurality of workpieces W) placed on the mounting device T. In this case, as shown in Figure 6A, the control device 3 may control the robot 1, the end effector 4, and at least one of the robot's movable devices so that the end effector 4 approaches the target object OBJ_tgt. In other words, the control device 3 may control the robot 1, the end effector 4, and at least one of the robot's movable devices so that the end effector 4 moves toward the target object OBJ_tgt. Subsequently, as shown in Figure 6B, the control device 3 may control at least the end effector 4 so that the end effector 4 holds the target object OBJ_tgt. Subsequently, as shown in Figure 6C, the control device 3 may control the robot 1, the end effector 4, and at least one of the robot's movable devices so that the end effector 4 holding the target object OBJ_tgt moves away from the mounting device T (in other words, moves away from it).

[0129] (2-1) Model Generation Process First, the model generation process performed by the model generation device 5 will be explained with reference to Figure 7. Figure 7 is a flowchart showing the flow of the model generation process. Note that the model generation process is the process of generating an object recognition model 50, which is an example of a learning model, and therefore can be considered to be the process equivalent to the learning phase in machine learning.

[0130] As shown in Figure 7, the computing unit 51 of the model generation device 5 generates a training dataset 560 used for machine learning to generate an object recognition model 50 (step S11).

[0131] An example of a training dataset 560 is shown in Figure 8A. As shown in Figure 8A, the training dataset 560 includes multiple training data 561. However, the training dataset 560 may also include a single training data 561. Each training data 561 includes training image data 562 and a ground truth label 563. Therefore, each training data 561 is data in which the training image data 562 and the ground truth label 563 are associated.

[0132] As shown in Figure 8B, the training image data 562 may show an image that includes an object OBJ that the robot 1 is scheduled to actually process using the end effector 4. In the following description, the image shown by the training image data 562 will be referred to as the training image 564. For example, the training image data 562 may show a training image 564 that includes a single object OBJ. For example, the training image data 562 may show a training image 564 that includes multiple objects OBJ. In other words, the training image data 562 may show a training image 564 that includes at least one object OBJ.

[0133] As shown in Figure 8B, the training image data 562 may show a training image 564 in which, in addition to or instead of the object OBJ, sample object OBJ_s that simulate the object OBJ are captured. For example, the training image data 562 may show a training image 564 in which a single sample object OBJ_s is captured. For example, the training image data 562 may show a training image 564 in which multiple sample object OBJ_s are captured. In other words, the training image data 562 may show a training image 564 in which at least one sample object OBJ_s is captured.

[0134] The sample object OBJ_s is not identical to the object OBJ that the robot 1 is scheduled to actually process using the end effector 4, but it is an object with the same or similar characteristics as object OBJ. In this case, the sample object OBJ_s may be considered equivalent to object OBJ. The sample object OBJ_s may also be considered substantially identical to object OBJ. For example, the sample object OBJ_s may be an object with the same or similar size as object OBJ. Note that the state in which the size of the sample object OBJ_s and the size of object OBJ are similar may mean that the difference between the size of the sample object OBJ_s and the size of object OBJ is less than or equal to a predetermined size threshold. For example, the sample object OBJ_s may be an object with the same or similar shape as object OBJ. Note that the state in which the shape of the sample object OBJ_s and the shape of object OBJ are similar may mean that the difference between the shape of the sample object OBJ_s and the shape of object OBJ is less than or equal to a predetermined shape threshold. For example, the sample object OBJ_s may be an object whose color is the same as or similar to that of object OBJ. Furthermore, the similarity in color between the sample object OBJ_s and object OBJ may mean that the colors of the sample object OBJ_s and object OBJ are adjacent or close to each other on the color wheel.

[0135] However, the sample object OBJ_s may not be identical to the object OBJ that the robot 1 is scheduled to actually process using the end effector 4, and may include objects with different characteristics from object OBJ. For example, the sample object OBJ_s may be an object of a different size from object OBJ. Note that the state in which the size of the sample object OBJ_s and the size of object OBJ are different may mean that the difference between the size of the sample object OBJ_s and the size of object OBJ exceeds a predetermined size threshold. For example, the sample object OBJ_s may be an object of a different shape from object OBJ. Note that the state in which the shape of the sample object OBJ_s and the shape of object OBJ are different may mean that the difference between the shape of the sample object OBJ_s and the shape of object OBJ exceeds a predetermined shape threshold. For example, the sample object OBJ_s may be an object of a different color from object OBJ. Furthermore, a state in which the color of sample object OBJ_s and the color of object OBJ are different may mean that the colors of sample object OBJ_s and object OBJ are not adjacent or close to each other in the color wheel.

[0136] If the training dataset 560 includes multiple training data 561, each of the multiple training data 561 may include multiple training image data 562 that are different from each other. The multiple training image data 562 may include at least two training image data 562 that each show at least two training images 564 in which the number of object OBJs captured in the training image 564 is different from each other. The multiple training image data 562 may include at least two training image data 562 that each show at least two training images 564 in which the number of sample object OBJ_s captured in the training image 564 is different from each other. The multiple training image data 562 may include at least two training image data 562 that each show at least two training images 564 in which the positions of the object OBJs captured in the training image 564 are different from each other. The multiple training image data 562 may include at least two training image data 562 that each show at least two training images 564 in which the positions of the sample object OBJ_s captured in the training image 564 are different from each other. Multiple training image data 562 may include at least two training image data 562 each showing at least two training images 564 in which the poses of the object OBJ reflected in the training image 564 are different from each other. Multiple training image data 562 may include at least two training image data 562 each showing at least two training images 564 in which the poses of the sample object OBJ_s reflected in the training image 564 are different from each other. Multiple training image data 562 may include at least two training image data 562 each showing at least two training images 564 in which the positional relationship between the object OBJ reflected in the training image 564 and other objects different from the object OBJ is different from each other. Multiple training image data 562 may include at least two training image data 562 each showing at least two training images 564 in which the positional relationship between the sample object OBJ_s reflected in the training image 564 and other objects different from the sample object OBJ_s is different from each other.

[0137] The correct label 563 indicates information about the object region WA in which object OBJ or sample object OBJ_s is captured within the training image 564 indicated by the training image data 562 associated with the correct label 563. For example, if object OBJ is captured in training image 564, the correct label 563 includes information about the object region WA in which object OBJ is captured within training image 564. For example, if sample object OBJ_s is captured in training image 564, the correct label 563 includes information about the object region WA in which sample object OBJ_s is captured within training image 564.

[0138] If the training image 564 contains multiple object OBJs, the correct label 563 may include information about multiple object regions WA in which each of the multiple object OBJs is captured within the training image 564. If the training image 564 contains multiple sample object OBJ_s, the correct label 563 may include information about multiple object regions WA in which each of the multiple sample object OBJ_s is captured within the training image 564.

[0139] The object region WA may include regions where an object OBJ or sample object OBJ_s exists, but may not include regions where an object OBJ or sample object OBJ_s does not exist. In this case, the size of each object region WA in the training image 564 may be the same as the size of the object OBJ or sample object OBJ_s captured in the training image 564. Furthermore, the shape of each object region WA in the training image 564 may be the same as the shape of the object OBJ or sample object OBJ_s captured in the training image 564.

[0140] The object region WA may include not only the region where the object OBJ or sample object OBJ_s exists, but also the region where the object OBJ or sample object OBJ_s does not exist. For example, the object region WA may include not only the region where the object OBJ or sample object OBJ_s exists, but also the region surrounding the region where the object OBJ or sample object OBJ_s exists. For example, the object region WA may include not only the region where the object OBJ or sample object OBJ_s exists, but also the region surrounding the region where the object OBJ or sample object OBJ_s exists. In this case, the size of each object region WA in the training image 564 may differ from the size of the object OBJ or sample object OBJ_s reflected in the training image 564. For example, the size of each object region WA in the training image 564 may be larger than the size of the object OBJ or sample object OBJ_s reflected in the training image 564. Furthermore, the shape of each object region WA in the training image 564 may differ from the shape of the object OBJ or sample object OBJ_s captured in the training image 564.

[0141] The object region WA may be a region defined based on the edges of an object OBJ or sample object OBJ_s reflected in the training image 564. As a first example, the object region WA may be a region that satisfies the condition that its outer edge coincides with the edges of an object OBJ or sample object OBJ_s reflected in the training image 564. In this case, the object region WA may be considered to include the region where the object OBJ or sample object OBJ_s exists, but does not include the region where the object OBJ or sample object OBJ_s does not exist. As a second example, the object region WA may be a region that satisfies the condition that its outer edge coincides with a virtual line obtained by expanding at least a portion of the edges of an object OBJ or sample object OBJ_s reflected in the training image 564 outwards. In this case, the object region WA may be considered to include the region where the object OBJ or sample object OBJ_s exists, as well as the region where the object OBJ or sample object OBJ_s does not exist.

[0142] Figure 8B shows object region WA#1 as an example of an object region WA that includes a region where an object OBJ or sample object OBJ_s exists, but does not include a region where an object OBJ or sample object OBJ_s does not exist, satisfying the condition that the outer edge of the object region WA coincides with the edge of an object OBJ or sample object OBJ_s captured in the training image 564. Furthermore, Figure 8B shows object region WA#2 as an example of an object region WA that includes a region where an object OBJ or sample object OBJ_s exists and a region where an object OBJ or sample object OBJ_s does not exist, satisfying the condition that the outer edge of the object region WA coincides with a virtual line obtained by extending at least a portion of the edge of an object OBJ or sample object OBJ_s outward. Furthermore, Figure 8B shows an example of an object region WA that includes both the region where object OBJ or sample object OBJ_s exists and the region where object OBJ or sample object OBJ_s does not exist. This is an object region WA#3 of a predetermined shape (for example, rectangular) that includes the region where object OBJ or sample object OBJ_s exists within its interior. Note that object region WA#3 may be considered equivalent to a so-called bounding box.

[0143] The correct label 563 may include any information that allows the object region WA to be identified within the training image 564 shown by the training image data 562, as information relating to the object region WA. The information that allows the object region WA to be identified may include any information that allows the position of the object region WA to be identified. The information that allows the object region WA to be identified may include any information that allows the shape of the object region WA to be identified. The information that allows the object region WA to be identified may include any information that allows the orientation of the object region WA to be identified.

[0144] An example of information capable of identifying at least one of the position, shape, and orientation of an object region WA is information indicating the coordinates of at least a portion of the object region WA. Information indicating the coordinates of at least a portion of the object region WA may include at least one of the following: information indicating the coordinates of at least a portion of the outer edge (i.e., edge) of the object region WA, information indicating the coordinates of the vertices of the object region WA, and information indicating the coordinates of a predetermined point inside the object region WA.

[0145] The arithmetic unit 51 may generate training data 561 containing the created training image data 562 and correct labels 563 by creating training image data 562 and correct labels 563. The arithmetic unit 51 may generate multiple training data 561 by repeating the process of generating training data 561. In other words, the arithmetic unit 51 may generate a training dataset 560 containing multiple training data 561 by repeating the process of generating training data 561.

[0146] The computing unit 51 may generate training data 561 by performing a simulation in which an object model OM is placed in a virtual simulation space SIM. The virtual simulation space SIM may be a simulation space that virtually represents a three-dimensional space.

[0147] The object model OM may be a model that represents an object OBJ (specifically, a model that represents at least a part of the object OBJ; the same applies hereinafter). In this case, the computing unit 51 may place the object model OM representing the object OBJ in the three-dimensional simulation space SIM. A model that represents the shape of the object OBJ may be used as the object model OM representing the object OBJ. In particular, a three-dimensional model that represents the three-dimensional shape of the object OBJ may be used as the object model OM representing the object OBJ. An example of a model that represents the shape of the object OBJ is a CAD (Computer-Aided Design) model of the object OBJ. An example of a model that represents the shape of the object OBJ is a model generated from a CAD model of the object OBJ (for example, a point cloud model). An example of a model that represents the shape of the object OBJ is a model of the object OBJ obtained by actually measuring the shape of the object OBJ using a three-dimensional shape measuring device such as a 3D scanner.

[0148] The object model OM may be a model that represents the sample object OBJ_s (specifically, a model that represents at least a part of the sample object OBJ_s; the same applies hereinafter). In this case, the computing unit 51 may place the object model OM representing the sample object OBJ_s in the three-dimensional simulation space SIM. A model that represents the shape of the sample object OBJ_s may be used as the object model OM representing the sample object OBJ_s. In particular, a three-dimensional model that represents the three-dimensional shape of the sample object OBJ_s may be used as the object model OM representing the sample object OBJ_s. An example of a model that represents the shape of the sample object OBJ_s is a CAD model of the sample object OBJ_s. An example of a model that represents the shape of the sample object OBJ_s is a model generated from the CAD model of the sample object OBJ_s (for example, a point cloud model). One example of a model representing the shape of the sample object OBJ_s is a model of the sample object OBJ_s obtained by actually measuring its shape using a three-dimensional shape measurement device such as a 3D scanner.

[0149] An example of a simulation space (SIM) is shown in Figure 9. As shown in Figure 9, the computing unit 51 places object models OM in the simulation space (SIM). For example, the computing unit 51 may place a single object model OM in the simulation space (SIM). Alternatively, for example, the computing unit 51 may place multiple object models OM in the simulation space (SIM).

[0150] When multiple object models OM are placed in the simulation space SIM, the computing unit 51 may arrange the multiple object models OM such that at least two of the multiple object models OM overlap at least partially. For example, the computing unit 51 may arrange the multiple object models OM haphazardly, haphazardly, or randomly. Alternatively, the computing unit 51 may arrange the multiple object models OM in a regular manner.

[0151] When the robot 1 performs processing on an object OBJ placed on the mounting device T, the computing device 51 may place a mounting device model TM representing the mounting device T in the simulation space SIM. In this case, the computing device 51 may place the mounting device model TM and the object model OM in the simulation space SIM such that the object model OM is placed on the mounting device model TM in the simulation space SIM.

[0152] As the mounting device model TM representing the mounting device T, a model showing the shape of the mounting device T may be used. In particular, as the mounting device model TM representing the mounting device T, a three-dimensional model showing the three-dimensional shape of the mounting device T may be used. An example of a model showing the shape of the mounting device T is a CAD model of the mounting device T. Another example of a model showing the shape of the mounting device T is a model of the mounting device T obtained by actually measuring the shape of the mounting device T using a three-dimensional shape measuring device such as a 3D scanner.

[0153] After the object model OM is placed in the simulation space SIM, the computing unit 51 may perform a simulation in which the simulation space SIM is imaged by a virtual imaging device 21 (hereinafter referred to as imaging device 21v). For example, the computing unit 51 may place a virtual imaging device 21v in the simulation space SIM. Subsequently, the computing unit 51 may perform a simulation in which the simulation space SIM is imaged using the imaging device 21v. In this case, the computing unit 51 may perform a simulation in which the position of the virtual imaging device 21v is moved and the simulation space SIM is imaged. As a result, the computing unit 51 may virtually generate image data that is virtually generated by imaged by the simulation space SIM using the imaging device 21v as training image data 562.

[0154] The computing device 51 may generate the correct labels 563 in parallel with or before / after the generation of the training image data 562. For example, because the computing device 51 itself places the object model OM in the simulation space SIM, the position and orientation of the object model OM in the simulation space SIM are known information to the computing device 51. Furthermore, because the computing device 51 itself performs a simulation to image the simulation space SIM using the imaging device 21v, the positional relationship and orientation relationship between the object model OM and the imaging device 21v in the simulation space SIM are known information to the computing device 51. Therefore, based on information regarding the position and orientation of the object model OM in the simulation space SIM, and information regarding the positional and orientation relationship between the object model OM and the imaging device 21v in the simulation space SIM, the computing device 51 can identify the object region WA in which the object model OM is captured (i.e., in which object OBJ or sample object OBJ_s is captured) within the training image 564 shown by the training image data 562 virtually generated by the imaging device 21v capturing the simulation space SIM. As a result, the computing device 51 can generate a correct label 563 indicating information about the object region WA in the training image 564. In other words, the computing device 51 can generate a correct label 563 indicating information about the identification result of the object region WA in the training image 564. As a result, the computing device 51 can generate training data 561 including the training image data 562 and the correct label 563.

[0155] However, the computing device 51 may identify the object region WA in which the object model OM is reflected in the simulation space SIM, based on instructions from the operator of the robot system SYS. In other words, the operator of the robot system SYS may manually change (in other words, set) the object region WA in which the object model OM is reflected in the simulation space SIM.

[0156] The computing device 51 may change the arrangement of the object model OM in the simulation space SIM each time it generates training data 561. Specifically, the computing device 51 may automatically change (in other words, set) the arrangement of the object model OM in the simulation space SIM. For example, the computing device 51 may automatically change the arrangement of the object model OM in the simulation space SIM according to predetermined rules. For example, the computing device 51 may randomly change the arrangement of the object model OM in the simulation space SIM. However, the computing device 51 may change the arrangement of the object model OM in the simulation space SIM based on instructions from the operator of the robot system SYS. In other words, the operator of the robot system SYS may manually change (in other words, set) the arrangement of the object model OM in the simulation space SIM.

[0157] The arrangement of object models OM may include the number of object models OM placed in the simulation space SIM. The arrangement of object models OM may include the position of at least one object model OM placed in the simulation space SIM. The arrangement of object models OM may include the orientation of at least one object model OM placed in the simulation space SIM. The arrangement of object models OM may include the positional relationship between at least one object model OM placed in the simulation space SIM and other objects (e.g., other object models OM, or other models such as mounting device models TM).

[0158] After the arrangement of the object model OM in the simulation space SIM is changed, the computing unit 51 may perform another simulation in which the simulation space SIM is imaged by the imaging device 21v. As a result, new training image data 562 is generated, which includes training image data 562 that is different from the training image data 562 that has already been generated, and a ground truth label 563 associated with the training image data 562. As an example, the computing unit 51 may arrange the object model OM in the simulation space SIM in the first arrangement. Subsequently, the computing unit 51 may perform a simulation in which the simulation space SIM, in which the object model OM is arranged in the first arrangement, is imaged by the imaging device 21v. As a result, the computing unit 51 can generate first training data 561 which includes first training image data 562 virtually generated by imaging a simulation space SIM in which the object model OM is arranged in a first arrangement using the imaging device 21v, and first correct label 563 indicating information about the object region WA within the training image 564 shown by the first training image data 562. Subsequently, the computing unit 51 may change the arrangement of the object model OM in the simulation space SIM from the first arrangement to a second arrangement different from the first arrangement. Subsequently, the computing unit 51 may perform a simulation in which the simulation space SIM in which the object model OM is arranged in the second arrangement is imaged using the imaging device 21v. As a result, the computing unit 51 can generate second training data 561 which includes a second training image data 562 virtually generated by imaging the simulation space SIM in which the object model OM is arranged in a second arrangement using the imaging device 21v, and a second correct label 563 indicating information about the object region WA within the training image 564 shown by the second training image data 562.

[0159] In this way, the computing device 51 can easily increase the number of generated training image data 562 (and furthermore, the number of training data 561 including the training image data 562) by alternately repeating the process of generating training data 561 and the process of changing the arrangement of the object model OM in the simulation space SIM.

[0160] The computing unit 51 may change the position of the imaging device 21v in the simulation space SIM each time it generates training data 561. In other words, the computing unit 51 may change the imaging position in which the imaging device 21v virtually images the simulation space SIM each time it generates training data 561. Specifically, the computing unit 51 may automatically change (in other words, set) the imaging position of the imaging device 21v. For example, the computing unit 51 may automatically change the imaging position of the imaging device 21v according to a predetermined rule. For example, the computing unit 51 may randomly change the imaging position of the imaging device 21v. However, the computing unit 51 may change the imaging position of the imaging device 21v based on instructions from the operator of the robot system SYS. In other words, the operator of the robot system SYS may manually change (in other words, set) the imaging position of the imaging device 21v.

[0161] After the imaging position of the imaging device 21v is changed, the computing device 51 may perform the simulation again to image the simulation space SIM with the imaging device 21v. As a result, new training data 561 is generated, which includes training image data 562 that is different from the training image data 562 that has already been generated, and a ground truth label 563 associated with the training image data 562. As an example, the computing device 51 may perform the simulation to image the simulation space SIM using the imaging device 21v located at a first imaging position in the simulation space. As a result, the computing device 51 can generate third training data 561 which includes a third training image data 562 that is virtually generated by image-taking the simulation space SIM using the imaging device 21v located at the first imaging position, and a third ground truth label 563 that indicates information about the object region WA in the training image 564 shown by the third training image data 562. Subsequently, the computing device 51 may change the imaging position of the imaging device 21v from the first imaging position to a second imaging position different from the first imaging position. Subsequently, the computing unit 51 may perform a simulation to image the simulation space SIM using the imaging device 21v located at a second imaging position within the simulation space SIM. As a result, the computing unit 51 can generate a fourth training data 561 which includes a fourth training image data 562 virtually generated by imaging the simulation space SIM using the imaging device 21v located at the second imaging position, and a fourth correct label 563 indicating information about the object region WA within the training image 564 shown by the fourth training image data 562.

[0162] In this way, the arithmetic unit 51 can easily increase the number of generated training image data 562 (and furthermore, the number of training data 561 that include the training image data 562) by alternately repeating the process of generating training data 561 and the process of changing the imaging position of the imaging device 21v.

[0163] The computing unit 51 may change the orientation of the imaging device 21v in the simulation space SIM each time it generates training data 561. In other words, the computing unit 51 may change the imaging orientation in which the imaging device 21v virtually images the simulation space SIM each time it generates training data 561. Specifically, the computing unit 51 may automatically change (in other words, set) the imaging orientation of the imaging device 21v. For example, the computing unit 51 may automatically change the imaging orientation of the imaging device 21v according to a predetermined rule. For example, the computing unit 51 may randomly change the imaging orientation of the imaging device 21v. However, the computing unit 51 may change the imaging orientation of the imaging device 21v based on instructions from the operator of the robot system SYS. In other words, the operator of the robot system SYS may manually change (in other words, set) the imaging orientation of the imaging device 21v.

[0164] After the imaging orientation of the imaging device 21v is changed, the computing device 51 may perform the simulation again to image the simulation space SIM with the imaging device 21v. As a result, new training data 561 is generated, which includes training image data 562 that is different from the training image data 562 that has already been generated, and a ground truth label 563 associated with the training image data 562. As an example, the computing device 51 may perform the simulation to image the simulation space SIM using the imaging device 21v positioned in the simulation space in a first imaging orientation. As a result, the computing device 51 can generate fifth training data 561 which includes a fifth training image data 562 that is virtually generated by imaging the simulation space SIM using the imaging device 21v positioned in the first imaging orientation, and a fifth ground truth label 563 that indicates information about the object region WA in the training image 564 shown by the fifth training image data 562. Subsequently, the computing unit 51 may change the imaging orientation of the imaging device 21v from the first imaging orientation to a second imaging orientation different from the first imaging orientation. After that, the computing unit 51 may perform a simulation in which it images the simulation space SIM using the imaging device 21v positioned in the second imaging orientation within the simulation space SIM. As a result, the computing unit 51 can generate sixth training data 561 which includes sixth training image data 562 virtually generated by imaging the simulation space SIM using the imaging device 21v positioned in the second imaging orientation, and sixth correct label 563 indicating information about the object region WA within the training image 564 shown by the sixth training image data 562.

[0165] In this way, the arithmetic unit 51 can easily increase the number of training image data 562 generated (and furthermore, the number of training data 561 including the training image data 562) by alternately repeating the process of generating training data 561 and the process of changing the imaging orientation of the imaging device 21v.

[0166] Furthermore, the computing device 51 may acquire image data IMD_2D, generated by the imaging device 21 capturing the real space in which at least one of the object OBJ and sample object OBJ_s is actually located, as training image data 562. In this case, the computing device 51 may generate the correct label 563 by analyzing the image shown by the image data IMD_2D acquired as training image data 562. For example, the computing device 51 may generate the correct label 563 indicating information about the object region WA in which the detected object OBJ or sample object OBJ_s exists by performing an object detection process to detect the object OBJ or sample object OBJ_s within the 2D image IMG_2D shown by the image data IMD_2D acquired as training image data 562. For example, the computing device 51 may generate the correct label 563 indicating information about the object region WA in which the detected object OBJ or sample object OBJ_s exists by performing the 2D matching process described later using the image data IMD_2D acquired as training image data 562.

[0167] Alternatively, a user such as an operator of the robot system SYS may manually identify the object region WA in which object OBJ or sample object OBJ_s is captured within the 2D image IMG_2D shown by the image data IMD_2D acquired as training image data 562. Furthermore, a user such as an operator of the robot system SYS may manually identify the object region WA in which object OBJ or sample object OBJ_s is captured within the training image 564 shown by the training image data 562 virtually generated by the imaging device 21v capturing the simulation space SIM. In other words, the annotation that generates the correct label 563 to the training image data 562 may be performed manually by the user.

[0168] Even when the image data IMD_2D is acquired as the training image data 562, the acquisition of the training image data 562 (and further, the assignment of the correct label 563) and the modification of at least one arrangement of object OBJ and sample object OBJ_s in real space, the imaging position of the imaging device 21, and the imaging orientation of the imaging device 21 may be repeated alternately. As a result, the number of training data 561 can be increased.

[0169] Furthermore, if the training dataset 560 has already been generated, the arithmetic unit 51 does not need to generate the training dataset 560 again. In other words, the arithmetic unit 51 does not need to perform the process in step S11 of Figure 7.

[0170] Again in Figure 7, after the training dataset 560 is generated in step S11, the computing unit 51 uses the training dataset 560 generated in step S11 to perform machine learning to generate the object recognition model 50 (steps S12 to S13).

[0171] Specifically, first, as shown in Figure 10, the computing device 51 inputs the training image data 562 included in the training dataset 560 into the object recognition model 50 (step S12 in Figure 7). For example, the computing device 51 may input the training image data 562 included in the training dataset 560 into the object recognition model 50 in its initial state. Alternatively, for example, the computing device 51 may input the training image data 562 included in the training dataset 560 into the object recognition model 50 that has already been generated.

[0172] As described above, the object recognition model 50 is a learning model that, when an image is input, can output object region information OAI regarding the object region WA in which an object OBJ exists within the input image. Therefore, when training image data 562 is input to the object recognition model 50, as shown in Figure 10, the object recognition model 50 outputs object region information OAI regarding the object region WA in which an object OBJ or sample object OBJ_s exists within the training image 564 indicated by the input training image data 562.

[0173] The object region information OAI output by the object recognition model 50 may include, like the correct label 563, any information that allows for the identification of the object region WA within the training image 564 shown by the training image data 562, as information regarding the object region WA. In this case, the explanation of the information regarding the object region WA shown by the correct label 563 is omitted, and the explanation of the information regarding the object region WA shown by the object region information OAI output by the object recognition model 50 is omitted, and therefore, in order to avoid redundant explanations, a detailed explanation of the object region information OAI is omitted.

[0174] If the training dataset 560 contains multiple training image data 562, the computing unit 51 may sequentially input the multiple training image data 562 contained in the training dataset 560 to the object recognition model 50. As a result, the object recognition model 50 sequentially outputs multiple object region information OAIs corresponding to each of the multiple training image data 562. The computing unit 51 may input all of the multiple training image data 562 contained in the training dataset 560 to the object recognition model 50, or it may input only some of the multiple training image data 562 contained in the training dataset 560 to the object recognition model 50.

[0175] In Figure 7 again, the computing device 51 then updates the object recognition model 50 by performing machine learning based on the object region information OAI output by the object recognition model 50 (i.e., the output of the object recognition model 50) and the correct labels 563 included in the training dataset 560 (step S13). For example, the computing device 51 may update the object recognition model 50 by performing machine learning based on information regarding the error between the object region information OAI and the correct labels 563, as shown in Figure 10. For example, the computing device 51 may update the object recognition model 50 by performing machine learning based on a loss function relating to the error between the object region information OAI and the correct labels 563, as shown in Figure 10. Examples of loss functions include at least one of the mean squared error, mean absolute error, and cross-entropy error. In this case, the computing device 51 may update the object recognition model 50 so that the loss calculated using the loss function is minimized.

[0176] Updating the object recognition model 50 may include updating the parameters of the object recognition model 50. For example, if the object recognition model includes a neural network, updating the object recognition model 50 may include updating at least one of the parameters of the object recognition model 50, namely the weights and biases.

[0177] Thereafter, the computing unit 51 may repeat machine learning to generate the object recognition model 50 until the learning termination condition is met (step S14). An example of a learning termination condition is that the accuracy of the object recognition model calculated by cross-validation is equal to or greater than a predetermined accuracy threshold.

[0178] (2-2) Robot Control Processing Next, the robot control processing performed by the control device 3 will be explained. Note that the robot control processing is a process that uses the object recognition model 50 generated by the model generation processing, which is a process equivalent to the learning phase in machine learning, and therefore may be considered to be a process equivalent to the inference phase in machine learning.

[0179] (2-2-1) Overall flow of robot control processing First, the overall flow of robot control processing will be explained with reference to Figure 11. Figure 11 is a flowchart showing the overall flow of robot control processing.

[0180] As shown in Figure 11, the position and attitude calculation unit 311 of the control device 3 acquires image data IMD from the imaging system 2 using the communication device 33 (step S21).

[0181] For example, the position and orientation calculation unit 311 may acquire image data IMD_2D from the imaging device 21 using the communication device 33 (step S21). Specifically, the imaging device 21 may image the object OBJ placed on the mounting device T at a predetermined imaging rate (specifically, at least a part of the workpiece W is imaged; the same applies hereinafter). For example, the imaging device 21 may image the object OBJ at an imaging rate of several tens to several hundreds of times per second (for example, 500 times). As a result, the imaging device 21 generates image data IMD_2D at a period corresponding to the predetermined imaging rate. The position and orientation calculation unit 311 acquires the image data IMD_2D each time the imaging device 21 generates the image data IMD_2D.

[0182] For example, the position and orientation calculation unit 311 may acquire image data IMD_3D from the imaging device 22 using the communication device 33 (step S21). Specifically, the imaging device 22 may image the object OBJ placed on the mounting device T at a predetermined imaging rate (specifically, at least a part of the workpiece W is imaged; the same applies hereinafter). For example, the imaging device 22 may image the object OBJ at an imaging rate of several tens to several hundreds of times per second (for example, 500 times). As a result, the imaging device 22 generates image data IMD_3D at a period corresponding to the predetermined imaging rate. The position and orientation calculation unit 311 acquires the image data IMD_3D each time the imaging device 22 generates the image data IMD_3D.

[0183] However, at least one of the imaging devices 21 and 22 does not have to periodically image the object OBJ at a desired imaging rate. For example, at least one of the imaging devices 21 and 22 may image the object OBJ when it receives a control signal from the control device 3 (or another control device different from the control device 3) to control at least one of the imaging devices 21 and 22 to image the object OBJ.

[0184] Each of the imaging devices 21 and 22 may image objects other than the object OBJ that the end effector 4 performs a predetermined process on (in this case, holds) (specifically, at least a part of the other object may be imaged; the same applies hereinafter). Since the other object is not the object that the robot 1 performs a predetermined process on, in the following description, objects other than the object OBJ will be referred to as non-target objects. An example of a non-target object is at least a part of the end effector 4. An example of a non-target object is at least a part of the robot arm 12. An example of a non-target object is at least a part of the peripheral objects that are located around the robot 1. An example of a peripheral object is the mounting device T. For example, if both the object OBJ and the non-target object are included in the imaging range (field of view) of the imaging device 21, each of the imaging devices 21 and 22 may image both the object OBJ and the non-target object. As a result, the imaging devices 21 and 22 may generate image data IMD_2D and IMD_3D, respectively, showing images in which both the object OBJ and the non-target object are captured. However, each of the imaging devices 21 and 22 may capture the object OBJ but not the non-target object. In other words, the imaging devices 21 and 22 may generate image data IMD_2D and IMD_3D, respectively, showing images in which the object OBJ is captured but the non-target object is not.

[0185] If multiple objects OBJ are placed on the mounting device T, each of the imaging devices 21 and 22 may image at least a portion of the multiple objects OBJ placed on the mounting device T. In other words, each of the imaging devices 21 and 22 may image a group of target objects that includes at least a portion of the multiple objects OBJ placed on the mounting device T. The group of target objects imaged by imaging device 21 may include multiple objects OBJ (or, in some cases, one object OBJ included in the imaging field of imaging device 21) from among the multiple objects OBJ placed on the mounting device T. The group of target objects imaged by imaging device 22 may include multiple objects OBJ (or, in some cases, one object OBJ included in the imaging field of imaging device 22) from among the multiple objects OBJ placed on the mounting device T.

[0186] In step S21, each time the position and orientation calculation unit 311 acquires image data IMD, the position and orientation calculation unit 311 calculates at least one of the position and orientation of the object OBJ based on the image data IMD acquired in step S21 and the object recognition model 50 generated by the model generation process described above (step S22). In particular, the position and orientation calculation unit 311 calculates at least one of the position and orientation of the target object OBJ_tgt, which is selected as the object OBJ that the end effector 4 actually performs the predetermined processing on. As a result, the position and orientation calculation unit 311 generates position and orientation data (in other words, position and orientation information) POI that indicates at least one of the position and orientation of the object OBJ (especially the target object OBJ_tgt). The process of calculating at least one of the position and orientation of the object OBJ in step S22 will be described in detail later with reference to Figure 12, etc.

[0187] Subsequently, the signal generation unit 312 of the control device 3 generates a robot control signal to control the robot 1, the end effector 4, and at least one of the robot movable devices to perform processing on the target object OBJ_tgt using the end effector 4, based on the position and orientation data POI generated in step S22 (step S23). Alternatively, the signal generation unit 312 may generate a control signal to control the mounting device T in addition to, or instead of, generating a robot control signal to control the robot 1, the end effector 4, and at least one of the robot movable devices (step S23). For example, if the mounting device T is movable relative to the support surface S as described above, the signal generation unit 312 may generate a control signal to control the movement of the mounting device T relative to the support surface S. As described above, in this embodiment, an example is described in which the robot 1 performs the holding process shown in Figures 6A to 6C using the end effector 4. Therefore, the signal generation unit 312 may generate a robot control signal to control the robot 1, the end effector 4, and at least one of the robot movable devices so as to hold the target object OBJ_tgt using the end effector 4. Subsequently, the signal generation unit 312 outputs the robot control signal generated in step S23 to the robot control device 13 using the communication device 33. As a result, the robot control device 13 controls the robot 1, the end effector 4, and at least one of the robot movable devices based on the robot control signal.

[0188] For example, the signal generation unit 312 may generate a robot control signal to control at least one of the robot 1 and the robot movable device so that the end effector 4 moves toward a position where the end effector 4 can process the target object OBJ_tgt which is located at the position calculated in step S22 and / or which is in the posture calculated in step S22. For example, the signal generation unit 312 may generate a robot control signal to control at least one of the robot 1 and the robot movable device so that the end effector 4 takes a posture where the end effector 4 can process the target object OBJ_tgt which is located at the position calculated in step S22 and / or which is in the posture calculated in step S22.

[0189] During the period when robot 1, end effector 4, and at least one of the robot's movable devices are controlled based on the robot control signal (for example, during the period when end effector 4 is approaching the target object OBJ_tgt), the imaging system 2 may image the object OBJ (in particular, the target object OBJ_tgt) again, and the position and orientation calculation unit 311 may again acquire the image data IMD generated again by the imaging system 2 image the object OBJ (in particular, the target object OBJ_tgt) again (step S21 in Figure 11). The position and orientation calculation unit 311 may then generate position and orientation data POI again based on the newly acquired image data IMD and the object recognition model 50 (step S22 in Figure 11). In other words, the position and orientation calculation unit 311 may update at least one of the position and orientation of the object OBJ (in particular, the target object OBJ_tgt). Subsequently, the signal generation unit 312 may generate the robot control signal again based on the updated position and orientation data POI (step S23 in Figure 11). In other words, during the period when the robot 1, the end effector 4, and at least one of the robot movable devices are controlled based on the robot control signal (for example, during the period when the end effector 4 is approaching the target object OBJ_tgt), the signal generation unit 312 may repeatedly generate the robot control signal (i.e., repeatedly generate the movement path of the end effector 4). To put it another way, during the period when the robot 1, the end effector 4, and at least one of the robot movable devices are controlled based on the robot control signal (for example, during the period when the end effector 4 is approaching the target object OBJ_tgt), the signal generation unit 312 may update the robot control signal (i.e., update the movement path of the end effector 4). In other words, during the period when robot 1, end effector 4, and at least one of the robot's movable devices are controlled based on the robot control signal (for example, the period when end effector 4 is approaching the target object OBJ_tgt), the control device 3 may repeat the process from step S21 to step S23 in Figure 11.

[0190] In particular, if the object OBJ (especially the target object OBJ_tgt) moves during a period when at least one of the robot 1, the end effector 4, and the robot movable device is controlled based on the robot control signal (for example, during a period when the end effector 4 is approaching the target object OBJ_tgt), the control device 3 may repeat the process from steps S21 to S23 in Figure 11. In this case, the imaging system 2 may repeatedly image the moving object OBJ (especially the target object OBJ_tgt). As a result, if the object OBJ (especially the target object OBJ_tgt) moves during a period when the end effector 4 is moving, the position and orientation calculation unit 311 can appropriately update the position and orientation data POI so that the movement of the object OBJ (especially the target object OBJ_tgt) is reflected in the position and orientation data POI. As a result, the signal generation unit 312 can generate a robot control signal based on the updated position and orientation data POI. Therefore, the signal generation unit 312 can control the robot 1, the end effector 4, and at least one of the robot's movable devices so that the end effector 4 performs a predetermined process on the moving object OBJ (in particular, the target object OBJ_tgt).

[0191] (2-2-2) Process for generating position and orientation data POI Next, with reference to Figure 12, the process of calculating at least one of the position and orientation of the object OBJ based on the image data IMD and the object recognition model 50 in step S22 of Figure 11 will be described. Figure 12 is a flowchart showing the flow of the process of calculating at least one of the position and orientation of the object OBJ based on the image data IMD and the object recognition model 50 in step S22 of Figure 11.

[0192] (2-2-2-1) Identification of object region WA As shown in Figure 12, the position and orientation calculation unit 311 inputs the image data IMD_2D acquired in step S21 of Figure 11 to the object recognition model 50 (step S221). As a result, as shown in Figure 13, the object recognition model 50 outputs object region information OAI relating to the object region WA in which the object OBJ exists within the 2D image IMG_2D indicated by the input image data IMD_.

[0193] The position and orientation calculation unit 311 acquires the object region information OAI output by the object recognition model 50. In other words, the position and orientation calculation unit 311 acquires the object region information OAI which indicates the result of the object region WA identification by the object recognition model 50. Thus, in this embodiment, the position and orientation calculation unit 311 uses the object recognition model 50 to identify the object region WA in which the object OBJ exists within the 2D image IMG_2D shown by the image data IMD_.

[0194] Furthermore, the characteristics of the object region WA in which object OBJ is captured within the 2D image IMG_2D may be the same as the characteristics of the object region WA in which object OBJ or sample object OBJ_s is captured within the training image 564 described above. In this case, the description of the object region WA in which object OBJ or sample object OBJ_s is captured within the training image 564 described above can be reused as the description of the object region WA in which object OBJ is captured within the 2D image IMG_2D. Specifically, the explanation of the object region WA in which object OBJ is captured within the training image 564 described above can be reused as an explanation of the object region WA in which object OBJ is captured within the 2D image IMG_2D shown by the image data IMD_2D, by replacing the phrase "training image data 562" with the phrase "image data IMD_2D", replacing the phrase "training image 564" with the phrase "2D image IMG_2D", and replacing the phrase "object OBJ or sample object OBJ_s" with the phrase "object OBJ". For this reason, in order to avoid redundant explanations, a detailed explanation of the object region WA in which object OBJ is captured within the 2D image IMG_2D is omitted.

[0195] Furthermore, the explanation of the object region information OAI regarding the object region WA in which an object OBJ or sample object OBJ_s exists within the training image 564 indicated by the input training image data 562, which is output by the object recognition model 50 that receives the aforementioned training image data 562, can be reused as an explanation of the object region information OAI regarding the object region WA in which an object OBJ exists within the 2D image IMG_2D indicated by the input image data IMD_2D, which is output by the object recognition model 50 that receives the image data IMD_2D. Specifically, the explanation of the object region information OAI output by the object recognition model 50 that receives the aforementioned training image data 562 can be reused as an explanation of the object region information OAI output by the object recognition model 50 that receives the image data IMD_2D, by replacing the phrase "training image data 562" with "image data IMD_2D", the phrase "training image 564" with "2D image IMG_2D", and the phrase "object OBJ or sample object OBJ_s" with "object OBJ". For this reason, in order to avoid redundant explanations, a detailed explanation of the object region information OAI output by the object recognition model 50 that receives the image data IMD_2D will be omitted.

[0196] (2-2-2-2) Generation of three-dimensional position data WSD Again in Figure 12, in parallel with or before / after the processing of step S221, the position and orientation calculation unit 311 generates three-dimensional position data WSD based on the image data IMD_3D (i.e., the two images shown by the image data IMD_3D) each time the position and orientation calculation unit 311 acquires image data IMD_3D from the imaging device 22 (step S222).

[0197] The three-dimensional position data WSD is data indicating the three-dimensional position of at least a portion of the object OBJ captured in the two images shown by the image data IMD_3D. For example, the three-dimensional position data WSD may be data indicating the three-dimensional position of at least a portion of the surface of the object OBJ. In particular, the three-dimensional position data WSD is data indicating the three-dimensional position of multiple points on the object OBJ. That is, the three-dimensional position data WSD is data indicating the three-dimensional position of multiple points on the object OBJ captured by the imaging device 22. For example, the three-dimensional position data WSD may be data indicating the three-dimensional position of multiple points on the surface of the object OBJ. For example, the three-dimensional position data WSD may be data indicating the three-dimensional position of multiple points corresponding to multiple parts of the surface of the object OBJ. In the following explanation, unless otherwise specified, the three-dimensional position of object OBJ may mean at least one of the following: the three-dimensional position of at least a part of the surface of object OBJ, the three-dimensional position of each of multiple points on object OBJ, the three-dimensional position of each of multiple points on the surface of object OBJ, and the three-dimensional position of each of multiple points corresponding to each of multiple parts on the surface of object OBJ.

[0198] To generate three-dimensional position data WSD, the position and orientation calculation unit 311 may calculate parallax by associating each part (e.g., each pixel) between two images represented by the image data IMD_3D. For example, the position and orientation calculation unit 311 may calculate parallax by associating each part of the projection pattern captured in the two images (i.e., each part of the projection pattern captured in each image). Subsequently, the position and orientation calculation unit 311 may calculate the three-dimensional positions of multiple points of the object OBJ using a well-known method based on the principle of triangulation with the calculated parallax. As a result, three-dimensional position data WSD indicating the three-dimensional positions of multiple points of the object OBJ is generated. In this case, associating each part of images with the projection pattern (i.e., each part of the captured projection pattern) results in higher accuracy in calculating parallax than associating each part of images without the projection pattern. Therefore, the accuracy of the generated three-dimensional position data WSD (i.e., the accuracy of calculating the three-dimensional position of each of the multiple points on the object OBJ) becomes higher.

[0199] Depth image data is an example of three-dimensional position data (WSD). Depth image data is image data in which depth information is associated with each pixel of the depth image shown by the depth image data, either in addition to or instead of brightness information. Depth information is information that indicates the distance (i.e., depth) between each part of the object OBJ captured in each pixel and the imaging device 22. The distance (i.e., depth) between each part of the object OBJ captured in each pixel and the imaging device 22 can be calculated from the parallax described above. Point cloud data is another example of three-dimensional position data (WSD). Point cloud data is data that shows the set of points in three-dimensional space corresponding to each part of the object OBJ captured in the image shown by the image data IMD_3D. For the sake of explanation, in the following explanation, an example in which point cloud data is used as three-dimensional position data (WSD) will be described.

[0200] The position and orientation calculation unit 311 may perform a resizing process to reduce the data size of the generated three-dimensional position data WSD after it has been generated. For example, if the three-dimensional position data WSD is point cloud data, the position and orientation calculation unit 311 may perform a resizing process that involves decimating (i.e., deleting or reducing) some of the points included in the point cloud indicated by the three-dimensional position data WSD. For example, if the three-dimensional position data WSD is depth image data, the position and orientation calculation unit 311 may perform a resizing process that involves decimating (i.e., deleting or reducing) some of the pixels of the depth image data indicated by the depth image data. After that, the position and orientation calculation unit 311 may use the three-dimensional position data WSD, whose data size has been reduced by the resizing process, to perform subsequent processing.

[0201] Alternatively, the position and orientation calculation unit 311 may perform a resizing process to reduce the data size of the image data IMD_3D before generating the three-dimensional position data WSD. For example, as a resizing process, the position and orientation calculation unit 311 may perform a process to decimate (i.e., delete or reduce) some of the pixels of the image shown in the image data IMD_3D. After that, the position and orientation calculation unit 311 may generate the three-dimensional position data WSD using the image data IMD_3D whose data size has been reduced by the resizing process. As a result, the position and orientation calculation unit 311 can generate three-dimensional position data WSD whose data size has been reduced by the resizing process. After that, the position and orientation calculation unit 311 may perform subsequent processing using the three-dimensional position data WSD generated based on the image data IMD_3D whose data size has been reduced by the resizing process.

[0202] However, the position and orientation calculation unit 311 does not have to perform resizing. In this case, the position and orientation calculation unit 311 may perform subsequent processing using three-dimensional position data WSD that has not been resized or three-dimensional position data WSD generated based on image data IMD_3D that has not been resized.

[0203] (2-2-2-3) Generation of position and orientation data POI1 (rough search) Subsequently, the position and orientation calculation unit 311 calculates at least one of the position and orientation of object OBJ based on the three-dimensional position data WSD generated in step S222 and the object region information OAI acquired in step S221 (i.e., the result of identifying the object region WA) (step S223). Specifically, the position and orientation calculation unit 311 calculates at least one of the position and orientation of an object OBJ located in an object region WA indicated by the object region information OAI, based on the three-dimensional position data WSD and the object region information OAI, as at least one of the position and orientation of the target object OBJ_tgt (step S223). As a result, the position and orientation calculation unit 311 generates position and orientation data (in other words, position and orientation information) POI1 that indicates at least one of the position and orientation of the object OBJ as position and orientation data POI1 that indicates at least one of the position and orientation of the target object OBJ_tgt (step S223). The process of generating the position and orientation data POI1 in step S223 may be called a rough search (rough search process).

[0204] In order to generate position and orientation data POI1, the position and orientation calculation unit 311 may first identify (in other words, select) a part of the three-dimensional position data WSD as the target data portion WSD_ROI based on the object region information OAI. For example, if the three-dimensional position data WSD is point cloud data as described above, the position and orientation calculation unit 311 may identify a part of the point cloud data (i.e., a part of the point cloud indicated by the three-dimensional position data) as the target data portion WSD_ROI based on the object region information OAI.

[0205] The position and orientation calculation unit 311 may identify the data portion corresponding to a single object region WA indicated by the object region information OAI in the three-dimensional position data WSD as the target data portion WSD_ROI. For example, as shown in Figure 14, the object region information OAI indicates a single object region WA in which an object OBJ exists within the 2D image IMG_2D. In this case, as shown in Figure 14, the position and orientation calculation unit 311 may identify a single object region (in other words, an object space) WA_3D in which an object OBJ exists within a single object region WA, based on the object region information OAI. For example, the position and orientation calculation unit 311 may identify a single object region (in other words, an object space) WA_2D in the three-dimensional space within the 2D imaging coordinate system where an object OBJ located in a single object region WA is located, based on the object region information OAI (i.e., the object region WA in which an object OBJ exists within the 2D image IMG_2D). Subsequently, as shown in Figure 14, the position and orientation calculation unit 311 may identify a single object region WA_3D in the three-dimensional space within the 3D imaging coordinate system corresponding to the three-dimensional position data WSD where an object OBJ located in a single object region WA_2D is located, based on the single object region WA_2D. For example, the position and orientation calculation unit 311 may generate an object region WA_3D by converting an object region WA_2D to an object region WA_3D using a transformation matrix (typically a rigid body transformation matrix) for converting coordinates in either the 2D or 3D imaging coordinate system to coordinates in the other of the 2D or 3D imaging coordinate system. The transformation matrix can be calculated by mathematical methods from external parameters indicating the positional relationship between the imaging device 21 and the imaging device 22. As an example, the transformation matrix can be calculated by mathematical methods that solve the PnP (Perspective n Point) problem based on external parameters indicating the positional relationship between the imaging device 21 and the imaging device 22. Subsequently, the position and orientation calculation unit 311 may identify the data portion of the three-dimensional position data WSD corresponding to an object region WA_3D as the data portion corresponding to an object region WA indicated by the object region information OAI (i.e., the target data portion WSD_ROI).However, the method for identifying the target data portion WSD_ROI corresponding to a single object region WA is not limited to the method described herein.

[0206] Subsequently, the position and orientation calculation unit 311 generates position and orientation data POI1 using the target data portion WSD_ROI. Here, the process of generating position and orientation data POI1 using the target data portion WSD_ROI may include data processing to extract (in other words, cut out) the target data portion WSD_ROI from the three-dimensional position data WSD, and then generating position and orientation data POI1 using the extracted target data portion WSD_ROI. The process of generating position and orientation data POI1 using the target data portion WSD_ROI may also include data processing to delete (in other words, remove) other data portions from the three-dimensional position data WSD other than the target data portion WSD_ROI, and then generating position and orientation data POI1 using the three-dimensional position data WSD from which the other data portions have been removed (i.e., the remaining target data portion WSD_ROI). The process of generating position and orientation data POI1 using the target data portion WSD_ROI may include a process of generating position and orientation data POI1 by selectively using the target data portion WSD_ROI, which is a part of the three-dimensional position data WSD, without performing data processing on the three-dimensional position data WSD to extract the target data portion WSD_ROI or to delete other data portions other than the target data portion WSD_ROI.

[0207] In order to generate position and orientation data POI1 using the target data portion WSD_ROI, the position and orientation calculation unit 311 may perform a matching process using the target data portion WSD_ROI. Specifically, the position and orientation calculation unit 311 may perform a matching process using the target data portion WSD_ROI and a template model which is a model representing the object OBJ (specifically, representing at least a part of the object OBJ; the same applies hereinafter). In the following description, the matching process using at least a part of the three-dimensional position data WSD (for example, the target data portion WSD_ROI) and a template model will be referred to as the 3D matching process.

[0208] The template model used for 3D matching processing may be a three-dimensional model (three-dimensional template model TMP_3D) that shows the three-dimensional shape of at least a part of the object OBJ. The three-dimensional template model TMP_3D may show the three-dimensional shape of at least a part of the object OBJ using both at least a part of the contour of the object OBJ and at least a part of the part of the object OBJ enclosed by the contour. An example of a three-dimensional model showing the three-dimensional shape of an object OBJ is a CAD model of the object OBJ. An example of a three-dimensional model showing the three-dimensional shape of an object OBJ is a CAD model of a sample object OBJ_s that simulates the object OBJ. An example of a three-dimensional model showing the three-dimensional shape of an object OBJ is a model (e.g., a point cloud model) generated from a CAD model of the object OBJ or a sample object OBJ_s. An example of a model showing the shape of an object OBJ is a model of the object OBJ obtained by actually measuring the shape of the object OBJ using a three-dimensional shape measuring device such as a 3D scanner. One example of a model representing the shape of an object's oblique joint (OBJ) is a model of an object's OBJ obtained by actually measuring the shape of a sample object OBJ_s, which simulates an object's OBJ, using a three-dimensional shape measurement device such as a 3D scanner.

[0209] The position and orientation calculation unit 311 may perform an object detection process as a 3D matching process to detect an object OBJ (in other words, a set of points corresponding to the three-dimensional template model TMP_3D) indicated by the three-dimensional template model TMP_3D within the point cloud indicated by the target data portion WSD_ROI. The 3D matching process (in this case, the object detection process) itself may be the same as an existing 3D matching process. For example, the position and orientation calculation unit 311 may perform the 3D matching process using a well-known method that includes at least one of RANSAC (Random Sample Consensus), SIFT (Scale-Invariant Feature Transform), ICP (Iterative Closest Point), and DSO (Direct Spare Odometry).

[0210] When 3D matching processing is performed, the position and orientation calculation unit 311 may translate, enlarge, reduce, and / or rotate the three-dimensional template model TMP_3D within the 3D imaging coordinate system corresponding to the target data portion WSD_ROI so that the feature locations of the three-dimensional template model TMP_3D (e.g., at least one of feature points and edges) approach (e.g., coincide with) the feature locations of the object OBJ (e.g., the point cloud corresponding to the object OBJ indicated by the target data portion WSD_ROI) that the target data portion WSD_ROI indicates in three dimensions. As a result, the position and orientation calculation unit 311 can determine the positional relationship between the coordinate system of the three-dimensional template model TMP_3D and the 3D imaging coordinate system. Subsequently, the position and orientation calculation unit 311 may calculate at least one of the position and orientation of the object OBJ in the 3D imaging coordinate system from at least one of the position and orientation of the object OBJ in the coordinate system of the three-dimensional template model TMP_3D (i.e., at least one of the position and orientation of the three-dimensional template model TMP_3D) based on the positional relationship between the coordinate system of the three-dimensional template model TMP_3D and the 3D imaging coordinate system. Subsequently, the position and orientation calculation unit 311 may convert at least one of the position and orientation of the object OBJ in the 3D imaging coordinate system to at least one of the position and orientation of the object OBJ in the reference coordinate system based on a transformation matrix for converting the three-dimensional coordinates in the 3D imaging coordinate system to three-dimensional coordinates in the reference coordinate system. However, if the 3D imaging coordinate system is used as the reference coordinate system, the position and orientation calculation unit 311 does not need to convert at least one of the position and orientation of the object OBJ in the 3D imaging coordinate system to at least one of the position and orientation of the object OBJ in the reference coordinate system. As a result, the position and orientation calculation unit 311 generates position and orientation data POI1 that indicates at least one of the position and orientation of the object OBJ in the reference coordinate system.

[0211] Furthermore, the position and orientation calculation unit 311 may calculate at least one of the position and orientation of the object OBJ using an existing method other than matching processing (3D matching processing) with respect to at least one of the target data portion WSD_ROI, object region WA, object region WA_2D, and object region WA_3D. For example, the position and orientation calculation unit 311 may calculate at least one of the position and orientation of the object OBJ based on at least one of the target data portion WSD_ROI, object region WA, object region WA_2D, and object region WA_3D using an inference model generated by machine learning. The inference model may be generated by machine learning so as to output at least one of the position and orientation of the object OBJ that indicates the three-dimensional position of the target data portion WSD_ROI (or at least a part of the three-dimensional position data WSD) when the target data portion WSD_ROI (or at least a part of the three-dimensional position data WSD) is input. The inference model may be generated by machine learning so that, when an object region WA is input, it outputs at least one of the position and orientation of the object OBJ contained in the object region WA. The inference model may be generated by machine learning so that, when an object region WA_2D is input, it outputs at least one of the position and orientation of the object OBJ contained in the object region WA_2D. The inference model may be generated by machine learning so that, when an object region WA_3D is input, it outputs at least one of the position and orientation of the object OBJ contained in the object region WA_3D.

[0212] The position and orientation calculation unit 311 may calculate, by performing a matching process, at least one of the following as the position of the object OBJ in the reference coordinate system: the position Tx of the object OBJ in the X-axis direction parallel to the X-axis of the reference coordinate system, the position Ty of the object OBJ in the Y-axis direction parallel to the Y-axis of the reference coordinate system, and the position Tz of the object OBJ in the Z-axis direction parallel to the Z-axis of the reference coordinate system. The position and orientation calculation unit 311 may also calculate, by performing a matching process, at least one of the following as the orientation of the object OBJ in the reference coordinate system: the amount of rotation Rx of the object OBJ around the X-axis of the reference coordinate system, the amount of rotation Ry of the object OBJ around the Y-axis of the reference coordinate system, and the amount of rotation Rz of the object OBJ around the Z-axis of the reference coordinate system. This is because the rotation amount Rx of the object OBJ around the X axis, the rotation amount Ry of the object OBJ around the Y axis, and the rotation amount Rz of the object OBJ around the Z axis are equivalent to the parameters representing the orientation of the object OBJ around the X axis, the parameters representing the orientation of the object OBJ around the Y axis, and the parameters representing the orientation of the object OBJ around the Z axis, respectively. For this reason, in the following explanation, the rotation amount Rx of the object OBJ around the X axis, the rotation amount Ry of the object OBJ around the Y axis, and the rotation amount Rz of the object OBJ around the Z axis will be referred to as the orientation Rx of the object OBJ around the X axis, the orientation Ry of the object OBJ around the Y axis, and the orientation Rz of the object OBJ around the Z axis, respectively.

[0213] Furthermore, the orientation Rx of object OBJ around the X-axis, the orientation Ry of object OBJ around the Y-axis, and the orientation Rz of object OBJ around the Z-axis may be considered to represent the position of object OBJ in the rotational direction around the X-axis, the position of object OBJ in the rotational direction around the Y-axis, and the position of object OBJ in the rotational direction around the Z-axis, respectively. In other words, the orientation Rx of object OBJ around the X-axis, the orientation Ry of object OBJ around the Y-axis, and the orientation Rz of object OBJ around the Z-axis may all be considered parameters that represent the position of object OBJ.

[0214] The position and orientation calculation unit 311 may calculate a matching similarity, which is the similarity between the template model and the detected object OBJ, when performing the matching process. In this case, the position and orientation calculation unit 311 may select (in other words, determine) an object OBJ as the target object OBJ_tgt for which the end effector 4 should actually perform the predetermined processing, if the object OBJ detected by the matching process has a calculated matching similarity that is equal to or greater than a predetermined matching judgment threshold used in the matching process. On the other hand, the position and orientation calculation unit 311 does not have to select an object OBJ as the target object OBJ_tgt for which the end effector 4 should actually perform the predetermined processing, if the object OBJ detected by the matching process has a calculated matching similarity that is less than the matching judgment threshold.

[0215] The matching threshold is a threshold used to detect object OBJ through the matching process. Specifically, the matching threshold is a threshold used to distinguish between a state in which the object detected by the matching process is highly likely to be object OBJ and a state in which the object detected by the matching process is unlikely to be object OBJ, based on the matching similarity of the object detected by the matching process. The matching threshold can also be said to be a threshold used to distinguish between a state in which the accuracy of at least one of the position and orientation of object OBJ calculated by the matching process is high and a state in which the accuracy of at least one of the position and orientation of object OBJ calculated by the matching process is low, based on the matching similarity. In this case, the matching similarity may be considered to be substantially equivalent to the probability (i.e., likelihood) that the object corresponding to the matching similarity is object OBJ. In other words, the matching similarity may be considered to be substantially equivalent to the likelihood of object OBJ corresponding to the matching similarity.

[0216] As described above, the imaging system 2 may image multiple objects OBJ placed on the mounting device T. In this case, the image data IMD generated by the imaging system 2 may contain multiple objects OBJ. As a result, in step S221 of Figure 12, the position and orientation calculation unit 311 may identify multiple object regions WA. In this case, the position and orientation calculation unit 311 may select (in other words, determine) one of the multiple objects OBJ present in each of the multiple object regions WA as the target object OBJ_tgt. For example, the position and orientation calculation unit 311 may perform a process to identify a target data portion WSD_ROI based on one object region WA for each of the multiple object regions WA detected in step S221 of Figure 12, and a matching process (3D matching process) to detect one object OBJ present in one object region WA using the target data portion WSD_ROI. As a result, the position and orientation calculation unit 311 can calculate the matching similarity of multiple object OBJs present in each of the multiple object regions WA. Subsequently, the position and orientation calculation unit 311 may select one of the multiple object OBJs present in each of the multiple object regions WA as the target object OBJ_tgt based on the matching similarity of the multiple object OBJs. In other words, the position and orientation calculation unit 311 may select one of the multiple object regions WA as the object region WA where the target object OBJ_tgt exists based on the matching similarity of the multiple object OBJs, and then select an object OBJ present in the selected object region WA as the target object OBJ_tgt.

[0217] As an example, the position and orientation calculation unit 311 may select one object OBJ from among a plurality of object OBJs that corresponds to a matching similarity that is equal to or greater than the matching judgment threshold, as the target object OBJ_tgt. As another example, the position and orientation calculation unit 311 may select one object OBJ from among a plurality of object OBJs that corresponds to the largest matching similarity that exceeds the matching judgment threshold, as the target object OBJ_tgt. As yet another example, the position and orientation calculation unit 311 may select one object OBJ from among a plurality of object OBJs that corresponds to a matching similarity that exceeds the matching judgment threshold and is the Nth (where N is a constant representing an integer of 2 or more) largest, as the target object OBJ_tgt. As yet another example, the position and orientation calculation unit 311 may select one object OBJ from among a plurality of object OBJs that corresponds to a matching similarity that is equal to or greater than the matching judgment threshold and is closest to the end effector 4, as the target object OBJ_tgt. As another example, if multiple objects OBJ are placed on the mounting device T (for example, if multiple objects are stacked loosely), the position and orientation calculation unit 311 may select as the target object OBJ_tgt one of the multiple objects OBJ that corresponds to a matching similarity of a value equal to or greater than the matching judgment threshold and has the largest Z coordinate along the Z axis (located at the highest position). As yet another example, the position and orientation calculation unit 311 may select as the target object OBJ_tgt one of the multiple objects OBJ that corresponds to a matching similarity of a value equal to or greater than the matching judgment threshold and that the end effector 4 can perform predetermined processing on.

[0218] Furthermore, the object recognition model 50 may be a learning model that, in addition to outputting object region information OAI regarding the object region WA in which an object OBJ exists when an image is input, outputs information regarding the matching similarity of the object OBJ present in the object region WA. For example, the object recognition model 50 may be a learning model that outputs numerical information indicating the matching similarity itself as information regarding the matching similarity. In this case, the position and orientation calculation unit 311 does not need to calculate the matching similarity during the matching processing. The position and orientation calculation unit 311 may use the information regarding the matching similarity output by the object recognition model 50 to select one object OBJ from among multiple object OBJs present in each of the multiple object region WAs as the target object OBJ_tgt.

[0219] When the object recognition model 50 outputs matching similarity, the training data 561 used to generate the object recognition model 50 may include, in addition to or instead of, a ground truth label 563 containing information about the object region WA, a ground truth label 565 containing information about matching similarity, as shown in Figure 15. The ground truth label 565 contains information about the matching similarity between the template model and the object OBJ or sample object OBJ_s captured in the training image 564 indicated by the training image data 562 associated with the ground truth label 565. For example, if the training image 564 contains an object OBJ, the ground truth label 565 contains information about the matching similarity between the template model and the object OBJ captured in the training image 564. For example, if the training image 564 contains sample object OBJ_s, the ground truth label 565 contains information about the matching similarity between the template model and the sample object OBJ_s captured in the training image 564.

[0220] If the training image 564 contains multiple object OBJs, the ground truth label 565 may include information regarding the matching similarity between each of the multiple object OBJs contained in the training image 564 and the template model. If the training image 564 contains multiple sample object OBJ_s, the ground truth label 565 may include information regarding the matching similarity between each of the multiple sample object OBJ_s contained in the training image 564 and the template model.

[0221] The correct label 565 may include any information that identifies the matching similarity, as information regarding the matching similarity. For example, the correct label 565 may include numerical information that indicates the matching similarity itself, as information regarding the matching similarity.

[0222] As described above, when training data 561 is generated by performing a simulation in which an object model OM is placed in a virtual simulation space SIM, the computing unit 51 may calculate the matching similarity between the object model OM placed in the simulation space and the template model, and generate a ground truth label 565 indicating information about the calculated matching similarity. For example, the computing unit 51 may calculate the matching similarity between the object model OM captured in the training image 564 shown by the training image data 562 obtained by imaging the simulation space with a virtual imaging device 21v, and the template model, and generate a ground truth label 565 indicating information about the calculated matching similarity. Alternatively, if image data IMD_2D, generated by the imaging device 21 capturing a real space in which at least one of the object OBJ and sample object OBJ_s is actually placed, is acquired as training image data 562, the computing device 51 may perform matching processing using the image data IMD_2D to calculate the matching similarity between at least one of the object OBJ and sample object OBJ_s captured in the 2D image IMG_2D shown by the image data IMD_2D and the template model, and generate a ground truth label 565 indicating information about the calculated matching similarity. Alternatively, the annotation to generate the ground truth label 565 and attach it to the training image data 562 may be performed manually by the user.

[0223] Furthermore, one example of matching processing using the image data IMD_2D is contour matching processing (in other words, edge matching processing). The template model used for performing contour matching processing may be a contour template model (edge ​​template model) that shows at least a part of the contour of the object OBJ. The contour template model does not have to show the part of the object OBJ that is enclosed by the contour. In this case, the arithmetic unit 51 may translate, enlarge, reduce and / or rotate the contour template model within the 2D imaging coordinate system based on the imaging device 21 so that the contour template model approaches (for example, matches) the contour of the object OBJ captured in the 2D image IMG_2D shown by the image data IMD_2D. As a result, the arithmetic unit 51 can calculate the matching similarity between the contour template model and at least one of the object OBJ captured in the 2D image IMG_2D shown by the image data IMD_2D and the sample object OBJ_s.

[0224] Another example of matching processing using image data IMD_2D is 2D matching processing. The template model used for 2D matching processing may be a two-dimensional model (two-dimensional template model) that shows the two-dimensional shape of at least a part of the object OBJ. The two-dimensional template model may show the two-dimensional shape of at least a part of the object OBJ using both at least a part of the contour of the object OBJ and at least a part of the portion of the object OBJ enclosed by the contour. In this case, the arithmetic unit 51 may translate, enlarge, reduce and / or rotate the two-dimensional template model in a 2D imaging coordinate system based on the imaging device 21 so that the feature locations of the two-dimensional template model (e.g., at least one of feature points and edges) approach (e.g., match) the feature locations of the object OBJ captured in the 2D image IMG_2D shown by the image data IMD_2D. As a result, the computing device 51 can calculate the matching similarity between the two-dimensional template model and at least one of the object OBJ and sample object OBJ_s captured in the 2D image IMG_2D shown by the image data IMD_2D.

[0225] Subsequently, the arithmetic unit 51 may update the object recognition model 50 by performing machine learning based on information regarding the error between the object region information OAI output by the object recognition model 50 and the correct label 563, and information regarding the error between the matching similarity output by the object recognition model 50 and the correct label 565. For example, the arithmetic unit 51 may update the object recognition model 50 by performing machine learning based on a loss function relating to the error between the object region information OAI output by the object recognition model 50 and the correct label 563, and the error between the matching similarity output by the object recognition model 50 and the correct label 565. In this case, the arithmetic unit 51 may update the object recognition model 50 so that the loss calculated using the loss function is minimized.

[0226] (2-2-2-4) Generation of position and orientation data POI2 (fine search) Again in Figure 12, after the position and orientation data POI1 is generated, the position and orientation calculation unit 311 calculates at least one of the position and orientation of the object OBJ (in particular, one object OBJ selected as the target object OBJ_tgt in step S223) based on the three-dimensional position data WSD generated in step S222 and the position and orientation data POI1 generated in step S223 (step S224). As a result, the position and orientation calculation unit 311 generates position and orientation data (in other words, position and orientation information) POI2 that shows at least one of the position and orientation of the target object OBJ_tgt (step S224). The position and orientation data POI2 generated in step S224 is used as the position and orientation data POI used to generate the robot control signal in step S23 of Figure 11. Furthermore, the process of generating position and orientation data POI2 in step S224 may be called a fine search (fine search process).

[0227] To generate position and orientation data POI2, the position and orientation calculation unit 311 performs a 3D matching process using at least a portion of the three-dimensional position data WSD and a template model (three-dimensional template model TMP_3D). Note that the 3D matching process in step S224 may be the same as the 3D matching process in step S223. For this reason, in order to avoid redundant explanations, a detailed explanation of the 3D matching process in step S224 is omitted.

[0228] However, in step S224, unlike in step S223, the position and orientation calculation unit 311 may generate position and orientation data POI2 by performing a 3D matching process using the three-dimensional position data WSD instead of the target data portion WSD_ROI. Alternatively, in step S224, similar to step S223, the position and orientation calculation unit 311 may generate position and orientation data POI2 by performing a 3D matching process using the target data portion WSD_ROI. In this case, the position and orientation calculation unit 311 may identify the target data portion WSD_ROI used in the 3D matching process in step S224 based on the object region information OAI relating to the object region WA where the object OBJ (in particular, the target object OBJ_tgt selected in step S223) exists, similar to how the target data portion WSD_ROI used in the 3D matching process in step S223 is identified. Alternatively, the position and orientation data POI1 generated in step S223 indicates at least one of the position and orientation of the target object OBJ_tgt. Therefore, it can be said that the position and orientation data POI1 indicates the region in which the target object OBJ_tgt exists. For this reason, the position and orientation calculation unit 311 may identify the target data portion WSD_ROI used in the 3D matching process in step S224 based on the position and orientation data POI1 which indicates at least one of the position and orientation of the target object OBJ_tgt.

[0229] In step S224, the position and orientation calculation unit 311 may generate position and orientation data POI2 by performing a 3D matching process using at least one of the position and orientation determined based on the position and orientation data POI1 and the three-dimensional position data WSD. Specifically, in step S224, the position and orientation calculation unit 311 may generate position and orientation data POI2 by performing a 3D matching process using a three-dimensional template model TMP_3D arranged to match at least one of the position and orientation determined based on the position and orientation data POI1 and the three-dimensional position data WSD.

[0230] For example, the position and orientation calculation unit 311 may determine at least one of the initial position and initial orientation of the three-dimensional template model TMP_3D based on the position and orientation data POI1. For instance, the 3D matching process includes a process of bringing the feature locations of the three-dimensional template model TMP_3D closer to the feature locations of the object OBJ indicated by the three-dimensional position data WSD within the 3D imaging coordinate system. For this reason, the position and orientation calculation unit 311 may determine at least one of the initial position and initial orientation of the three-dimensional template model TMP_3D within the 3D imaging coordinate system.

[0231] To determine at least one of the initial position and initial orientation of the three-dimensional template model TMP_3D, first, as shown in Figure 16, the position and orientation calculation unit 311 converts position and orientation data POI1, which represents at least one of the position and orientation of the target object OBJ_tgt in the reference coordinate system, into position and orientation data POI1_t, which represents the position and orientation of the target object OBJ_tgt in the 3D imaging coordinate system. To convert position and orientation data POI1 into position and orientation data POI1_t, the position and orientation calculation unit 311 may use a transformation matrix (typically a rigid body transformation matrix) to convert coordinates in either the reference coordinate system or the 3D imaging coordinate system to coordinates in the other of the reference coordinate system or the 3D imaging coordinate system. The transformation matrix can be calculated by mathematical methods from external parameters that represent the positional relationship between the reference coordinate system and the imaging device 22. For example, the transformation matrix can be calculated using a mathematical method that solves the PnP (Perspective n Point) problem based on external parameters that indicate the positional relationship between the reference coordinate system and the imaging device 22.

[0232] Subsequently, as shown in Figure 16, the position and orientation calculation unit 311 may set the position of the target object OBJ_tgt indicated by the position and orientation data POI1_t to the initial position of the three-dimensional template model TMP_3D. Alternatively, the position and orientation calculation unit 311 may set the orientation of the target object OBJ_tgt indicated by the position and orientation data POI1_t to the initial orientation of the three-dimensional template model TMP_3D.

[0233] Subsequently, the position and orientation calculation unit 311 may place the three-dimensional template model TMP_3D at the determined initial position within the 3D imaging coordinate system. The position and orientation calculation unit 311 may also place the three-dimensional template model TMP_3D at the determined initial orientation within the 3D imaging coordinate system. Subsequently, the position and orientation calculation unit 311 starts processing to bring the feature locations of the three-dimensional template model TMP_3D, which is placed at the initial position and / or at the initial orientation, closer to the feature locations of the object OBJ indicated by the three-dimensional position data WSD. As a result, the position and orientation calculation unit 311 generates position and orientation data POI2 as a result of the 3D matching process.

[0234] However, in step S224, the position and orientation calculation unit 311 may generate position and orientation data POI2 without using at least one of the position determined based on position and orientation data POI1 and the orientation determined based on position and orientation data POI1.

[0235] Furthermore, in step S224, the position and orientation calculation unit 311 generates position and orientation data POI2 using the three-dimensional position data WSD used to generate position and orientation data POI1 in step S223. In other words, the three-dimensional position data WSD used to generate position and orientation data POI2 is the same as the three-dimensional position data WSD used to generate position and orientation data POI1. To put it another way, the image data IMD_3D used to generate the three-dimensional position data WSD for generating position and orientation data POI2 is the same as the image data IMD_3D used to generate the three-dimensional position data WSD for generating position and orientation data POI1. However, in step S224, the position and orientation calculation unit 311 may generate position and orientation data POI2 using a different three-dimensional position data WSD than the three-dimensional position data WSD used to generate position and orientation data POI1 in step S223. In other words, the three-dimensional position data WSD used to generate the position and orientation data POI2 may be different from the three-dimensional position data WSD used to generate the position and orientation data POI1. To put it another way, the image data IMD_3D used to generate the three-dimensional position data WSD for generating the position and orientation data POI2 may be different from the image data IMD_3D used to generate the three-dimensional position data WSD for generating the position and orientation data POI1. As an example, in step S223, the position and orientation calculation unit 311 may generate the first three-dimensional position data WSD based on the first image data IMD_3D generated when the imaging device 22 images the object OBJ at a first time, and generate the position and orientation data POI1 based on the first three-dimensional position data WSD. Subsequently, in step S224, the position and orientation calculation unit 311 may generate a second three-dimensional position data WSD based on the second image data IMD_3D generated by the imaging device 22 re-imaging the object OBJ at a second time different from the first time, and generate position and orientation data POI2 based on the second three-dimensional position data WSD.

[0236] Furthermore, the three-dimensional template model TMP_3D used in the 3D matching process (i.e., the fine search process) in step S224 may be the same as the three-dimensional template model TMP_3D used in the 3D matching process (i.e., the rough search process) in step S223. In other words, in step S224, the position and orientation calculation unit 311 may perform a 3D matching process using the three-dimensional template model TMP_3D used in the 3D matching process in step S223. Alternatively, the three-dimensional template model TMP_3D used in the 3D matching process in step S224 may be different from the three-dimensional template model TMP_3D used in the 3D matching process in step S223. In other words, in step S224, the position and orientation calculation unit 311 may perform a 3D matching process using a template model TMP_3D different from the three-dimensional template model TMP_3D used in the 3D matching process in step S223.

[0237] (3) Technical Effects As described above, in this embodiment, the control device 3 can identify the object region WA in which an object OBJ exists using the object recognition model 50, which is a learning model that can be generated by machine learning. For this reason, the control device 3 can identify the object region WA more easily than when the object region WA is identified without using the object recognition model 50. Furthermore, the control device 3 can identify an object region WA in which there is less surrounding area where no object OBJ exists (i.e., an area surrounding the region in which an object OBJ exists, in which no object OBJ exists) compared to when the object region WA is identified without using the object recognition model 50. For example, the control device 3 can identify an object region WA in which the proportion of surrounding area where no object OBJ exists is smaller compared to when the object region WA is identified without using the object recognition model 50.

[0238] In particular, the control device 3 can identify the object region WA using image data IMD_2D, which represents a two-dimensional image suitable for processing in the learning model (e.g., matrix operations), and the object recognition model 50. Therefore, the control device 3 can identify the object region WA more easily and quickly compared to the case where the object region WA is identified without using the image data IMD_2D and the object recognition model 50.

[0239] Furthermore, in this embodiment, the control device 3 uses the target data portion WSD_ROI, which corresponds to the object region WA of the three-dimensional position data WSD, to perform a 3D matching process (i.e., a rough search process) to generate position and orientation data POI1 that indicates at least one of the position and orientation of the target object OBJ_tgt. Here, the data size of the target data portion WSD_ROI is smaller than the data size of the three-dimensional position data WSD. Therefore, the time required to perform the rough search process using the target data portion WSD_ROI is shorter than the time required to perform the rough search process using the three-dimensional position data WSD. As a result, compared to the case where the rough search process is performed using the three-dimensional position data WSD, the control device 3 can reduce the time required to generate the position and orientation data POI1.

[0240] In this embodiment, the control device 3 can use the three-dimensional position data WSD, whose data size has been reduced by resizing, in order to generate the position and orientation data POI1 (i.e., perform a rough search). Therefore, the time required to perform a rough search using the target data portion WSD_ROI identified from the three-dimensional position data WSD, whose data size has been reduced by resizing, is shorter than the time required to perform a rough search using the target data portion WSD_ROI identified from the original three-dimensional position data WSD, whose data size has not been reduced by resizing. Therefore, compared to the case where a rough search is performed using the target data portion WSD_ROI identified from the original three-dimensional position data WSD, whose data size has not been reduced by resizing, the control device 3 can reduce the time required to generate the position and orientation data POI1.

[0241] Furthermore, the smaller the object region WA identified by the control device 3 using the object recognition model 50 (i.e., the object region WA identified by the object recognition model 50), the smaller the data size of the target data portion WSD_ROI becomes. For this reason, if prioritizing the reduction of the time required to generate the position and orientation data POI1, the object recognition model 50 may be generated by machine learning to identify an object region WA that includes the region where an object OBJ exists but does not include the region where an object OBJ does not exist (for example, object region WA#1 in Figure 8), instead of identifying an object region WA that includes the region where an object OBJ exists and the region where an object OBJ does not exist (for example, object region WA#2 or WA#3 in Figure 8).

[0242] Furthermore, in this embodiment, after performing a rough search, the control device 3 performs a 3D matching process (i.e., a fine search process) to generate position and orientation data POI2 that indicates at least one of the position and orientation of the target object OBJ_tgt using the three-dimensional position data WSD. As a result, the position and orientation data POI2 is used as a deterministic position and orientation data POI to generate the robot control signal. Here, the position and orientation data POI2 generated by the fine search process using the original three-dimensional position data WSD, whose data size has not been reduced by the resizing process, is highly likely to indicate at least one of the position and orientation of the target object OBJ_tgt with higher accuracy compared to the position and orientation data POI1 generated by the rough search process using the three-dimensional position data WSD, whose data size has been reduced by the resizing process. Therefore, the control device 3 can generate the robot control signal using the position and orientation data POI2 that indicates at least one of the position and orientation of the target object OBJ_tgt with higher accuracy. As a result, the robot 1 can perform predetermined processing more appropriately using the end effector 4.

[0243] Furthermore, in this embodiment, the control device 3 can perform a fine search using at least one of the position and orientation determined based on the position and orientation data POI1 generated by the rough search process. Specifically, the control device 3 can perform a 3D matching process using a three-dimensional template model TMP_3D positioned to match at least one of the position and orientation determined based on the position and orientation data POI1 generated by the rough search process, as a fine search process. Here, the position and orientation data POI1 indicates at least one of the position and orientation of the target object OBJ_tgt. Therefore, compared to the case where the three-dimensional template model TMP_3D is positioned without using the position and orientation determined based on the position and orientation data POI1, there is a high probability that the three-dimensional template model TMP_3D will be positioned at the same or close to the position of the target object OBJ_tgt and / or that the orientation of the three-dimensional template model TMP_3D will be the same as or close to the orientation of the target object OBJ_tgt. As a result, the possibility of not being able to detect the object OBJ (in this case, the target object OBJ_tgt detected by the rough search process) corresponding to the three-dimensional template model TMP_3D becomes lower. In other words, the possibility of detection errors occurring becomes lower. This is because the greater the discrepancy between the initial position and orientation of the three-dimensional template model TMP_3D and the actual position and orientation of the object OBJ (in this case, the target object OBJ_tgt detected by the rough search process), the higher the possibility that an object different from the object OBJ corresponding to the three-dimensional template model TMP_3D will be detected by the 3D matching process, or that the object OBJ corresponding to the three-dimensional template model TMP_3D will not be detected. Therefore, the control device 3 can generate position and orientation data POI2 that shows at least one of the position and orientation of the target object OBJ_tgt with higher accuracy.

[0244] (4) Modifications Next, we will explain modifications of the robot system SYS.

[0245] (4-1) First Modification In the above description, in step S223 of Figure 12, the position and orientation calculation unit 311 calculates at least one of the position and orientation of the object OBJ (in particular, the target object OBJ_tgt) by performing a 3D matching process. In the first modification, the position and orientation calculation unit 311 may calculate at least one of the position and orientation of the object OBJ (in particular, the target object OBJ_tgt) using the object recognition model 50 without performing a 3D matching process.

[0246] Specifically, in the first modified example, the object recognition model 50 may be a learning model that, when an image is input, outputs position and orientation information relating to at least one of the position and orientation of the object OBJ present in the object region WA (i.e., the object OBJ captured in the input image), in addition to or instead of object region information OAI relating to the object region WA in which the object OBJ exists within the input image. For example, the object recognition model 50 may be a learning model that outputs numerical information (e.g., coordinate information) indicating at least one of the position and orientation of the object OBJ present in the object region WA as position and orientation information.

[0247] In this case, the position and orientation calculation unit 311 may use the position and orientation information output by the object recognition model 50 as at least one of the position and orientation of the object OBJ. The position and orientation calculation unit 311 may use the position and orientation information output by the object recognition model 50 as position and orientation data POI1 indicating at least one of the position and orientation of the object OBJ. The position and orientation calculation unit 311 may use the information regarding at least one of the position and orientation of the object OBJ output by the object recognition model 50 as the result of a rough search process. In this case, the position and orientation calculation unit 311 does not need to generate three-dimensional position data WSD in step S222 of Figure 12, nor does it need to perform 3D matching processing in step S223 of Figure 12. As a result, the time required to generate the position and orientation data POI1 can be shortened. Furthermore, the control device 3 does not need to have a logical processing block (so-called software module) necessary to generate the position and orientation data POI1.

[0248] As described above, when the imaging system 2 images multiple objects OBJ placed on the mounting device T, the object recognition model 50 may output multiple position and orientation information, each indicating at least one of the positions and orientations of the multiple objects OBJ. In this case, the position and orientation calculation unit 311 may use one of the multiple position and orientation information outputs from the object recognition model 50 as position and orientation data POI1. For example, when the object recognition model 50 outputs a matching similarity score as described above, the position and orientation calculation unit 311 may use one of the multiple position and orientation information outputs from the object recognition model 50, indicating at least one of the positions and orientations of one object OBJ selected as the target object OBJ_tgt based on the matching similarity score described above, as position and orientation data POI1.

[0249] When the object recognition model 50 outputs position and orientation information, the training data 561 used to generate the object recognition model 50 may include, as shown in Figure 17, a ground truth label 563 containing information about the object region WA (or at least one of the ground truth labels 563 and the ground truth label 565 containing information about matching similarity shown in Figure 15), as well as a ground truth label 566 containing position and orientation information. The ground truth label 566 includes position and orientation information relating to at least one of the position and orientation of an object OBJ or sample object OBJ_s captured in the training image 564 shown by the training image data 562 associated with the ground truth label 566. For example, if an object OBJ is captured in the training image 564, the ground truth label 566 includes position and orientation information relating to at least one of the position and orientation of the object OBJ captured in the training image 564. For example, if the training image 564 contains the sample object OBJ_s, the correct label 566 includes positional and orientation information relating to at least one of the position and orientation of the sample object OBJ_s contained in the training image 564.

[0250] If the training image 564 contains multiple object OBJs, the correct label 566 may include multiple positional orientation information corresponding to each of the multiple object OBJs contained in the training image 564. If the training image 564 contains multiple sample object OBJ_s, the correct label 566 may include multiple positional orientation information corresponding to each of the multiple sample object OBJ_s contained in the training image 564.

[0251] The correct label 566 may include any information that can identify at least one of the position and orientation of the object OBJ or the sample object OBJ_s as position and orientation information. For example, the correct label 566 may include numerical information (e.g., coordinate information) that indicates at least one of the position and orientation of the object OBJ or the sample object OBJ_s itself as position and orientation information.

[0252] As described above, when training data 561 is generated by performing a simulation in which an object model OM is placed in a virtual simulation space SIM, the computing unit 51 may calculate at least one of the position and orientation of the object model OM placed in the simulation space and generate a ground truth label 566 indicating information about at least one of the calculated position and orientation of the object model OM. When image data IMD_2D, generated by the imaging device 21 capturing images of the real space in which at least one of the object OBJ and sample object OBJ_s is actually placed, is acquired as training image data 562, the computing unit 51 may perform a matching process using the image data IMD_2D to calculate at least one of the position and orientation of at least one of the object OBJ and sample object OBJ_s captured in the 2D image IMG_2D shown by the image data IMD_2D, and generate a ground truth label 565 indicating information about at least one of the calculated position and orientation. An example of a matching process using image data IMD_2D is at least one of the contour matching process and 2D matching process described above. Alternatively, the annotation process, which involves assigning the correct label 566 to the training image data 562, may be performed manually by the user.

[0253] Subsequently, the arithmetic unit 51 may update the object recognition model 50 by performing machine learning based on information regarding the error between the object region information OAI output by the object recognition model 50 and the correct label 563, and information regarding the error between the position and orientation information output by the object recognition model 50 and the correct label 566. For example, the arithmetic unit 51 may update the object recognition model 50 by performing machine learning based on a loss function relating to the error between the object region information OAI output by the object recognition model 50 and the correct label 563, and the error between the position and orientation information output by the object recognition model 50 and the correct label 566. In this case, the arithmetic unit 51 may update the object recognition model 50 so that the loss calculated using the loss function is minimized.

[0254] Furthermore, the position and orientation calculation unit 311 may calculate at least one of the position and orientation of the object OBJ (especially the target object OBJ_tgt) by performing 3D matching processing, and may also calculate at least one of the position and orientation of the object OBJ (especially the target object OBJ_tgt) using the object recognition model 50. In this case, the position and orientation calculation unit 311 may generate position and orientation data POI1 used in the fine search processing of step S224 in Figure 12 using first position and orientation data POI1 indicating at least one of the position and orientation of the object OBJ calculated by 3D matching processing, and second position and orientation data POI1 indicating at least one of the position and orientation of the object OBJ calculated using the object recognition model 50. For example, the position and orientation calculation unit 311 may calculate the average value (e.g., a simple average or a weighted average) of the position of the object OBJ indicated by the first position and orientation data POI1 and the position of the object OBJ indicated by the second position and orientation data POI1, and generate position and orientation data POI1 that indicates the calculated position as the position of the object OBJ, as the position and orientation data POI1 used in the fine search process. For example, the position and orientation calculation unit 311 may calculate the average value (e.g., a simple average or a weighted average) of the orientation of the object OBJ indicated by the first position and orientation data POI1 and the orientation of the object OBJ indicated by the second position and orientation data POI1, and generate position and orientation data POI1 that indicates the calculated orientation as the orientation of the object OBJ, as the position and orientation data POI1 used in the fine search process.

[0255] (4-2) Second Modification In the second modification, in step S223 of Figure 12 described above, the position and orientation calculation unit 311 may interpolate the target data portion WSD_ROI, which is a part of the three-dimensional position data WSD, before performing the 3D matching process using the target data portion WSD_ROI. After that, the position and orientation calculation unit 311 may generate the position and orientation data POI1 by performing a 3D matching process using the interpolated target data portion WSD_ROI and the three-dimensional template model TMP_3D.

[0256] Furthermore, in step S224 of Figure 12 described above, if the position and orientation calculation unit 311 performs 3D matching processing using the target data portion WSD_ROI, the position and orientation calculation unit 311 may interpolate the target data portion WSD_ROI before performing 3D matching processing using the target data portion WSD_ROI. Subsequently, the position and orientation calculation unit 311 may generate position and orientation data POI2 by performing 3D matching processing using the interpolated target data portion WSD_ROI and the three-dimensional template model TMP_3D.

[0257] Furthermore, in step S224 of Figure 12 described above, if the position and orientation calculation unit 311 performs 3D matching processing using the three-dimensional position data WSD, the position and orientation calculation unit 311 may interpolate the three-dimensional position data WSD before performing 3D matching processing using the three-dimensional position data WSD. Subsequently, the position and orientation calculation unit 311 may generate position and orientation data POI2 by performing 3D matching processing using the interpolated three-dimensional position data WSD and the three-dimensional template model TMP_3D.

[0258] Furthermore, the process of interpolating the three-dimensional position data WSD may be the same as the process of interpolating the target data portion WSD_ROI. For this reason, the following explanation will describe the process of interpolating the target data portion WSD_ROI. The explanation of the process of interpolating the target data portion WSD_ROI can be reused as an explanation of the process of interpolating the three-dimensional position data WSD by replacing the phrase "target data portion WSD_ROI" with the phrase "three-dimensional position data WSD".

[0259] As described above, this embodiment describes an example in which the point cloud data described above is used as the three-dimensional position data WSD. In this case, the target data portion WSD_ROI is also point cloud data. In this case, the position and orientation calculation unit 311 may interpolate the target data portion WSD_ROI by interpolating the point cloud indicated by the target data portion WSD_ROI. In other words, the position and orientation calculation unit 311 may interpolate the target data portion WSD_ROI by adding new points that constitute the point cloud to the point cloud indicated by the target data portion WSD_ROI (i.e., interpolating). The position and orientation calculation unit 311 may interpolate the target data portion WSD_ROI such that the number of points indicated by the interpolated target data portion WSD_ROI is greater than the number of points indicated by the uninterpolated target data portion WSD_ROI.

[0260] The position and orientation calculation unit 311 may interpolate the target data portion WSD_ROI using an interpolation model IM representing the object OBJ, as shown in Figure 18. Figure 18 shows an example of interpolating the first target data portion WSD_ROI using the interpolation model IM, and an example of interpolating the second target data portion WSD_ROI using the interpolation model IM. The interpolation model IM is a three-dimensional model representing the three-dimensional shape of the object OBJ. An example of the interpolation model IM is a CAD (Computer-Aided Design) model of the object OBJ. An example of the interpolation model IM is a CAD model of a sample object OBJ_s that simulates the object OBJ. An example of the interpolation model IM is a model (e.g., a point cloud model) generated from a CAD model of the object OBJ or sample object OBJ_s. One example of an interpolation model IM is a model of an object OBJ obtained by actually measuring the shape of the object OBJ using a three-dimensional shape measurement device such as a 3D scanner. Another example of an interpolation model IM is a model of an object OBJ obtained by actually measuring the shape of a sample object OBJ_s that simulates an object OBJ using a three-dimensional shape measurement device such as a 3D scanner.

[0261] In order to interpolate the target data portion WSD_ROI using the interpolation model IM, the position and orientation calculation unit 311 may align the interpolation model IM with respect to the target data portion WSD_ROI, as shown in Figure 18. For example, the position and orientation calculation unit 311 may align the interpolation model IM with respect to the target data portion WSD_ROI by adjusting the positional relationship between the object OBJ, which represents the three-dimensional position of the target data portion WSD_ROI, and the interpolation model IM. For example, the position and orientation calculation unit 311 may align the interpolation model IM with respect to the target data portion WSD_ROI by adjusting the orientation relationship between the object OBJ, which represents the three-dimensional position of the target data portion WSD_ROI, and the interpolation model IM. As an example, the position and orientation calculation unit 311 may translate, enlarge, shrink, and / or rotate the interpolation model IM so that the feature locations of the interpolation model IM (for example, at least one of feature points and edges) approach (for example, coincide with) the feature locations of the object OBJ (for example, the point cloud corresponding to the object OBJ indicated by the target data portion WSD_ROI) where the target data portion WSD_ROI indicates a three-dimensional position. In this case, the position and orientation calculation unit 311 may align the interpolation model IM with respect to the target data portion WSD_ROI so that the similarity between the interpolation model IM and the object OBJ where the target data portion WSD_ROI indicates a three-dimensional position (for example, the matching similarity described above) is maximized or exceeds a predetermined interpolation threshold.

[0262] Subsequently, the position and orientation calculation unit 311 may interpolate the target data portion WSD_ROI with the interpolation model IM. For example, the position and orientation calculation unit 311 may generate new data obtained by combining the target data portion WSD_ROI with at least a part of the interpolation model IM as the interpolated target data portion WSD_ROI. For example, the position and orientation calculation unit 311 may generate new data obtained by combining the target data portion WSD_ROI with data indicating the three-dimensional position of the object OBJ using at least a part of the interpolation model IM as the interpolated target data portion WSD_ROI. For example, the position and orientation calculation unit 311 may generate new point cloud data obtained by combining the target data portion WSD_ROI, which is point cloud data, with point cloud data indicating points included in at least a part of the interpolation model IM, which is a point cloud model, as the interpolated target data portion WSD_ROI. In other words, the position and orientation calculation unit 311 may generate new point cloud data obtained by combining point cloud data indicating the three-dimensional positions of multiple points of an object OBJ using at least a part of the target data portion WSD_ROI, which is point cloud data, and the interpolation model IM, which is a point cloud model, as the interpolated target data portion WSD_ROI. For example, if the interpolation model IM is not a point cloud model, the position and orientation calculation unit 311 may convert the interpolation model IM to a point cloud model and generate new point cloud data obtained by combining it with point cloud data indicating points included in at least a part of the target data portion WSD_ROI, which is point cloud data, as the interpolated target data portion WSD_ROI.

[0263] As a result, as shown in Figure 18, the number of points indicated by the interpolated target data portion WSD_ROI is greater than the number of points indicated by the uninterpolated target data portion WSD_ROI. In other words, the proportion of the object OBJ where the interpolated target data portion WSD_ROI indicates the three-dimensional position is greater than the proportion of the object OBJ where the uninterpolated target data portion WSD_ROI indicates the three-dimensional position is greater than the proportion of the object OBJ where the uninterpolated target data portion WSD_ROI indicates the three-dimensional position. Therefore, it becomes more likely that the interpolated target data portion WSD_ROI indicates the three-dimensional position of a part of the object OBJ that the uninterpolated target data portion WSD_ROI did not indicate the three-dimensional position of. As a first example, if the first part of the object OBJ is located in a blind spot of the imaging range of the imaging device 22, the first part of the object OBJ may not be captured in the image shown by the image data IMD_3D generated by the imaging device 22. In this case, the uninterpolated target data portion WSD_ROI may not indicate the three-dimensional position of the first part of the object OBJ. However, even in this case, if the target data portion WSD_ROI is interpolated by the interpolation model IM that represents the first part of the object OBJ, the interpolated target data portion WSD_ROI is more likely to indicate the three-dimensional position of the first part of the object OBJ. As a second example, if the second part of the object OBJ is dark, the image data IMD_3D generated by the imaging device 22 may not clearly show the second part of the object OBJ. In this case, the uninterpolated target data portion WSD_ROI may not indicate the three-dimensional position of the second part of the object OBJ. However, even in this case, if the target data portion WSD_ROI is interpolated by the interpolation model IM that represents the second part of the object OBJ, the interpolated target data portion WSD_ROI is more likely to indicate the three-dimensional position of the second part of the object OBJ. As a third example, if the imaging device 22 images a third portion of a thin, flat object OBJ from a direction intersecting the thickness direction of the third portion, the image data IMD_3D generated by the imaging device 22 may only show a small portion of the third portion OP3, which corresponds to the edge of the thin part of the flat object OBJ.In other words, the image data IMD_3D generated by the imaging device 22 may not capture the fourth part OP4 of the object OBJ (i.e., most of the object OBJ), other than the third part OP3, which corresponds to the edge of the thin part of the flat plate. In this case, as shown in the right-hand diagram of Figure 18, the uninterpolated target data portion WSD_ROI may not indicate the three-dimensional position of the fourth part of the object OBJ other than the edge of the thin part of the flat plate (i.e., most of the object OBJ). However, even in this case, if the target data portion WSD_ROI is interpolated by the interpolation model IM which represents the fourth part OP4 of the object OBJ, then, as shown in the right-hand diagram of Figure 18, the interpolated target data portion WSD_ROI is more likely to indicate the three-dimensional position of the fourth part OP4 of the object OBJ. As a result, compared to the case where 3D matching processing is performed using the uninterpolated target data portion WSD_ROI, the likelihood of detecting object OBJs by 3D matching processing using the interpolated target data portion WSD_ROI increases. In other words, the position and orientation calculation unit 311 can appropriately detect object OBJs that the end effector 4 should perform predetermined processing on by performing 3D matching processing. Conversely, the possibility of a situation occurring where the position and orientation calculation unit 311 fails to detect any object OBJs that the end effector 4 should perform predetermined processing on by performing 3D matching processing is low. For this reason, the control device 3 can appropriately control the robot 1, the end effector 4, and at least one of the robot's movable devices to perform predetermined processing on the object OBJs.

[0264] Furthermore, compared to the case where 3D matching processing is performed using the uninterpolated target data portion WSD_ROI, the matching similarity of the object OBJ detected by 3D matching processing using the interpolated target data portion WSD_ROI may be higher. The higher the matching similarity of the object OBJ, the higher the calculation accuracy of at least one of the position and orientation of the object OBJ. This is because, as the matching similarity of the object OBJ increases, the proportion of the object OBJ parts that match the three-dimensional template model TMP_3D increases, and therefore the position and orientation of the three-dimensional template model TMP_3D calculated as the position and orientation of the object OBJ become closer to the actual position and orientation of the object OBJ, respectively. For this reason, compared to the case where 3D matching processing is performed using the uninterpolated target data portion WSD_ROI, the position and orientation calculation unit 311 can calculate at least one of the position and orientation of the object OBJ with higher accuracy by performing 3D matching processing using the interpolated target data portion WSD_ROI.

[0265] Furthermore, the interpolation model IM used to interpolate the target data portion WSD_ROI may be the same as the three-dimensional template model TMP_3D used in the 3D matching process (i.e., at least one of the rough search process and the fine search process) in at least one of steps S223 and S224. In other words, the position and orientation calculation unit 311 may interpolate the target data portion WSD_ROI by using the three-dimensional template model TMP_3D used in the 3D matching process in at least one of steps S223 and S224 as the interpolation model IM. Alternatively, the interpolation model IM used to interpolate the target data portion WSD_ROI may be different from the three-dimensional template model TMP_3D used in the 3D matching process in at least one of steps S223 and S224. In other words, the position and orientation calculation unit 311 may interpolate the target data portion WSD_ROI using an interpolation model IM that is different from the three-dimensional template model TMP_3D used in the 3D matching process in at least one of steps S223 and S224.

[0266] When interpolating the target data portion WSD_ROI used in the fine search process in step S224 of Figure 12, the position and orientation calculation unit 311 may interpolate the target data portion WSD_ROI using the position and orientation data POI1 generated in the rough search process. For example, the position and orientation calculation unit 311 may place the interpolated model IM at the position indicated by the position and orientation data POI1 before aligning the interpolated model IM with respect to the target data portion WSD_ROI. For example, the position and orientation calculation unit 311 may place the interpolated model IM in the orientation indicated by the position and orientation data POI1 before aligning the interpolated model IM with respect to the target data portion WSD_ROI. As a result, there is a high probability that the interpolated model IM will be placed at the same or close position as the object OBJ that the target data portion WSD_ROI represents in three dimensions, and / or that the orientation of the interpolated model IM will be the same as or close to the orientation of the object OBJ that the target data portion WSD_ROI represents in three dimensions. As a result, compared to the case where the interpolation model IM is positioned at a location significantly different from the object OBJ that indicates the three-dimensional position of the target data portion WSD_ROI, and / or the orientation of the interpolation model IM is significantly different from the orientation of the object OBJ that indicates the three-dimensional position of the target data portion WSD_ROI, the position and orientation calculation unit 311 can shorten the time required to align the interpolation model IM with respect to the target data portion WSD_ROI.

[0267] Furthermore, even when interpolating the target data portion WSD_ROI used in the rough search process in step S223 of Figure 12, the position and orientation calculation unit 311 may use at least one of the position and orientation of the object OBJ, which the target data portion WSD_ROI represents, to align the interpolation model IM with respect to the target data portion WSD_ROI. For example, the position and orientation calculation unit 311 may use the position and orientation information output by the object recognition model 50 in step S221 of Figure 12 (i.e., information regarding at least one of the position and orientation of the object OBJ) to align the interpolation model IM with respect to the target data portion WSD_ROI. For example, the position and orientation calculation unit 311 may calculate at least one of the position and orientation of the object OBJ by performing a matching process using at least one of the image data IMD_2D acquired in step S21 of Figure 11 and the three-dimensional position data WSD generated in step S222 of Figure 12, and then use the calculation result of at least one of the position and orientation of the object OBJ to align the interpolation model IM with the target data portion WSD_ROI. As an example of a matching process using the image data IMD_2D, at least one of the contour matching process and 2D matching process described above can be given. As an example of a matching process using the three-dimensional position data WSD, the 3D matching process described above can be given.

[0268] (4-3) Third Modification As described above, the end effector 4 may perform a predetermined process on one object OBJ selected as the target object OBJ_tgt from among a plurality of object OBJs, which include at least two object OBJs that are at least partially overlapping. Here, as shown in Figure 19, an example will be described in which the end effector 4 performs the above-described holding process in a situation where the first object OBJ#1 and the second object OBJ#2 among the plurality of object OBJs are overlapping such that the first object OBJ#1 is located on top of the second object OBJ#2. In this example, if the matching similarity of the first object OBJ#1 is higher than the matching similarity of the second object OBJ#2, the first object OBJ#1 is selected as the target object OBJ_tgt. In this case, the end effector 4 can properly hold the first object OBJ#1 without being significantly affected by the second object OBJ#2 located below the first object OBJ#1. On the other hand, if the matching similarity of the second object OBJ#2 is higher than that of the first object OBJ#1, the second object OBJ#2 is selected as the target object OBJ_tgt. In this case, when the end effector 4 holds the second object OBJ#2, it may be affected by the first object OBJ#1 which is located above the second object OBJ#2. For example, when the end effector 4 holds the second object OBJ#2, the first object OBJ#1 which is located above the second object OBJ#2 may also be unintentionally lifted by the end effector 4. However, because the first object OBJ#1 is not properly held by the end effector 4, the first object OBJ#1 which is unintentionally lifted by the end effector 4 may fall.

[0269] Therefore, in the third modified example, in step S223 of Figure 12, the position and orientation calculation unit 311 may, in addition to or instead of selecting one of the multiple object OBJs as the target object OBJ_tgt based on matching similarity, select one of the multiple object OBJs as the target object OBJ_tgt based on the priority of the multiple object OBJs. Below, an example will be described in which the position and orientation calculation unit 311 selects either the first object OBJ#1 or the second object OBJ#2 as the target object OBJ_tgt based on the priority of the first object OBJ#1 and the priority of the second object OBJ#2.

[0270] The priority of an object OBJ may be an index value determined based on the degree of overlap of the object OBJs. For example, the priority of one object OBJ may be an index value determined based on the degree of overlap between that object OBJ and other object OBJs that are different from that object OBJ. For example, the priority of one object OBJ may be an index value that increases as the degree of overlap between one object OBJ and other object OBJs decreases. In particular, the priority of one object OBJ may be an index value that increases as the degree of overlap between one object OBJ and other object OBJs decreases when other object OBJs are located on top of one object OBJ. Furthermore, the degree of overlap between one object OBJ and another object OBJ may be an index value equivalent to the ratio of the size (e.g., area) of the portion of one object OBJ that is overlapped by another object OBJ to the size (e.g., area) of one object OBJ included in the imaging field of the imaging system 2, assuming that the other object OBJ is not overlapping one object OBJ. For this reason, the smaller the portion of one object OBJ that is overlapped by another object OBJ, the lower the degree of overlap between one object OBJ and another object OBJ. In other words, the smaller the portion of one object OBJ that is overlapped by another object OBJ, the higher the priority of one object OBJ.

[0271] The position and orientation calculation unit 311 may calculate the priority of multiple object OBJs in order to select one of the multiple object OBJs as the target object OBJ_tgt based on the priority of the multiple object OBJs.

[0272] As a first example for calculating the priority of an object OBJ, the position and orientation calculation unit 311 may calculate the priority by performing the matching process described above. As an example, the position and orientation calculation unit 311 may calculate the priority by performing the contour matching process described above. For example, if no other object OBJ is superimposed on one object OBJ, it is highly likely that the 2D image IMG_2D shown by the image data IMD_2D used for contour matching will capture all (or most) of the contour of the object OBJ. For this reason, the similarity (i.e., matching similarity) between the contour of the object calculated by the edge matching process and the contour template model will be high. On the other hand, if other object OBJs are superimposed on one object OBJ, it is highly likely that the 2D image IMG_2D shown by the image data IMD_2D used for contour matching will not capture part of the contour of the object OBJ. Therefore, the proportion of the contour of an object OBJ that is not captured in the 2D image IMG_2D increases, and the lower the matching similarity becomes. For this reason, the matching similarity of an object OBJ calculated by the contour matching process essentially indicates the degree of overlap between one object OBJ and other object OBJs. Accordingly, the position and orientation calculation unit 311 may calculate the matching similarity of an object OBJ by performing contour matching processing and calculate the priority of an object OBJ based on the matching similarity of the object OBJ. For example, the position and orientation calculation unit 311 may calculate the priority of an object OBJ such that the priority of an object OBJ increases as the matching similarity of the object OBJ increases.

[0273] Furthermore, the matching similarity of an object OBJ calculated by a matching process different from the contour matching process (for example, at least one of 2D matching and 3D matching processes) can also be said to substantially indicate the degree of overlap between one object OBJ and other object OBJs. For this reason, the position and orientation calculation unit 311 may calculate the priority of object OBJs by performing a matching process different from the contour matching process (for example, at least one of 2D matching and 3D matching processes) in addition to or instead of the contour matching process. Alternatively, the position and orientation calculation unit 311 may calculate the priority of object OBJs based on at least one of the image data IMD acquired in step S21 of Figure 11 and the three-dimensional position data WSD generated in step S222 of Figure 12, in addition to or instead of performing the matching process.

[0274] As a second example for calculating the priority of an object OBJ, the position and orientation calculation unit 311 may calculate the priority using the object recognition model 50 described above. In this case, the object recognition model 50 may be a learning model that, in addition to outputting object region information OAI relating to the object region WA in which an object OBJ exists when an image is input, outputs information regarding the priority of an object OBJ existing in the object region WA. For example, the object recognition model 50 may be a learning model that outputs numerical information indicating the priority of the object OBJ itself existing in the object region WA as information regarding priority.

[0275] If the object recognition model 50 outputs a priority, the training data 561 used to generate the object recognition model 50 may include, in addition to or instead of, the ground truth label 563 (or at least one of the ground truth labels 563, 565, and 566 described above) which includes information about the object region WA, as shown in Figure 20, a ground truth label 567 which includes information about the priority. The ground truth label 567 includes information about the priority (e.g., an index value that depends on the degree of overlap) of the object OBJ or sample object OBJ_s captured in the training image 564 indicated by the training image data 562 associated with the ground truth label 567. For example, if the training image 564 captures an object OBJ, the ground truth label 567 includes information about the priority of the object OBJ captured in the training image 564. For example, if the training image 564 captures sample object OBJ_s, the ground truth label 567 includes information about the priority of the sample object OBJ_s captured in the training image 564.

[0276] If multiple object OBJs are captured in the training image 564, the correct label 567 may include information regarding the priority of each of the multiple object OBJs captured in the training image 564. If multiple sample object OBJ_s are captured in the training image 564, the correct label 567 may include information regarding the priority of each of the multiple sample object OBJ_s captured in the training image 564.

[0277] The correct label 567 may include any information that identifies the priority as information regarding priority. For example, the correct label 567 may include numerical information that indicates the priority itself as information regarding priority.

[0278] As described above, when training data 561 is generated by performing a simulation in which an object model OM is placed in a virtual simulation space SIM, the computing unit 51 may calculate an index value that depends on the degree of overlap between one object model OM placed in the simulation space and other models (for example, other object models OM), as the priority of one object model OM, and generate a ground truth label 567 that indicates information about the calculated priority. For example, the computing unit 51 may calculate the degree of overlap between one object model OM and other models (for example, other object models OM) that are captured in the training image 564 shown by the training image data 562 obtained by imaging the simulation space with a virtual imaging device 21v, and generate a ground truth label 567 that indicates information about an index value (i.e., the priority of one object model OM) that depends on the calculated degree of overlap. Alternatively, if the image data IMD_2D generated by the imaging device 21 capturing a real space in which at least one of the object OBJ and sample object OBJ_s is actually located is acquired as the training image data 562, the arithmetic unit 51 may perform matching processing using the image data IMD_2D to calculate the priority of at least one of the object OBJ and sample object OBJ_s captured in the 2D image IMG_2D shown by the image data IMD_2D, and generate a correct label 567 indicating information about the calculated priority. Alternatively, the annotation to assign the correct label 567 to the training image data 562 may be performed manually by the user.

[0279] After the priorities of multiple object OBJs have been calculated, the position and orientation calculation unit 311 may select one of the multiple object OBJs as the target object OBJ_tgt based on the priorities. For example, the position and orientation calculation unit 311 may select the object OBJ with the highest priority among the multiple object OBJs as the target object OBJ_tgt. For example, the position and orientation calculation unit 311 may select one object OBJ whose priority among the multiple object OBJs is equal to or greater than a predetermined priority threshold as the target object OBJ_tgt. As an example, Figure 19 shows an example where the priority of the first object OBJ #1 (= 0.922) is higher than the priority of the second object OBJ (= 0.752) because the first object OBJ #1 is superimposed on the second object OBJ #2. In this case, the position and orientation calculation unit 311 may select the first object OBJ #1 as the target object OBJ_tgt. On the other hand, if the priority of the second object OBJ#2 is higher than the priority of the first object OBJ#1, the position and orientation calculation unit 311 may select the second object OBJ#2 as the target object OBJ_tgt.

[0280] In the example shown in Figure 19, the matching similarity of the first object OBJ#1 (=0.93) is lower than that of the second object OBJ (=0.95). Therefore, if the target object OBJ_tgt is selected based on matching similarity without using priority, the position and orientation calculation unit 311 will select the second object OBJ#2, which is superimposed on the first object OBJ#1, as the target object OBJ_tgt. As a result, the aforementioned technical problem may occur in which the second object OBJ#2 is unintentionally lifted. However, in the third modified example, the target object OBJ_tgt is selected based on priority in addition to or instead of matching similarity, thus reducing the likelihood of such a technical problem occurring.

[0281] Furthermore, when the position and orientation calculation unit 311 calculates the priority of an object OBJ using the object recognition model 50, the object recognition model 50 may also output information regarding the certainty (in other words, accuracy) of the priority in addition to the priority. In this case, the position and orientation calculation unit 311 may select the target object OBJ_tgt based on the priority of the object OBJ and the certainty of that priority. For example, the position and orientation calculation unit 311 may select as the target object OBJ_tgt one object OBJ among a plurality of object OBJs that has the highest priority and whose certainty of priority is above a certain value. For example, the position and orientation calculation unit 311 may select as the target object OBJ_tgt one object OBJ among a plurality of object OBJs whose priority is above a predetermined priority threshold and whose certainty of priority is above a certain value.

[0282] If the object recognition model 50 outputs information regarding the likelihood of priority, the object recognition model 50 may be a learning model that outputs information regarding the likelihood of priority of an object OBJ present in the object region WA when a certain image is input. If the object recognition model 50 outputs the likelihood of priority, the training data 561 used to generate the object recognition model 50, and the correct labels 567, may include information regarding the likelihood of priority.

[0283] The position and orientation calculation unit 311 may select one of the multiple object OBJs as the target object OBJ_tgt based on both the matching similarity and priority.

[0284] As a first example of selecting a target object OBJ_tgt based on both matching similarity and priority, the position and orientation calculation unit 311 may select at least one object OBJ from among a plurality of object OBJs based on matching similarity, and then select one object OBJ as the target object OBJ_tgt from among the at least one object OBJ selected based on matching similarity, based on priority. For example, the position and orientation calculation unit 311 may select at least one object OBJ from among a plurality of object OBJs whose matching similarity is equal to or greater than a matching determination threshold, and then select one object OBJ as the target object OBJ_tgt from among the at least one object OBJ selected based on matching similarity, which has the highest priority or is equal to or greater than a priority threshold.

[0285] As a second example of selecting the target object OBJ_tgt based on both matching similarity and priority, the position and orientation calculation unit 311 may select at least one object OBJ from among a plurality of object OBJs based on priority, and then select one object OBJ as the target object OBJ_tgt from among the at least one object OBJ selected based on priority, based on matching similarity. For example, the position and orientation calculation unit 311 may select at least one object OBJ from among a plurality of object OBJs whose priority is equal to or greater than the priority threshold, and then select one object OBJ as the target object OBJ_tgt from among the at least one object OBJ selected based on priority, which has the highest matching similarity or is equal to or greater than the matching determination threshold.

[0286] As a third example of selecting the target object OBJ_tgt based on both matching similarity and priority, the position and orientation calculation unit 311 may perform a predetermined calculation using the matching similarity and priority, and select the target object OBJ_tgt based on the result of the predetermined calculation. The position and orientation calculation unit 311 may select the target object OBJ_tgt from among a plurality of object OBJs, where the result of the predetermined calculation indicates that it should be selected as the target object OBJ_tgt. For example, the position and orientation calculation unit 311 may select the object OBJ that yields the highest calculated value as a result of the predetermined calculation as the target object OBJ_tgt. For example, the position and orientation calculation unit 311 may select the object OBJ that yields a calculated value equal to or greater than a predetermined calculation threshold as the target object OBJ_tgt.

[0287] For example, if the calculated value obtained as a result of a predetermined calculation performed on the first object OBJ#1 and the second object OBJ#2 indicates that the first object OBJ#1 should be selected as the target object OBJ_tgt rather than the second object OBJ#2, the position and orientation calculation unit 311 may select the first object OBJ#1 as the target object OBJ_tgt. On the other hand, if the calculated value obtained as a result of a predetermined calculation performed on the first object OBJ#1 and the second object OBJ#2 indicates that the second object OBJ#2 should be selected as the target object OBJ_tgt rather than the first object OBJ#1, the position and orientation calculation unit 311 may select the second object OBJ#2 as the target object OBJ_tgt.

[0288] One example of a predetermined calculation is a first calculation that multiplies the matching similarity by the priority (i.e., a first calculation expressed by the formula "matching similarity × priority"). Another example of a predetermined calculation is a second calculation that calculates the simple average of the matching similarity and the priority (i.e., a second calculation expressed by the formula "(matching similarity + priority) / 2"). In this case, the contribution rate of the matching similarity to the selection of the target object OBJ_tgt and the contribution rate of the priority to the selection of the target object OBJ_tgt will be approximately the same.

[0289] One example of a predetermined operation is a third operation that calculates a weighted average of matching similarity and priority (that is, a third operation expressed by the formula "α × matching similarity + (1 - α) × priority (where α is a weight representing a real number greater than 0 and less than 1)"). In this case, the contribution rate of matching similarity to the selection of target object OBJ_tgt and the contribution rate of priority to the selection of target object OBJ_tgt can be adjusted by the weight α. For example, if the contribution rate of matching similarity to the selection of target object OBJ_tgt is given more weight than the contribution rate of priority to the selection of target object OBJ_tgt, the weight α may be set to a relatively large value (for example, a value greater than 0.5). For example, if the contribution rate of priority to the selection of target object OBJ_tgt is given more weight than the contribution rate of matching similarity to the selection of target object OBJ_tgt, the weight α may be set to a relatively small value (for example, a value less than 0.5).

[0290] The weight α may be set in advance before the robot control process starts. The weight α may be set as appropriate after the robot control process starts. The weight α may be set by the position and orientation calculation unit 311. The weight α may be set by the operator of the robot system SYS. The weight α may be set for each type of object OBJ. For example, when robot control processing is performed to control the robot 1 etc. to perform a predetermined process on a first type of object OBJ, the weight α may be set to a first value corresponding to the first type of object OBJ. For example, when robot control processing is performed to control the robot 1 etc. to perform a predetermined process on a second type of object OBJ that is different from the first type of object OBJ, the weight α may be set to a second value (for example, a value different from the first value) corresponding to the second type of object OBJ. As a result, the contribution rate of the matching similarity to the selection of the target object OBJ_tgt and the contribution rate of the priority to the selection of the target object OBJ_tgt can be adjusted for each type of object OBJ. For example, the contribution rates of matching similarity and priority can be adjusted to enable high-precision detection of object OBJs according to the shapes of the first type of object OBJ and the second type of object OJT, respectively.

[0291] (4-4) Other Modifications (4-4-1) Modification of the process for calculating at least one of the position and orientation of the object OBJ In the above description, the position and orientation calculation unit 311 generates position and orientation data POI1 indicating at least one of the position and orientation of the object OBJ by performing a rough search process in step S223 of Figure 12, and then generates position and orientation data POI2 indicating at least one of the position and orientation of the object OBJ as a robot control signal used to generate a robot control signal in step S23 of Figure 11 by performing a fine search process. However, the position and orientation calculation unit 311 generates position and orientation data POI1 as a robot control signal used to generate a robot control signal in step S23 of Figure 11 by performing a rough search process in step S223 of Figure 12. In this case, the position and orientation calculation unit 311 does not need to perform a fine search process in step S224 of Figure 12.

[0292] In the above description, the position and orientation calculation unit 311 generates position and orientation data POI1 in step S223 of Figure 12 by performing a 3D matching process using the target data portion WSD_ROI and the template model. However, the position and orientation calculation unit 311 may also generate position and orientation data POI1 in step S223 of Figure 12 by performing a matching process (for example, at least one of contour matching and 2D matching) using at least a part of the image data IMD_2D and the template model. In this case, the position and orientation calculation unit 311 may identify a data portion of the image data IMD_2D corresponding to one object region WA as part of the image data IMD_2D used for matching, similar to when selecting a data portion of the three-dimensional position data WSD corresponding to one object region WA as the target data portion WSD_ROI.

[0293] Similarly, in the above description, the position and orientation calculation unit 311 generates position and orientation data POI2 in step S224 of Figure 12 by performing a 3D matching process using the three-dimensional position data WSD and the template model. However, the position and orientation calculation unit 311 may also generate position and orientation data POI2 in step S224 of Figure 12 by performing a matching process using at least a part of the image data IMD_2D and the template model.

[0294] (4-4-2) Modification of data input to object recognition model 50 (4-4-2-1) First modification of data input to object recognition model 50 In the above description, the arithmetic unit 51 inputs the image data IMD_2D to the object recognition model 50 in step S221 of Figure 12, thereby obtaining object region information OAI relating to the object region WA in which object OBJ exists within the 2D image IMG_2D shown by the image data IMD_2D. However, in step S221 of Figure 12, the arithmetic unit 51 may input the image data IMD_3D to the object recognition model 50 in addition to or instead of the image data IMD_2D, thereby obtaining object region information OAI relating to the object region WA in which object OBJ exists within the image shown by the image data IMD_3D.

[0295] The image data IMD_3D input to the object recognition model 50 may be the same as at least one of the image data IMD_3D used to generate the three-dimensional position data WSD used to generate the position and orientation data POI1 in step S223 of Figure 12, and the image data IMD_3D used to generate the three-dimensional position data WSD used to generate the position and orientation data POI2 in step S224 of Figure 12. Alternatively, the image data IMD_3D input to the object recognition model 50 may be different from the image data IMD_3D used to generate the three-dimensional position data WSD used to generate at least one of the position and orientation data POI1 and POI2.

[0296] When the image data IMD_3D is input to the object recognition model 50, the training image data 562 used for machine learning of the object recognition model 50 may include image data (so-called stereo image data) that includes two images in which at least two of the object OBJ and sample object OBJ_s are captured. As a result, using such training image data 562, when image data IMD_3D is input, an object recognition model 50 is generated that can output at least one of the following: (i) object region information OAI relating to the object region WA in which an object OBJ exists within the image shown by the input image data IMD_3D; (ii) information relating to the matching similarity of an object OBJ captured in the image shown by the input image data IMD_3D; (iii) position and orientation information relating to at least one of the position and orientation of an object OBJ captured in the image shown by the input image data IMD_3D; and (iv) information relating to the priority of an object OBJ captured in the image shown by the input image data IMD_3D.

[0297] (4-4-2-2) The second modification calculation device 51 for data input to the object recognition model 50 may, in step S221 of Figure 12, input the three-dimensional position data WSD generated in step S222 of Figure 12 to the object recognition model 50 in addition to or instead of at least one of the image data IMD_2D and IMD_3D, thereby obtaining object region information OAI relating to the object region WA in which the object OBJ exists in the space corresponding to the three-dimensional position data WSD. However, when the three-dimensional position data WSD is input to the object recognition model 50, the object region WA may indicate the space in which the object OBJ or sample object OBJ_s exists, as the region in which the object OBJ or sample object OBJ_s exists. For this reason, the object region WA may be called the object space.

[0298] The three-dimensional position data WSD input to the object recognition model 50 may be the same as at least one of the three-dimensional position data WSD used to generate position and orientation data POI1 in step S223 of Figure 12 and the three-dimensional position data WSD used to generate position and orientation data POI2 in step S224 of Figure 12. In other words, the image data IMD_3D used to generate the three-dimensional position data WSD input to the object recognition model 50 may be the same as the image data IMD_3D used to generate the three-dimensional position data WSD for generating at least one of the position and orientation data POI1 and POI2. Alternatively, the three-dimensional position data WSD input to the object recognition model 50 may be different from the three-dimensional position data WSD used to generate at least one of the position and orientation data POI1 and POI2. In other words, the image data IMD_3D used to generate the three-dimensional position data WSD input to the object recognition model 50 may be different from the image data IMD_3D used to generate the three-dimensional position data WSD for generating at least one of the position and orientation data POI1 and POI2.

[0299] When three-dimensional position data WSD is input to the object recognition model 50, the training dataset 560 used for machine learning of the object recognition model 50 may include, in addition to or instead of the training data 561 which includes the training image data 562 described above and at least one of the correct labels 563 and 565 to 567, training data 568 which includes training three-dimensional position data 5690 and at least one of the correct labels 5691 to 5694.

[0300] The learned three-dimensional position data 5690 may be three-dimensional position data (for example, at least one of point cloud data and depth image data) that indicates the three-dimensional position of an object OBJ that the robot 1 is scheduled to actually process using the end effector 4. For example, the learned three-dimensional position data 5690 may indicate the three-dimensional position of a single object OBJ. For example, the learned three-dimensional position data 5690 may indicate the three-dimensional positions of multiple object OBJs. In other words, the learned three-dimensional position data 5690 may indicate the three-dimensional position of at least one object OBJ.

[0301] The training three-dimensional position data 5690 may be three-dimensional position data that shows the three-dimensional position of sample object OBJ_s that simulates the object OBJ, in addition to or instead of the object OBJ. For example, the training three-dimensional position data 5690 may show the three-dimensional position of a single sample object OBJ_s. For example, the training three-dimensional position data 5690 may show the three-dimensional positions of multiple sample object OBJ_s. In other words, the training three-dimensional position data 5690 may show the three-dimensional position of at least one sample object OBJ_s.

[0302] If the training dataset 560 includes multiple training data 568, each training data 568 may include multiple training three-dimensional position data 5690 that are different from each other. Each training three-dimensional position data 5690 may include at least two training three-dimensional position data 5690 that have different numbers of object OBJs indicating three-dimensional positions. Each training three-dimensional position data 5690 may include at least two training three-dimensional position data 5690 that have different numbers of sample object OBJ_s. Each training three-dimensional position data 5690 may include at least two training three-dimensional position data 5690 that have different positions and orientations of object OBJs. Each training three-dimensional position data 5690 may include at least two training three-dimensional position data 5690 that have different positions and orientations of sample object OBJ_s. The multiple training three-dimensional position data 5690 may include at least two training three-dimensional position data 5690 in which the positional relationship between an object OBJ and another object different from the object OBJ is different from that object OBJ. The multiple training three-dimensional position data 5690 may include at least two training three-dimensional position data 5690 in which the positional relationship between a sample object OBJ_s and another object different from the sample object OBJ_s is different from that object OBJ_s.

[0303] The correct label 5691, like the correct label 563 described above, indicates information about the object region WA. Specifically, the correct label 5691 indicates information about the object region WA in which the object OBJ or sample object OBJ_s exists within the space corresponding to the learned three-dimensional position data 5690 associated with the correct label 5691. Note that other features of the correct label 5691 may be the same as other features of the correct label 563 described above. For this reason, in order to avoid redundant explanations, a detailed explanation of the correct label 5691 is omitted.

[0304] The ground truth label 5692, like the ground truth label 565 described above, indicates information regarding matching similarity. Specifically, the ground truth label 5692 indicates information regarding the matching similarity between the object OBJ or sample object OBJ_s whose three-dimensional position is indicated by the training three-dimensional position data 5690 associated with the ground truth label 5692 and the template model. Note that other features of the ground truth label 5692 may be the same as other features of the ground truth label 565 described above. For this reason, in order to avoid redundant explanations, a detailed explanation of the ground truth label 5692 is omitted.

[0305] The correct label 5693, like the correct label 566 described above, indicates position and orientation information. Specifically, the correct label 5693 indicates position and orientation information relating to at least one of the position and orientation of the object OBJ or sample object OBJ_s whose three-dimensional position is indicated by the learned three-dimensional position data 5690 associated with the correct label 5693. Note that other features of the correct label 5693 may be the same as other features of the correct label 566 described above. For this reason, in order to avoid redundant explanations, a detailed explanation of the correct label 5693 is omitted.

[0306] The correct label 5694, like the correct label 567 described above, indicates information regarding priority. Specifically, the correct label 5694 indicates positional orientation information regarding the priority of the object OBJ or sample object OBJ_s whose three-dimensional position is indicated by the learned three-dimensional position data 5690 associated with the correct label 5694. Note that other features of the correct label 5694 may be the same as other features of the correct label 567 described above. For this reason, in order to avoid redundant explanations, a detailed explanation of the correct label 5694 is omitted.

[0307] The arithmetic unit 51 may generate training data 568 containing the created training three-dimensional position data 5690 and at least one of the correct labels 5691 to 5694 by creating training three-dimensional position data 5690 and at least one of the correct labels 5691 to 5694. The arithmetic unit 51 may generate multiple training data 568 by repeating the process of generating training data 568.

[0308] The computing unit 51 may generate training data 568 by performing a simulation in which the object model OM is placed in a virtual simulation space SIM. The process of generating training data 568 by performing a simulation in which the object model OM is placed in the simulation space SIM may be the same as the process of generating training data 561 by performing a simulation in which the object model OM is placed in the simulation space SIM. Therefore, in order to avoid redundant explanations, the detailed setup of the process of generating training data 568 by performing a simulation in which the object model OM is placed in the simulation space SIM will be omitted. However, a brief overview will be given below.

[0309] The computing unit 51 may perform a simulation in which it images the simulation space SIM, in which the object model OM is placed, using a virtual imaging device 22 (hereinafter referred to as imaging device 22v). As a result, the computing unit 51 may generate image data (stereo image data) that is virtually generated by imaging the simulation space SIM using the imaging device 22v. Subsequently, the computing unit 51 may generate training three-dimensional position data 5690 from the virtually generated image data (stereo image data). The computing unit 51 may generate three-dimensional position data WSD, which is used as training three-dimensional position data 5690, from image data IMD_3D generated by the imaging device 22 imaging the real space in which at least one of the object OBJ and sample object OBJ_s is actually placed. Furthermore, the computing unit 51 may generate at least one of the correct labels 5691 to 5694 in parallel with or before or after the generation of the training three-dimensional position data 5690.

[0310] The computing unit 51 may change the arrangement of the object model OM in the simulation space SIM each time it generates training data 568. The computing unit 51 may change the position of the imaging device 22v in the simulation space SIM each time it generates training data 568. The computing unit 51 may change the position of the imaging device 22v in the simulation space SIM each time it generates training data 568. As a result, the computing unit 51 can easily increase the number of training three-dimensional position data 5690 (and furthermore, the number of training data 568 containing training three-dimensional position data 5690).

[0311] It is expected that the accuracy of the information output by the object recognition model 50 will increase as the amount of data input to the object recognition model 50 increases. On the other hand, the less data input to the object recognition model 50, the shorter the time required for machine learning of the object recognition model 50. For this reason, the type of data input to the object recognition model 50 may be set appropriately, taking into consideration both the effect of improving the accuracy of the information output by the object recognition model 50 and the effect of shortening the time required for machine learning of the object recognition model 50. For example, if the effect of improving the accuracy of the information output by the object recognition model 50 is prioritized over the effect of shortening the time required for machine learning of the object recognition model 50, the object recognition model 50 may be input with at least one of the image data IMD_3D and three-dimensional position data WSD in addition to the image data IMD_2D. As another example, if the effect of reducing the time required for machine learning of the object recognition model 50 is prioritized over the effect of improving the accuracy of the information output by the object recognition model 50, then the object recognition model 50 may be input with image data IMD_2D, but may not be input with image data IMD_3D and three-dimensional position data WSD.

[0312] Furthermore, if the object recognition model 50 outputs positional orientation information relating to at least one of the position and orientation of the object OBJ, the object recognition model 50 may be input with at least one of the image data IMD_3D and the three-dimensional position data WSD. In this case, because at least one of the image data IMD_3D and the three-dimensional position data WSD includes depth information of the object OBJ, the accuracy of the positional orientation information output by the object recognition model 50 that is input with at least one of the image data IMD_3D and the three-dimensional position data WSD will be higher than the accuracy of the positional orientation information output by the object recognition model 50 that is not input with at least one of the image data IMD_3D and the three-dimensional position data WSD.

[0313] (4-4-3) Modification of the method for identifying the object region WA In the above description, in step S221 of Figure 12, the position and orientation calculation unit 311 uses the object recognition model 50 to identify the object region WA in which the object OBJ exists. However, the position and orientation calculation unit 311 may identify the object region WA without using the object recognition model 50. For example, in step S221 of Figure 12, the position and orientation calculation unit 311 may perform a matching process using the image data IMD_2D (for example, at least one of contour matching and 2D matching) to detect the object OBJ reflected in the 2D image IMG_2D shown by the image data IMD_2D, and identify the region containing the detected object OBJ as the object region WA.

[0314] (4-4-4) Modification of the Imaging System 2 In the above description, the imaging system 2 is equipped with an imaging device 22 and an illumination device 23 in order to generate the image data IMD_3D. However, the imaging system 2 does not need to be equipped with an illumination device 23 in order to generate the image data IMD_3D. This is because, as described above, the imaging device 22 is a stereo camera, and therefore, three-dimensional position data WSD, which indicates the three-dimensional position of multiple points of the object OBJ, can be generated from the image data IMD_3D, which shows the two images generated by the two image sensors of the stereo camera.

[0315] The imaging device 22 does not have to be a stereo camera. For example, the imaging device 22 may be a monocular camera that uses a single image sensor to image the object OBJ. Even in this case, the image data IMD_3D shows the object OBJ onto which the projection pattern from the illumination device 23 is projected. In this case, the shape of the projection pattern shown in the image data IMD_3D reflects the three-dimensional shape of the object OBJ onto which the projection pattern is projected. Therefore, even if the imaging device 22 is not a stereo camera, the position and orientation calculation unit 311 can generate three-dimensional position data WSD using a well-known process based on the projection pattern shown in the image data IMD_3D.

[0316] The imaging system 2 includes an imaging device 21 which is a monocular camera, but does not necessarily include an imaging device 22 which is a stereo camera. In this case, image data generated by the imaging device 21 imaging the object OBJ during the period when the illumination device 23 is not projecting a desired projection pattern onto the object OBJ may be used as image data IMD_2D. On the other hand, image data generated by the imaging device 21 imaging the object OBJ during the period when the illumination device 23 is projecting a desired projection pattern onto the object OBJ may be used as image data IMD_3D.

[0317] The imaging system 2 includes an imaging device 22 which is a stereo camera, but does not necessarily include an imaging device 21. In this case, the image data generated by either of the two monocular cameras of the imaging device 22 capturing an object OBJ may be used as the image data IMD_2D. Alternatively, the image data showing two images generated by both of the two monocular cameras of the imaging device 22 capturing an object OBJ may be used as the image data IMD_3D.

[0318] The imaging system 2 may include an imaging device 21 that includes cameras other than monocular cameras and stereo cameras. For example, the imaging system 2 may include an imaging device 21 that includes three or more monocular cameras. For example, the imaging system 2 may include an imaging device 21 that includes at least one of a light field camera, a plenoptic camera, and a multispectral camera.

[0319] The imaging system 2 may include an imaging device 22 that includes cameras other than monocular cameras and stereo cameras. For example, the imaging system 2 may include an imaging device 22 that includes three or more monocular cameras. For example, the imaging system 2 may include an imaging device 22 that includes at least one of a light field camera, a plenooptic camera, and a multispectral camera.

[0320] If the imaging device 22 includes a light field camera, the imaging system 2 may include the imaging device 22, which is a light field camera, but may not include the imaging device 21. The light field camera included in the imaging device 22 may generate multiple images obtained by imaging the object OBJ from multiple different viewpoints each time the object OBJ is imaged. In this case, image data showing one of the multiple images may be used as image data IMD_2D. Image data showing at least two of the multiple images may be used as image data IMD_3D.

[0321] In the above description, the imaging system 2 is provided on the robot arm 12. However, the imaging system 2 does not have to be attached to the robot arm 12. For example, the imaging system 2 may be provided on a robot that does not have a robot arm 12. For example, the imaging system 2 may be positioned at any position from which it can image an object OBJ. Also, for example, the imaging system 2 may be positioned at any position from which it can illuminate the object OBJ with illumination light. For example, the imaging system 2 may be attached to a structure such as a column. As an example, the imaging system 2 may be provided on a support device capable of suspending the imaging system 2 from above the object OBJ. The support device may include a plurality of leg members extending upward from a support surface S, and a beam member connecting the plurality of leg members via the upper ends or vicinity thereof. The robot 1 may be able to move the support device. As another example, if the robot 1 is installed on a robotic movable device as described above, the imaging system 2 may be installed on the robotic movable device on which the robot 1 is installed. Furthermore, at least one of the imaging device 21, imaging device 22, and illumination device 23 may be attached to the robot arm 12, and at least one of the imaging device 21, imaging device 22, and illumination device 23 may be attached to a location other than the robot arm 12 (for example, a structure such as a column or a robotic movable device). For example, the imaging device 21 may be attached to a location other than the robot arm 12, and the imaging device 22 may be attached to the robot arm 12.

[0322] The support device (or any structure, hereafter the same in this paragraph) to which the imaging system 2 is attached may be fixed. In other words, the imaging system 2 may be attached to a fixed support device. In this case, the imaging system 2 may image the object OBJ while stationary. As a result, the imaging system 2 can image the object OBJ more accurately compared to when the imaging system 2 is moving and images the object OBJ. Furthermore, when the imaging system 2 is attached to a fixed support device, the position of the imaging system 2 in the reference coordinate system is fixed, making it easier to anticipate occlusion. In other words, it becomes easier to anticipate the positional relationship between the imaging system 2 and obstacles that create blind spots for the imaging system 2. As a result, the imaging system 2 can appropriately image the object OBJ at a time when occlusion does not occur.

[0323] (4-4-5) Alternative Means for Imaging System 2 As shown in Figure 22, the robot system SYS may include, in addition to or instead of, the imaging system 2 capable of imaging an object OBJ, a measurement system 6 capable of measuring at least one of the distance to the object OBJ and its three-dimensional shape. In this case, the control device 3 may use the measurement results of the measurement system 6 as the imaging results of the imaging system 2 (i.e., image data IMD). In this case, the control device 3 may generate three-dimensional position data WSD based on the measurement results of the measurement system 6. Alternatively, the control device 3 may use the measurement results of the measurement system 6 as three-dimensional position data WSD. Thereafter, similar to the case where three-dimensional position data WSD is generated from image data IMD generated by the imaging system 2, the control device 3 (in particular, the position and orientation calculation unit 311) may perform 3D matching processing using at least a part of the three-dimensional position data WSD. In other words, the measurement results of the measurement system 6 may be used in the same way as the imaging results (image data IMD) of the imaging system 2. The measurement system 6 may be used in the same way as the imaging system 2.

[0324] The measurement system 6 may be an existing system capable of optically measuring at least one of the distance to the object OBJ and its three-dimensional shape. For example, the measurement system 6 may be an existing system capable of non-contact measuring at least one of the distance to the object OBJ and its three-dimensional shape. For example, the measurement system 6 may be capable of optically measuring at least one of the distance to the object OBJ and its three-dimensional shape by irradiating the object OBJ with measurement electromagnetic waves and detecting the electromagnetic waves returning from the irradiated object OBJ towards the measurement system 6. The measurement electromagnetic waves may be light (in this case, the light irradiated onto the object OBJ may be called measurement light) or radio waves (in this case, the radio waves irradiated onto the object OBJ may be called measurement waves). Examples of such measurement systems 6 include at least one of the following: a Time Of Flight (TOF) interferometer, a Frequency Modulated Continuous Wave (FMCW) interferometer, an optical comb interferometer that uses light containing frequency components arranged at equal intervals on the frequency axis (optical frequency comb) as measurement light, a triangulation measuring instrument, and a millimeter-wave radar that uses millimeter waves as measurement waves. At least one of the TOF measuring instrument and the FMCW measuring instrument may be referred to as a Light Detection and Ranging (LiDAR). The control device 3 (in particular, the position and attitude calculation unit 311) may generate three-dimensional position data WSD, which corresponds to point cloud data, based on the measurement results of the object OBJ by such measurement systems 6. Subsequently, the control device 3 (particularly the position and orientation calculation unit 311) may perform 3D matching processing using the three-dimensional position data WSD, similar to the case where three-dimensional position data WSD is generated from the image data IMD generated by the imaging system 2. The measurement system 6 may also be referred to as a measurement device or a measurement sensor.

[0325] Furthermore, the imaging system 2 also images the object OBJ by detecting the reflected light from the object OBJ using the image sensors provided in each of the imaging devices 21. In other words, both the imaging system 2 and the measurement system 6 can be said to be detecting the reflected light from the object OBJ. For this reason, the imaging system 2 and the measurement system 6 may be collectively referred to as the light detection system. The light detection system may also be a system that includes at least one of the imaging system 2 and the measurement system 6.

[0326] (4-4-6) Modified Version of Robot 1 In the above description, the end effector 4 is attached to the robot arm 12 (for example, the tip of the robot arm 12). However, as shown in Figure 23, a side view showing the appearance of a modified version of robot 1, robot 1 may be equipped with a moving device 14, the moving device 14 may be attached to the robot arm 12 (for example, the tip of the robot arm 12), and the end effector 4 may be attached to the moving device 14.

[0327] The moving device 14 is equipped with an actuator (in other words, a motor), and the power output by the actuator can be used to move the end effector 4 attached to the moving device 14. The moving device 14 may also be able to move the end effector 4 relative to the robot arm 12 to which the moving device 14 is attached. In other words, the moving device 14 may be able to move the end effector 4 so as to change the positional relationship between the robot arm 12 and the end effector 4. The moving device 14 may also be able to move the end effector 4 relative to an object OBJ on which the robot 1 performs a predetermined process. In other words, the moving device 14 may be able to move the end effector 4 so as to change the positional relationship between the object OBJ and the end effector 4.

[0328] The moving device 14 may be capable of moving the end effector 4 along a predetermined translation axis (i.e., a linear axis extending in the linear direction). For example, the moving device 14 may be capable of moving the end effector 4 along a first translation axis. For example, in addition to or instead of moving the end effector 4 along the first translation axis, the moving device 14 may be capable of moving the end effector 4 along a second translation axis that intersects (typically perpendicular to) the first translation axis and is different from the first translation axis. For example, in addition to or instead of moving the end effector 4 along at least one of the first and second translation axes, the moving device 14 may be capable of moving the end effector 4 along a third translation axis that intersects (typically perpendicular to) the first and second translation axes and is different from the first and second translation axes. Furthermore, since the translation axis is an axis along the direction in which the end effector 4 moves, it may also be called a moving axis.

[0329] The moving device 14 may be capable of rotating (i.e., rotationally moving) the end effector 4 around a predetermined axis of rotation. For example, the moving device 14 may be capable of rotating the end effector 4 around a first axis of rotation. For example, in addition to or instead of rotating the end effector 4 around a first axis of rotation, the moving device 14 may be capable of rotating the end effector 4 around a second axis of rotation that intersects (typically orthogonal to) the first axis of rotation and is different from the first axis of rotation. For example, in addition to or instead of rotating the end effector 4 around at least one of the first and second axes of rotation, the moving device 14 may be capable of rotating the end effector 4 around a third axis of rotation that intersects (typically orthogonal to) the first and second axes of rotation and is different from the first and second axes of rotation.

[0330] In this case, if the robot 1 is equipped with a moving device 14, the control device 3 may move the end effector 4 by controlling the moving device 14 in addition to or instead of controlling the robot arm 12. In other words, the control device 3 may move the end effector 4 by controlling the moving device 14 in addition to or instead of moving the robot arm 12. In this case, the robot control process may include a process for generating a robot control signal for controlling the moving device 14 equipped with the robot 1. For example, the robot control process may include a process for generating a robot control signal for controlling the moving device 14 to move the end effector 4.

[0331] (5) Additional Notes Regarding the embodiments described above, the following additional notes are disclosed.

[0332] [Appendix A1] A robot equipped with a processing device for processing at least one object and an imaging system capable of generating image data by imaging the at least one object, and a control device for generating a control signal for controlling at least one of the processing devices, wherein the control device comprises a calculation device for generating the control signal and an output device for outputting the control signal generated by the calculation device, the calculation device identifies an object region in which the at least one object exists using at least one of first image data which is the image data generated by the imaging system imaging the at least one object, and first three-dimensional position data generated from second image data which is the image data generated by the imaging system imaging the at least one object, and generates first position and orientation information indicating at least one of the position and orientation of the target object using the result of identifying the object region and second three-dimensional position data indicating the three-dimensional position generated from third image data which is the image data generated by the imaging system imaging the at least one object, A control device that generates second position and orientation information indicating at least one of the position and orientation of the target object using the first position and orientation information and third three-dimensional position data indicating the three-dimensional position generated from the fourth image data which is image data generated when the imaging system images the at least one object, and generates the control signal for executing the processing on the target object based on the second position and orientation information.

[0333] [Appendix A2] The control device according to Appendix A1, wherein the object region corresponds to the region in the first image data and the first three-dimensional position data in which the target object exists.

[0334] [Note A3] The control device according to Note A1 or A2, wherein the object region is set based on the shape of the edge of the target object.

[0335] [Appendix A4] The control device according to any one of Appendix A1 to A3, wherein the calculation device identifies a first data portion of the second three-dimensional position data corresponding to the object region and generates the first position and orientation information using the first data.

[0336] [Appendix A5] The control device according to Appendix A4, wherein the calculation device generates the first position and orientation information by performing a matching process using the first data portion and the three-dimensional model of the object.

[0337] [Appendix A6] The control device according to any one of Appendix A1 to A5, wherein the calculation device generates the second position and orientation information by performing a matching process using a three-dimensional model of the object positioned in accordance with at least one of the position and orientation determined based on the first position and orientation information and the third three-dimensional position data.

[0338] [Appendix A7] The control device according to any one of Appendix A1 to A6, wherein the computing device identifies the object region using at least one of the first image data and the first three-dimensional position data and a learning model generated by machine learning.

[0339] [Appendix A8] The control device according to Appendix A7, which is generated by machine learning using first training data, in which training image data showing a first sample object that simulates the object and a first training image showing at least one of the object is associated with a first ground truth label indicating information about the region in which the first sample object and at least one of the object exist.

[0340] [Appendix A9] The control device described in Appendix A8, which generates the first learning data by performing a simulation in which the first sample object and a first object model representing at least one of the objects are placed in a virtual simulation space.

[0341] [Appendix A10] The control device according to Appendix A9, wherein at least one of the number, position, orientation, and positional relationship between the first object models and other objects is automatically set.

[0342] [Appendix A11] The control device according to Appendix A9 or A10, wherein the training image data includes image data virtually generated by imaging the simulation space in which the first object model is located with a virtual imaging device, and the first correct label includes information relating to the region in which the first object model exists within the simulation space.

[0343] [Appendix A12] The control device according to Appendix A11, wherein the learning model is generated by machine learning using a first learning dataset which includes a plurality of first learning data, and the plurality of first learning data include: first learning image data which is learning image data virtually generated by imaging the simulation space in which the first object model is arra...

Claims

1. A robot equipped with a processing device that processes at least one of a plurality of objects and an imaging system capable of generating image data by imaging at least one of the plurality of objects, and a control device that generates a control signal for controlling at least one of the processing devices, wherein the control device comprises a calculation device that generates the control signal and an output device that outputs the control signal generated by the calculation device, the calculation device identifies an object region in which an object to be processed by the processing device exists, using at least one of first image data which is the image data generated by the imaging system imaging at least one of the plurality of objects, and first three-dimensional position data generated from second image data which is the image data generated by the imaging system imaging at least one of the plurality of objects, and generates first position and orientation information indicating at least one of the position and orientation of the object using the result of identifying the object region and second three-dimensional position data indicating the three-dimensional position generated from third image data which is the image data generated by the imaging system imaging at least one of the plurality of objects, A control device that generates second position and orientation information indicating at least one of the position and orientation of the target object using the first position and orientation information and third three-dimensional position data indicating the three-dimensional position generated from the fourth image data which is image data generated when the imaging system captures at least one of the plurality of objects, and generates the control signal for executing the processing on the target object based on the second position and orientation information.

2. The control device according to claim 1, wherein the object region corresponds to the region in the first image data and the first three-dimensional position data in which the target object exists.

3. The control device according to claim 1 or 2, wherein the object region is set based on the shape of the edge of the target object.

4. The control device according to any one of claims 1 to 3, wherein the calculation device identifies a first data portion of the second three-dimensional position data that corresponds to the object region, and generates the first position and orientation information using the first data.

5. The control device according to claim 4, wherein the calculation device generates first position and orientation information by performing a matching process using the first data portion and the three-dimensional model of the object.

6. The control device according to claim 5, wherein the calculation device interpolates the first data portion using a three-dimensional model of the object, and generates at least one of the first position and orientation information and the second position and orientation information using the interpolated first data and the third three-dimensional position data.

7. The control device according to any one of claims 1 to 6, wherein the calculation device generates the second position and orientation information by performing a matching process using a three-dimensional model of the object positioned to match at least one of the position and orientation determined based on the first position and orientation information, and the third three-dimensional position data.

8. The control device according to any one of claims 1 to 7, wherein the computing device identifies the object region using at least one of the first image data and the first three-dimensional position data and a learning model generated by machine learning.

9. The control device according to claim 8, wherein the learning model is generated by machine learning using first learning data, which is associated with first sample object that simulates the object and first learning image showing at least one of the object, and first ground truth label indicating information about the region in which the first sample object and at least one of the object exist.

10. The control device according to claim 9, wherein the first training data is generated by performing a simulation in which the first sample object and a first object model representing at least one of the objects are placed in a virtual simulation space.

11. The control device according to claim 10, wherein at least one of the number, position, orientation, and positional relationship between the first object models and other objects is automatically set.

12. The control device according to claim 10 or 11, wherein the training image data includes image data virtually generated by capturing the simulation space in which the first object model is located with a virtual imaging device, and the first correct label includes information relating to the region in which the first object model exists within the simulation space.

13. The control device according to any one of claims 8 to 12, wherein the learning model is generated by machine learning using second learning data, which is associated with second learning data, in which a second sample object that simulates the object and a second learning image showing at least one of the object are second sample objects and a second ground truth label indicating information about the region in which the second sample object and at least one of the object exist.

14. The control device according to claim 13, wherein the second training data is generated by performing a simulation in which the second sample object and a second object model representing at least one of the objects are placed in a virtual simulation space.

15. The control device according to claim 14, wherein at least one of the number, position, orientation, and positional relationship between the second object model and other objects is automatically set.

16. The control device according to claim 14 or 15, wherein the learned three-dimensional position data includes three-dimensional position data generated from image data virtually generated by imaging the simulation space in which the second object model is placed with a virtual imaging device, and the second correct label includes information relating to the region in which the second object model exists within the simulation space.

17. The control device according to any one of claims 8 to 16, wherein the computing device generates third position and orientation information indicating at least one of the position and orientation of the target object using the learning model and at least one of the first image data and the first three-dimensional position data, and generates second position and orientation information using the third position and orientation information and the third three-dimensional position data.

18. The control device according to claim 17, wherein the calculation device generates the second position and orientation information by performing a matching process using a three-dimensional model of the object positioned in accordance with at least one of the position determined based on the third position and orientation information and the orientation determined based on the third position and orientation information, and the third three-dimensional position data.

19. The control device according to claim 17 or 18, wherein the computing device identifies a first data portion of the second three-dimensional position data that corresponds to the object region, interpolates the first data portion using the third position and orientation information and the three-dimensional model of the object, and generates the second position and orientation information using the interpolated first data portion.

20. The control device according to any one of claims 1 to 19, wherein the computing device selects the target object on which to perform the processing from among the plurality of objects, and generates first position and orientation information using the result of identifying the object region and the second three-dimensional position data.

21. The control device according to claim 20, wherein the arithmetic unit generates priority levels for a first object and a second object included in the plurality of objects, and the selection is such that, based on the priority levels, the first object is selected as the target object on which processing should be performed.

22. The control device according to claim 21, wherein the imaging system images the first object and the second object, each of which is the object; the computing device determines the priority of the first object and the priority of the second object based on the imaging results of the imaging system; the processing device selects the first object as the target object on which the processing device should perform the processing based on the priority; and generates first position and orientation information indicating at least one of the position and orientation of the selected first object using the identification result of the object region in which the selected first object exists and the second three-dimensional position data.

23. The control device according to claim 22, wherein the calculation device generates first position and orientation information indicating the position and orientation of the first object using the result of identifying the object region in which the first object exists and the second three-dimensional position data, when the priority of the first object is higher than the priority of the second object.

24. The control device according to any one of claims 20 to 22, wherein the calculation device calculates the priority based on at least one of the first image data and the first three-dimensional position data.

25. The control device according to any one of claims 21 to 24, wherein the priority is determined based on the degree of overlap between the first object and the second object.

26. The control device according to claim 25, wherein the priority of the first object increases as the degree of overlap between the first object and other objects different from the first object decreases, and the priority of the second object increases as the degree of overlap between the second object and other objects different from the second object decreases.

27. The control device according to any one of claims 21 to 26, wherein the arithmetic unit calculates the priority using at least one of the first image data and the first three-dimensional position data and a learning model generated by machine learning.

28. The control device according to claim 27, wherein the arithmetic unit generates the first position and attitude information based on the priority and the likelihood of the priority calculated using the learning model.

29. The control device according to any one of claims 21 to 28, wherein the computing device generates first position and orientation information based on the similarity between the target object and the three-dimensional model of the target object, indicated by at least one of the first image data and the first three-dimensional position data, and the priority.

30. The control device according to claim 29, wherein the imaging system images the first object and the second object, each of which is the object; the computing device performs a first calculation based on the imaging results of the imaging system using the similarity of the first object and the priority of the first target object; performs a second calculation using the similarity of the second object and the priority of the second object; selects the first object as the target object on which the processing device will perform the processing based on the results of the first and second calculations; and generates first position and orientation information indicating at least one of the position and orientation of the selected first object using the identification result of the object region in which the selected first object exists and the second three-dimensional position data.

31. The control device according to claim 29 or 30, wherein the arithmetic device calculates the similarity using a learning model generated by machine learning to output the similarity when at least one of the first image data and the first three-dimensional position data is input, and at least one of the first image data and the first three-dimensional position data.

32. The control device according to any one of claims 20 to 31, wherein the calculation device calculates the priority based on the likelihood of the object calculated using the learning model and the degree of overlap of the objects.

33. The control device according to any one of claims 1 to 32, wherein the fourth image data is identical to the third image data, and the third three-dimensional position data is identical to the second three-dimensional position data.

34. The control device according to any one of claims 1 to 33, wherein the second image data, the third image data, and the fourth image data are identical, and the first three-dimensional position data, the second three-dimensional position data, and the third three-dimensional position data are identical.

35. The control device according to any one of claims 1 to 34, wherein at least one of the plurality of objects in the third image data and the fourth image data includes the target object.

36. A robot equipped with a processing device for processing at least one object, and a control device for generating a control signal for controlling at least one of the processing devices, wherein the control device comprises a calculation device for generating the control signal and an output device for outputting the control signal generated by the calculation device, the calculation device identifies a first data portion in which the object exists from first three-dimensional position data generated by an imaging system capable of generating image data by imaging the object, generates position and orientation information of the object based on a second data portion that interpolates the first data portion calculated using a three-dimensional model of the object and the first data portion, and generates the control device for executing the processing on the object based on the position and orientation information.

37. The control device according to claim 36, wherein the imaging system is provided on the robot.

38. The control device according to claim 36 or 37, wherein the three-dimensional model of the object is a first model, the computing device adjusts the positional relationship and orientation relationship between the first model and a second model which is a three-dimensional model of the object indicated by the first data portion, and at least a part of the second model whose positional relationship and orientation relationship have been adjusted is used as the second data portion.

39. The control device according to any one of claims 36 to 38, wherein the three-dimensional position data is second three-dimensional position data, the position and orientation information is second position and orientation information, and the calculation device generates first position and orientation information indicating the position and orientation of the object using at least one of first image data which is image data generated by the imaging system capturing the object, and first three-dimensional position data generated from second image data which is image data generated by the imaging system capturing the target object, and calculates the second data portion using the first position and orientation information and a three-dimensional model of the object.

40. The control device according to claim 39, wherein the three-dimensional model of the object is a first model, the computing device adjusts the positional relationship and orientation relationship between the first model and a second model which is a three-dimensional model of the object indicated by the first data portion, and data indicating the three-dimensional position of each of multiple points of the object using at least a part of the second model whose positional relationship and orientation relationship have been adjusted is used as the second data portion.

41. The control device according to claim 39 or 40, wherein the computing device generates first position and orientation information using a learning model generated by machine learning to output first position and orientation information when at least one of first image data, which is image data generated when the imaging system images the object, and second image data, which is image data generated when the imaging system images the target object, is input, and at least one of the first image data and the first three-dimensional position data.

42. The control device according to claim 41, wherein the computing device generates first position and orientation information by inputting at least one of the first image data and the first three-dimensional position data into the learning model.

43. The control device according to claim 41 or 42, wherein the learning model is generated by machine learning using learning data associated with learning image data showing a sample object that simulates the object and a learning image showing at least one of the objects, and a ground truth label indicating the result of identifying a region in which the sample object and at least one of the objects exist.

44. The control device according to any one of claims 39 to 43, wherein the arithmetic unit generates first position and orientation information by performing a matching process using at least one of the first image data and the first three-dimensional position data.

45. A robot equipped with a processing device for performing processing on at least one of a plurality of objects, and a control device for generating a control signal for controlling at least one of the processing devices, wherein the control device comprises a calculation device for generating the control signal and an output device for outputting the control signal generated by the calculation device, the calculation device uses at least one of first image data, which is image data generated when an imaging system capable of generating image data by imaging the plurality of objects images at least one of the plurality of objects, and first three-dimensional position data generated from second image data, which is image data generated when the imaging system images at least one of the plurality of objects, and a learning model generated by machine learning to select a target object from the plurality of objects to be processed by the processing device, and generates the control signal for performing the processing on the selected target object.

46. ​​The control device according to claim 45, wherein the imaging system is provided on the robot.

47. The control device according to claim 45 or 46, wherein the computing device calculates the priority of a first object and a second object included in the plurality of objects using at least one of the first image data and the first three-dimensional position data and the learning model, and the selection includes selecting the first object as the target object based on the priority.

48. The control device according to claim 47, wherein the imaging system images the first object and the second object, which are each the object, the computing device determines the priority of the first object and the priority of the second object based on the imaging results of the imaging system, and selects the first object as the target object based on the priority.

49. The control device according to claim 48, wherein the arithmetic unit selects the first object as the target object when the priority of the first object is higher than the priority of the second object.

50. The control device according to any one of claims 47 to 49, wherein the priority is determined based on the degree of overlap between the first object and the second object.

51. The control device according to claim 50, wherein the priority of the first object increases as the degree of overlap between the first object and other objects different from the first object decreases, and the priority of the second object increases as the degree of overlap between the second object and other objects different from the second object decreases.

52. The control device according to any one of claims 47 to 51, wherein the arithmetic unit calculates the priority using at least one of the first image data and the first three-dimensional position data and the learning model.

53. The control device according to claim 52, wherein the computing device selects the target object based on the priority and the likelihood of the priority calculated using the learning model.

54. The control device according to any one of claims 47 to 53, wherein the computing device selects the target object based on the similarity between the target object and the three-dimensional model of the target object, indicated by at least one of the first image data and the first three-dimensional position data, and the priority.

55. The control device according to claim 54, wherein the imaging system images the first object and the second object, each of which is the object; the computing device performs a first calculation based on the imaging results of the imaging system using the similarity of the first object and the priority of the first target object; performs a second calculation using the similarity of the second object and the priority of the second object; and selects the first object as the target object based on the results of the first and second calculations.

56. The control device according to claim 55, wherein the calculation device selects the first object as the target object when the results of the first and second calculations indicate that the first object should be selected as the target object rather than the second object.

57. The control device according to any one of claims 54 to 56, wherein the arithmetic device calculates the similarity using a learning model generated by machine learning to output the similarity when at least one of the first image data and the first three-dimensional position data is input, and at least one of the first image data and the first three-dimensional position data.

58. The control device according to any one of claims 47 to 57, wherein the calculation device calculates the priority based on the likelihood of the object calculated using the learning model and the degree of overlap of the objects.

59. A robot provided with a processing device for processing at least one object, and a control device for generating a control signal for controlling at least one of the processing devices, wherein the control device comprises a calculation device for generating the control signal and an output device for outputting the control signal generated by the calculation device, the calculation device generates first position and orientation information indicating the position and orientation of the object using first three-dimensional position data generated from first image data which is image data generated by an imaging system capable of generating image data by imaging the object, second position and orientation information indicating the position and orientation of the object based on second three-dimensional position data generated from second image data which is image data generated by the imaging system for imaging the object, and at least one of the position determined based on the first position and orientation information and the orientation determined based on the first position and orientation information, and the control device generates the control signal based on the second position and orientation information.

60. The control device according to claim 59, wherein the imaging system is provided on the robot.

61. The control device according to claim 59 or 60, wherein the calculation device generates the second position and orientation information by performing a matching process using a three-dimensional model of the object positioned in accordance with at least one of the position determined based on the first position and orientation information and the orientation determined based on the first position and orientation information, and the second three-dimensional position data.

62. A control system comprising the control device according to any one of claims 1 to 61 and the imaging system.

63. A robot system comprising the control device according to any one of claims 1 to 61, the imaging system, and the robot.

64. A robot equipped with a processing device for processing at least one of a plurality of objects and an imaging system capable of generating image data by imaging at least one of the plurality of objects, and a control method for generating a control signal for controlling at least one of the processing devices, comprising: identifying an object region in which a target object to be processed by the processing device exists, using at least one of first image data which is the image data generated by the imaging system imaging at least one of the plurality of objects, and first three-dimensional position data generated from second image data which is the image data generated by the imaging system imaging at least one of the plurality of objects; generating first position and orientation information indicating at least one of the position and orientation of the target object using the result of identifying the object region and second three-dimensional position data indicating a three-dimensional position generated from third image data which is the image data generated by the imaging system imaging at least one of the plurality of objects; and generating second position and orientation information indicating at least one of the position and orientation of the target object using the first position and orientation information and third three-dimensional position data indicating a three-dimensional position generated from fourth image data which is the image data generated by the imaging system imaging at least one of the plurality of objects. A control method comprising generating the control signal for performing the processing on the target object based on the second position and orientation information.

65. A robot equipped with a processing device for processing at least one object, and a control method for generating a control signal for controlling at least one of the processing devices, comprising: identifying a first data portion in which the object exists from first three-dimensional position data generated by an imaging system capable of generating image data by imaging the object; generating position and orientation information of the object based on a second data portion that interpolates the first data portion calculated using a three-dimensional model of the object and the first data portion; and generating the control signal for the object to perform the processing based on the position and orientation information.

66. A robot equipped with a processing device that performs processing on at least one of a plurality of objects, and a control method for generating a control signal for controlling at least one of the processing devices, comprising: selecting a target object from the plurality of objects to be processed by the processing device using at least one of first image data which is image data generated by an imaging system capable of generating image data by imaging at least one of the plurality of objects, and first three-dimensional position data generated from second image data which is image data generated by the imaging system by imaging at least one of the plurality of objects, and a learning model generated by machine learning; and generating the control signal for performing the processing on the selected target object.

67. A robot provided with a processing device for processing at least one object, and a control method for generating a control signal for controlling at least one of the processing devices, comprising: generating first position and orientation information indicating the position and orientation of an object using first three-dimensional position data generated from first image data which is image data generated by an imaging system capable of generating image data by imaging the object; generating second position and orientation information indicating the position and orientation of an object based on second three-dimensional position data generated from second image data which is image data generated by the imaging system by imaging the object, and at least one of a position determined based on the first position and orientation information and an orientation determined based on the first position and orientation information; and generating the control signal based on the second position and orientation information.

68. A computer program that causes a computer to execute the control method described in any one of claims 64 to 67.

Citation Information

Patent Citations

  • Robot control system, robot control method and program

    JP2022169255A

  • Control system, driving device, input device and control program

    JP2023003602A

  • Control device, control system, robot system, control method, and computer program

    WO2023209974A1