Calculation device, calculation system, robot system, calculation method, and computer program

By using the control unit and the learning unit of the computing device in the robot system, the control signals and models are generated to correct the robot coordinate system, and the processing object image data is generated by shooting and generating the problem of accurately correcting the coordinate system in the prior art, and high-precision object position and posture recognition are achieved.

CN119947870APending Publication Date: 2025-05-06NIKON CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202280100172.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2022-09-29
Publication Date
2025-05-06

Smart Images

  • Figure CN119947870A_ABST
    Figure CN119947870A_ABST
Patent Text Reader

Abstract

The arithmetic device includes: a control unit that outputs a control signal for controlling an imaging unit and a robot provided with the imaging unit; and a learning unit that generates a model that determines a parameter of the calculation process in the control unit by learning the imaging results of the object to be learned by the imaging unit. The control unit outputs a first control signal for driving the robot such that the imaging unit has a predetermined positional relationship with respect to the object to be learned, and causes the imaging unit to image the object to be learned in the predetermined positional relationship. The learning unit generates a model by learning using learning image data generated by causing the imaging unit to capture an image of an object to be learned in a predetermined positional relationship by control based on the first control signal. The control unit calculates the position and / or orientation of the object to be processed using a parameter determined using the model and image data to be processed generated by imaging the object to be processed having substantially the same shape as the object to be learned by the imaging unit.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the technical field of computing devices, computing systems, robot systems, computing methods, and computer programs. Background Technology

[0002] For example, a method has been proposed that, in a hand-eye robotic system where a vision sensor is fixed to the tip of the robot's hand, corrects the robot system's coordinate system based on the result of the vision sensor observing a mark from multiple observation positions (see Patent Document 1). Patent Document 2 is cited as another related technology.

[0003] [Existing Technical Documents]

[0004] [Patent Literature]

[0005] Patent Document 1: Japanese Patent Application Publication No. 2012-91280

[0006] Patent Document 2: Japanese Patent Application Publication No. 2010-188439 Summary of the Invention

[0007] According to a first approach, a computing device is provided, comprising: a control unit that outputs a control signal for controlling a camera unit and a robot equipped with the camera unit; and a learning unit that generates a model that determines parameters for computational processing in the control unit by learning from the camera unit's acquisition of a learning object. The control unit outputs a first control signal for driving the robot such that the camera unit is in a predetermined position relative to the learning object, and that the camera unit acquires a picture of the learning object under the predetermined position. The learning unit generates the model by learning from learning image data, wherein the learning image data is generated by: acquiring a picture of the learning object under the predetermined position by controlling the first control signal; and the control unit performing the computational processing using parameters determined by the model and processing object image data generated by the camera unit acquiring a processing object with approximately the same shape as the learning object, to calculate at least one of the position and posture of the processing object.

[0008] According to a second method, a computing device is provided that calculates the position and posture of an object. In the computing device, the object is disposed on a platform provided with at least one mark. The computing device includes a computing unit that calculates at least one of the object's position and object posture based on image data generated by capturing images of the object and the at least one mark by an imaging unit.

[0009] According to a third approach, a computing system is provided, which includes a computing device provided by the first approach and the imaging unit.

[0010] According to a fourth approach, a robot system is provided, which includes a computing device provided by a first approach, the imaging unit, and the robot.

[0011] According to a fifth method, a computational method is provided, which is a computational method in a computing device. The computing device includes: a control unit that outputs control signals for controlling a camera unit and a robot equipped with the camera unit; and a learning unit that generates a model that determines parameters for computational processing in the control unit by learning from the camera unit's acquisition of a learning object. The computational method includes: the control unit outputting a first control signal for driving the robot such that the camera unit is in a predetermined position relative to the learning object, and that the camera unit acquires an image of the learning object under the predetermined position; the learning unit generating the model by learning from learning image data, the learning image data being generated by: acquiring an image of the learning object under the predetermined position by controlling the first control signal; and the control unit performing the computational processing using parameters determined by the model and processing object image data generated by the camera unit acquiring an image of a processing object with approximately the same shape as the learning object, to calculate at least one of the position and posture of the processing object.

[0012] According to the sixth method, a computer program is provided that enables a computer to perform the arithmetic methods provided by the fifth method.

[0013] According to a seventh method, a computing device is provided, which includes a recording medium storing a computer program provided by a sixth method, and is capable of executing the computer program. Attached Figure Description

[0014] Figure 1 This is a block diagram showing the structure of the robot system according to the first embodiment.

[0015] Figure 2 This is a side view showing the appearance of the robot according to the first embodiment.

[0016] Figure 3 This is a block diagram showing the structure of the imaging unit in the first embodiment.

[0017] Figure 4 This is a block diagram showing the structure of the control device in the first embodiment.

[0018] Figure 5 This diagram is a summary of the operation of the decision unit and the object recognition unit in the first embodiment.

[0019] Figure 6 This diagram illustrates an example of a method for constructing a parameter-determined model according to the first embodiment.

[0020] Figure 7 This is a diagram illustrating an example of the positional relationship between the imaging unit and the learning object in the first embodiment.

[0021] Figure 8 This is a diagram showing an example of an input screen displayed by the output device of the control device according to the first embodiment.

[0022] Figure 9 This is a diagram showing an outline of a robot system according to a first variation of the first embodiment.

[0023] Figure 10 This is a block diagram showing the structure of the control device in a first variation of the first embodiment.

[0024] Figure 11 This is a block diagram showing the structure of other devices in a first variation of the first embodiment.

[0025] Figure 12 This is a block diagram showing the structure of the control device in a second variation of the first embodiment.

[0026] Figure 13 This is a block diagram showing the structure of other devices in a first variation of the first embodiment.

[0027] Figure 14 This is a conceptual diagram of the robot system according to the second embodiment.

[0028] Figure 15 This is an example of a screen displayed as a user interface.

[0029] Figure 16 This is another example of a screen displayed as a user interface.

[0030] Figure 17 This is another example of a screen displayed as a user interface.

[0031] Figure 18 This is another example of a screen displayed as a user interface.

[0032] Figure 19 This is another example of a screen displayed as a user interface.

[0033] Explanation of reference numerals in the attached figures

[0034] 1, 2: Robot System

[0035] 10, 31, 32, 33: Robots

[0036] 12: Robotic Arm

[0037] 13: End effector

[0038] 14: Robot control device

[0039] 20: Filming Unit

[0040] 21, 22: Filming equipment

[0041] 100, 101, 102, 301, 302, 303: Control devices

[0042] 110: Computing device

[0043] 111, 211: Data Generation Department

[0044] 112, 212: Learning Object Recognition Department

[0045] 113, 213: Study Department

[0046] 114: Decision-Making Department

[0047] 115: Object Recognition Processing Unit

[0048] 116: Signal generation department

[0049] 120, 220: Storage devices

[0050] 130, 230: Communication device

[0051] 140, 240: Input devices

[0052] 150, 250: Output devices

[0053] 160, 260: Data bus

[0054] 400: Management Device

[0055] 450: Monitor Detailed Implementation

[0056] The implementation methods of the computing device, computing system, robot system, computing method, and computer program are described.

[0057] <First Implementation>

[0058] In the first embodiment, an example of applying a computing device, a computing system, a robot system, a computing method, and a computer program in robot system 1 will be described.

[0059] (1) Overview of Robot Systems

[0060] Reference Figure 1 and Figure 2 An overview of robot system 1 is provided. Figure 1 This is a block diagram representing the structure of robot system 1. Figure 2 This is a side view showing the appearance of robot 10.

[0061] exist Figure 1 In this system, robot system 1 includes robot 10, imaging unit 20, and control device 100. Robot 10 is a device capable of performing prescribed processing on a workpiece W, which is an object. For example, ... Figure 2 As shown, the robot 10 may include a base 11, a robot arm 12, and a robot control device 14.

[0062] The base 11 is a component that serves as the foundation of the robot 10. The base 11 is disposed on a supporting surface such as the ground. The base 11 may also be fixed to the supporting surface. The base 11 may also be movable relative to the supporting surface. In the case where the base 11 is movable relative to the supporting surface, the base 11 may also be able to move independently on the supporting surface. In this case, the base 11 may also be mounted on an Automatic Guided Vehicle (AGV). That is, the robot 10 may also be mounted on an AGV.

[0063] Robotic arm 12 is assembled on base 11. Robotic arm 12 is a device consisting of multiple links connected by joints. An actuator is built into the joint. The links are rotatable about an axis defined by the joint via the actuator built into the joint. In addition, at least one link is also telescopic along the direction in which the link extends. Furthermore, the device including the multiple links connected by joints and base 11 can be referred to as robotic arm 12.

[0064] An end effector 13 is assembled on the robot arm 12. That is, an end effector 13 is assembled on the robot 10. Figure 2 In the example shown, the end effector 13 is assembled at the tip of the robotic arm 12. The end effector 13 is movable by the movement of the robotic arm 12. That is, the robotic arm 12 moves the end effector 13. In other words, the robot 10 moves the end effector 13. Furthermore, the robot 10 may include the end effector 13 as part of the robot 10. In other words, the end effector 13 may constitute part of the robot 10.

[0065] The end effector 13 is a device that performs a prescribed process (in other words, a prescribed action) on the workpiece W. Furthermore, since the end effector 13 performs a prescribed process on the workpiece W, it can be called a processing device.

[0066] The end effector 13 can perform a gripping process for gripping a workpiece W. The end effector 13 can also perform a configuration process for positioning the workpiece W gripped by the end effector 13 at a desired position. The end effector 13 can also perform an embedding process for embedding the workpiece W (a first object) gripped by the end effector 13 into another object (a second object) that is different from the first object. In this case, the second object into which the first object, the workpiece W, is embedded can also be referred to as workpiece W.

[0067] As an example of an end effector 13 performing at least one of the grasping, configuration, and embedding processes, examples include a hand gripper capable of grasping the workpiece W by physically clamping it, and a vacuum gripper capable of grasping the workpiece W by vacuum suction. Furthermore, Figure 2 This example shows an end effector 13 as a hand gripper. The number of fingers or claws in the hand gripper is not limited to 2, but can be 3 or more (e.g., 3, 6, etc.).

[0068] Furthermore, the desired position for arranging workpiece W (the first object) can be the desired position for another object (the second object) different from workpiece W. In this case, the other object (the second object) can be a arranging device for arranging workpiece W, which is the first object. A pallet can be cited as an example of a arranging device. The arranging device can be arranged on a support surface such as the ground. The arranging device can also be fixed to the support surface. Alternatively, the arranging device can be movable relative to the support surface. As an example, the arranging device can also be movable on the support surface. In this case, the arranging device can also be called an automated guided vehicle (AGV). When the arranging device is movable relative to the support surface, the workpiece W arranged on the arranging device also moves relative to the support surface as the arranging device moves. Therefore, a movable arranging device can also function as a moving device that moves workpiece W. Furthermore, a belt conveyor can be used as a arranging device. Furthermore, the arranging device, which is the second object and on which the workpiece W, which is the first object, is arranged, can also be called workpiece W. Furthermore, the workpiece W, which is the first object, can be arranged on the arranging device before being grasped by the end effector 13.

[0069] Furthermore, the gripping process performed by the end effector 13 to grip the workpiece W can be referred to as a holding process to hold the workpiece W. Additionally, the configuration process performed by the end effector 13 to position the workpiece W gripped by the end effector 13 at a desired position can be referred to as a release process (in other words, a release action) to release the workpiece W gripped by the end effector 13 to the desired position (i.e., separation). The object gripped by the end effector 13 is not limited to the object itself; at least one of the object held by the end effector 13 and the object positioned (in other words, released) by the end effector 13 can also be the object (i.e., the workpiece W) for which the end effector 13 performs its prescribed processing. In the case where the first object gripped by the end effector 13 as the workpiece W is embedded in a second object different from the first object, the end effector 13 embeds the first object into the second object, therefore the second object can be considered the object for which the end effector 13 performs its prescribed processing. That is, in this case, the second object in which the first object is embedded can also be referred to as the workpiece W. When the first object, which is a workpiece W, grasped by the end effector 13 is placed in a second object different from the first object, the end effector 13 places the first object in the second object. Therefore, the second object can be considered the object to which the end effector 13 performs its prescribed processing. That is, in this case, the second object in which the first object is placed can also be referred to as workpiece W.

[0070] Furthermore, the end effector 13 is not limited to a hand gripper or a vacuum gripper, but can also be a machining device for machining workpiece W. The machining device can perform at least one of the following: additional machining to add new shapes to workpiece W, removal machining to remove a portion of workpiece W, welding machining to join two workpieces W (a first object and a second object), and cutting machining to cut off workpiece W. The machining device can use a tool to machine workpiece W. In this case, the machining device including the tool can be assembled on the robot arm 12. Alternatively, the machining device can machine workpiece W by irradiating it with an energy beam (e.g., light, electromagnetic waves, and charged particle beams). In this case, the machining device including an irradiation device to irradiate the workpiece W with an energy beam can be assembled on the robot arm 12.

[0071] As an example of an end effector 13, a processing device can perform welding processing, welding parts onto a workpiece W. The processing device can use a soldering iron to weld the parts onto the workpiece W. In this case, the processing device including the soldering iron can be assembled into a robotic arm 12. Alternatively, the processing device can also weld parts onto the workpiece W by irradiating the solder with an energy beam (e.g., light, electromagnetic waves, and charged particle beams). In this case, the processing device including an irradiation device that irradiates the workpiece W with an energy beam can be assembled into the robotic arm 12. Furthermore, the energy beam irradiated for processing the workpiece W can be referred to as processing light.

[0072] As another example, a measuring device for measuring workpiece W, serving as an end effector 13, can be assembled on robot arm 12. The measuring device can be capable of measuring characteristics of workpiece W. Examples of characteristics of workpiece W include at least one of the following: shape, size, and temperature. The measuring device can use a contact probe to measure workpiece W. In this case, the measuring device including the contact probe can be assembled on robot arm 12. Alternatively, the measuring device can measure workpiece W by irradiating it with an energy beam (e.g., light, electromagnetic waves, and charged particle beams). In this case, the measuring device including an irradiation device for irradiating the workpiece W with an energy beam can be assembled on robot arm 12.

[0073] also, Figure 2 The robot 10 is an example of a robot including robotic arms 12 (i.e., a vertical articulated robot). However, robot 10 can also be a robot different from a vertical articulated robot. Robot 10 can also be a selective compliance assembly robot (i.e., a horizontal articulated robot). Robot 10 can also be a parallel linkage robot. Robot 10 can also be a biarm robot including two robotic arms 12. Robot 10 can also be an orthogonal coordinate robot. Robot 10 can also be a cylindrical coordinate robot. Robot 10 can also be referred to as a movable device. A movable device can replace robot 10 or include at least one of automated guided vehicles (AGVs) and unmanned aerial vehicles (UAVs) on its basis. Robot 10 can be installed in at least one of automated guided vehicles and UAVs. Furthermore, robot system 1 includes robot 10, which can be referred to as a movable device. Therefore, robot system 1 can also be referred to as a movable system. Furthermore, a movable system can replace robot 10 or include at least one of, for example, automated guided vehicles and UAVs on its basis. In addition, the movable system may also include a robot 10 installed in at least one of the automated guided vehicles and unmanned aerial vehicles.

[0074] The robot control device 14 controls the movements of the robot 10. As an example, the robot control device 14 can control the movements of the robot arm 12. The robot control device 14 can control the movements of the robot arm 12 such that a desired link rotates about an axis specified by a desired joint. The robot control device 14 can control the movements of the robot arm 12 such that the end effector 13 assembled on the robot arm 12 is positioned in a desired position (in other words, moved). As another example, the robot control device 14 can also control the movements of the end effector 13. The robot control device 14 can also control the movements of the end effector 13 such that the end effector 13 grasps the workpiece W at a desired time. That is, the robot control device 14 can also control the movements of the end effector 13 such that the end effector 13 performs grasping processing at a desired time. The robot control device 14 can also control the movements of the end effector 13 to position the workpiece W grasped by the end effector 13 at the desired time in the desired position (i.e., release the grasped workpiece W). That is, the robot control device 14 can also control the movement of the end effector 13, so that the end effector 13 is configured at the desired time. Furthermore, the robot control device 14 can also control the movement of the end effector 13, so that the end effector 13 is engaged at the desired time. When the end effector 13 is a gripper, the robot control device 14 can control the timing of the gripper's opening and closing. When the end effector 13 is a vacuum gripper, the robot control device 14 can control the timing of turning the vacuum of the vacuum gripper on / off.

[0075] The imaging unit 20 is a device for taking pictures of an object (e.g., workpiece W). The imaging unit 20 can be assembled onto the robot arm 12. For example, as... Figure 2 As shown, the imaging unit 20 can also be assembled at the front end of the robot arm 12. In this case, the imaging unit 20 assembled on the robot arm 12 can be moved by the movement of the robot arm 12. That is, the robot arm 12 can move the imaging unit 20. The imaging unit 20 may also not be assembled on the robot arm 12. For example, a bracket can be assembled above the robot 10, and the imaging unit 20 can be assembled on the bracket.

[0076] Reference Figure 3 An example of the shooting unit 20 will be explained. Figure 3 This is a block diagram representing the structure of the imaging unit 20. For example... Figure 3As shown, the imaging unit 20 may include an imaging device 21, an imaging device 22, and a projection device 23. The imaging device 21 is a single monocular camera. In other words, the imaging device 21 has one imaging element. The imaging device 22 is a stereo camera with two monocular cameras. In other words, the imaging device 22 has two imaging elements. Furthermore, at least one of the imaging devices 21 and 22 may include three or more monocular cameras. At least one of the imaging devices 21 and 22 may be at least one of a light field camera, a plenoptic camera, and a multispectral camera. Additionally, the imaging unit 20 may not include the projection device 23. Furthermore, the imaging unit 20 may not include either the imaging device 21 or the imaging device 22.

[0077] Imaging device 21 generates image data IMG_2D by photographing an object (e.g., workpiece W). The generated image data IMG_2D is output from imaging device 21 to control device 100. Imaging device 22 generates image data IMG_3D by photographing an object (e.g., workpiece W). Here, image data IMG_3D includes two image data generated by the two monocular cameras of the stereo camera of imaging device 22. Image data IMG_3D is output from imaging device 22 to control device 100.

[0078] Projection device 23 is a device capable of irradiating an object (e.g., workpiece W) with projection light. Projection device 23 can be a device capable of projecting a desired projection pattern onto workpiece W by irradiating it with projection light. Projection device 23 can be a projector. The desired projection pattern can include a random pattern. The random pattern can be a projection pattern with a different pattern for each unit illumination area. The random pattern can include a random dot pattern. The desired projection pattern is not limited to a random pattern; it can include a one-dimensional or two-dimensional lattice pattern, or a pattern different from a lattice pattern. Furthermore, the projection light irradiating the object from projection device 23 can be called patterned light or structured light because it can project the projection pattern onto workpiece W. Furthermore, projection device 23 can be called a light projection device because it can project the desired projection pattern.

[0079] Furthermore, the projection device 23 may be a projection light that is different from the projection light that can project a desired projection pattern onto an object (e.g., workpiece W). The projection light that is different from the projection light that can project a desired projection pattern may be called illumination light. When the projection device 23 irradiates illumination light as projection light, the projection device 23 may be called an illumination device.

[0080] Furthermore, at least one of the shooting devices 21, 22 and 23 in the shooting unit 20 can be assembled in the robot arm 12, and at least one other shooting device 21, 22 and 23 in the shooting unit 20 can be assembled in a different location from the robot arm 12.

[0081] The imaging devices 21 and 22 can simultaneously and synchronously photograph an object (e.g., workpiece W). Specifically, the imaging devices 21 and 22 can photograph the object at the same time. Alternatively, they can photograph the object at different times.

[0082] Here, the state of "simultaneous imaging of an object by both imaging devices 21 and 22" can, as the name suggests, include the state where "the imaging time of imaging device 21 and the imaging time of imaging device 22 are exactly the same." The state of "simultaneous imaging of an object by both imaging devices 21 and 22" can also include the state where "although the imaging time of imaging device 21 and the imaging time of imaging device 22 are not exactly the same, the time deviation between their imaging times is less than the allowable upper limit, and therefore they are considered to be substantially the same." Here, if there is a time deviation between the imaging times of imaging device 21 and 22, the control of the robotic arm 12 will have an error. The "allowable upper limit" can be an allowable upper limit set based on the control error of the robotic arm 12 caused by the time deviation between the imaging times of imaging devices 21 and 22.

[0083] The imaging device 22 can photograph an object onto which a desired projection pattern is projected. That is, the imaging device 22 can photograph an object when the projection device 23 illuminates it with projection light capable of projecting the desired pattern. In this case, the object with the desired projection pattern can be captured in the image data IMG_3D generated by the imaging device 22. On the other hand, the imaging device 21 may not photograph an object onto which a desired projection pattern is projected. That is, the imaging device 21 may not photograph an object when the projection device 23 illuminates it with projection light capable of projecting the desired pattern. In this case, an object without the desired projection pattern can be captured in the image data IMG_2D generated by the imaging device 21.

[0084] Furthermore, when the projection device 23 illuminates an object with projection light capable of projecting a desired projection pattern, the imaging devices 21 and 22 can simultaneously photograph the object with the desired projection pattern. In this case, the projection device 23 can illuminate the object with projection light containing a light component of a first wavelength band (e.g., the wavelength band of blue light). The imaging device 21 may have a filter capable of attenuating the light component of the first wavelength band. Here, the imaging device 21 can photograph the object by receiving light from the object through the filter via the imaging element. As described above, when the projection light from the projection device 23 contains a light component of the first wavelength band, the filter provided by the imaging device 21 attenuates the projection light. Therefore, the return light (e.g., at least one of reflected and scattered light from the object illuminated by the projection light (i.e., the object with the projection pattern) is attenuated by the filter provided by the imaging device 21. As a result, even when the projection device 23 illuminates the object with projection light, the imaging device 21 can photograph the object without being affected by the projection light emitted from the projection device 23. On the other hand, the imaging device 22 may not have a filter capable of attenuating the light components of the first wavelength band. Therefore, the imaging device 22 can photograph an object illuminated by projected light (i.e., an object with a projected pattern).

[0085] Furthermore, the imaging unit 20 may include only the imaging device 21. The imaging unit 20 may also include only the imaging device 22. The imaging unit 20 may also include both the imaging device 22 and the projection device 23 (in other words, the imaging unit 20 may not include the imaging device 21). As described above, the imaging unit 20 captures images of an object (e.g., workpiece W). Therefore, the imaging unit 20 may be referred to as an imaging unit.

[0086] return Figure 2 The control device 100 performs robot control processing. Robot control processing is the process of generating robot control signals for controlling the robot 10. Specifically, the control device 100 generates robot control signals based on at least one of image data IMG_2D and image data IMG_3D output from the imaging unit 20. The control device 100 calculates at least one of the position and pose of an object (e.g., workpiece W) within the global coordinate system of the robot system 1 based on, for example, at least one of the image data IMG_2D and image data IMG_3D. The control device 100 generates robot control signals based on at least one of the calculated position and pose of the object. The generated robot control signals are output to the robot control device 14 of the robot 10. Furthermore, the robot control signals may include signals for controlling the movements of the robot 10.

[0087] In addition to performing robot control processing, the control device 100 can also perform end effector control processing. End effector control processing may include generating end effector control signals for controlling the end effector 13. Specifically, the control device 100 may generate end effector control signals based on at least one of the calculated position and orientation of the object. End effector control processing may or may not be included in robot control processing. That is, the end effector control signals generated by the control device 100 may or may not be included in the robot control signals.

[0088] In the following description, for ease of explanation, an example is given where the end effector control processing is included in the robot control processing (i.e., the end effector control signal is included in the robot control signal). Therefore, in the following description, robot control processing may refer to the process of generating at least one of a robot control signal and an end effector control signal. Furthermore, in the following description, robot control signal may refer to at least one of a signal used to control robot 10 and a signal used to control end effector 13. Additionally, at least one of the robot control signal and the end effector control signal may be referred to as a second control signal.

[0089] The robot control device 14 can control the drive of the actuators built into the joint of the robot arm 12 based on robot control signals. As described above, the robot control signals may include signals for controlling the end effector 13 (i.e., end effector control signals). That is, the end effector 13 can be controlled via the robot control signals. In this case, the robot control device 14 can control the drive of the actuators built into the end effector 13 based on the robot control signals. At this time, the robot control device 14 can convert the robot control signals into control signals expressed in the language inherent to the robot 10. In this case, it can be said that the control device 100 indirectly controls the robot 10 via the robot control device 14. Furthermore, the control device 100 may also control the robot 10 without using the robot control device 14. In this case, the robot control signals may be expressed in the language inherent to the robot 10. In this case, the robot 10 may also not include the robot control device 14.

[0090] Furthermore, the "global coordinate system" can be considered a coordinate system based on the robot 10. The control device 100 can calculate at least one of the position and pose of the object in a coordinate system different from the global coordinate system based on image data IMG_2D and image data 3D, and generate robot control signals. As a coordinate system different from the global coordinate system, at least one of the following can be listed: a coordinate system based on the imaging device 21 (the two-dimensional (2D) imaging coordinate system described later) and a coordinate system based on the imaging device 22 (the three-dimensional (3D) imaging coordinate system described later).

[0091] Thus, the control device 100 and the imaging unit 20 are used to control the robot 10. Therefore, the system including the control device 100 and the imaging unit 20 can also be called a robot control system or a control system.

[0092] Reference Figure 4 The control device 100 is described below. Figure 4 This is a block diagram showing the structure of the control device 100. Figure 4 In this configuration, the control device 100 includes a processing unit 110, a storage unit 120, and a communication unit 130. The control device 100 may also include an input unit 140 and an output unit 150. Alternatively, the control device 100 may not include at least one of the input unit 140 and the output unit 150. The processing unit 110, storage unit 120, communication unit 130, input unit 140, and output unit 150 can be connected via a data bus 160.

[0093] The computing device 110 may include at least one of, for example, a central processing unit (CPU), a graphics processing unit (GPU), and a field programmable gate array (FPGA).

[0094] Storage device 120 may include at least one of, for example, random access memory (RAM), read-only memory (ROM), hard disk drive, magneto-optical disk drive, solid-state drive (SSD), and hard disk array. That is, storage device 120 may also include non-temporary storage media.

[0095] The communication device 130 is capable of communicating with both the robot 10 and the camera unit 20. The communication device 130 may also be capable of communicating with other devices different from the robot 10 and the camera unit 20 via an unmapped network.

[0096] Input device 140 may also include at least one of a keyboard, mouse, and touchscreen. Input device 140 may also include a recording medium reading device capable of reading information recorded on a removable recording medium such as a Universal Serial Bus (USB) memory. Furthermore, when information is input to control device 100 via communication device 130 (in other words, when control device 100 obtains information via communication device 130), communication device 130 may function as an input device.

[0097] The output device 150 may include at least one of a display, a speaker, and a printer. The output device 150 may also be capable of outputting information to a removable storage medium such as a USB flash drive. Furthermore, when outputting information from the control device 100 via the communication device 130, the communication device 130 may function as an output device.

[0098] The arithmetic unit 110 may include a data generation unit 111, a learning object recognition unit 112, a learning unit 113, a decision unit 114, a processing object recognition unit 115, and a signal generation unit 116 as logically implemented functional blocks. These functional blocks can be implemented by the arithmetic unit 110 executing a computer program. Furthermore, detailed descriptions of each of the data generation unit 111, the learning object recognition unit 112, the learning unit 113, the decision unit 114, the processing object recognition unit 115, and the signal generation unit 116 will be provided later, and therefore are omitted here.

[0099] The arithmetic unit 110 can read computer programs stored in the storage device 120. The arithmetic unit 110 can use a recording medium reading device (not shown) included in the control device 100 to read computer programs that are readable by a computer and stored on a non-temporary recording medium. The arithmetic unit 110 can acquire (i.e., download or read) computer programs from a device (not shown) located outside the control device 100 via a communication device 130 (or other communication device). The arithmetic unit 110 can execute the read computer programs. As a result, logical function blocks for performing the processes (e.g., the robot control processes described above) that the control device 100 should perform can be implemented within the arithmetic unit 110. That is, the arithmetic unit 110 can function as a controller for implementing logical function blocks for performing the processes that the control device 100 should perform.

[0100] Furthermore, as the recording medium for recording the computer program executed by the arithmetic unit 110, the following can be used: Compact Disc ROM (CD-ROM), Compact Disc Recordable (CD-R), Compact Disc Rewritable (CD-RW), Flexible Disk, Magnet Optical Disc (MO), Digital Versatile Disc ROM (DVD-ROM), Digital Versatile Disc RAM (DVD-RAM), Recordable Digital Versatile Disc Recordable (DVD-R), Writeable Digital Video Disc (DVD+R), Rewritable Digital Versatile Disc Rewritable (DVD-RW), and Reproducible Digital Versatile Disc (DVD-RW). At least one of the following: optical discs such as ReWritable (DVD+RW) and Blu-ray (registered trademark), magnetic media such as magnetic tape, optical disks, semiconductor storage such as USB storage, and any medium capable of storing other programs. The recording medium may include a device capable of recording computer programs (e.g., a general-purpose or special-purpose device installed in a state where a computer program can be executed in the form of at least one of software and firmware). Furthermore, the various processes or functions contained in the computer program may be implemented by logic processing blocks executed by the arithmetic unit 110 (i.e., computer) within the arithmetic unit 110, or by hardware such as a defined gate array (FPGA (Field Programmable Gate Array), Application Specific Integrated Circuit (ASIC)) included in the arithmetic unit 110, or by a mixture of logic processing blocks and partial hardware modules that implement a portion of the hardware components.

[0101] (2) Robot control processing

[0102] The robot control processing will now be explained. Here, we will primarily describe the processing performed by the object recognition unit 115 and the signal generation unit 116, which are implemented within the computing unit 110 of the control device 100. Additionally, the end effector 13 assembled in the robot 10 performs prescribed processing on a workpiece W, which is an example of an object. Since the end effector 13 performs prescribed processing on the workpiece W, the workpiece W can be referred to as the processing object (or processing object).

[0103] Furthermore, the object to be processed may also be an object that is substantially the same as the object to be learned, as described later. That is, the object to be processed is not limited to an object with the same shape as the object to be learned, but may also be an object that is similar to the object to be learned to the same extent as it can be considered. The state that "the object to be processed is an object that is similar to the object to be learned to the same extent as it can be considered" may include at least one of the following states: (i) the difference between the shape of the object to be processed and the shape of the object to be learned is a state where ... object to be processed and the shape of the object to be learned are different, and the object to be processed and the object to be learned are different, and the object to be processed and the object to be learned are different, and the object to be processed and the object to be learned are different, and the object to be processed and the object to be learned are different, and the object to be processed and the object to be learned are different, and the object to be processed and the object to be learned are different, and the object to be processed and the object to be learned are different, and the object to be processed and the object to be learned are different, and the object to be processed and the object to be learned are different, and the object to be processed and the object to be learned are different, and the object

[0104] In the image represented by image data (e.g., image data IMG_2D, image data IMG_3D) generated by the imaging unit 20, for example, by photographing the workpiece W, the workpiece W, which can be referred to as the object to be processed, is captured. Therefore, the image data generated by the imaging unit 20, for example, by photographing the workpiece W, can be referred to as image data of the object to be processed.

[0105] The processed object image data may include image data (e.g., image data IMG_2D) generated by a single monocular camera (which serves as the imaging device 21) within the imaging unit 20, which captures images of the workpiece W. The processed object image data may also include image data (e.g., image data IMG_3D) generated by a stereo camera (which serves as the imaging device 22) within the imaging unit 20, which captures images of the workpiece W. Furthermore, when the stereo camera (which serves as the imaging device 22) captures images of the workpiece W, a desired projection pattern (e.g., random dots) from the projection device 23 included in the imaging unit 20 may be projected onto the workpiece W. Alternatively, when the stereo camera (which serves as the imaging device 22) captures images of the workpiece W, the desired projection pattern may not be projected onto the workpiece W.

[0106] The control device 100 acquires image data IMG_2D from the imaging device 21 using the communication device 130. Specifically, the imaging device 21 captures images of the workpiece W at a predetermined 2D imaging rate. The imaging device 21 can capture images of the workpiece W at a rate of tens to hundreds (for example, 500 times) per second. As a result, the imaging device 21 generates image data IMG_2D at a period corresponding to the predetermined 2D imaging rate. The imaging device 21 can generate tens to hundreds (for example, 500) images of IMG_2D per second. Whenever the imaging device 21 generates image data IMG_2D, the control device 100 acquires the image data IMG_2D. That is, the control device 100 can acquire tens to hundreds (for example, 500) images of IMG_2D per second.

[0107] The control device 100 then uses the communication device 130 to acquire image data IMG_3D from the imaging device 22. Specifically, the imaging device 22 captures images of the workpiece W at a predetermined 3D imaging rate. The 3D imaging rate can be the same as the 2D imaging rate, or it can be different from the 2D imaging rate. The imaging device 22 can capture images of the workpiece W at a 3D imaging rate of tens to hundreds (for example, 500 times) per second. As a result, the imaging device 22 generates image data IMG_3D at a period corresponding to the predetermined 3D imaging rate. For example, the imaging device 22 can generate tens to hundreds (for example, 500) image data IMG_3D per second. Whenever the imaging device 22 generates image data IMG_3D, the control device 100 acquires the image data IMG_3D. That is, the control device 100 can acquire tens to hundreds (for example, 500) image data IMG_3D per second.

[0108] Whenever the control device 100 acquires image data IMG_3D, the object recognition processing unit 115 generates three-dimensional position data WSD based on the acquired image data IMG_3D, representing the positions of points corresponding to each part of the workpiece W in three-dimensional space. Furthermore, the term "positions of points corresponding to each part of the workpiece W in three-dimensional space" will henceforth be appropriately referred to as "the individual three-dimensional positions of multiple points on the workpiece W".

[0109] In an image represented by image data IMG_3D, for example, a workpiece W with a projected pattern is captured. In this case, the three-dimensional shape of the workpiece W with the projected pattern is reflected in the projected pattern captured in the image represented by image data IMG_3D. Therefore, the object recognition unit 115 generates three-dimensional position data WSD based on the projected pattern captured in the image represented by image data IMG_3D.

[0110] The object recognition unit 115 can calculate the disparity by establishing correspondences between the parts (e.g., each pixel) of the images represented by the two image data contained in the image data IMG_3D. Specifically, in the correspondence establishment, the object recognition unit 115 can calculate the disparity by establishing correspondences between the parts of the projection patterns captured in the images represented by the two image data (i.e., the parts of the projection patterns captured in each image). The object recognition unit 115 can generate three-dimensional position data WSD (i.e., the three-dimensional positions of multiple points of the workpiece W can be calculated) using a well-known method based on the triangulation principle using the calculated disparity. Thus, compared to establishing correspondences between the parts of the images with projection patterns captured (i.e., the parts of the projected projection patterns captured), the accuracy of disparity calculation becomes higher. Therefore, the accuracy of the generated three-dimensional position data WSD (i.e., the accuracy of calculating the three-dimensional positions of multiple points of the workpiece W) becomes higher. Furthermore, the method for establishing the correspondence (i.e., matching method) between the parts represented by the two image data contained in the image data IMG_3D can be, for example, a known method of at least one of semi-global block matching (SGBM) and sum of absolute difference (SAD). Additionally, the parallax can be calculated by establishing the correspondence between the parts of the images not included in the projection pattern, and the three-dimensional position data WSD can be generated using a known method based on the triangulation principle using the calculated parallax. Furthermore, since the three-dimensional position data WSD represents the three-dimensional shape of the workpiece W, it can be called three-dimensional shape data. Furthermore, since the three-dimensional position data WSD represents the distance from the imaging unit 20 (e.g., the imaging device 22) to the workpiece W, it can be called distance data.

[0111] As long as the three-dimensional position data WSD can represent the three-dimensional positions of multiple points of the workpiece W, it can be any data. As an example of three-dimensional position data WSD, depth image data can be listed. Depth image data is an image that, in addition to luminance information, establishes a relationship between depth information and each pixel of the depth image represented by the depth image data. Depth information represents the distance (i.e., depth) between each part of the object captured in each pixel and the imaging device 22. Furthermore, depth image data can be an image where the luminance information of each pixel represents the depth of each part of the object (the distance between each part of the object and the imaging device 22). The object recognition processing unit 115 can calculate the distance between each part of the object captured in the image represented by the image data IMG_3D and the imaging device 22 based on the projected pattern captured in the image represented by the image data IMG_3D, and establish a relationship between the calculated distance as depth information and each pixel of the image data IMG_3D, thereby generating a depth image. As another example of three-dimensional position data WSD, point group data can be listed. Point group data is data representing the positions of points corresponding to each part of an object (e.g., workpiece W) captured in an image represented by image data IMG_3D in three-dimensional space. The object recognition processing unit 115 can generate point group data based on depth image data and camera parameters of the imaging device 22. Furthermore, three-dimensional position data WSD, representing the three-dimensional positions of multiple points of workpiece W, can also be referred to as data representing the position of workpiece W. Since the object recognition processing unit 115 uses image data IMG_3D to generate three-dimensional position data WSD, it can also be said that the position of workpiece W is calculated using image data IMG_3D. Thus, the process of calculating the positions of points corresponding to each part of an object (e.g., workpiece W) in three-dimensional space using image data IMG_3D (in other words, the three-dimensional positions of multiple points of the object) can be called "position calculation processing".

[0112] Then, the object recognition unit 115 calculates at least one of the position and orientation of the workpiece W based on the image data IMG_2D and the three-dimensional position data WSD. For example, the object recognition unit 115 can calculate at least one of the position and orientation of a representative point of the workpiece W. Examples of representative points of the workpiece W include at least one of the center of the workpiece W, the center of gravity of the workpiece W, the vertex of the workpiece W, the center of the surface of the workpiece W, and the center of gravity of the surface of the workpiece W. The representative point of the workpiece W can be referred to as a feature point of the workpiece W.

[0113] At this time, the object recognition unit 115 can calculate at least one of the position and orientation of the workpiece W in the global coordinate system. As described above, the global coordinate system is the coordinate system that serves as the reference for the robot system 1. Specifically, the global coordinate system is the coordinate system used to control the robot 10. The robot control device 14 can control the robot arm 12 so that the end effector 13 is located at the desired position in the global coordinate system. The global coordinate system can be a coordinate system defined by mutually orthogonal X-axis (GL), Y-axis (GL), and Z-axis (GL). Furthermore, "GL" is a symbol representing the global coordinate system.

[0114] The object recognition unit 115 can calculate at least one of the following: the position Tx(GL) of the workpiece W in the X-axis direction (GL) parallel to the X-axis (GL); the position Ty(GL) of the workpiece W in the Y-axis direction (GL) parallel to the Y-axis (GL); and the position Tz(GL) of the workpiece W in the Z-axis direction (GL) parallel to the Z-axis (GL), as the position of the workpiece W in the global coordinate system. The object recognition unit 115 can also calculate at least one of the following: the rotation amount Rx(GL) of the workpiece W about the X-axis (GL); the rotation amount Rey(GL) of the workpiece W about the Y-axis (GL); and the rotation amount Rz(GL) of the workpiece W about the Z-axis (GL), as the pose of the workpiece W in the global coordinate system. The rotation amounts Rx(GL) of the workpiece W about the X-axis (GL), Ry(GL) of the workpiece W about the Y-axis (GL), and Rz(GL) of the workpiece W about the Z-axis (GL) are equivalent to the parameters representing the orientation of the workpiece W about the X-axis (GL), the orientation of the workpiece W about the Y-axis (GL), and the orientation of the workpiece W about the Z-axis (GL), respectively. Therefore, in the following description, the rotation amounts Rx(GL) of the workpiece W about the X-axis (GL), Ry(GL) of the workpiece W about the Y-axis (GL), and Rz(GL) of the workpiece W about the Z-axis (GL) are referred to as the orientations Rx(GL) of the workpiece W about the X-axis (GL), Ry(GL) of the workpiece W about the Y-axis (GL), and Rz(GL) of the workpiece W about the Z-axis (GL), respectively.

[0115] Furthermore, the orientations Rx(GL) of the workpiece W about the X-axis (GL), Ry(GL) of the workpiece W about the Y-axis (GL), and Rz(GL) of the workpiece W about the Z-axis (GL) can be considered as representing the positions of the workpiece W in the rotational directions about the X-axis (GL), the Y-axis (GL), and the Z-axis (GL), respectively. That is, the orientations Rx(GL) of the workpiece W about the X-axis (GL), Ry(GL) of the workpiece W about the Y-axis (GL), and Rz(GL) of the workpiece W about the Z-axis (GL) can all be considered as parameters representing the position of the workpiece W. Therefore, orientations Rx(GL), Ry(GL), and Rz(GL) can be called positions Rx(GL), Ry(GL), and Rz(GL), respectively.

[0116] Furthermore, the object recognition unit 115 can calculate the position or orientation of the workpiece W in the global coordinate system. That is, the object recognition unit 115 can also calculate both the position and orientation of the workpiece W in the global coordinate system. In this case, the object recognition unit 115 can calculate at least one of position Tx(GL), position Ty(GL), and position Tz(GL) as the position of the workpiece W in the global coordinate system. That is, the position of the workpiece W in the global coordinate system can be represented by at least one of position Tx(GL), position Ty(GL), and position Tz(GL). Additionally, the object recognition unit 115 can also calculate at least one of orientation Rx(GL), orientation Ry(GL), and orientation Rz(GL) as the orientation of the workpiece W in the global coordinate system. That is, the orientation of the workpiece W in the global coordinate system can be represented by at least one of orientation Rx(GL), orientation Ry(GL), and orientation Rz(GL).

[0117] Then, the signal generation unit 116 generates a robot control signal based on at least one of the position and orientation of the workpiece W calculated as described above. The signal generation unit 116 can generate a robot control signal (in other words, an end effector control signal) that causes the end effector 13 assembled on the robot 10 to perform a predetermined process on the workpiece W. The signal generation unit 116 can also generate a robot control signal that causes the positional relationship between the end effector 13 assembled on the robot 10 and the workpiece W to become a desired positional relationship. The signal generation unit 116 can also generate a robot control signal for controlling the movement of the robot arm 12, causing the positional relationship between the end effector 13 assembled on the robot 10 and the workpiece W to become a desired positional relationship. The signal generation unit 116 can also generate a robot control signal (in other words, an end effector control signal) that causes the end effector 13 to perform a predetermined process on the workpiece W at the point in time when the positional relationship between the end effector 13 assembled on the robot 10 and the workpiece W becomes the desired positional relationship. The signal generation unit 116 can also generate robot control signals (i.e., end effector control signals) for controlling the movement of the end effector 13, so that the workpiece W is subjected to specified processing at the time point when the positional relationship between the end effector 13 assembled on the robot 10 and the workpiece W becomes the desired positional relationship.

[0118] Furthermore, the robot control signal may include signals that can be directly used by the robot control device 14 to control the actions of the robot 10. The robot control signal may also include signals that can be directly used as robot drive signals for controlling the actions of the robot 10 by the robot control device 14. In this case, the robot control signal may also include signals that can be directly used as robot drive signals for controlling the actions of the robot 10 by the robot control device 14. In this case, the robot control device 14 can directly use the robot control signal to control the actions of the robot 10. For example, the control device 100 may generate a drive signal for an actuator built into the joint of the robot arm 12 as a robot control signal, and the robot control device 14 can directly use the robot control signal generated by the control device 100 to control the actuator built into the joint of the robot arm 12. The robot control signal may include signals that can be directly used by the robot control device 14 to control the actions of the end effector 13. The robot control signal may also include signals that can be directly used as end effector drive signals for controlling the actions of the end effector 13 by the robot control device 14. In such cases, for example, the control device 100 may generate a drive signal (end-effector drive signal) for the actuator that moves the gripper constituting the end-effector 13 as a robot control signal, and the robot control device 14 may directly use the robot control signal generated by the control device 100 to control the actuator of the end-effector 13. For example, the control device 100 may generate a drive signal (end-effector drive signal) for driving the vacuum device of the vacuum gripper constituting the end-effector 13 as a robot control signal, and the robot control device 14 may directly use the robot control signal generated by the control device 100 to control the vacuum device of the end-effector 13.

[0119] Furthermore, as described above, if the robot control signals include signals that can be directly used by the robot control device 14 to control the actions of the robot 1 and signals that can be directly used by the robot control device 14 to control the actions of the end effector 13, the robot 10 may not include the robot control device 14.

[0120] Furthermore, the robot control signals may include signals necessary for generating robot drive signals required by the robot control device 14 to control the actions of the robot 10. Additionally, the signals necessary for generating robot drive signals required by the robot control device 14 to control the actions of the robot 10 can also be considered signals that form the basis for the robot control device 14 to control the actions of the robot 10. Furthermore, the signals necessary for generating signals required by the robot control device 14 to control the actions of the robot 10 may be signals representing at least one of the position and orientation of the workpiece W in the global coordinate system. Furthermore, the signals necessary for generating signals required by the robot control device 14 to control the actions of the robot 10 may also be signals representing the desired positional relationship between the robot 10 and the workpiece W in the global coordinate system. Furthermore, the signals necessary for generating signals required by the robot control device 14 to control the actions of the robot 10 may include signals representing at least one of the desired position and orientation of the end effector 13 in the global coordinate system. Furthermore, the signals that can be used to generate the signals required by the robot control device 14 to control the actions of the robot 10 can be, for example, signals representing at least one of the desired position and posture of the desired front end of the robot arm 12 in the global coordinate system, or signals representing at least one of the desired position and posture of the imaging unit 20 in the global coordinate system.

[0121] The signal generation unit 116 uses the communication device 130 to output the generated robot control signal to the robot 10 (e.g., the robot control device 14). As a result, the robot control device 14 controls the actions of the robot 10 (e.g., the actions of at least one of the robot arm 12 and the end effector 13) based on the robot control signal.

[0122] Subsequently, the control device 100 may repeat the aforementioned series of processes until it is determined that the robot control process is to end. That is, the control device 100 may also continue to acquire image data IMG_2D and image data IMG_3D from the imaging device 21 and imaging device 22 respectively during the period when the robot 10's movements are controlled based on the robot control signals. Since the robot 10's movements are controlled based on the robot control signals, the workpiece W moves relative to the imaging device 21 and imaging device 22 (in other words, at least one of the imaging device 21 and imaging device 22 and the workpiece W moves). At this time, the imaging device 21 and imaging device 22 may each capture images of the workpiece W during the relative movement between the workpiece W and the imaging device 21 and imaging device 22. In other words, the control device 100 may continue to perform robot control processing during the relative movement between the workpiece W and the imaging device 21 and imaging device 22. As a result, the control device 100 can also, during the period of controlling the movement of the robot 10 based on the robot control signals, calculate (i.e., update) at least one of the position and orientation of the workpiece W based on the newly acquired image data IMG_2D and image data IMG_3D. The imaging devices 21 and 22 can each capture images of the workpiece W while it is stationary and both imaging devices 21 and 22 are stationary. In other words, the control device 100 can perform robot control processing while both imaging devices 21 and 22 and the workpiece W are stationary.

[0123] Furthermore, the object recognition unit 115 can calculate at least one of the position and orientation of the workpiece W as at least one of the position and orientation of the workpiece W in the global coordinate system. The object recognition unit 115 can also calculate at least one of the position and orientation of the workpiece W as at least one of the position and orientation of the workpiece W in a coordinate system different from the global coordinate system. In this case, the signal generation unit 116 can generate a robot control signal based on at least one of the position and orientation of the workpiece W in a coordinate system different from the global coordinate system calculated by the object recognition unit 115.

[0124] (2-1) Calculation and processing of the position and orientation of the object

[0125] The object recognition unit 115 will describe the process for calculating at least one of the position and orientation of an object (e.g., workpiece W). Furthermore, for ease of explanation, the process for calculating at least one of the position and orientation of workpiece W, which is an example of an object, will be described below. The object recognition unit 115 can calculate at least one of the position and orientation of any object by performing the same operation as the process for calculating at least one of the position and orientation of workpiece W. That is, the following description of the process for calculating at least one of the position and orientation of workpiece W can be used as a description of the process for calculating at least one of the position and orientation of an object by replacing the phrase "workpiece W" with the phrase "object".

[0126] The object recognition processing unit 115 can calculate at least one of the position and orientation of the workpiece W by performing matching processing using image data IMG_2D, image data IMG_3D, and three-dimensional position data WSD, and tracking processing using image data IMG_2D, image data IMG_3D, and three-dimensional position data WSD.

[0127] (2-2) 2D Matching Processing

[0128] Matching using image data IMG_2D can be referred to as 2D matching. Furthermore, the matching process itself can be the same as existing matching processes. Therefore, detailed descriptions of the matching process are omitted, and its overview is provided below.

[0129] The object recognition unit 115 can, for example, perform matching processing on the image data IMG_2D using the workpiece W captured in a two-dimensional image represented by the two-dimensional model data IMG_2M as a template. Here, the two-dimensional model data IMG_2M is data representing a two-dimensional model of the workpiece W. The two-dimensional model data IMG_2M is data representing a two-dimensional model having a reference two-dimensional shape of the workpiece W. In this embodiment, the two-dimensional model data IMG_2M is image data representing a reference two-dimensional image of the workpiece W. More specifically, the two-dimensional model data IMG_2M is image data representing a two-dimensional image containing a two-dimensional model of the workpiece W. The two-dimensional model data IMG_2M is image data containing a two-dimensional image of a two-dimensional model having a reference two-dimensional shape of the workpiece W. For example, the two-dimensional model data IMG_2M can be two-dimensional image data representing multiple two-dimensional models of the workpiece W generated by virtually projecting a three-dimensional model of the workpiece W (e.g., a CAD model created using Computer-Aided Design, CAD) from multiple different directions onto virtual planes orthogonal to multiple different directions. Furthermore, the two-dimensional model data IMG_2M can also be image data representing a two-dimensional image obtained by photographing the actual workpiece W in advance. In this case, the two-dimensional model data IMG_2M can also be image data representing multiple two-dimensional images generated by photographing the actual workpiece W from multiple different shooting directions. Furthermore, the three-dimensional model of the workpiece W can be a three-dimensional model having the same shape as the three-dimensional shape of the workpiece W obtained by measuring the actual workpiece W in advance. Furthermore, the actual workpiece W for which shape measurement has been performed in advance can also be a reference or a good workpiece W. The two-dimensional model data IMG_2M can also be image data representing a two-dimensional image generated by photographing the actual workpiece W in advance. In this case, the two-dimensional model data IMG_2M can also be image data representing multiple two-dimensional images generated by photographing the actual workpiece W from multiple different shooting directions. The workpiece W photographed in the two-dimensional image represented by the image data IMG_2D, which is the two-dimensional model data IMG_2M, can also be called the two-dimensional model of the workpiece W. Furthermore, the actual workpiece W photographed in advance can also be a reference or a good workpiece W.

[0130] The object recognition unit 115 can, for example, move, enlarge, reduce, and / or rotate the workpiece W (a two-dimensional model of workpiece W) captured in the two-dimensional image represented by the two-dimensional model data IMG_2M, so that the feature parts (e.g., at least one of feature points and edges) of the workpiece W (a two-dimensional model of workpiece W) captured in the two-dimensional image represented by the two-dimensional model data IMG_2M are close to (typically identical to) the feature parts of the workpiece W captured in the image represented by the image data IMG_2D. That is, the object recognition unit 115 can change the positional relationship between the coordinate system of the two-dimensional model data IMG_2M (e.g., the coordinate system of the CAD model) and the 2D shooting coordinate system based on the shooting device 21 that captures the workpiece W, so that the feature parts of the workpiece W (a two-dimensional model of workpiece W) captured in the two-dimensional image represented by the two-dimensional model data IMG_2M are close to (typically identical to) the feature parts of the workpiece W captured in the image represented by the image data IMG_2D.

[0131] As a result, the object recognition unit 115 can determine the positional relationship between the coordinate system of the two-dimensional model data IMG_2M and the 2D shooting coordinate system. Then, based on the positional relationship between the coordinate system of the two-dimensional model data IMG_2M and the 2D shooting coordinate system, the object recognition unit 115 can calculate at least one of the position and orientation of the workpiece W in the 2D shooting coordinate system, according to at least one of the position and orientation of the workpiece W in the coordinate system of the two-dimensional model data IMG_2M. Here, the 2D shooting coordinate system is a coordinate system defined by mutually orthogonal X-axis (2D), Y-axis (2D), and Z-axis (2D). At least one of the X-axis (2D), Y-axis (2D), and Z-axis (2D) can be an axis along the optical axis of the optical system (especially terminal optical elements such as objective lenses) included in the shooting device 21. Furthermore, the optical axis of the optical system included in the shooting device 21 can be regarded as the optical axis of the shooting device 21. In addition, "2D" is a symbol representing the 2D shooting coordinate system.

[0132] The object recognition unit 115 can calculate at least one of the following: the position Tx (2D) of the workpiece W in the X-axis direction (2D), the position Ty (2D) of the workpiece W in the Y-axis direction, and the position Tz (2D) of the workpiece W in the Z-axis direction, as the position of the workpiece W in the 2D imaging coordinate system. The object recognition unit 115 can also calculate at least one of the following: the pose Rx (2D) of the workpiece W around the X-axis (2D), the pose Ry (2D) of the workpiece W around the Y-axis (2D), and the pose Rz (2D) of the workpiece W around the Z-axis (2D), as the pose of the workpiece W in the 2D imaging coordinate system. Furthermore, the object recognition unit 115 can calculate either the position or the pose of the workpiece W in the 2D imaging coordinate system. That is, the object recognition unit 115 can calculate at least one of the position and the pose of the workpiece W in the 2D imaging coordinate system.

[0133] Furthermore, the object recognition processing unit 115 can change the positional relationship between the coordinate system of the two-dimensional model data IMG_2M (e.g., the coordinate system of the CAD model) and the 2D shooting coordinate system based on the shooting device 21 that takes pictures of the workpiece W, so that the feature parts in a part of the workpiece W (two-dimensional model of workpiece W) captured in the two-dimensional image represented by the two-dimensional model data IMG_2M are close to (typically consistent with) the feature parts in a part of the workpiece W captured in the image represented by the image data IMG_2D.

[0134] Furthermore, the method for calculating at least one of the position and orientation of workpiece W is not limited to the method described above, and may also be other well-known methods for calculating at least one of the position and orientation of workpiece W based on image data IMG_2D. As a matching process using image data IMG_2D, well-known methods such as Scale-Invariant Feature Transform (SIFT) and Speed-Upped Robust Feature (SURF) may be used.

[0135] Furthermore, the object recognition unit 115 can calculate at least one of the position and orientation of the workpiece W without using the two-dimensional model data IMG_2M. The object recognition unit 115 can also calculate at least one of the position and orientation of the workpiece W based on feature parts (feature points, edges, and at least one mark on the workpiece) of the workpiece W captured in the image represented by image data IMG_2D. Furthermore, the object recognition unit 115 can also calculate at least one of the position and orientation of the workpiece W based on multiple feature parts (multiple feature points, multiple edges, and at least one mark on the workpiece) of the workpiece W captured in the image represented by image data IMG_2D.

[0136] (2-3) 3D Matching Processing

[0137] Matching processes using WSD (3D position data) can be referred to as 3D matching processes. Furthermore, the matching process itself can be identical to existing matching processes. Therefore, detailed descriptions of the matching process are omitted, and its overview is provided below.

[0138] The object recognition unit 115 can also perform matching processing on the three-dimensional position data WSD using the workpiece W represented by the three-dimensional model data WMD as a template. Here, the three-dimensional model data WMD is data representing the three-dimensional model of the workpiece W. That is, the three-dimensional model data WMD is data representing the reference three-dimensional shape of the workpiece W. The three-dimensional model can be a CAD model of the workpiece W, which is an example of the three-dimensional model of the workpiece W. The three-dimensional model can be a three-dimensional model having the same shape as the three-dimensional shape of the workpiece W obtained by measuring the three-dimensional shape of the actual workpiece W in advance. In this case, the three-dimensional model data WMD can be generated in advance based on the image data IMG_3D generated by the imaging device 22 photographing the workpiece W on which a projection pattern from the projection device 23 is projected. Alternatively, the three-dimensional model data WMD can also be generated in advance by shape measurement using a well-known three-dimensional shape measuring device different from the robot system 1. In this case, the three-dimensional model data WMD can be depth image data representing the three-dimensional model of the workpiece W. The three-dimensional model data WMD can be point group data representing the three-dimensional model of the workpiece W. In addition, the actual workpiece W, which was photographed or measured in advance to generate the 3D model data WMD, can also be a reference or a good workpiece W.

[0139] The object recognition unit 115 can, for example, move, enlarge, reduce, and / or rotate the workpiece W represented by the 3D model data WMD, so that the feature parts of the workpiece W represented by the 3D model data WMD are close to (typically identical to) the feature parts of the workpiece W represented by the 3D position data WSD. That is, the object recognition unit 115 can change the positional relationship between the coordinate system of the 3D model data WMD (e.g., the coordinate system of the CAD model) and the 3D imaging coordinate system based on the imaging device 22 that has photographed the workpiece W, so that the feature parts of the workpiece W represented by the 3D position data WSD are close to (typically identical to) the feature parts of the workpiece W represented by the 3D model data WMD.

[0140] As a result, the object recognition unit 115 can determine the positional relationship between the coordinate system of the 3D model data WMD and the 3D shooting coordinate system. Then, based on the positional relationship between the coordinate system of the 3D model data WMD and the 3D shooting coordinate system, the object recognition unit 115 can calculate at least one of the position and orientation of the workpiece W in the 3D shooting coordinate system, according to at least one of the position and orientation of the workpiece W in the coordinate system of the 3D model data WMD. Here, the 3D shooting coordinate system is a coordinate system defined by mutually orthogonal X-axis (3D), Y-axis (3D), and Z-axis (3D). At least one of the X-axis (3D), Y-axis (3D), and Z-axis (3D) can be an axis along the optical axis of the optical system (especially terminal optical elements such as objective lenses) included in the shooting device 22. Furthermore, the optical axis of the optical system included in the shooting device 22 can be regarded as the optical axis of the shooting device 22. In the case that the shooting device 22 is a stereo camera including two monocular cameras, the optical axis can be the optical axis of the optical system included in either of the two monocular cameras. That is, the optical axis can be the optical axis of either of the two SLR cameras. Furthermore, "3D" is a symbol representing the coordinate system used for 3D photography.

[0141] The object recognition unit 115 can calculate at least one of the following: the position Tx(3D) of the workpiece W in the X-axis direction (3D), the position Ty(3D) of the workpiece W in the Y-axis direction (3D), and the position Tz(3D) of the workpiece W in the Z-axis direction (3D) as the position of the workpiece W in the 3D imaging coordinate system. The object recognition unit 115 can also calculate at least one of the following: the pose Rx(3D) of the workpiece W around the X-axis (3D), the pose Ry(3D) of the workpiece W around the Y-axis (3D), and the pose Rz(3D) of the workpiece W around the Z-axis (3D) as the pose of the workpiece W in the 3D imaging coordinate system. Furthermore, the object recognition unit 115 can calculate either the position or the pose of the workpiece W in the 3D imaging coordinate system. That is, the object recognition unit 115 can calculate at least one of the position and the pose of the workpiece W in the 3D imaging coordinate system.

[0142] Furthermore, the object recognition processing unit 115 can change the positional relationship between the coordinate system of the three-dimensional model data WMD (e.g., the coordinate system of the CAD model) and the 3D shooting coordinate system based on the shooting device 22 that has taken a picture of the workpiece W, so that the feature part in a part of the workpiece W represented by the three-dimensional position data WSD is close to (typically consistent with) the feature part in a part of the workpiece W represented by the three-dimensional model data WMD.

[0143] Furthermore, the method for calculating the position of workpiece W is not limited to the method described above, and other well-known methods for calculating the position of workpiece W based on three-dimensional position data WSD can also be used. As a matching process using three-dimensional position data WSD, well-known methods such as at least one of Random Sample Consensus (RANSAC), SIFT (Scale Invariant Feature Transform), Iterative Closest Point (ICP), and Direct Sparse Odometry (DSO) can be used.

[0144] Furthermore, the object recognition unit 115 can calculate at least one of the position and orientation of the workpiece W without using the three-dimensional model data WMD. For example, the object recognition unit 115 can also calculate at least one of the position and orientation of the workpiece W based on the feature parts (feature points, edges, and marks on the workpiece) represented by the three-dimensional position data WSD. Furthermore, the object recognition unit 115 can also calculate at least one of the position and orientation of the workpiece W based on multiple feature parts (multiple feature points, multiple edges, and multiple marks on the workpiece) represented by the three-dimensional position data WSD.

[0145] (2-4) 2D tracking processing

[0146] The tracking process using two image data points, IMG_2D#t1 and IMG_2D#t2, generated by capturing images of the workpiece W at different times t1 and t2 by the imaging device 21, can be referred to as 2D tracking processing. Furthermore, time t2 is set to a time after time t1. Moreover, the 2D tracking process itself can be the same as existing tracking processes. Therefore, a detailed description of the tracking process is omitted, and its outline is described below.

[0147] The object recognition unit 115 can track at least one feature part (e.g., at least one feature point and edge) that is identical to at least one feature part of the workpiece W captured in image data IMG_2D#t1 within image data IMG_2D#t2. That is, the object recognition unit 115 can calculate the change in at least one of the position and orientation of the at least one feature part in the 2D imaging coordinate system between time t1 and time t2. Then, the object recognition unit 115 can calculate the change in at least one of the position and orientation of the workpiece W in the 2D imaging coordinate system between time t1 and time t2 based on the change in the position and orientation of the at least one feature part in the 2D imaging coordinate system.

[0148] The object recognition unit 115 can calculate at least one of the following as the change in position of workpiece W in the 2D imaging coordinate system: the change in position Tx (2D) of workpiece W in the X-axis direction (2D) ΔTx (2D), the change in position Ty (2D) of workpiece W in the Y-axis direction ΔTy (2D), and the change in position Tz (2D) of workpiece W in the Z-axis direction ΔTz (2D). The object recognition unit 115 can also calculate at least one of the following as the change in position of workpiece W in the 2D imaging coordinate system: the change in posture Rx (2D) of workpiece W around the X-axis (2D) ΔRx (2D), the change in posture Ry (2D) of workpiece W around the Y-axis (2D) ΔRy (2D), and the change in posture Rz (2D) of workpiece W around the Z-axis (2D).

[0149] Furthermore, the method for calculating the change in at least one of the position and orientation of workpiece W is not limited to the method described above, but may also be other well-known methods for calculating the change in at least one of the position and orientation of workpiece W using two image data IMG_2D#t1 and IMG_2D#t2.

[0150] (2-5) 3D tracking processing

[0151] The tracking process using two three-dimensional position data WSDPD#s1 and WSDPD#s2, corresponding to different times s1 and s2 respectively, can be called 3D tracking process. Here, the two three-dimensional position data WSDPD#s1 and WSDPD#s2 are generated from two image data IMG_3D#s1 and IMG_3D#s2 generated by the imaging device 22 capturing images of the workpiece W at different times s1 and s2 respectively. Furthermore, time s2 is set to a time after time s1. In addition, the tracking process itself can be the same as existing tracking processes. Therefore, a detailed description of the tracking process is omitted, and its outline is described below.

[0152] The object recognition unit 115 can track at least one feature part (e.g., at least one of feature points and edges) that is identical to at least one feature part of the workpiece W represented by the three-dimensional position data WSDPD#s2. That is, the object recognition unit 115 can calculate the change in at least one of the position and orientation of the at least one feature part in the 3D imaging coordinate system between time s1 and time s2. Then, the object recognition unit 115 can calculate the change in at least one of the position and orientation of the workpiece W in the 3D imaging coordinate system between time s1 and time s2 based on the change in the position and orientation of the at least one feature part in the 3D imaging coordinate system.

[0153] The object recognition unit 115 can calculate at least one of the following as the change in position Tx(3D) of workpiece W in the X-axis direction (3D), the change in position Ty(3D) of workpiece W in the Y-axis direction (3D), and the change in position Tz(3D) of workpiece W in the Z-axis direction (3D): ΔTx(3D), as the change in position of workpiece W in the 3D imaging coordinate system. The object recognition unit 115 can also calculate at least one of the following as the change in posture Rx(3D) of workpiece W around the X-axis (3D), the change in posture Ry(3D) of workpiece W around the Y-axis (3D), and the change in posture Rz(3D) of workpiece W around the Z-axis (3D): ΔRz(3D), as the change in posture of workpiece W in the 3D imaging coordinate system.

[0154] Furthermore, the method for calculating the change in at least one of the position and orientation of workpiece W is not limited to the method described above, but may also be other well-known methods for calculating the change in at least one of the position and orientation of workpiece W using two three-dimensional position data WSDPD#s1 and WSDPD#s2.

[0155] The method for calculating the change in at least one of the position and orientation of workpiece W is not limited to the method described above, and may also be other well-known methods for calculating the change in at least one of the position and orientation of workpiece W using two three-dimensional position data WSDPD#s1 and WSDPD#s2. As a tracking process using two three-dimensional position data WSDPD#s1 and WSDPD#s2, at least one of the well-known methods RANSAC (Random Sample Consensus), SIFT (Scale Invariant Feature Transform), ICP (Iterative Closest Point), and DSO (Direct Sparse Odometry) can be used.

[0156] (2-6) Processing the actions of the object recognition unit

[0157] As an example, the object recognition unit 115 calculates at least one of the position and pose of the workpiece W in the global coordinate system based on the results of 2D matching processing, 3D matching processing, 2D tracking processing, and 3D tracking processing. For example, the object recognition unit 115 can calculate positions Tx(GL), Ty(GL), and Tz(GL), as well as poses Rx(GL), Ry(GL), and Rz(GL), as the position and pose of the workpiece W in the global coordinate system. That is, the object recognition unit 115 can calculate all six degrees of freedom (DoF) positions and poses of the workpiece W in the global coordinate system as the position and pose of the workpiece W.

[0158] Furthermore, the object recognition unit 115 may not perform all of the 2D matching, 3D matching, 2D tracking, and 3D tracking processes. For example, the object recognition unit 115 may calculate at least one of the position and orientation of the workpiece W based on the results of the 2D matching and 2D tracking processes. For example, the object recognition unit 115 may calculate at least one of the position and orientation of the workpiece W based on the results of the 3D matching and 3D tracking processes. For example, the object recognition unit 115 may calculate at least one of the position and orientation of the workpiece W based on either the results of the 2D matching or 3D matching processes. For example, the object recognition unit 115 may calculate at least one of the position and orientation of the workpiece W based on the results of the 2D matching and 3D matching processes.

[0159] Furthermore, the object recognition unit 115 can also calculate all of the 6DoF positions and poses of the workpiece W in the global coordinate system. That is, the object recognition unit 115 can also calculate at least one of the 6DoF positions and poses of the workpiece W in the global coordinate system. The object recognition unit 115 can calculate at least one of position Tx(GL), position Ty(GL), and position Tz(GL) as the position of the workpiece W in the global coordinate system. The object recognition unit 115 can also calculate at least one of pose Rx(GL), pose Ry(GL), and pose Rz(GL) as the pose of the workpiece W in the global coordinate system.

[0160] Furthermore, position Tx(GL) is the position of workpiece W in the X-axis direction, parallel to the X-axis of the global coordinate system. Position Ty(GL) is the position of workpiece W in the Y-axis direction, parallel to the Y-axis of the global coordinate system. Position Tz(GL) is the position of workpiece W in the Z-axis direction, parallel to the Z-axis of the global coordinate system. Pose Rx is the pose of workpiece W about the X-axis of the global coordinate system. Pose Ry is the pose of workpiece W about the Y-axis of the global coordinate system. Pose Rz is the pose of workpiece W about the Z-axis of the global coordinate system.

[0161] In order to calculate at least one of the position and orientation of the workpiece W in the global coordinate system, the object recognition unit 115 can correct the result of the 2D matching process based on the result of the 2D tracking process. Similarly, the object recognition unit 115 can correct the result of the 3D matching process based on the result of the 3D tracking process.

[0162] Specifically, the object recognition unit 115 can calculate the change amount ΔTx(2D), the change amount ΔTy(2D), and the change amount ΔRz(2D) during 2D tracking processing. As an example of the result of 2D tracking processing, the object recognition unit 115 can generate information about the change amount ΔTx(2D), the change amount ΔTy(2D), and the change amount ΔRz(2D).

[0163] The object recognition unit 115 can calculate the position Tx (2D), position Ty (2D), and pose Rz (2D) from the result of the 2D matching process. As an example of the result of the 2D matching process, the object recognition unit 115 can generate information about the position Tx (2D), position Ty (2D), and pose Rz (2D).

[0164] The object recognition unit 115 can calculate the position Tx'(2D) of the workpiece W in the X-axis direction, which is parallel to the X-axis of the 2D shooting coordinate system, by correcting the position Tx(2D) based on the change amount ΔTx(2D). The object recognition unit 115 can calculate the position Ty'(2D) of the workpiece W in the Y-axis direction, which is parallel to the Y-axis of the 2D shooting coordinate system, by correcting the position Ty(2D) based on the change amount ΔTy(2D). The object recognition unit 115 can calculate the posture Rz'(2D) of the workpiece W around the Z-axis of the 2D shooting coordinate system by correcting the posture Rz(2D) based on the change amount ΔRz(2D). Furthermore, the process of correcting the result of the 2D matching process based on the result of the 2D tracking process may also include adding the result of the 2D tracking process to the result of the 2D matching process.

[0165] The object recognition unit 115 can calculate the change amount ΔTz(3D), change amount ΔRx(3D), and change amount ΔRy(3D) during 3D tracking processing. As an example of the result of 3D matching processing, the object recognition unit 115 can generate information about the change amount ΔTz(3D), change amount ΔRx(3D), and change amount ΔRy(3D).

[0166] The object recognition unit 115 can calculate the position Tz (3D), pose Rx (3D), and pose Ry (3D) in the 3D matching process. As an example of the result of the 3D matching process, the object recognition unit 115 can generate information about the position Tz (3D), pose Rx (3D), and pose Ry (3D).

[0167] The object recognition unit 115 can calculate the position Tz'(3D) of the workpiece W in the Z-axis direction, which is parallel to the Z-axis of the 3D shooting coordinate system, by correcting the position Tz(3D) based on the change amount ΔTz(3D). The object recognition unit 115 can calculate the posture Rx'(3D) of the workpiece W around the X-axis of the 3D shooting coordinate system by correcting the posture Rx(3D) based on the change amount ΔRx(3D). The object recognition unit 115 can calculate the posture Ry'(3D) of the workpiece W around the Y-axis of the 3D shooting coordinate system by correcting the posture Ry(3D) based on the change amount ΔRy(3D). Furthermore, the process of correcting the result of the 3D matching process based on the result of the 3D tracking process may also include the process of adding the result of the 3D tracking process to the result of the 3D matching process.

[0168] The object recognition unit 115 can calculate the positions Tx(GL), Ty(GL), Tz(GL), Rx(GL), Ry(GL), and Rz(GL) of the workpiece W in the global coordinate system based on the positions Tx'(2D), Ty'(2D), Rz'(2D), Tz'(3D), Rx'(3D), and Ry'(3D). Specifically, firstly, the object recognition unit 115 converts the positions Tx'(2D), Ty'(2D), Rz'(2D), Tz'(3D), Rx'(3D), and Ry'(3D) into positions within a common coordinate system that serves as either the 2D or 3D shooting coordinate system. Furthermore, any coordinate system different from the 2D or 3D shooting coordinate system can be used as the common coordinate system.

[0169] As an example, when the 2D shooting coordinate system is used as the common coordinate system, the position Tx' (2D), position Ty' (2D), and pose Rz' (2D) already represent the position and pose within the 2D shooting coordinate system. Therefore, the object recognition unit 115 can avoid transforming the position Tx' (2D), position Ty' (2D), and pose Rz' (2D). On the other hand, the object recognition unit 115 transforms the position Tz' (3D), pose Rx' (3D), and pose Ry' (3D) into the position Tz' (2D) along the Z-axis of the 2D shooting coordinate system, the pose Rx' (2D) around the X-axis of the 2D shooting coordinate system, and the pose Ry' (2D) around the Y-axis of the 2D shooting coordinate system. The object recognition unit 115 can use a transformation matrix M for converting the position and pose within the 3D shooting coordinate system to the position and pose within the 2D shooting coordinate system. 32 The positions Tz' (3D), poses Rx' (3D), and poses Ry' (3D) are transformed into positions Tz' (2D), poses Rx' (2D), and poses Ry' (2D). Furthermore, the transformation matrix M... 32 It can be calculated using mathematical methods based on external parameters representing the positional relationship between the shooting device 21 and the shooting device 22.

[0170] Furthermore, in the example described above, position Tx (2D), position Ty (2D), and pose Rz (2D) are calculated through 2D matching processing, and the changes in position Tx (2D) ΔTx (2D), position Ty (2D) ΔTy (2D), and pose Rz (2D) ΔRz (2D) are calculated through 2D tracking processing. However, the 2D matching processing is not limited to position Tx (2D), position Ty (2D), and pose Rz (2D); at least one of position Tx (2D), position Ty (2D), position Tz (2D), pose Rx (2D), pose Ry (2D), and pose Rz (2D) can be calculated. Not limited to the changes ΔTx(2D), ΔTy(2D), and ΔRz(2D), at least one of the changes ΔTx(2D) of position Tx(2D), ΔTy(2D) of position Ty(2D), ΔTz(2D) of position Tz(2D), ΔRx(2D) of posture Rx(2D), ΔRy(2D) of posture Ry(2D), and ΔRz(2D) of posture Rz(2D) can be calculated in 2D tracking processing.

[0171] In the example described, position Tz(3D), pose Rx(3D), and pose Ry(3D) are calculated through 3D matching processing, and the changes in position Tz(3D), pose Rx(3D), and pose Ry(3D) are calculated through 3D tracking processing. However, the 3D matching processing is not limited to position Tz(3D), position Rx(3D), and pose Ry(3D); at least one of position Tx(3D), position Ty(3D), position Tz(3D), pose Rx(3D), pose Ry(3D), and pose Rz(3D) can be calculated. In 3D tracking processing, not limited to the change amount ΔTz(3D), change amount ΔRx(3D), and change amount ΔRy(3D), at least one of the change amount ΔTx(3D) of position Tx(3D), change amount ΔTy(3D) of position Ty(3D), change amount ΔTz(3D) of position Tz(3D), change amount ΔRx(3D) of pose Rx(3D), change amount ΔRy(3D) of pose Ry(3D), and change amount ΔRz(3D) of pose Rz(3D) can be calculated.

[0172] The object recognition unit 115 can calculate the 6DoF position and pose of the workpiece W in the global coordinate system based on the 6DoF position and pose of the workpiece W in the 2D imaging coordinate system, which is a common coordinate system. That is, the object recognition unit 115 can calculate the position Tx(GL), position Ty(GL), position Tz(GL), pose Rx(GL), pose Ry(GL), and pose Rz(GL) of the workpiece W in the global coordinate system based on the 6DoF position and pose of the workpiece W in the 2D imaging coordinate system.

[0173] The object recognition unit 115 can also use a transformation matrix M for converting the position and pose in the 2D shooting coordinate system to the position and pose in the global coordinate system. 2G The 6DoF position and pose of the workpiece W in the 2D imaging coordinate system are converted into the 6DoF position and pose of the workpiece W in the global coordinate system. The object recognition unit 115 can use the position Tx' (2D) in the 2D imaging coordinate system and the transformation matrix M 2G According to "Tx(GL)=M" 2G The position Tx(GL) in the global coordinate system is calculated using the formula ·Tx'(2D)”.

[0174] Transformation matrix M 2G For example, it may include a product of transformation matrices reflecting the changes in the position coordinates of the imaging device 21 or imaging device 22 caused by the rotation of the links about the axis specified by the joints of the robotic arm 12. The transformation matrix may be a so-called rotation matrix, a rotation matrix containing a parallel component, or a matrix based on Euler angles. Furthermore, regarding the coordinate transformation of the robotic arm itself using the transformation matrix, existing transformation methods can be used, and therefore their detailed description is omitted.

[0175] Furthermore, the object recognition unit 115 may not need to be converted to a reference coordinate system (e.g., a 2D camera coordinate system). The object recognition unit 115 can use a transformation matrix M. 2G The positions Tx' (2D), Ty' (2D), and pose Rz' (2D) are converted into the positions Tx(GL), Ty(GL), and pose Rz(GL) of the workpiece W in the global coordinate system. The object recognition unit 115 can, for example, use a transformation matrix M for converting the positions and poses in the 3D imaging coordinate system to the positions and poses in the global coordinate system. 3G The positions Tz' (3D), poses Rx' (3D), and poses Ry' (3D) are converted into the positions Tz (GL), poses Rx (GL), and poses Ry (GL) of the workpiece W in the global coordinate system.

[0176] Furthermore, the 2D shooting coordinate system, the 3D shooting coordinate system, and the common coordinate system are examples of the aforementioned "coordinate system different from the global coordinate system".

[0177] Furthermore, the object recognition unit 115 may not perform 3D matching processing and 3D tracking processing (i.e., the object recognition unit 115 may perform 2D matching processing and 2D tracking processing).

[0178] Specifically, the object recognition unit 115 can calculate position Tx (2D), position Ty (2D), position Tz (2D), pose Rx (2D), pose Ry (2D), and pose Rz (2D) from the result of 2D matching processing. As an example of the result of 2D matching processing, the object recognition unit 115 can generate information about position Tx (2D), position Ty (2D), position Tz (2D), pose Rx (2D), pose Ry (2D), and pose Rz (2D).

[0179] In 2D tracking processing, the object recognition unit 115 can calculate the changes in position Tx (2D) ΔTx (2D), position Ty (2D) ΔTy (2D), position Tz (2D) ΔTz, pose Rx (2D) ΔRx (2D), pose Ry (2D) ΔRy (2D), and pose Rz (2D) ΔRz (2D). As an example of the result of 2D tracking processing, the object recognition unit 115 can generate information about the changes ΔTx (2D), ΔTy (2D), ΔTz (2D), ΔRx (2D), ΔRy (2D), and ΔRz (2D).

[0180] The object recognition unit 115 can calculate the position Tx'(2D) of the workpiece W in the X-axis direction, which is parallel to the X-axis of the 2D shooting coordinate system, by correcting the position Tx(2D) based on the change amount ΔTx(2D). The object recognition unit 115 can calculate the position Ty'(2D) of the workpiece W in the Y-axis direction, which is parallel to the Y-axis of the 2D shooting coordinate system, by correcting the position Ty(2D) based on the change amount ΔTy(2D). The object recognition unit 115 can calculate the position Tz'(2D) of the workpiece W in the Z-axis direction, which is parallel to the Z-axis of the 2D shooting coordinate system, by correcting the position Tz(2D) based on the change amount ΔTz(2D). The object recognition unit 115 can calculate the posture Rx'(2D) of the workpiece W around the X-axis of the 2D shooting coordinate system by correcting the posture Rx(2D) based on the change amount ΔRx(2D). The object recognition unit 115 can calculate the pose Ry'(2D) of the workpiece W about the Y-axis of the 2D shooting coordinate system by correcting the pose Ry(2D) based on the change amount ΔRy(2D). The object recognition unit 115 can calculate the pose Rz'(2D) of the workpiece W about the Z-axis of the 2D shooting coordinate system by correcting the pose Rz(2D) based on the change amount ΔRz(2D).

[0181] The object recognition unit 115 can use a transformation matrix M 2G (That is, a transformation matrix used to convert the position and pose in the 2D shooting coordinate system into the position and pose in the global coordinate system), for example, converting the position Tx'(2D), position Ty'(2D), position Tz'(2D), pose Rx'(2D), pose Ry'(2D), and pose Rz'(2D) of the workpiece W in the global coordinate system into the position Tx(GL), position Ty(GL), position Tz(GL), pose Rx(GL), pose Ry(GL), and pose Rz(GL) of the workpiece W.

[0182] Furthermore, the object recognition unit 115 can calculate at least one of the following in the global coordinate system: position Tx(GL), position Ty(GL), position Tz(GL), pose Rx(GL), pose Ry(GL), and pose Rz(GL) of the workpiece W: based on the results of the 2D matching process and the 2D tracking process, respectively.

[0183] Thus, the object recognition unit 115 can calculate at least one of the position and orientation of the workpiece W in the global coordinate system based on the results of 2D matching processing and 2D tracking processing.

[0184] Furthermore, the object recognition unit 115 may not perform 2D tracking processing in addition to not performing 3D matching processing and 3D tracking processing (i.e., the object recognition unit 115 may perform 2D matching processing).

[0185] Specifically, the object recognition unit 115 can calculate the position Tx (2D), position Ty (2D), position Tz (2D), pose Rx (2D), pose Ry (2D), and pose Rz (2D) from the result of the 2D matching process. As an example of the result of the 2D matching process, the object recognition unit 115 can generate information about the position Tx (2D), position Ty (2D), position Tz (2D), pose Rx (2D), pose Ry (2D), and pose Rz (2D).

[0186] The object recognition unit 115 can use a transformation matrix M 2GThe positions Tx (2D), Ty (2D), Tz (2D), and poses Rx (2D), Ry (2D), and Rz (2D) are converted into the positions Tx (GL), Ty (GL), Tz (GL), Rx (GL), Ry (GL), and Rz (GL) of the workpiece W in the global coordinate system. Furthermore, the object recognition unit 115 can calculate at least one of the positions Tx (GL), Ty (GL), Tz (GL), Rx (GL), Ry (GL), and Rz (GL) of the workpiece W in the global coordinate system based on the result of the 2D matching process. Thus, the object recognition unit 115 can calculate at least one of the positions and poses of the workpiece W in the global coordinate system based on the result of the 2D matching process. In this case, the imaging unit 20 may not include the imaging device 22.

[0187] Furthermore, the object recognition unit 115 may not perform 2D matching processing and 2D tracking processing (i.e., the object recognition unit 115 may perform 3D matching processing and 3D tracking processing).

[0188] Specifically, the object recognition unit 115 can calculate position Tx (3D), position Ty (3D), position Tz (3D), pose Rx (3D), pose Ry (3D), and pose Rz (3D) in the 3D matching process. As an example of the result of the 3D matching process, the object recognition unit 115 can generate information about position Tx (3D), position Ty (3D), position Tz (3D), pose Rx (3D), pose Ry (3D), and pose Rz (3D).

[0189] In 3D tracking processing, the object recognition unit 115 can calculate the change in position Tx (3D) ΔTx (3D), the change in position Ty (3D) ΔTy (3D), the change in position Tz ΔTz (3D), the change in pose Rx (3D) ΔRx (3D), the change in pose Ry (3D) ΔRy (3D), and the change in pose Rz (3D) ΔRz (3D). As an example of the result of 3D tracking processing, the object recognition unit 115 can generate information about the changes ΔTx (3D), ΔTy (3D), ΔTz (3D), ΔRx (3D), ΔRy (3D), and Rz (3D).

[0190] The object recognition unit 115 can calculate the position Tx'(3D) of the workpiece W in the X-axis direction, which is parallel to the X-axis of the 3D shooting coordinate system, by correcting the position Tx(3D) based on the change amount ΔTx(3D). The object recognition unit 115 can calculate the position Ty'(3D) of the workpiece W in the Y-axis direction, which is parallel to the Y-axis of the 3D shooting coordinate system, by correcting the position Ty(3D) based on the change amount ΔTy(3D). The object recognition unit 115 can calculate the position Tz'(3D) of the workpiece W in the Z-axis direction, which is parallel to the Z-axis of the 3D shooting coordinate system, by correcting the position Tz(3D) based on the change amount ΔTz(3D). The object recognition unit 115 can calculate the posture Rx'(3D) of the workpiece W around the X-axis of the 3D shooting coordinate system by correcting the posture Rx(3D) based on the change amount ΔRx(3D). The object recognition unit 115 can calculate the pose Ry'(3D) of the workpiece W about the Y-axis of the 3D imaging coordinate system by correcting the pose Ry(3D) based on the change amount ΔRy(3D). The object recognition unit 115 can calculate the pose Rz'(3D) of the workpiece W about the Z-axis of the 3D imaging coordinate system by correcting the pose Rz(3D) based on the change amount ΔRz(3D).

[0191] The object recognition unit 115 can use a transformation matrix M 3G (That is, the transformation matrix used to convert the position and pose in the 3D shooting coordinate system into the position and pose in the global coordinate system), converting the position Tx'(3D), position Ty'(3D), position Tz'(3D), pose Rx'(3D), pose Ry'(3D), and pose Rz'(3D) of the workpiece W in the global coordinate system into the position Tx(GL), position Ty(GL), position Tz(GL), pose Rx(GL), pose Ry(GL), and pose Rz(GL).

[0192] Furthermore, the object recognition unit 115 can calculate at least one of the following positions of the workpiece W in the global coordinate system: position Tx(GL), position Ty(GL), position Tz(GL), pose Rx(GL), pose Ry(GL), and pose Rz(GL) based on the results of the 3D matching process and the 3D tracking process, respectively.

[0193] Thus, the object recognition unit 115 can calculate at least one of the position and orientation of the workpiece W in the global coordinate system based on the results of 3D matching processing and 3D tracking processing.

[0194] Furthermore, the object recognition unit 115 may not perform 3D tracking processing in addition to not performing 2D matching processing and 2D tracking processing (i.e., the object recognition unit 115 may perform 3D matching processing).

[0195] Specifically, the object recognition unit 115 can calculate position Tx (3D), position Ty (3D), position Tz (3D), pose Rx (3D), pose Ry (3D), and pose Rz (3D) in the 3D matching process. As an example of the result of the 3D matching process, the object recognition unit 115 can generate information about position Tx (3D), position Ty (3D), position Tz (3D), pose Rx (3D), pose Ry (3D), and pose Rz (3D).

[0196] The object recognition unit 115 can use a transformation matrix M 3G The positions Tx (3D), Ty (3D), Tz (3D), and poses Rx (3D), Ry (3D), and Rz (3D) are converted into the positions Tx (GL), Ty (GL), Tz (GL), Rx (GL), Ry (GL), and Rz (GL) of the workpiece W in the global coordinate system. Furthermore, the object recognition unit 115 can calculate at least one of the positions Tx (GL), Ty (GL), Tz (GL), Rx (GL), Ry (GL), and Rz (GL) of the workpiece W in the global coordinate system based on the result of the 3D matching process. Thus, the object recognition unit 115 can calculate at least one of the positions and poses of the workpiece W in the global coordinate system based on the result of the 3D matching process.

[0197] Furthermore, the object recognition unit 115 may omit 2D matching processing, 2D tracking processing, and 3D tracking processing, and may also omit 3D matching processing. In this case, the object recognition unit 115 can generate three-dimensional position data WSD based on image data IMG_3D. As described above, the three-dimensional position data WSD can be referred to as data representing the position of workpiece W. Therefore, the object recognition unit 115 generating three-dimensional position data WSD based on image data IMG_3D can be considered as the object recognition unit 115 calculating the position of workpiece W based on image data IMG_3D. That is, the object recognition unit 115 can calculate the position of workpiece W by generating three-dimensional position data WSD based on image data IMG_3D. In this case, the imaging unit 20 may also omit the imaging device 21.

[0198] Furthermore, before performing the matching and tracking processes, the object recognition unit 115 may perform at least one of the following preprocessing operations on the image represented by at least one of image data IMG_2D and image data IMG_3D: automatic exposure (AE) processing, gamma correction processing, high dynamic range (HDR) processing, and filtering processing (e.g., smoothing filtering, sharpening filtering, differential filtering, median filtering, dilation filtering, edge hand filtering, bandpass filtering, etc.). Regarding the filtering processing, only one filtering operation may be performed, or multiple filtering operations may be performed.

[0199] AE processing is the automatic adjustment of the exposure of at least one of the imaging devices 21 and 22. The object recognition unit 115 can adjust the exposure of the imaging device 22 based on the luminance values ​​(e.g., individual luminance values ​​of multiple pixels or the average of luminance values ​​of multiple pixels) of the image data IMG_3D acquired by the imaging device 22. Furthermore, the object recognition unit 115 can also adjust the exposure of the imaging device 22 based on the luminance values ​​(e.g., individual luminance values ​​of multiple pixels or the average of luminance values ​​of multiple pixels) of the respective images (IMG_3D) acquired using the imaging device 22 at different exposures. Alternatively, the object recognition unit 115 may adjust the exposure of the imaging device 22 without relying on the luminance values ​​of the image data IMG_3D acquired by the imaging device 22. The object recognition unit 115 can also adjust the exposure of the imaging device 22 based on the measurement results of a metering device (not shown) capable of measuring brightness. In this case, the metering device may also be included in the imaging unit 20. Furthermore, when the object recognition unit 115 adjusts the exposure of the shooting device 21, the explanation is the same as that for adjusting the exposure of the shooting device 22 by the object recognition unit 115, so its explanation is omitted. Gamma correction processing is a process that adjusts the hue or color of an image represented by at least one of image data IMG_2D and image data IMG_3D. If the pixel value before gamma correction processing is set to V... in Set the pixel value after γ correction to V. out Then the γ correction process is represented as "V out =A·V in γFurthermore, "A" is a constant. When γ is less than 1, the image after γ correction is darker than the image before γ correction. When γ is greater than 1, the image after γ correction is brighter than the image before γ correction. Additionally, γ correction can be a process of adjusting the contrast of at least one of image data IMG_2D and image data IMG_3D. γ correction can also include emphasizing the edges of objects such as workpiece W captured in at least one of image data IMG_2D and image data IMG_3D. HDR processing is a process of generating an HDR image by combining multiple images (e.g., two images) of workpiece W taken with different exposures.

[0200] As described above, the computing device 110 and the imaging unit 20, which have an object recognition unit 115, are used to calculate (calculate) at least one of, for example, the position and orientation of the workpiece W. Therefore, a system including the computing device 110 and the imaging unit 20 can be called a computing system.

[0201] (3) Processing parameters are determined

[0202] To improve the accuracy of at least one of the position and pose (e.g., 6DoF) of an object (e.g., workpiece W) calculated from image data such as IMG_2D and IMG_3D output from the imaging unit 20, various parameters in at least one of the matching and tracking processes can be changed. Regarding 2D matching processing, one or more parameters that can be changed may exist. Regarding 3D matching processing, one or more parameters that can be changed may exist. Regarding 2D tracking processing, one or more parameters that can be changed may exist. Regarding 3D tracking processing, one or more parameters that can be changed may exist. Regarding position calculation processing, one or more parameters that can be changed may exist.

[0203] As described above, in the matching process (e.g., at least one of 2D matching and 3D matching), at least one of the position and orientation of the object (e.g., workpiece W) is calculated. In the tracking process (e.g., at least one of 2D tracking and 3D tracking), the change in at least one of the position and orientation of the object (e.g., workpiece W) is calculated. In the position calculation process, the three-dimensional positions of multiple points of the object (e.g., workpiece W) are calculated. Therefore, at least one of the matching process, tracking process, and position calculation process can be referred to as a computational process. Therefore, the multiple parameters in at least one of the matching process, tracking process, and position calculation process can be considered as multiple parameters in the computational process. Hereafter, these parameters (i.e., the parameters of the computational process) will be referred to as "processing parameters".

[0204] Furthermore, the computational processing can include the concepts of 2D matching processing, 2D tracking processing, 3D matching processing, 3D tracking processing, and position calculation processing. Specifically, the computational processing can include at least one of 2D matching processing, 2D tracking processing, 3D matching processing, 3D tracking processing, and position calculation processing. That is, the computational processing can include 2D matching processing and 2D tracking processing, but on the other hand, exclude position calculation processing, 3D matching processing, and 3D tracking processing. The computational processing can also include 2D matching processing, but on the other hand, exclude 2D tracking processing, position calculation processing, 3D matching processing, and 3D tracking processing. The computational processing can also include position calculation processing, 3D matching processing, but on the other hand, exclude 3D tracking processing, 2D matching processing, and 2D tracking processing.

[0205] Furthermore, the computational processing can be a concept that includes at least one of the preprocessing steps performed as AE processing, γ correction processing, HDR processing, and filtering processing. That is, the computational processing is not limited to processing that calculates (computes) at least one of the following: the position and pose of an object, the three-dimensional positions of multiple points of the object (e.g., the three-dimensional positions represented by 3D data WSD), and the amount of change in the position and pose of the object. It can be a concept that includes processing (e.g., preprocessing) performed on data (e.g., at least one of image data IMG_2D and image data IMG_3D) used to calculate at least one of the amount of change in the position and pose of the object.

[0206] The processing parameters are values ​​used to determine the content of the computational processing. When the processing parameters change, the content of the computational processing changes. Therefore, deciding whether to appropriately execute the computational processing based on the processing parameters will affect the accuracy of at least one of the calculated position and orientation of the object (e.g., workpiece W) obtained through the computational processing. As described above, since robot control signals are generated based on at least one of the calculated position and orientation of the object (e.g., workpiece W), the accuracy of the calculated position and orientation of the object (e.g., workpiece W) will affect the control error of at least one of the robot 10 and the end effector 13. Specifically, the higher the accuracy of the calculated position and orientation of the object (e.g., workpiece W), the smaller the control error of at least one of the robot 10 and the end effector 13.

[0207] Here, it is extremely difficult for the user of the robot system 1 to individually change (adjust) and appropriately set (determine) the various parameters (many processing parameters) of the computational processing in order to achieve the desired accuracy in calculating at least one of the position and orientation of the object (e.g., workpiece W).

[0208] Furthermore, processing parameters can also be values ​​of at least one of the threshold or target value for the arithmetic operation. Furthermore, processing parameters can also be arguments of the executed program. Furthermore, processing parameters can also be the type of function of the program performing the arithmetic operation. Furthermore, processing parameters can also be information used to determine whether to perform the arithmetic operation itself (the program performing the arithmetic operation). Furthermore, processing parameters can also be information used to enable or disable the function of the program performing the arithmetic operation.

[0209] Here, refer to Figure 5 The method for determining one or more processing parameters in the computational processing will be described. The computational unit 110 of the control device 100 has a determination unit 114. When image data (e.g., image data IMG_2D, image data IMG_3D) of an object (e.g., workpiece W) is input, the determination unit 114 can determine one or more processing parameters using a parameter determination model that determines one or more processing parameters in the computational processing (e.g., at least one of matching processing and tracking processing). That is, by inputting image data (e.g., at least one of image data IMG_2D and image data IMG_3D) of an object into the parameter determination model, the determination unit 114 can automatically determine one or more processing parameters (in other words, it can automatically change one or more processing parameters).

[0210] Furthermore, the decision unit 114 may read the parameter determination model from the storage device 120 (in other words, it may be accessible). Alternatively, the decision unit 114 may not read the parameter determination model from the storage device 120. The decision unit 114 may store the parameter determination model. Furthermore, whenever a decision is made on one or more processing parameters, the decision unit 114 may read the parameter determination model from the storage device 120.

[0211] As described above, the workpiece W can be referred to as the object to be processed. Furthermore, the image data generated by the imaging unit 20 by photographing the workpiece W can be referred to as the image data of the object to be processed.

[0212] Therefore, it can be said that the decision unit 114 can use the processing object image data and the parameter determination model to determine one or more processing parameters in the computational processing of the processing object image data. Specifically, the decision unit 114 can also input image data IMG_3D, which is an example of processing object image data, into the parameter determination model, and determine one or more processing parameters in the processing using the position of the input image data IMG_3D. The decision unit 114 can also input image data IMG_2D, which is an example of processing object image data, into the parameter determination model, and determine one or more processing parameters in the 2D matching processing using the input image data IMG_2D. The decision unit 114 can also input image data IMG_3D, which is an example of processing object image data, into the parameter determination model, and determine one or more processing parameters in the 3D matching processing using the input image data IMG_3D. Furthermore, the decision unit 114 can also input image data IMG_2D, which is an example of processing object image data, into the parameter determination model, and determine one or more processing parameters in the 2D tracking processing using the input image data IMG_2D. The decision unit 114 may also input image data IMG_3D, which is an example of image data of the processing object, into the parameter decision model to determine one or more processing parameters in the 3D tracking process using the input image data IMG_3D.

[0213] Furthermore, the decision unit 114 may also use a parameter determination model to determine one processing parameter. In other words, the decision unit 114 may also determine multiple processing parameters without using a parameter determination model. Even in the case described above, the technical effects described later can be obtained. In addition, the decision unit 114 may also be referred to as a reasoning unit because it uses a parameter determination model to reason about one or more processing parameters in the computational process.

[0214] The object recognition unit 115 of the computing device 110 can perform at least one of matching processing and tracking processing based on one or more processing parameters determined by the decision unit 114 (in other words, determined by the parameter determination model) to calculate at least one of the position and posture of the object (e.g., workpiece W).

[0215] Specifically, the object recognition unit 115 may also perform 2D matching processing on image data IMG_2D#1, which is an example of image data of the processing object, using one or more processing parameters in the 2D matching processing determined by the determination unit 114. As a result, the object recognition unit 115 can calculate at least one of the position and pose of the object (e.g., workpiece W) captured in the image represented by the image data IMG_2D#1 in the 2D shooting coordinate system. Furthermore, the determination unit 114 may determine one or more processing parameters in the preprocessing performed on the image data IMG_2D#1 used in the 2D matching processing, based on or instead of one or more processing parameters in the 2D matching processing. Here, the one or more processing parameters in the preprocessing may include at least a portion of the processing parameters that determine whether filtering processing is performed, such as ON / OFF.

[0216] Furthermore, the object recognition unit 115 may also use one or more processing parameters in the 2D matching process determined by the decision unit 114 to change at least one of the processing parameters in the 2D matching process. In this case, the object recognition unit 115 may also use the changed processing parameters in the 2D matching process to perform 2D matching processing and calculate at least one of the position and orientation of the object (e.g., workpiece W) captured in the image represented by image data IMG_2D#1 in the 2D shooting coordinate system.

[0217] The object recognition unit 115 can also perform position calculation processing using image data IMG_3D#1, which is an example of image data of the processing object, by calculating one or more processing parameters in the position calculation process determined by the determination unit 114. As a result, the object recognition unit 115 can calculate the three-dimensional positions of multiple points of the object (e.g., workpiece W) captured in the image represented by the image data IMG_3D#1. Furthermore, the determination unit 114 can determine one or more processing parameters in the preprocessing performed on the image data IMG_3D#1 used in the position calculation process, based on or instead of one or more processing parameters in the position calculation process. Here, the one or more processing parameters in the preprocessing may include at least a portion of the processing parameters for determining whether to perform filtering processing, such as turning on / off preprocessing.

[0218] Furthermore, the object recognition unit 115 may also use the position determined by the determination unit 114 to calculate one or more processing parameters in the processing, and change at least one of the one or more processing parameters in the position calculation processing. In this case, the object recognition unit 115 may also perform position calculation processing using one or more processing parameters in the changed position calculation processing, and calculate the three-dimensional positions of multiple points of the object (e.g., workpiece W) captured in the image represented by image data IMG_3D#1. The object recognition unit 115 may also generate three-dimensional position data (e.g., three-dimensional data WSD) representing the three-dimensional positions of the multiple points of the calculated object.

[0219] The object recognition unit 115 can also perform 3D matching processing using image data IMG_3D#1 (an example of image data of the processing object) and three-dimensional position data (e.g., three-dimensional data WSD) generated by position calculation processing, using one or more processing parameters in the 3D matching process determined by the determination unit 114. As a result, the object recognition unit 115 can calculate at least one of the position and pose of the object (e.g., workpiece W) captured in the image represented by the image data IMG_3D#1 in the 3D shooting coordinate system. Furthermore, the determination unit 114 can determine one or more processing parameters in the preprocessing performed on the image data IMG_3D#1 used in the 3D matching process, based on or instead of one or more processing parameters in the 3D matching process. Here, the one or more processing parameters in the preprocessing may include at least a portion of processing parameters that determine whether to perform filtering processing, such as turning on / off preprocessing.

[0220] Furthermore, the object recognition unit 115 may also use one or more processing parameters in the 3D matching process determined by the decision unit 114 to change at least one of the one or more processing parameters in the 3D matching process. In this case, the object recognition unit 115 may also use the changed one or more processing parameters in the 3D matching process to perform 3D matching processing and calculate at least one of the position and orientation of the object (e.g., workpiece W) captured in the image represented by image data IMG_3D#1 in the 3D shooting coordinate system.

[0221] The object recognition unit 115 can also perform 2D tracking processing using image data IMG_2D#1 and image data IMG_2D#2, which are examples of image data of the processing object, using one or more processing parameters in the 2D tracking processing determined by the determination unit 114. As a result, the object recognition unit 115 can calculate at least one of the position and posture of the object (e.g., workpiece W) captured in the image represented by image data IMG_2D#1 in the 2D shooting coordinate system, and the amount of change of at least one of the position and posture of the object (e.g., workpiece W) captured in the image represented by image data IMG_2D#2 in the 2D shooting coordinate system. In addition, the determination unit 114 can determine one or more processing parameters in the preprocessing performed on the image data IMG_2D#1 and image data IMG_2D#2 used in the 2D tracking processing, based on or instead of one or more processing parameters in the 2D tracking processing. Here, the one or more processing parameters in the preprocessing may include at least a portion of the processing parameters that determine whether to perform filtering processing, such as turning on / off preprocessing.

[0222] Furthermore, the object recognition unit 115 may also use one or more processing parameters in the 2D tracking process determined by the decision unit 114 to change at least one of the one or more processing parameters in the 2D tracking process. In this case, the object recognition unit 115 may also perform 2D tracking processing using the changed one or more processing parameters in the 2D tracking process, and calculate at least one of the position and orientation of the object (e.g., workpiece W) captured in the image represented by image data IMG_2D#1 in the 2D shooting coordinate system, and the change in at least one of the position and orientation of the object (e.g., workpiece W) captured in the image represented by image data IMG_2D#2 in the 2D shooting coordinate system.

[0223] The object recognition unit 115 can also perform 3D tracking processing using image data IMG_3D#1 and image data IMG_3D#2, which are examples of image data of the processing object, using one or more processing parameters in the 3D tracking processing determined by the determination unit 114. As a result, the object recognition unit 115 can calculate at least one of the position and pose of the object (e.g., workpiece W) captured in the image represented by image data IMG_3D#1 in the 3D shooting coordinate system, and the amount of change of at least one of the position and pose of the object (e.g., workpiece W) captured in the image represented by image data IMG_3D#2 in the 3D shooting coordinate system. In addition, the determination unit 114 can determine one or more processing parameters in the preprocessing performed on the image data IMG_3D#1 and image data IMG_3D#2 used in the 3D tracking processing, based on or instead of one or more processing parameters in the 3D tracking processing. Here, the one or more processing parameters in the preprocessing may include at least a portion of the processing parameters that determine whether to perform filtering processing, such as turning on / off preprocessing.

[0224] Furthermore, the object recognition unit 115 may also use one or more processing parameters in the 3D tracking process determined by the decision unit 114 to change at least one of the one or more processing parameters in the 3D tracking process. In this case, the object recognition unit 115 may also use the changed one or more processing parameters in the 3D tracking process to perform 3D tracking processing and calculate at least one of the position and orientation of the object (e.g., workpiece W) captured in the image represented by image data IMG_3D#1 in the 3D shooting coordinate system, and the change in at least one of the position and orientation of the object (e.g., workpiece W) captured in the image represented by image data IMG_3D#2 in the 3D shooting coordinate system.

[0225] Furthermore, image data IMG_3D, image data IMG_3D#1, and image data IMG_3D#2 can be generated by photographing an object (e.g., workpiece W) onto which a desired projection pattern (e.g., random dots) is projected by a projection device 23 included in the imaging unit 20, using a stereo camera that serves as the imaging device 22. Additionally, image data IMG_3D, image data IMG_3D#1, and image data IMG_3D#2 can be generated by photographing an object (e.g., workpiece W) onto which the desired projection pattern is not projected, using a stereo camera that serves as the imaging device 22 included in the imaging unit 20.

[0226] As described below, parametric determination models can be generated through learning using teacher data (i.e., through so-called teacher-led learning). Furthermore, parametric determination models are not limited to learning using teacher data; they can also be generated through learning without teacher data (i.e., through so-called teacherless learning). The learning used to generate parametric determination models can be machine learning or deep learning.

[0227] A parameter determination model outputs one or more processing parameters from image data containing objects (e.g., at least one of image data IMG_2D and image data IMG_3D). In other words, a parameter determination model is a structure that derives one or more processing parameters from image data containing objects. Specifically, a parameter determination model is a structure that performs at least one of some evaluation and judgment on image data containing objects and outputs one or more processing parameters as the result of that evaluation and judgment. The parameter determination model can be a mathematical model using, for example, decision trees, random forests, support vector machines, or Naive Bayes. Parameter determination models can be generated through deep learning. A parameter determination model generated through deep learning can be a mathematical model constructed using a multi-layered neural network with multiple intermediate layers (also called hidden layers). The neural network can be, for example, a convolutional neural network. Furthermore, since the parameter determination model can calculate one or more processing parameters based on at least one of image data IMG_2D and image data IMG_3D, it can also be called a computational model. Furthermore, the parameter determination model can also be said to infer one or more processing parameters based on at least one of the image data IMG_2D and image data IMG_3D, and therefore can be called an inference model or inferencer.

[0228] In addition, the following processing parameters can be listed as examples of processing parameters.

[0229] As an example of a processing parameter in 2D matching processing, it may be at least one of a threshold for determining the edge of an object (e.g., workpiece W) captured in the image (e.g., a threshold for the difference in luminance values ​​between adjacent pixels), and a range for parallel shifting, scaling, and / or rotation of model data (e.g., 2D model data IMG_2M). Furthermore, the processing parameter in 2D matching processing is not limited to the threshold for determining the edge of an object captured in the image, and at least one of the range for parallel shifting, scaling, and / or rotation of model data; it may also be other known parameters for 2D matching processing. As an example of a processing parameter in 3D matching processing, it may be at least one of the data division ratio of 3D position data (e.g., 3D position data WSD), and the proportion of feature quantities (e.g., normal vectors) calculated in the respective 3D positions of multiple points represented by the 3D position data (e.g., 3D position data WSD). In this case, the data division ratio of the 3D position data includes at least one of the data division ratio of the point group represented by the point group data and the pixel division ratio of the depth image represented by the depth image data. Furthermore, the proportion of feature quantities calculated from the respective 3D positions of multiple points represented by the 3D position data includes at least one of the proportion of point groups whose feature quantities are calculated from the point group data, and the proportion of pixels whose feature quantities are calculated from the multiple pixels of the depth image represented by the depth image data. Additionally, at least one of the point group and pixels from which feature quantities are calculated can be referred to as a feature region. Furthermore, the processing parameters in 3D matching processing are not limited to the data-to-data ratio of the 3D position data, and at least one of the proportion of feature quantities calculated from the respective 3D positions of multiple points represented by the 3D position data, but can also be other known 3D matching processing parameters. As an example of processing parameters in 2D or 3D tracking processing, it can be at least one of the number of feature regions (at least one of the number of feature points and the number of edges), and at least one of the threshold used to determine whether something is a feature region. For example, in the case of 2D tracking processing, the threshold used to determine whether something is a feature region can be a threshold for the luminance value of the image represented by the image data IMG_2D (e.g., two image data IMG_2D#t1 and IMG_2D#t2). For example, in the case of 3D tracking processing, the threshold used to determine whether something is a feature region can also be the threshold of the 3D position represented by the 3D position data WSD (e.g., two 3D position data WSDPD#s1 and WSDPD#s2). Furthermore, the processing parameters in 2D or 3D tracking processing are not limited to at least one of the number of feature regions and the threshold used to determine whether something is a feature region; they can also be other known parameters for 2D or 3D tracking processing.As an example of a processing parameter in the location calculation process, it may be at least one of the following: a threshold for determining whether the disparity is erroneous (e.g., the difference between adjacent disparities); a threshold for determining whether it is noise (e.g., at least one of the luminance value and disparity of pixels in each image represented by the two image data contained in image data IMG_3D); the width of the region of adjacent pixels where the luminance value of pixels in each image represented by the two image data contained in image data IMG_3D exceeds a predetermined threshold; and the width of the region where the disparity exceeds a predetermined threshold. Furthermore, the processing parameter in the location calculation process is not limited to the example described above, and may also be other known parameters for location calculation processes. As an example of a processing parameter in AE processing, it may be at least one of a target value such as on / off (in other words, whether it is applied) and brightness. Furthermore, the processing parameter in AE processing is not limited to at least one of a target value such as on / off and brightness, and may also be other known parameters for AE processing. As an example of a processing parameter in γ correction processing, it may be at least one of an on / off state and a γ value. Furthermore, the processing parameters in gamma correction processing are not limited to at least one of on / off and gamma value, but may also be other known gamma correction processing parameters. As an example of processing parameters in HDR processing, at least one of on / off, exposure, and number of image acquisitions may also be included. Similarly, as an example of processing parameters in filtering processing, at least one of on / off, the combination of filters used, the filter value, and the filter threshold may also be included.

[0230] Furthermore, the processing unit 110 of the control device 100 may not have a determination unit 114. In this case, the processing parameters may be determined by another device different from the control device 100 but having the same function as the determination unit 114 (e.g., another control device different from the control device 100 or another processing device different from the processing unit 110). The control device 100 may send image data output from the imaging unit 20, such as image data IMG_2D and image data IMG_3D, to the other device. The other device may determine one or more processing parameters by inputting the image data sent from the control device 100 into a model equivalent to a parameter determination model. The other device may send the determined one or more processing parameters to the control device 100. Furthermore, the object recognition unit 115 may not use one or more processing parameters determined by the determination unit 114. In this case, the object recognition unit 115 may determine the processing parameters. In this case, the object recognition unit 115 may determine one or more processing parameters by inputting image data into a model equivalent to a parameter determination model.

[0231] As described above, the signal generation unit 116 can generate robot control signals and end-effector control signals to control the movements of the robot 10 and the end-effector 13, causing the end-effector 13 to perform prescribed processing on an object (e.g., workpiece W). The object recognition unit 115 can also calculate at least one of the position and orientation of the object (e.g., workpiece W), causing the signal generation unit 116 to generate robot control signals and end-effector control signals. The determination unit 114 uses a parameter determination model to determine one or more processing parameters in the computational processing performed by the object recognition unit 115. Thus, the determination unit 114, the object recognition unit 115, and the signal generation unit 116 are components related to the prescribed processing of the object (e.g., workpiece W). Therefore, the determination unit 114, the object recognition unit 115, and the signal generation unit 116 can be referred to as a control unit. In addition, the computing device 110 having the determination unit 114, the object recognition unit 115, and the signal generation unit 116 can also be referred to as a control unit.

[0232] The decision unit 114 uses one or more processing parameters determined by the parameter decision model, for example, to generate robot control signals and end effector control signals that control the actions of the robot 10 and the end effector 13, so that the end effector 13 performs specified processing on the workpiece W.

[0233] Therefore, one or more processing parameters in the 2D matching process determined by the decision unit 114 can be referred to as "two-dimensional matching parameters for processing". Furthermore, the "two-dimensional matching parameters for processing" may also include one or more processing parameters in the preprocessing of the image data used in the 2D matching process, either on the basis of or in place of one or more processing parameters in the 2D matching process.

[0234] One or more processing parameters in the 3D matching process determined by the decision unit 114 may also be referred to as "processing 3D matching parameters". Furthermore, the "processing 3D matching parameters" may include one or more processing parameters in the preprocessing of the image data used in the 3D matching process, either on the basis of or in place of one or more processing parameters in the 3D matching process.

[0235] One or more processing parameters in the position calculation process determined by the determination unit 114 may be referred to as "position calculation parameters for processing". Furthermore, the "position calculation parameters for processing" may include one or more processing parameters in the preprocessing of the image data used in the position calculation process, either on the basis of or in place of one or more processing parameters in the position calculation process.

[0236] One or more processing parameters in the 2D tracking process determined by the decision unit 114 may also be referred to as "two-dimensional tracking parameters for processing". Furthermore, the "two-dimensional tracking parameters for processing" may include one or more processing parameters, such as those used in the preprocessing of the image data used in the 2D tracking process, based on or replacing one or more processing parameters in the 2D tracking process.

[0237] One or more processing parameters in the 3D tracking process determined by the decision unit 114 may also be referred to as "processing 3D tracking parameters". Furthermore, the "processing 3D tracking parameters" may include one or more processing parameters, such as those used in the preprocessing of the image data used in the 3D tracking process, based on or replacing one or more processing parameters in the 3D tracking process.

[0238] One or more processing parameters in the preprocessing performed on the image data (e.g., at least one of image data IMG_2D and image data IMG_3D) determined by the decision unit 114 may be referred to as "preprocessing parameters". (4) Method for generating the parameter determination model

[0239] (4-1)Overview

[0240] Here, refer to Figure 6 This section provides an overview of the methods for generating parameter-determined models. Figure 6 This is a diagram illustrating an example of how parameters determine the generation method of a model.

[0241] exist Figure 6 In the middle, the data generation unit 111 of the arithmetic device 110 (refer to Figure 4Based on the image data relating to the learning object and the model data corresponding to the learning object, the data generation unit 111 calculates at least one of the position and pose of the learning object captured in the image represented by the learning image data, within the shooting coordinate system (e.g., at least one of a 2D shooting coordinate system and a 3D shooting coordinate system). The data generation unit 111 generates position and pose data representing at least one of the calculated position and pose of the learning object. Furthermore, based on the learning image data and the model data corresponding to the learning object, the data generation unit 111 calculates the three-dimensional positions (i.e., the three-dimensional shape of the learning object) of multiple points captured in the image represented by the learning image data. The data generation unit 111 generates three-dimensional (3D) position data representing the calculated three-dimensional positions of the multiple points of the learning object.

[0242] Here, the three-dimensional position data can be point group data representing the three-dimensional positions of multiple points on the learning object, or depth image data representing the three-dimensional positions of multiple points on the learning object. Furthermore, since the three-dimensional position data represents the three-dimensional shape of the workpiece W, it can also be called three-dimensional shape data. Additionally, since the three-dimensional position data represents the distance from the imaging unit 20 (e.g., the imaging device 22) to the workpiece W, it can also be called distance data.

[0243] Furthermore, the position of the learning object represented by the position and pose data can also be the position of a representative point of the learning object. Examples of representative points of the learning object include at least one of the following: the center of the learning object, the center of gravity of the learning object, the vertex of the learning object, the center of the surface of the learning object, and the center of gravity of the surface of the learning object. The representative point of the learning object can also be called a feature point of the learning object.

[0244] Furthermore, the data generation unit 111 can calculate at least one of the position and pose of the learning object in the image represented by the learning image data, within the shooting coordinate system, based on the learning image data. That is, the data generation unit 111 may calculate at least one of the position and pose of the learning object in the image represented by the learning image data, within the shooting coordinate system, without using model data corresponding to the learning object. Similarly, the data generation unit 111 can also calculate the three-dimensional positions of multiple points of the learning object in the image represented by the learning image data, each based on the learning image data. That is, the data generation unit 111 may calculate the three-dimensional positions of multiple points of the learning object in the image represented by the learning image data, each based on the learning image data, without using model data corresponding to the learning object.

[0245] Furthermore, the data generation unit 111 may not calculate one of the following: at least one of the position and pose of the learning object in the shooting coordinate system, and one of the three-dimensional positions of each of the multiple points of the learning object. That is, the data generation unit 111 may calculate at least one of the position and pose of the learning object in the shooting coordinate system, but on the other hand, it may not calculate the three-dimensional positions of each of the multiple points of the learning object. In this case, the data generation unit 111 may generate position and pose data representing at least one of the calculated position and pose of the learning object in the shooting coordinate system, but on the other hand, it may not generate three-dimensional position data. Alternatively, the data generation unit 111 may calculate the three-dimensional positions of each of the multiple points of the learning object, but on the other hand, it may not calculate the position and pose of the learning object in the shooting coordinate system. In this case, the data generation unit 111 may generate three-dimensional position data representing the calculated three-dimensional positions of each of the multiple points of the learning object, but on the other hand, it may not generate position and pose data.

[0246] Furthermore, the data generation unit 111 can include the three-dimensional positions (i.e., the three-dimensional shape of the learning object) of multiple points of the learning object in the position and pose data. The data generation unit 111 can also calculate the three-dimensional positions of the multiple points of the learning object as the position of the learning object captured in the image represented by the learning image data. Moreover, the data generation unit 111 can also generate position and pose data representing the calculated three-dimensional positions of the multiple points of the learning object. This is because the position of the learning object represented by the position and pose data, and the three-dimensional positions of the multiple points of the learning object, can both represent the same position.

[0247] However, the data generation unit 111 may also calculate the three-dimensional positions of each of the multiple points of the learning object as the position of the learning object captured in the image represented by the learning image data. In this case, the position of the learning object may refer to the position of the representative point of the learning object. That is, the data generation unit 111 may also calculate at least one of the position and pose of the representative point of the learning object captured in the image represented by the learning image data based on the learning image data and the model data corresponding to the learning object.

[0248] The individual three-dimensional positions of multiple points on a learning object can be considered as the individual three-dimensional positions of a non-specific set of points on the learning object. Therefore, it can also be said that the position of a representative point on the learning object differs from the individual three-dimensional positions of the multiple points on the learning object.

[0249] Furthermore, the learning image data can be generated by photographing the learning object using the imaging unit 20. The learning image data can also be generated by photographing the learning object using an imaging device different from the imaging unit 20. Additionally, as described later, the learning image data can also be generated as virtual data.

[0250] The learning object is the object used in the learning process to generate the parameter determination model described later. Furthermore, the learning object can also be described as an object captured in an image represented by the learning image data. The learning object is an object that is substantially the same as the object (e.g., workpiece W) that is being processed by the end effector 13. Here, "the object that is being processed by the end effector 13" refers to the object being processed by the end effector 13, and therefore can also be called the processing object. Therefore, the learning object can be described as an object that is substantially the same as the processing object. That is, the learning object is not limited to an object with the same shape as the processing object, but can be an object that is similar to the processing object to a degree that can be considered the same as the processing object. The state that “the learning object is an object that is similar to the processing object to the same extent as the processing object” may include at least one of the following states: (i) the difference between the shape of the learning object and the processing object is a manufacturing error; (ii) the difference between the shape of the learning object and the processing object is such that the shapes in the image captured by the imaging unit 20 are considered to be the same; (iii) the difference between the shape of the learning object and the processing object is such that the learning object is slightly deformed due to contact with another object (e.g., another learning object); (iv) the difference between the shape of the learning object and the processing object is such that the shapes are slightly deformed due to being placed or held on a platform, etc.; (v) a part of the learning object and a part of the processing object have different shapes when the imaging unit 20 is not used or cannot be used to capture images.

[0251] Furthermore, the model data corresponding to the learning object can also be two-dimensional model data representing a two-dimensional model of the learning object. In this case, the model data corresponding to the learning object can be two-dimensional image data representing multiple two-dimensional models of the learning object generated by virtually projecting a three-dimensional model (e.g., a CAD model) of the learning object from multiple different directions onto virtual planes orthogonal to each of the multiple different directions. Alternatively, the model data corresponding to the learning object can also be image data representing two-dimensional images of the actual learning object taken beforehand. The model data corresponding to the learning object can also be image data representing multiple two-dimensional images generated by taking pictures of the actual learning object from multiple different shooting directions. Furthermore, the two-dimensional model data corresponding to the learning object can also be the two-dimensional model data IMG_2M. In this case, the model data corresponding to the learning object can be the same as the model data corresponding to the workpiece W. Alternatively, the model data corresponding to the learning object can be different from the model data corresponding to the workpiece W.

[0252] Furthermore, the model data corresponding to the learning object can also be three-dimensional model data representing a three-dimensional model of the learning object. In this case, the model data corresponding to the learning object can be CAD data of the learning object, or it can be three-dimensional model data having the same shape as the three-dimensional shape of the learning object obtained by measuring the actual three-dimensional shape of the learning object beforehand. Furthermore, the three-dimensional model data corresponding to the learning object can also be the three-dimensional model data WMD. In this case, the model data corresponding to the learning object can be the same as the model data corresponding to the workpiece W. Furthermore, the model data corresponding to the learning object can also be different from the model data corresponding to the workpiece W.

[0253] At least one of the learned image data and the position / pose data and three-dimensional position data generated by the data generation unit 111 based on the learned image data can be stored in the storage device 120 of the control device 100 in association with each other. That is, the learned image data and the position / pose data and three-dimensional position data generated based on the learned image data can also be stored in the storage device 120 in association with each other. The learned image data and the position / pose data generated based on the learned image data are stored in the storage device 120 in association with each other, but on the other hand, the three-dimensional position data generated based on the learned image data may not be associated with the learned image data. The learned image data and the three-dimensional position data generated based on the learned image data are stored in the storage device 120 in association with each other, but on the other hand, the position / pose data generated based on the learned image data may not be associated with the learned image data.

[0254] Position and pose data and three-dimensional position data can be generated based on or instead of learning image data generated by photographing the learning object by the imaging unit 20, based on virtual data serving as learning image data. Virtual data serving as learning image data can be generated as follows: After arranging a three-dimensional model of the learning object (e.g., a three-dimensional model of the learning object represented by CAD data of the learning object) in virtual space, virtual photography is performed on the three-dimensional model by a virtual imaging unit (e.g., virtual imaging unit 20), thereby generating virtual data serving as learning image data. At this time, since the positional relationship between the three-dimensional model in virtual space and the virtual imaging unit (e.g., virtual imaging unit 20) is known, at least one of the position and pose of representative points of the three-dimensional model (i.e., the learning object) captured in the image represented by the virtual data, or the three-dimensional positions of multiple points of the three-dimensional model (i.e., the learning object), are also known.

[0255] Furthermore, virtual data can also be generated by the data generation unit 111. In this case, the data generation unit 111 can also generate at least one of position and pose data and three-dimensional position data from the virtual data. As described above, at least one of the position and pose of the representative point of the captured three-dimensional model in the image represented by the virtual data is known, so the data generation unit 111 can generate position and pose data based on the known position and pose of the representative point of the three-dimensional model. In addition, the three-dimensional positions of each of the multiple points of the three-dimensional model are also known, so the data generation unit 111 can generate three-dimensional position data based on the known three-dimensional positions of each of the multiple points of the three-dimensional model. The data generation unit 111 can also generate virtual data based on input from the user via the input device 140 (e.g., at least one of the input indicating the configuration of the three-dimensional model in the virtual space and the configuration of the virtual shooting unit in the virtual space). Furthermore, the data generation unit 111 may not generate virtual data. In this case, the virtual data can be generated by other devices different from the control device 100 (e.g., at least one of other control devices and other computing devices). The virtual data generated by the other devices can also be stored in the storage device 120 of the control device 100. In this case, the data generation unit 111 can also read the virtual data stored in the storage device 120 and generate at least one of position and pose data and three-dimensional position data based on the read virtual data. Alternatively, the data generation unit 111 can also acquire virtual data generated by the other devices via the communication device 130. The data generation unit 111 can also generate at least one of position and pose data and three-dimensional position data based on the acquired virtual data. Furthermore, other devices different from the control device 100 can generate virtual data and generate at least one of position and pose data and three-dimensional position data based on the generated virtual data. In this case, the control device 100 can acquire virtual data generated by other devices, as well as at least one of position and pose data and three-dimensional position data, via the communication device 120.

[0256] Furthermore, the virtual imaging unit can be a single virtual monocular camera or a virtual stereo camera with two virtual monocular cameras. The virtual imaging unit may include a single virtual monocular camera and a virtual stereo camera. Additionally, a projection pattern identical to the projection pattern (e.g., random dots) that can be projected by the projection device 23 included in the imaging unit 20 can virtually project onto the 3D model in virtual space. In this case, the 3D model with the projected projection pattern can be captured in the image represented by the virtual data. Alternatively, the projection pattern may not be virtually projected onto the 3D model in virtual space. In this case, the 3D model without a projected projection pattern can be captured in the image represented by the virtual data.

[0257] Furthermore, the data generation unit 111 may also use multiple learning image data captured at different times to generate displacement data involved in tracking processing (e.g., at least one of 2D tracking processing and 3D tracking processing) (see "(4-3-3) Others" described later). The multiple learning image data and the displacement data involved in tracking processing generated by the data generation unit 111 based on the multiple learning image data can be stored in the storage device 120 in association with each other.

[0258] The learning object recognition unit 112 of the computing device 110 can use the learning image data to perform matching processing such as CAD matching (e.g., at least one of 2D matching processing and 3D matching processing) to calculate at least one of the position and posture of the learning object in the image represented by the learning image data in the shooting coordinate system (e.g., at least one of 2D shooting coordinate system and 3D shooting coordinate system).

[0259] At this time, the learning object recognition unit 112 can modify one or more processing parameters in the matching process, so that at least one of the position and posture of the learning object in the shooting coordinate system (at least one of the position and posture calculated by the learning object recognition unit 112) and at least one of the position and posture of the learning object in the shooting coordinate system represented by the position and posture data associated with the learning image data used in the matching process of the learning object recognition unit 112 (the position and posture data generated by the data generation unit 111) are close (typically consistent). In addition, the learning object recognition unit 112 can modify one or more processing parameters in the preprocessing performed on the learning image data used in the matching process based on or instead of one or more processing parameters in the matching process, so that at least one of the position and posture of the learning object in the shooting coordinate system (at least one of the position and posture calculated by the learning object recognition unit 112) and at least one of the position and posture of the learning object in the shooting coordinate system represented by the position and posture data associated with the learning image data used in the matching process (the position and posture data generated by the data generation unit 111) are close (typically consistent). Furthermore, at least one of the position and pose of the learning object may not be at least one of the position and pose in the shooting coordinate system. At least one of the position and pose of the learning object may be at least one of the position and pose in the global coordinate system, or at least one of the position and pose in a coordinate system different from both the shooting coordinate system and the global coordinate system.

[0260] The learning object recognition unit 112 can use learning image data to perform position calculation processing using well-known correspondence establishment methods such as SGBM or SAD, calculating the three-dimensional positions of multiple points of the learning object captured in the image represented by the learning image data. The learning object recognition unit 112 can modify one or more processing parameters in the position calculation process so that the calculated three-dimensional positions of the multiple points of the learning object (the three-dimensional positions of the multiple points calculated by the learning object recognition unit 112) are close to (typically consistent with) the three-dimensional positions of the multiple points of the learning object represented by the three-dimensional position data associated with the learning image data used in the position calculation process of the learning object recognition unit 112 (the three-dimensional position data generated by the data generation unit 111). Furthermore, the learning object recognition unit 112 can modify one or more processing parameters in the preprocessing performed on the learning image data used in the position calculation process, based on or instead of one or more processing parameters in the position calculation process.

[0261] Furthermore, the learning object recognition unit 112 can use the learning image data to perform only one of the matching process and the position calculation process. That is, the learning object recognition unit 112 can perform at least one of the matching process and the position calculation process.

[0262] Furthermore, the learning object recognition unit 112 can perform tracking processing (e.g., at least one of 2D tracking processing and 3D tracking processing) using multiple learning image data (e.g., one learning image data and another learning image data captured at different times) (see "(4-4-4) Others" described later). The learning object recognition unit 112 can track at least one feature part (e.g., at least one of feature points and edges) that is identical to at least one feature part of the learning object captured in the image represented by one learning image data within another learning image data. That is, the learning object recognition unit 112 can calculate the change in at least one of the position and pose of the learning object within the shooting coordinate system (e.g., a 2D shooting coordinate system or a 3D shooting coordinate system) between one learning image data and another learning image data.

[0263] The learning object recognition unit 112 can modify one or more processing parameters in the tracking process, making the change in at least one of the position and pose of the learning object in the shooting coordinate system (the change calculated by the learning object recognition unit 112) close (typically consistent) to the change in at least one of the position and pose of the learning object in the shooting coordinate system represented by the displacement data involved in the tracking process associated with one learning image data and another learning image data (the displacement data generated by the data generation unit 111). Furthermore, the learning object recognition unit 112 can modify one or more processing parameters in the preprocessing performed on the learning image data used in the tracking process, based on or replacing one or more processing parameters in the tracking process. Moreover, the change in at least one of the position and pose of the learning object may not be the change in the position and pose of the learning object in the shooting coordinate system. The change in the position and pose of the learning object may be the change in the position and pose of the learning object in the global coordinate system, or it may be the change in the position and pose of the learning object in a coordinate system different from the shooting coordinate system and the global coordinate system.

[0264] Furthermore, regarding the modification of at least one of the processing parameters in the matching process (e.g., at least one of 2D matching process and 3D matching process) by the learning object recognition unit 112 when using matching processing of the learning image data, the modification of at least one of the processing parameters in the matching process and the preprocessing of the learning image data used in the matching process can be described as follows: The learning object recognition unit 112 optimizes at least one of the processing parameters in the matching process and the preprocessing of the learning image data used in the matching process, such that at least one of the calculated position and pose of the learning object in the shooting coordinate system (at least one of the position and pose calculated by the learning object recognition unit 112) and at least one of the position and pose of the learning object in the shooting coordinate system represented by the position and pose data associated with the learning image data (position and pose data generated by the generation unit 111) are close (typically consistent).

[0265] Furthermore, at least one of the position and pose of the learning object may not be at least one of the position and pose in the shooting coordinate system. At least one of the position and pose of the learning object may be at least one of the position and pose in the global coordinate system, or at least one of the position and pose in a coordinate system different from both the shooting coordinate system and the global coordinate system.

[0266] Furthermore, regarding the position calculation process using the learning image data, the learning object recognition unit 112 changes at least one of the processing parameters in the position calculation process and at least one of the processing parameters in the preprocessing performed on the learning image data used in the position calculation process. This can be described as follows: The learning object recognition unit 112 optimizes at least one of the processing parameters in the position calculation process and at least one of the processing parameters in the preprocessing performed on the learning image data used in the position calculation process, so that the calculated three-dimensional positions of each of the multiple points of the learning object (the three-dimensional positions of each of the multiple points calculated by the learning object recognition unit 112) and the three-dimensional positions of each of the multiple points of the learning object represented by the three-dimensional position data associated with the learning image data (the three-dimensional position data generated by the generation unit 111) are close (typically consistent).

[0267] Furthermore, regarding tracking processing (e.g., at least one of 2D tracking processing and 3D tracking processing) using multiple learning image data captured at different times, the learning object recognition unit 112 changes at least one of the processing parameters in the tracking processing and at least one of the processing parameters in the preprocessing of the learning image data used in the tracking processing. This can be described as follows: The learning object recognition unit 112 optimizes at least one of the processing parameters in the tracking processing and at least one of the processing parameters in the preprocessing of the learning image data used in the tracking processing, such that the change in at least one of the calculated position and pose of the learning object (the change calculated by the learning object recognition unit 112) and the change in at least one of the position and pose of the learning object represented by the displacement data involved in the tracking processing associated with the multiple learning image data (the displacement data generated by the data generation unit 111) are close (typically consistent).

[0268] Based on this, "changing the processing parameters" can also be called "optimizing the processing parameters".

[0269] As a result of matching processing using the learned image data, the learned object recognition unit 112 can output output data representing at least one of one or more processing parameters in the matching process, including the consistency involved in the matching process and the optimized matching process. The consistency involved in the matching process is an index representing the degree of consistency between at least one of the position and pose of the learned object in the shooting coordinate system calculated by the learned object recognition unit 112 and at least one of the position and pose of the learned object in the shooting coordinate system represented by the position and pose data generated by the data generation unit 111. In addition to including the consistency involved in the matching process and one or more processing parameters in the matching process, the output data may also include the time required for the matching process using one or more processing parameters. Furthermore, the output data may also include one or more processing parameters from the preprocessing performed on the learned image data used in the matching process, based on or instead of one or more processing parameters in the matching process.

[0270] Furthermore, at least one of the position and pose of the learning object may not be at least one of the position and pose in the shooting coordinate system. At least one of the position and pose of the learning object may be at least one of the position and pose in the global coordinate system, or at least one of the position and pose in a coordinate system different from both the shooting coordinate system and the global coordinate system.

[0271] In this case, the learning image data and the output data output by the learning object recognition unit 112 after matching the learning image data can be stored in the storage device 120 of the control device 100 in association with each other. Furthermore, the learning object recognition unit 112 can perform one matching process using one learning image data and output one output data. In this case, one output data can be associated with one learning image data. The learning object recognition unit 112 can perform multiple matching processes using one learning image data and output multiple output data. In this case, multiple output data can be associated with one learning image data.

[0272] Furthermore, the "consistency involved in the matching process" can be an indicator that varies according to at least one of the following: the difference between at least one of the position and posture of the learning object in the shooting coordinate system calculated by the learning object recognition unit 112 and the position and posture of the learning object in the shooting coordinate system represented by the position and posture data generated by the data generation unit 111; and the ratio of at least one of the position and posture of the learning object in the shooting coordinate system calculated by the learning object recognition unit 112 and the position and posture of the learning object in the shooting coordinate system represented by the position and posture data generated by the data generation unit 111. When the consistency involved in the matching process varies according to the difference between at least one of the position and posture of the learning object in the shooting coordinate system calculated by the learning object recognition unit 112 and the position and posture of the learning object in the shooting coordinate system represented by the position and posture data, the smaller the difference (in other words, the closer to 0), the higher the consistency involved in the matching process. When the consistency involved in the matching process varies based on the ratio of at least one of the position and pose of the learning object in the shooting coordinate system calculated by the learning object recognition unit 112 to at least one of the position and pose of the learning object in the shooting coordinate system represented by the position and pose data, the closer the ratio is to 1, the higher the consistency involved in the matching process can be.

[0273] As a result of the position calculation processing using the learned image data, the learned object recognition unit 112 can output output data representing at least one of the consistency involved in the position calculation processing and one or more processing parameters in the optimized position calculation processing. The consistency involved in the position calculation processing is an index representing the degree of consistency between the three-dimensional positions of multiple points of the learned object calculated by the learned object recognition unit 112 and the three-dimensional positions of multiple points of the learned object represented by the three-dimensional position data generated by the data generation unit 111. In addition to including the consistency involved in the position calculation processing and one or more processing parameters in the position calculation processing, the output data may also include the time required for the position calculation processing using one or more processing parameters. Furthermore, the output data may also include one or more processing parameters from the preprocessing performed on the learned image data used in the position calculation processing, based on or instead of one or more processing parameters in the position calculation processing.

[0274] In this case, the learned image data and the output data output by the learning object recognition unit 112 after performing position calculation processing using the learned image data can be stored in the storage device 120 of the control device 100 in association with each other. Furthermore, the learning object recognition unit 112 can perform position calculation processing once using one piece of learned image data and output one piece of output data. In this case, one piece of output data can be associated with one piece of learned image data. The learning object recognition unit 112 can perform position calculation processing multiple times using one piece of learned image data and output multiple pieces of output data. In this case, multiple pieces of output data can be associated with one piece of learned image data.

[0275] Furthermore, the "consistency involved in the position calculation process" can be an indicator that varies according to at least one of the following: the difference between the three-dimensional positions of each point of the learning object calculated by the learning object recognition unit 112 and the three-dimensional positions of each point of the learning object represented by the three-dimensional position data generated by the data generation unit 111; and the ratio of the three-dimensional positions of each point of the learning object calculated by the learning object recognition unit 112 to the three-dimensional positions of each point of the learning object represented by the three-dimensional position data generated by the data generation unit 111. When the consistency involved in the position calculation process varies according to the difference between the three-dimensional positions of each point of the learning object calculated by the learning object recognition unit 112 and the three-dimensional positions of each point of the learning object represented by the three-dimensional position data generated by the data generation unit 111, the smaller the difference (in other words, the closer to 0), the higher the consistency involved in the position calculation process. When the consistency involved in the position calculation process varies according to the ratio of the three-dimensional positions of each of the multiple points of the learning object calculated by the learning object recognition unit 112 to the three-dimensional positions of each of the multiple points of the learning object represented by the three-dimensional position data generated by the data generation unit 111, the closer the ratio is to 1, the higher the consistency involved in the position calculation process can be.

[0276] Furthermore, when the learning object recognition unit 112 performs tracking processing using multiple learning image data captured at different times, the learning object recognition unit 112 can output output data representing the consistency involved in the tracking processing and one or more processing parameters in the optimized tracking processing. The consistency involved in the tracking processing is an index representing the degree of consistency between the change in at least one of the position and pose of the learning object calculated based on the result of the tracking processing performed by the learning object recognition unit 112 (the change calculated by the learning object recognition unit 112) and the change in at least one of the position and pose of the learning object represented by the displacement data involved in the tracking processing associated with the multiple learning image data (displacement data generated by the data generation unit 111). In addition to including the consistency involved in the tracking processing and one or more processing parameters in the tracking processing, the output data may also include the time required for the tracking processing using one or more processing parameters. Furthermore, the output data may also include one or more processing parameters from preprocessing performed on at least one of the multiple learning image data used in the tracking processing, based on or instead of the one or more processing parameters in the tracking processing.

[0277] In this case, the multiple learning image data and the output data output by the learning object recognition unit 112 after tracking processing using the multiple learning image data can be stored in the storage device 120 of the control device 100 in association with each other. Furthermore, the learning object recognition unit 112 can perform a single tracking process using multiple learning image data and output one output data. In this case, one output data can be associated with multiple learning image data. The learning object recognition unit 112 can perform multiple tracking processes using multiple learning image data and output multiple output data. In this case, multiple output data can be associated with multiple learning image data.

[0278] Furthermore, the "consistency involved in the tracking process" can be an indicator that varies according to at least one of the following: the change in at least one of the position and posture of the learning object calculated based on the result of the tracking process performed using the learning object recognition unit 112 (the change calculated by the learning object recognition unit 112), and the difference between the change in at least one of the position and posture of the learning object represented by the displacement data involved in the tracking process associated with multiple learning image data (the displacement data generated by the data generation unit 111); and the ratio of the change in at least one of the position and posture of the learning object calculated based on the result of the tracking process performed using the learning object recognition unit 112 (the change calculated by the learning object recognition unit 112), and the change in at least one of the position and posture of the learning object represented by the displacement data involved in the tracking process associated with multiple learning image data (the displacement data generated by the data generation unit 111). When the consistency involved in the tracking process varies based on the difference between the change in at least one of the position and posture of the learning object calculated by the learning object recognition unit 112 (the change calculated by the learning object recognition unit 112) and the change in at least one of the position and posture of the learning object represented by the displacement data (displacement data generated by the data generation unit 111) involved in the tracking process associated with multiple learning image data, the smaller the difference (in other words, the closer it is to 0), the higher the consistency involved in the tracking process can be. When the consistency involved in the tracking process varies based on the ratio of the change in at least one of the position and posture of the learning object calculated by the learning object recognition unit 112 (the change calculated by the learning object recognition unit 112) and the change in at least one of the position and posture of the learning object represented by the displacement data (displacement data generated by the data generation unit 111) involved in the tracking process associated with multiple learning image data, the closer the ratio is to 1, the higher the consistency involved in the tracking process can be.

[0279] Teacher data can be generated from learning image data and one or more processing parameters represented by output data associated with the learning image data. Specifically, teacher data can be generated that uses learning image data as input data (referred to as a "case") and one or more processing parameters contained in the output data associated with the learning image data as the positive solution data.

[0280] Here, as described later, at least one of the position and pose of the learning object in the shooting coordinate system represented by the position and pose data generated by the data generation unit 111 is infinitely close to the true value (refer to "(4-3-1) Position and Pose Data Generation Process"). As described above, one or more processing parameters, such as those in the matching process, contained in the output data associated with the learning image data are generated by changing one or more processing parameters in the matching process so that at least one of the position and pose of the learning object captured in the image represented by the learning image data, calculated by the learning object recognition unit 112, is close to (typically identical to) at least one of the position and pose of the learning object in the shooting coordinate system represented by the position and pose data associated with the learning image data (the position and pose data generated by the data generation unit 111), which is infinitely close to the true value. Therefore, if the learning object recognition unit 112 performs matching processing using the generated processing parameters (i.e., one or more processing parameters in the matching process), the learning object recognition unit 112 can calculate at least one of the position and pose of the learning object captured in the image represented by the learning image data, which is infinitely close to the true value, in the shooting coordinate system. Therefore, it can be said that the one or more processing parameters involved in the matching process contained in the output data are processing parameters that are infinitely close to the true value or are the true value.

[0281] Furthermore, in the output data, if one or more processing parameters from the preprocessing performed on the learning image data used in the matching process are included on the basis of or instead of one or more processing parameters from the matching process, it can be said that the same applies to one or more processing parameters from the preprocessing.

[0282] As described below, the three-dimensional positions of multiple points of the learning object represented by the three-dimensional position data generated by the data generation unit 111 are infinitely close to the true values ​​(refer to "(4-3-2) Generation Process of Three-Dimensional Position Data"). As described above, one or more processing parameters in the position calculation process contained in the output data associated with the learning image data are generated by changing one or more processing parameters in the position calculation process so that the three-dimensional positions of multiple points of the learning object captured in the image represented by the learning image data, calculated by the learning object recognition unit 112, are close to (typically consistent with) the three-dimensional positions of multiple points of the learning object represented by the three-dimensional position data (three-dimensional position data generated by the data generation unit 111) associated with the learning image data, which are infinitely close to the true values. Therefore, if the learning object recognition unit 112 performs position calculation processing using the generated processing parameters (i.e., one or more processing parameters in the position calculation process), the learning object recognition unit 112 can calculate the three-dimensional positions of multiple points of the learning object captured in the image represented by the learning image data, which are infinitely close to the true values. Therefore, it can be said that the one or more processing parameters involved in the position calculation process contained in the output data are infinitely close to the true value or are the true value.

[0283] Furthermore, in the output data, if one or more processing parameters are included in the preprocessing of the learned image data used in the location calculation process, or in place of one or more processing parameters in the location calculation process, it can be said that the same applies to one or more processing parameters in the preprocessing.

[0284] As described below, the change in at least one of the position and pose of the learning object represented by the displacement data involved in the tracking process (at least one of 2D tracking process and 3D tracking process) generated by the data generation unit 111 is infinitely close to the true value (see "(4-3-3) Other"). As described above, one or more processing parameters in the tracking process contained in the output data associated with multiple learning image data captured at different times are generated by modifying one or more processing parameters in the tracking process so that the change in at least one of the position and pose of the learning object calculated based on the result of the tracking process, and the change in at least one of the position and pose of the learning object represented by the displacement data (displacement data generated by the data generation unit 111) involved in the tracking process associated with the multiple learning image data are close (typically consistent). Therefore, if the learning object recognition unit 112 performs tracking processing using the generated processing parameters (i.e., one or more processing parameters in the tracking processing), the learning object recognition unit 112 can calculate the change in at least one of the position and pose of the learning object captured in the image represented by multiple learning image data, which is infinitely close to the true value. Therefore, it can be said that one or more processing parameters in the tracking processing contained in the output data are processing parameters that are infinitely close to the true value or are the true value.

[0285] Furthermore, in the output data, if one or more processing parameters are included in the preprocessing of at least one of the multiple learning image data used in the tracking process, based on or instead of one or more processing parameters in the tracking process, it can be said that the same applies to one or more processing parameters in the preprocessing.

[0286] Therefore, one or more processing parameters contained in the output data associated with the learning image data can also be called forward retrieval data. Furthermore, as mentioned above, changing the processing parameters can be called optimizing the processing parameters. Therefore, one or more processing parameters contained in the output data associated with the learning image data can also be called optimized parameters. Thus, it can be said that teacher data can be generated from the learning image data and the optimized parameters.

[0287] When the learning object recognition unit 112 performs matching processing and position calculation processing using a learning image data, one or more processing parameters from the matching processing and the position calculation processing may also be included in the output data associated with the learning image data. When the learning object recognition unit 112 performs matching processing using a learning image data but does not perform position calculation processing, one or more processing parameters from the matching processing may be included in the output data associated with the learning image data. When the learning object recognition unit 112 performs position calculation processing but does not perform matching processing, one or more processing parameters from the position calculation processing may be included in the output data associated with the learning image data. In either case, the output data may also include one or more processing parameters from the preprocessing performed on the learning image data.

[0288] Furthermore, when multiple output data are associated with a single learning image data, teacher data can be generated that uses one or more processing parameters contained in one of the multiple output data as the positive solution data. As a method for selecting one output data from multiple output data, at least one of the following methods can be listed: (i) selecting the output data representing the highest consistency among the consistency values ​​represented by each of the multiple output data; (ii) selecting the output data representing the shortest processing time among the processing times represented by each of the multiple output data; (iii) selecting the output data representing a consistency value above a consistency threshold among the consistency values ​​represented by each of the multiple output data; (iv) selecting the output data representing a processing time below a time threshold among the processing times represented by each of the multiple output data; and (v) selecting the output data with the largest product of consistency value and the inverse of processing time based on the consistency value and processing time represented by each of the multiple output data. Furthermore, when multiple output data are associated with a single learning image data, multiple sets of teacher data can be generated that use a single learning image data as input data.

[0289] In addition, teacher data can also be generated as follows: one learning image data and another learning image data captured at different times are used as input data, and one or more processing parameters (e.g., one or more processing parameters of at least one of 2D tracking processing and 3D tracking processing) represented by the output data associated with the one learning image data and the other learning image data are used as teacher data for forward solving. Furthermore, the one or more processing parameters represented by the output data associated with the one learning image data and the other learning image data may include, in addition to one or more processing parameters of at least one of 2D tracking processing and 3D tracking processing, one or more processing parameters such as preprocessing performed on at least one of the one learning image data and the other learning image data.

[0290] Furthermore, when virtual data exists as learning image data, the learning object recognition unit 112 can use the virtual data to perform matching processing to calculate at least one of the position and pose of the learning object captured in the image represented by the virtual data. At this time, the learning object recognition unit 112 can change one or more processing parameters in at least one of the matching processing and the preprocessing performed on the virtual data used in the matching processing, so that at least one of the calculated position and pose of the learning object (i.e., the output value of the learning object recognition unit 112) is close to (typically consistent with) the position and pose of the learning object represented by the position and pose data associated with the virtual data used in the matching processing (the position and pose data generated by the data generation unit 111).

[0291] In this case, the virtual data and the output data output by the learning object recognition unit 112 after matching the virtual data can be stored in the storage device 120 of the control device 100 in association with each other. Furthermore, the output data output after matching the virtual data can be at least one of the following: a value representing the consistency degree corresponding to the consistency degree, the time required for the matching process, and the value of one or more processing parameters when the learning object recognition unit 112 calculates at least one of the position and posture of the learning object.

[0292] Furthermore, the learning object recognition unit 112 can use virtual data to perform position calculation processing to generate three-dimensional position data. At this time, the learning object recognition unit 112 can change one or more processing parameters in at least one of the position calculation processing and the preprocessing performed on the virtual data used in the position calculation processing, so that the three-dimensional positions of the multiple points of the learning object represented by the generated three-dimensional position data (i.e., the output values ​​of the learning object recognition unit 112) are close to (typically consistent with) the three-dimensional positions of the multiple points of the learning object represented by the three-dimensional position data associated with the virtual data used in the position calculation processing (the three-dimensional position data generated by the data generation unit 111).

[0293] In this case, the virtual data and the output data output by the learning object recognition unit 112 after performing position calculation processing using the virtual data can be stored in the storage device 120 of the control device 100 in association with each other. Furthermore, the output data output after performing position calculation processing using the virtual data can be at least one of the following: a consistency value corresponding to the consistency, the time required for the position calculation processing, and a value of one or more processing parameters when the learning object recognition unit 112 calculates the three-dimensional positions of multiple points of the learning object.

[0294] In the presence of virtual data as learning image data, teacher data can be generated from the virtual data and the values ​​of one or more processing parameters represented by the output data associated with the virtual data. That is, teacher data can be generated that uses the virtual data as learning image data as input data and the values ​​of one or more processing parameters represented by the output data associated with the virtual data as the forward solution data.

[0295] The learning unit 113 of the computing device 110 generates (constructs) a parameter determination model by learning using teacher data. The learning unit 113 can generate the parameter determination model by performing machine learning using teacher data. In this case, as an example, a regression model can be generated as the parameter determination model by learning the relationship between the learning image data contained in the teacher data and one or more processing parameters that serve as positive solution data. This regression model is used to output (calculate) one or more processing parameters that are positive or close to positive solutions based on the image data (processing object image data) generated and input by the imaging unit 20 after (i.e., during actual operation after learning) of the object to be processed. Furthermore, the learning unit 113 is not limited to machine learning; it can generate the parameter determination model by performing deep learning using teacher data. Furthermore, "during actual operation" refers to the time when the control device 100 performs robot control processing to control at least one of the robot 10 and the end effector 13, causing the end effector 13 to perform prescribed processing on the object to be processed (e.g., workpiece W).

[0296] In the learning process using teacher data in the learning unit 113, the model 1131, which subsequently becomes a parameter-determining model, can determine the values ​​of one or more processing parameters from at least one of matching processing, position calculation processing, tracking processing, and preprocessing for the learning image data that serves as input data. The learning unit 113 can learn from the model 1131, which subsequently becomes a parameter-determining model, such that the values ​​of one or more processing parameters output from the model 1131 are close to (typically consistent with) the values ​​of the optimal parameters (i.e., the processing parameters as positive solution data) contained in the teacher data.

[0297] Furthermore, the object recognition unit 112 can perform at least one of matching processing and position calculation processing according to the first algorithm. Additionally, the parameter determination model generated by the learning unit 113 can be based on the structure of the second algorithm. The first and second algorithms can be algorithms involved in machine learning such as Differential Evolution (DE), nearest neighbor method, Naive Bayes method, decision tree, and support vector machine, or algorithms involved in deep learning that utilize neural networks to generate feature quantities and combine weighted coefficients. As a deep learning model, a Convolutional Neural Network (CNN) model can be used. Furthermore, the first and second algorithms can be the same or different.

[0298] (4-2) Learning image data acquisition and processing

[0299] Reference Figure 7 An example of acquiring and processing learning image data will be illustrated. Figure 7This diagram illustrates an example of the positional relationship between the imaging unit 20 and the learning object. Furthermore, the imaging unit 20 may include imaging devices 21 and 22, and a projection device 23; it may also include only imaging device 21, only imaging device 22, or both imaging device 22 and projection device 23. Hereinafter, for ease of explanation, it will be referred to as "imaging unit 20". By replacing "imaging unit 20" with at least one of "imaging device 21" and "imaging device 22", it can be used as a description of at least one of the following: a description of processing for acquiring learning image data using "imaging device 21", and a description of processing for acquiring learning image data using "imaging device 22".

[0300] exist Figure 7 In the process, the learning object is positioned on a platform ST marked with a marker M. Figure 7 In this configuration, four markers M are set on the platform ST. Alternatively, only one marker M may be set on the platform ST. Or, more than two markers M may be set on the platform ST. That is, at least one marker M may be set on the platform ST.

[0301] The marker M can be a marker representing information related to the stage ST, such as its position and orientation. The marker M can be an Augmented Reality (AR) marker (also called an AR tag). The marker M may not be an AR marker; it can be any other two-dimensional barcode. The marker M can also be one or more markers used to calculate the position and orientation of the stage ST. The marker M can be multiple markers (e.g., three markers) capable of calculating the position and orientation of the stage ST (e.g., at least one of a cross marker and a well-shaped marker). The marker M can be separate from the stage ST, or it can be formed on the stage ST by existing processing methods such as machining (i.e., the marker M can also be integrally formed with the stage ST).

[0302] Furthermore, the position and orientation of the stage ST can be in the global coordinate system, the 2D imaging coordinate system, or the 3D imaging coordinate system. The position and orientation of the stage ST can also be in a coordinate system different from the global coordinate system, the 2D imaging coordinate system, and the 3D imaging coordinate system.

[0303] When the imaging unit 20 photographs the learning object, the signal generation unit 116 of the control device 100 can generate a robot control signal to drive the robot 10, so that the imaging unit 20 is in a predetermined position relative to the learning object. Additionally, the signal generation unit 116 can generate a photographing control signal to cause the imaging unit 20 to photograph the learning object under the predetermined positional relationship. Furthermore, the learning object can be photographed by the imaging device 21, by the imaging device 22, or by both the imaging device 21 and the imaging device 22.

[0304] As an example of the positional relationship specified above, the following can be listed: Figure 7 The first positional relationship, the second positional relationship, and the third positional relationship in the text. Figure 7 The dashed arrows in the diagram represent the optical axis of the imaging unit 20 (e.g., at least one of imaging devices 21 and 22). The first positional relationship, the second positional relationship, and the third positional relationship can each be derived from a learning coordinate system (refer to...) with a point on the optical axis of the imaging unit 20 as the origin. Figure 7 The origin of the learning coordinate system can be a point on the learning object, a point on the platform ST, or a point in space. Hereafter, as an example, the case where the origin of the learning coordinate system is located on the learning object will be explained. Figure 7 As shown, the learning coordinate system can be a coordinate system defined by mutually orthogonal X-axis, Y-axis, and Z-axis. One axis of the learning coordinate system (e.g., the Z-axis) can be the axis along the optical axis of the imaging unit 20 in the first positional relationship. Furthermore, one axis of the learning coordinate system (e.g., the Z-axis) may not be along the optical axis of the imaging unit 20 in the first positional relationship, but may be along the optical axis of the imaging unit 20 in other positional relationships such as the second positional relationship. Furthermore, if the positional relationship of the imaging unit 20 relative to the learning object can be defined, then one axis of the learning coordinate system (e.g., the Z-axis) may not be along the optical axis of the imaging unit 20. One axis of the learning coordinate system can be an axis along the vertically upward direction. Furthermore, the relationship between the learning coordinate system and the global coordinate system is assumed to be known. That is, a transformation matrix for converting the position and pose within the learning coordinate system to the position and pose within the global coordinate system can be used to convert the defined positional relationship in the learning coordinate system to the defined positional relationship in the global coordinate system. The transformation matrix can be calculated based on the relationship between the learning coordinate system and the global coordinate system.

[0305] In the learning coordinate system, the first positional relationship can be defined as the pose of the imaging unit 20 about the X-axis (in other words, the rotation of the imaging unit 20 about the X-axis), the pose of the imaging unit 20 about the Y-axis (in other words, the rotation of the imaging unit 20 about the Y-axis), and the position of the imaging unit 20 in the Z-axis direction parallel to the Z-axis. Furthermore, the pose of the imaging unit 20 about the X-axis and the pose of the imaging unit 20 about the Y-axis can also be considered as parameters representing the position of the imaging unit 20. Therefore, the pose of the imaging unit 20 about the X-axis and the pose of the imaging unit 20 about the Y-axis can be respectively referred to as the position of the imaging unit 20 in the rotation direction about the X-axis and the position of the imaging unit 20 in the rotation direction about the Y-axis. The position of the imaging unit 20 in the Z-axis direction can be referred to as the distance from the learning object to the imaging unit 20.

[0306] Furthermore, when at least one of the imaging unit 20 and the learning object moves, the first positional relationship can also be defined as the relative posture of the imaging unit 20 and the learning object around the X-axis, the relative posture of the imaging unit 20 and the learning object around the Y-axis, and the relative position of the imaging unit 20 and the learning object in the Z-axis direction parallel to the Z-axis. Furthermore, when the imaging unit 20 is stationary (i.e., the imaging unit 20 does not move) and the learning object moves on the other hand, the first positional relationship can be defined as the posture of the learning object around the X-axis, the posture of the learning object around the Y-axis, and the position of the learning object in the Z-axis direction parallel to the Z-axis. In this case, the position of the learning object in the Z-axis direction can be referred to as the distance from the imaging unit 20 to the learning object.

[0307] As described above, the learning coordinate system can be a coordinate system defined by mutually orthogonal X-axis, Y-axis, and Z-axis. The Z-axis of the learning coordinate system can be the axis along the optical axis of the imaging unit 20 in the first positional relationship. If the Z-axis is called the first axis, the X-axis is called the second axis, and the Y-axis is called the third axis, then it can be said that the posture of the imaging unit 20 can be the posture of the learning object in the coordinate system defined by the first axis along the optical axis of the optical system of the imaging unit 20, the second axis orthogonal to the first axis, and the third axis orthogonal to the first and second axes, i.e., the posture about at least one of the first, second, and third axes in the learning coordinate system.

[0308] In the learning coordinate system, the second positional relationship can be defined as the posture of the imaging unit 20 around the X-axis, the posture of the imaging unit 20 around the Y-axis, and the position of the imaging unit 20 in the Z-axis direction parallel to the Z-axis. The positional relationship between the imaging unit 20 and the learning object in the learning coordinate system under the second positional relationship differs from the positional relationships of the first and third positional relationships. Furthermore, if one axis of the learning coordinate system (e.g., the Z-axis) is along the optical axis of the imaging unit 20 in the first positional relationship, then that axis (e.g., the Z-axis) may not be along the optical axis of the imaging unit 20 in the second positional relationship. Additionally, the first, second, and third positional relationships can also be defined in a learning coordinate system where one of the mutually orthogonal X-axis, Y-axis, and Z-axis (e.g., the Z-axis) is along the optical axis of the imaging unit 20 in the second positional relationship.

[0309] In the learning coordinate system, the third positional relationship can be defined as the posture of the imaging unit 20 around the X-axis, the posture of the imaging unit 20 around the Y-axis, and the position of the imaging unit 20 in the Z-axis direction parallel to the Z-axis. The positional relationship between the imaging unit 20 and the learning object in the learning coordinate system under the third positional relationship differs from the positional relationships of the first and second positional relationships. Furthermore, if one axis of the learning coordinate system (e.g., the Z-axis) is along the optical axis of the imaging unit 20 in the first positional relationship, then that axis (e.g., the Z-axis) may not be along the optical axis of the imaging unit 20 in the third positional relationship. Additionally, the first, second, and third positional relationships can also be defined in a learning coordinate system where one of the mutually orthogonal X-axis, Y-axis, and Z-axis (e.g., the Z-axis) is along the optical axis of the imaging unit 20 in the third positional relationship.

[0310] For example, when the imaging unit 20 is used to photograph the learning object under a first positional relationship, the signal generation unit 116 can use a transformation matrix to convert the position and posture in the learning coordinate system into the position and posture in the global coordinate system, and convert the first positional relationship defined in the learning object coordinate system into the first positional relationship in the global coordinate system. The signal generation unit 116 can generate a robot control signal for controlling the movement of the robot 10 (robot arm 12), so that the positional relationship of the imaging unit 20 relative to the learning object becomes the first positional relationship in the global coordinate system. The signal generation unit 116 can use the communication device 130 to output the generated robot control signal to the robot 10 (e.g., robot control device 14). As a result, the robot control device 14 can also control the movement of the robot 10 (e.g., the movement of the robot arm 12) based on the robot control signal, and set the positional relationship of the imaging unit 20 relative to the learning object as the first positional relationship.

[0311] Furthermore, the specified positional relationship is not limited to the positional relationship between the shooting unit 20 and the learning object, but can also refer to the positional relationship between the learning object and the shooting unit 20. That is, the specified positional relationship can also refer to the relative positional relationship between the learning object and the shooting unit 20. The specified positional relationship can include a first positional relationship, a second positional relationship, and a third positional relationship. Therefore, the specified positional relationship can also refer to multiple different positional relationships between the learning object and the shooting unit 20.

[0312] For example, the operation of the control device 100 when the shooting unit 20 shoots the learning object under the first, second, and third positional relationships will be explained. Furthermore, the shooting unit 20 may also shoot the learning object under at least one of the first, second, and third positional relationships. The shooting unit 20 may also shoot the learning object under positional relationships different from the first, second, and third positional relationships. In such cases, the shooting unit 20 may shoot the learning object under these different positional relationships based on the first, second, and third positional relationships, or it may shoot the learning object under these different positional relationships instead of the first, second, and third positional relationships. Furthermore, these different positional relationships may be one or multiple positional relationships.

[0313] The signal generation unit 116 can generate a shooting control signal for causing the shooting unit 20 to shoot the learning object under a first positional relationship. In this case, the signal generation unit 116 first uses a transformation matrix for converting the position and posture in the learning coordinate system to the position and posture in the global coordinate system, and converts the first positional relationship specified in the learning coordinate system into a first positional relationship in the global coordinate system. The signal generation unit 116 can generate a robot control signal representing the first positional relationship between the shooting unit 20 and the learning object based on the first positional relationship in the global coordinate system. The signal generation unit 116 can output the generated robot control signal to the robot 10 using the communication device 130. The signal generation unit 116 can generate a shooting control signal for causing the shooting unit 20 to shoot the learning object under the first positional relationship. The signal generation unit 116 can output the generated shooting control signal to the shooting unit 20 using the communication device 130.

[0314] As a result, after the relationship between the imaging unit 20 and the learning object becomes a first positional relationship, the imaging unit 20 (e.g., at least one of the imaging devices 21 and 22) can generate learning image data by photographing the learning object. Furthermore, the signal generation unit 116 can generate an imaging control signal, causing the imaging unit 20 to photograph the learning object while the positional relationship between the imaging unit 20 and the learning object becomes the first positional relationship, thereby generating learning image data.

[0315] The signal generation unit 116 can similarly generate a shooting control signal for causing the shooting unit 20 to shoot the learning object under the second positional relationship. In this case, the signal generation unit 116 can use a transformation matrix for converting the position and posture in the learning coordinate system to the position and posture in the global coordinate system, and convert the second positional relationship specified in the learning coordinate system to the second positional relationship in the global coordinate system. Furthermore, the transformation matrix may be different from the transformation matrix used to convert the first positional relationship specified in the learning coordinate system to the first positional relationship in the global coordinate system. The signal generation unit 116 can generate a robot control signal representing the second positional relationship between the shooting unit 20 and the learning object based on the second positional relationship in the global coordinate system. The signal generation unit 116 can output the generated robot control signal to the robot 10 using the communication device 130. The signal generation unit 116 can generate a shooting control signal for causing the shooting unit 20 to shoot the learning object under the second positional relationship. The signal generation unit 116 can output the generated shooting control signal to the shooting unit 20 using the communication device 130.

[0316] As a result, after the relationship between the imaging unit 20 and the learning object becomes a second positional relationship, the imaging unit 20 (e.g., at least one of the imaging devices 21 and 22) can generate learning image data by photographing the learning object. Furthermore, the signal generation unit 116 can generate an imaging control signal, causing the imaging unit 20 to photograph the learning object while the positional relationship between the imaging unit 20 and the learning object becomes a second positional relationship, thereby generating learning image data.

[0317] The signal generation unit 116 can similarly generate a shooting control signal for causing the shooting unit 20 to shoot the learning object under a third positional relationship. In this case, the signal generation unit 116 can use a transformation matrix for converting the position in the learning coordinate system to the position in the global coordinate system, and convert the third positional relationship specified in the learning coordinate system to the third positional relationship in the global coordinate system. Furthermore, the transformation matrix may be different from the transformation matrix used to convert the first positional relationship specified in the learning coordinate system to the first positional relationship in the global coordinate system, and the transformation matrix used to convert the second positional relationship specified in the learning coordinate system to the second positional relationship in the global coordinate system. The signal generation unit 116 can generate a robot control signal representing the third positional relationship between the shooting unit 20 and the learning object based on the third positional relationship in the global coordinate system. The signal generation unit 116 can output the generated robot control signal to the robot 10 using the communication device 130. The signal generation unit 116 can generate a shooting control signal for causing the shooting unit 20 to shoot the learning object under a third positional relationship. The signal generation unit 116 can use the communication device 130 to output the generated shooting control signal to the shooting unit 20.

[0318] As a result, after the relationship between the imaging unit 20 and the learning object becomes a third positional relationship, the imaging unit 20 (e.g., at least one of the imaging devices 21 and 22) can generate learning image data by photographing the learning object. Furthermore, the signal generation unit 116 can generate an imaging control signal, enabling the imaging unit 20 to photograph the learning object while the positional relationship between the imaging unit 20 and the learning object becomes a third positional relationship, thereby generating learning image data.

[0319] Through this series of actions, the imaging unit 20 captures images of the learning object under the first, second, and third positional relationships, thereby generating (acquiring) learning image data. Furthermore, as mentioned above, the positional relationship in which the imaging unit 20 captures the learning object is not limited to the first, second, and third positional relationships. Moreover, the defined positional relationships can include multiple different positional relationships such as the first, second, and third positional relationships. Therefore, it can be said that the defined positional relationships can be multiple different positional relationships of the imaging unit 20 relative to the learning object.

[0320] As described above, the signal generation unit 116 can generate robot control signals for controlling the robot 10 and shooting control signals for controlling the shooting unit 20. Thus, the signal generation unit 116, which generates both robot control signals for controlling the robot 10 and shooting control signals for controlling the shooting unit 20, can also be called a control unit. The signal generation unit 116 can output robot control signals to the robot 10 using the communication device 130. Additionally, the signal generation unit 116 can output shooting control signals to the shooting unit 20 using the communication device 130. Therefore, it can be said that the signal generation unit 116 can output control signals (e.g., shooting control signals and robot control signals) for controlling the shooting unit 20 and the robot 10 on which the shooting unit 20 is installed. Furthermore, the signal generation unit 116 can also output both shooting control signals for controlling the shooting unit 20 and robot control signals for controlling the robot 10. The signal generation unit 116 can also output common control signals to the shooting unit 20 and the robot 10.

[0321] As described above, the signal generation unit 116 can generate a robot control signal representing a predetermined positional relationship (e.g., at least one of a first positional relationship, a second positional relationship, and a third positional relationship) between the imaging unit 20 and the learning object. The signal generation unit 116 can output the robot control signal to the robot 10 using the communication device 130. The signal generation unit 116 can also generate an imaging control signal for causing the imaging unit 20 to image the learning object under the predetermined positional relationship. The signal generation unit 116 can output the imaging control signal to the imaging unit 20 using the communication device 130. Here, considering the control of the robot 10's movements based on the robot control signal, it can be said that the robot control signal is a control signal used to drive the robot 10. Therefore, it can be said that the signal generation unit 116 can output control signals (e.g., robot control signals and imaging control signals) for driving the robot 10 such that the imaging unit 20 is in a predetermined positional relationship relative to the learning object, and that the imaging unit 20 images the learning object under the predetermined positional relationship.

[0322] Furthermore, the control signal used to drive the robot 10, causing the imaging unit 20 to be in a predetermined position relative to the learning object and to capture images of the learning object under that predetermined position, is a signal that controls both the robot 10 and the imaging unit 20, and can therefore be called the first control signal. Thus, the first control signal can be a control signal used to drive the robot 10, causing the imaging unit 20 to be in a predetermined position relative to the learning object and to capture images of the learning object under that predetermined position. Therefore, it can be said that the first control signal (e.g., at least one of the robot control signal and the imaging control signal) can be a signal used to drive the robot 10 to change to various of a plurality of positional relationships, and to cause the imaging unit 20 to capture images of the learning object each time it changes to one of the plurality of positional relationships.

[0323] As described above, the learning image data is generated by photographing the learning object by the imaging unit 20 (e.g., at least one of the imaging devices 21 and 22). That is, the learning image data can be described as data representing an image of the learning object captured by the imaging unit 20. Furthermore, the learning image data may include both image data IMG_2D generated by the imaging device 21 and image data IMG_3D generated by the imaging device 22. Specifically, the learning image data may include, for example, image data generated by a single monocular camera of the imaging device 21, and two image data generated by the two monocular cameras of the stereo camera of the imaging device 22. Furthermore, the learning image data may also include image data IMG_2D but not image data IMG_3D. Additionally, the learning image data may also include image data IMG_3D but not image data IMG_2D.

[0324] Furthermore, the image data IMG_3D, which serves as learning image data, can also be generated by the imaging device 22 photographing a learning object onto which the desired projection pattern from the projection device 23 of the imaging unit 20 is projected. Alternatively, the image data IMG_3D, which serves as learning image data, can also be generated by the imaging device 22 photographing a learning object onto which the desired projection pattern is not projected.

[0325] That is, when the imaging unit 20 includes a stereo camera as an imaging device 22 and a projection device 23, the stereo camera as the imaging device 22 is used to photograph a learning object from which a desired pattern of light (e.g., a random dot pattern) is projected from the projection device 23 in a first positional relationship, thereby generating image data IMG_3D as learning image data. The stereo camera as the imaging device 22 is used to photograph a learning object from which patterned light is projected from the projection device 23 in a second positional relationship, thereby generating image data IMG_3D as learning image data. The stereo camera as the imaging device 22 is used to photograph a learning object from which patterned light is projected from the projection device 23 in a third positional relationship, thereby generating image data IMG_3D as learning image data.

[0326] On the other hand, when the imaging unit 20 includes a stereo camera (as imaging device 22) and a projection device 23, the stereo camera (as imaging device 22) is used to photograph the learning object in a first positional relationship where the desired pattern light is not projected from the projection device 23, thereby generating image data IMG_3D as learning image data. The stereo camera (as imaging device 22) is used to photograph the learning object in a second positional relationship where the pattern light is not projected from the projection device 23, thereby generating image data IMG_3D as learning image data. The stereo camera (as imaging device 22) is used to photograph the learning object in a third positional relationship where the pattern light is not projected from the projection device 23, thereby generating image data IMG_3D as learning image data. In this case, the imaging unit 20 may also not include the projection device 23.

[0327] The signal generation unit 116 can generate robot control signals and shooting control signals, so that the shooting unit 20 can shoot the learning object in the order of the first position relationship, the second position relationship, and the third position relationship. However, the shooting order is not limited to this.

[0328] As described above, the prescribed positional relationship is not limited to the first, second, and third positional relationships. The prescribed positional relationship can be four or more different positional relationships, or it can be a single positional relationship. Here, the prescribed positional relationship can be automatically set by the control device 100 (e.g., the signal generation unit 116). Specifically, the control device 100 (e.g., the signal generation unit 116) can automatically set one or more prescribed positional relationships based on three-dimensional model data, where the three-dimensional model data represents the three-dimensional model of the learning object represented by the CAD data corresponding to the learning object (e.g., the CAD model of the learning object). The control device 100 can automatically set one or more prescribed positional relationships based on image data captured by the imaging unit 20 that includes the entire learning object (so-called image data captured by scaling down).

[0329] In order to set a predetermined positional relationship, the control device 100 (e.g., signal generation unit 116) can display, as an output device 150, the following on a display screen: Figure 8 The input screen shown. The user of robot system 1 can input the range of the pose of the imaging unit 20 about the X-axis in the learning coordinate system (equivalent to...) via input device 140. Figure 8 The "roll angle" and the range of pose of the shooting unit 20 around the Y-axis in the learning coordinate system (equivalent to) Figure 8 The "pitch" angle and the range of pose of the shooting unit 20 around the Z-axis in the learning coordinate system (equivalent to...) Figure 8 The "yaw angle" and the range of distance from the learning object to the imaging unit 20 (equivalent to...) Figure 8 (The "Z width"). Furthermore, such as... Figure 8 As shown, even when the input screen contains input fields for "roll angle," "pitch angle," "yaw angle," and "Z-width," the user of robot system 1 can still input at least one of "roll angle," "pitch angle," "yaw angle," and "Z-width" (i.e., there may be uninputted items). Furthermore, the user can also input either the posture or distance in the learning coordinate system. The posture in the learning coordinate system that the user can input is not limited to "roll angle," "pitch angle," and "yaw angle," but can also be at least one of "roll angle," "pitch angle," and "yaw angle."

[0330] The control device 100 (e.g., signal generation unit 116) can automatically set one or more predetermined positional relationships within the range where the input positional relationship is changed. The control device 100 (e.g., signal generation unit 116) can set positional relationships at predetermined intervals within the range where the input positional relationship is changed (e.g., the range of at least one of posture change and distance change). Furthermore, in this case, if the predetermined interval becomes narrower, the number of predetermined positional relationships set increases; if the predetermined interval becomes wider, the number of predetermined positional relationships set decreases. Moreover, the predetermined interval can be input by the user via the input device 140 or preset.

[0331] As described above, by having the user input the range of positional relationship changes via the input device 140, the control device 100 (e.g., the signal generation unit 116) can automatically generate one or more positional relationships. Therefore, the user inputting the range of positional relationship changes via the input device 140 can be interpreted as the control device 100 (e.g., the signal generation unit 116) receiving input of the range of positional relationship changes. Furthermore, it can be said that the control device 100 (e.g., the signal generation unit 116) can determine (set) a predetermined positional relationship within the range of positional relationship changes based on the input (in other words, the received) range of positional relationship changes.

[0332] Here, "the range of changes in positional relationship" refers to the range within which the positional relationship between the shooting unit 20 and the learning object is changed when the shooting unit 20 shoots the learning object. In other words, "the range of changes in positional relationship" can also be described as the range of possible positional relationships between the shooting unit 20 and the learning object when the shooting unit 20 shoots the learning object. The positional relationship between the shooting unit 20 and the learning object can be represented by at least one of the pose and distance of the shooting unit 20 in the learning coordinate system, or by at least one of the pose and distance of the learning object in the learning coordinate system, or by at least one of the pose and distance of the shooting unit 20 and the learning object relative to the shooting unit 20 and the learning object in the learning coordinate system. Furthermore, "pose" is not limited to "roll angle," "pitch angle," and "yaw angle," but can also be at least one of "roll angle," "pitch angle," and "yaw angle."

[0333] The phrase “the extent to which the positional relationship is changed based on the input (in other words, the received)” is not limited to the meaning of “the extent to which the positional relationship is changed based solely on the input”, but may include the meaning of “at least based on the extent to which the positional relationship is changed (i.e., based on other information in addition to the extent to which the positional relationship is changed)” and “based on at least a portion of the extent to which the positional relationship is changed”.

[0334] As described above, the learning image data is used in the learning process to generate the parameter determination model. The parameter determination model outputs (calculates) one or more processing parameters that are correct or close to correct based on image data (hereinafter referred to as processing object image data) generated by the imaging unit 20 photographing the processing object (e.g., workpiece W) during actual operation after learning. Here, the predetermined positional relationship between the imaging unit 20 and the learning object when the imaging unit 20 photographs the learning object can be set to be independent of the positional relationship between the imaging unit 20 and the processing object during actual operation. When using learning image data—that is, learning image data generated by the imaging unit 20 photographing the learning object under the positional relationship where the imaging unit 20 does not photograph the processing object during actual operation—to learn and thereby generate the parameter determination model, there is a concern that one or more processing parameters output by the parameter determination model based on the image data generated by the imaging unit 20 photographing the processing object during actual operation may deviate from the correct solution. In other words, the following learning image data—that is, the learning image data generated by the shooting unit 20 taking pictures of the learning object under the positional relationship between the shooting unit 20 and the learning object, such as when the shooting unit 20 does not take pictures of the object being processed during actual operation—may become a factor (e.g., noise) that hinders the improvement of the accuracy of the parameter determination model in the learning of the parameter determination model.

[0335] As described above, by limiting the range of positional relationship changes by user input, the positional relationship between the imaging unit 20 and the learning target object can be prevented from being captured during actual operation, where the imaging unit 20 does not capture images of the object being processed. That is, by limiting the range of positional relationship changes, it is possible to prevent the acquisition of learning image data that becomes noise for generating the parameter determination model. Furthermore, the time required for acquiring learning image data (in other words, capturing images of the learning target object) can be suppressed. Moreover, the range of positional relationship changes input by the user via the input device 140 is not limited to the learning coordinate system and can also be the range of other coordinate systems. The range of positional relationship changes can be the range of the global coordinate system. That is, the changed range can be the range of the orientation of the imaging unit 20 around the X-axis (GL), the range of the orientation of the imaging unit 20 around the Y-axis (GL), the range of the orientation of the imaging unit 20 around the Z-axis (GL), and the range of the distance from the learning target object to the imaging unit 20. In the case described above, when the imaging unit 20 captures the learning object under one or more positional relationships, there is no need for coordinate system transformation based on the signal generation unit 116 (e.g., transformation from the learning coordinate system to the global coordinate system).

[0336] also, Figure 8 The input screen shown may not be displayed on the display, which is the output device 150. In this case, it may be displayed on a display included in another device different from the control device 100. Figure 8 The input screen shown indicates that the user can input the range within which the positional relationship is changed via a device other than the control device 100. This other device can automatically set one or more predetermined positional relationships within the input range. The other device can then send a signal representing the set one or more predetermined positional relationships to the control device 100.

[0337] Furthermore, the stage ST can be configured to change its position and orientation. That is, the stage ST can have a mechanism (not shown) capable of changing its position and orientation. In this case, the predetermined positional relationship between the imaging unit 20 and the learning object can be achieved by the action of at least one of the robot 10 (e.g., robot arm 12) and the stage ST. That is, the predetermined positional relationship can be achieved by the action of only the robot 10, by the action of only the stage ST, or by the action of both the robot 10 and the stage ST.

[0338] In this case, the signal generation unit 116 can replace the robot control signal representing the predetermined positional relationship between the imaging unit 20 and the learning object, or generate a platform control signal representing the predetermined positional relationship between the imaging unit 20 and the learning object based thereon. The signal generation unit 116 can use the communication device 130 to output the platform control signal to a mechanism capable of changing the position and orientation of the platform ST. Therefore, it can be said that the signal generation unit 116 can output control signals (e.g., robot control signals, platform control signals, and imaging control signals) to drive at least one of the robot 10 included in the imaging unit 20 and the platform ST on which the learning object is disposed, so that the imaging unit 20 and the learning object are in a predetermined positional relationship, and the imaging unit 20 captures images of the learning object under the predetermined positional relationship.

[0339] As described above, the defined positional relationships may include a first positional relationship, a second positional relationship, and a third positional relationship. In this case, the signal generation unit 116 can generate at least one of a robot control signal and a platform control signal, as well as a shooting control signal, so that the shooting unit 20 can shoot the learning object under each of the first, second, and third positional relationships. The robot control signal, platform control signal, and shooting control signal can be described as signals used to drive at least one of the robot 10 and the platform ST to change to various positions, and each time the positional relationship is changed, the shooting unit 20 can shoot the learning object.

[0340] Furthermore, the control signal that drives at least one of the robot 10 and the platform ST to make the imaging unit 20 and the learning object have a predetermined positional relationship, and enables the imaging unit 20 to photograph the learning object under the predetermined positional relationship, is a signal that controls at least one of the robot 10 and the mechanism that can change the position and posture of the platform ST, as well as the imaging unit 20, and is therefore called the first control signal.

[0341] (4-3) Generation and processing of target data

[0342] As described above, the learning object recognition unit 112 can change one or more processing parameters in the matching process using the learning image data, so that at least one of the position and posture of the learning object captured in the image represented by the learning image data (at least one of the position and posture calculated by the learning object recognition unit 112) is close to (typically identical to) the position and posture of the learning object represented by the position and posture data (the position and posture data generated by the data generation unit 111). Therefore, the position and posture of the learning object represented by the position and posture data generated by the data generation unit 111 can be said to be the target that the position and posture of the learning object calculated by the learning object recognition unit 112 in the matching process using the learning image data is close to. Therefore, the position and posture data generated by the data generation unit 111 can be in other words represented as target data representing the target in the matching process using the learning image data.

[0343] Furthermore, the learning object recognition unit 112 can modify one or more processing parameters during the position calculation process using the learning image data, so that the three-dimensional positions of multiple points of the learning object captured in the image represented by the learning image data (the three-dimensional positions of multiple points calculated by the learning object recognition unit 112) are close to (typically identical to) the three-dimensional positions of multiple points of the learning object represented by the three-dimensional position data (the three-dimensional position data generated by the data generation unit 111). Therefore, the three-dimensional positions of multiple points of the learning object represented by the three-dimensional position data generated by the data generation unit 111 can be considered as close to the target of the three-dimensional positions of multiple points of the learning object calculated by the learning object recognition unit 112 during the position calculation process using the learning image data. Therefore, the three-dimensional position data generated by the data generation unit 111 can be in other words represented as target data representing the target in the position calculation process using the learning image data.

[0344] That is, the position and pose data and the three-dimensional position data generated by the data generation unit 111 can also be referred to as target data. Furthermore, target data can be a concept that includes either position and pose data or three-dimensional position data. That is, target data can also be a concept that does not include the other of position and pose data and three-dimensional position data. Therefore, at least one of position and pose data and three-dimensional position data may not be referred to as target data.

[0345] Since the position and posture data generated by the data generation unit 111 can also be called target data, the position and posture data generated by the data generation unit 111 shall hereafter be referred to as "target position and posture data". Since the three-dimensional position data generated by the data generation unit 111 can also be called target data, the three-dimensional position data generated by the data generation unit 111 shall hereafter be referred to as "target three-dimensional position data".

[0346] When the imaging unit 20 is positioned relative to the learning object and captures a picture of the learning object, the imaging unit 20 captures the learning object such that the image represented by the learning image data generated after the capture of the learning object includes the platform ST (in particular at least one marker M set on the platform ST) and the learning object disposed on the platform ST. That is, the predetermined positional relationship between the imaging unit 20 and the learning object is set such that when the imaging unit 20 captures the learning object, the platform ST (in particular at least one marker M set on the platform ST) and the learning object disposed on the platform ST are captured.

[0347] Will be through Figure 7The learning image data generated by the imaging unit 20 photographing the learning object under the first positional relationship shown is called the first learning image data. In this case, it can be said that the first learning image data can be generated by the imaging unit 20 photographing the learning object under the first positional relationship. As described above, the learning object is disposed on the stage ST. At least one marker M is provided on the stage ST. Therefore, it can be said that the first learning image data can be generated by the imaging unit 20 photographing both the learning object and at least one marker M under the first positional relationship. Multiple markers M may also be provided on the stage ST. Therefore, it can be said that the first learning image data can be generated by the imaging unit 20 photographing both the learning object and at least one of the multiple markers M under the first positional relationship.

[0348] Will be through Figure 7 The learning image data generated by the imaging unit 20 photographing the learning object under the second positional relationship shown is called the second learning image data. In this case, it can be said that the second learning image data can be generated by the imaging unit 20 photographing the learning object under the second positional relationship. As described above, the learning object is disposed on the stage ST. At least one marker M is provided on the stage ST. Therefore, it can be said that the second learning image data can be generated by the imaging unit 20 photographing both the learning object and at least one marker M under the second positional relationship. Multiple markers M may also be provided on the stage ST. Therefore, it can be said that the second learning image data can be generated by the imaging unit 20 photographing both the learning object and at least one other marker M under the second positional relationship.

[0349] Furthermore, the phrase "at least one other of the plurality of marks M" refers to at least one mark M captured by the shooting unit 20 together with the learning object in the second positional relationship, which may be different from at least one mark M captured by the shooting unit 20 together with the learning object in the first positional relationship. Specifically, when the shooting unit 20 captures both the learning object and one mark M together in the second positional relationship, the one mark M may not be included in the at least one mark M captured by the shooting unit 20 together with the learning object in the first positional relationship. When the shooting unit 20 captures both the learning object and multiple marks M together in the second positional relationship, at least a portion of the multiple marks M may not be included in the at least one mark M captured by the shooting unit 20 together with the learning object in the first positional relationship.

[0350] When the platform ST has only one marker M, the first learning image data can be generated by having the imaging unit 20 simultaneously photograph the learning object and the marker M in a first positional relationship. The second learning image data can be generated by having the imaging unit 20 simultaneously photograph the learning object and the marker M in a second positional relationship. In this case, the same marker M is photographed along with the learning object in both the image represented by the first learning image data and the image represented by the second learning image data. Here, for the purpose of explaining the marker M photographed in the multiple images represented by the multiple learning image data, the first learning image data and the second learning image data are listed. However, the imaging unit 20 can generate more than three learning image data sets, or it can generate only one. Furthermore, the first positional relationship and the second positional relationship can be... Figure 7 The first positional relationship and the second positional relationship shown are different.

[0351] Image data (e.g., learned image data) generated by the imaging unit 20 (e.g., at least one of imaging device 21 and imaging device 22) can be assigned camera parameter information representing camera parameters at the time of shooting. Camera parameters are a concept that includes at least one of the camera parameters involved in the single monocular camera as imaging device 21 and the camera parameters involved in the stereo camera as imaging device 22. That is, the camera parameter information may include at least one of information representing the camera parameters involved in the single monocular camera as imaging device 21 and information representing the camera parameters involved in the stereo camera as imaging device 22.

[0352] Furthermore, camera parameters may include at least one of intrinsic parameters, extrinsic parameters, and distortion coefficients. Intrinsic parameters may include at least one of focal length, optical center, and clipping factor. Extrinsic parameters may be determined by rotation matrices and translation vectors.

[0353] Hereinafter, an example of a method by which the data generation unit 111 of the computing device 110 generates target position and pose data and target three-dimensional position data using learning image data will be described.

[0354] (4-3-1) Generation and processing of position and pose data

[0355] The data generation unit 111 detects at least one marker M set on the stage ST in the image represented by the learning image data. Based on the detected marker M, the data generation unit 111 calculates the position and orientation of the stage ST. Furthermore, the data generation unit 111 can also calculate the position and orientation of the stage ST based on the detected marker M and the camera parameter information assigned to the learning image data. The method for calculating the position and orientation of the stage ST can also be the same as existing methods. Therefore, a detailed description of the calculation method is omitted. Furthermore, the calculated position and orientation of the stage ST can be, for example, the position and orientation of the stage ST in a global coordinate system, the position and orientation of the stage ST in a 2D shooting coordinate system, or the position and orientation of the stage ST in a 3D shooting coordinate system. Additionally, the data generation unit 111 can also calculate only one of the position and orientation of the stage ST. That is, the data generation unit 111 can calculate at least one of the position and orientation of the stage ST. Furthermore, at least one of the position and orientation of the platform ST can also be at least one of the position and orientation of a representative point of the platform ST. The representative point of the platform ST can include at least one of the center of the platform ST, the centroid of the platform ST, the vertex of the platform ST, the center of the surface of the platform ST, and the centroid of the surface of the platform ST. The representative point of the platform ST can also be called a feature point of the platform ST. Furthermore, when the data generation unit 111 calculates at least one of the position and orientation of the platform ST, it can use the constraint that the platform ST is a plane.

[0356] As described above, the data generation unit 111 can calculate at least one of the position and orientation of the platform ST based on at least one marker M detected in the image represented by the learning image data. Here, multiple markers M may also be provided on the platform ST. When multiple markers M are provided on the platform ST, when the imaging unit 20 captures a learning object disposed on the platform ST, regardless of the positional relationship between the imaging unit 20 and the learning object, it is expected that at least one marker M will be included within the field of view of the imaging unit 20 (i.e., at least one marker M will be captured in the image represented by the learning image data). Furthermore, when multiple markers M are provided on the platform ST, it is possible to capture two or more markers M in the image represented by the learning image data. For example, the accuracy of at least one of the position and orientation of the platform ST calculated based on two or more markers M is expected to be higher than that of at least one of the position and orientation of the platform ST calculated based on only one marker M. In addition, the calculated position of the platform ST may also be a concept encompassing the three-dimensional positions of multiple points on the platform ST. Data representing the three-dimensional positions of the multiple points on the platform ST may be, for example, depth image data, point group data, etc.

[0357] Thus, multiple markers M can be set on the stage ST, ensuring that at least one marker M is included within the field of view of the imaging unit 20 regardless of the positional relationship between the imaging unit 20 and the learning object (i.e., at least one marker M is captured in the image represented by the learning image data). Furthermore, to accurately calculate at least one of the position and pose of the representative point (e.g., the center of gravity) of the stage ST, and at least one of the three-dimensional positions (e.g., depth image, point group) of multiple points on the stage ST, multiple markers M can be set on the stage ST.

[0358] The data generation unit 111 can calculate the relative position and pose of the platform ST and the learning object through matching processing using the learned image data. Furthermore, in addition to using the image represented by the learned image data, the data generation unit 111 can also use camera parameter information assigned to the learned image data for matching processing, thereby calculating the relative position and pose of the platform ST and the learning object. Furthermore, the data generation unit 111 can also calculate only one of the relative position and pose of the platform ST and the learning object. That is, the data generation unit 111 can calculate at least one of the relative position and pose of the platform ST and the learning object.

[0359] When performing 2D matching processing, the following processing can be performed. In this case, the learning image data used in the 2D matching processing can be image data generated by taking a picture of the learning object by a single monocular camera, which is the shooting device 21 included in the shooting unit 20, in a specified positional relationship. Here, the image data is referred to as "image data IMG_2D#p1".

[0360] Furthermore, typically, the desired projection pattern (e.g., a random dot pattern) is not projected onto the learning object captured in the image represented by image data IMG_2D#p1. That is, typically, image data IMG_2D#p1 is image data generated by capturing the learning object without projecting the desired projection pattern using a single monocular camera in a specified positional relationship.

[0361] Suppose that when the desired projection pattern is projected onto the learning object, the desired projection pattern is also projected onto at least a portion of the stage ST on which the learning object is disposed. Therefore, the desired projection pattern may overlap with the mark M captured in the image data generated by taking a picture of the learning object with the desired projection pattern projected by a single monocular camera. As a result, the data generation unit 111 may fail to detect the mark M captured in the image. In contrast, if the image data is generated by taking a picture of the learning object without the desired projection pattern projected by the single monocular camera, the mark M captured in the image data can be appropriately detected. Furthermore, the image data IMG_2D#p1 can be image data generated by taking a picture of the learning object with the desired projection pattern projected by the single monocular camera in a predetermined positional relationship.

[0362] The data generation unit 111 uses the CAD data corresponding to the learning object to generate reference image data IMG_2M. The reference image data IMG_2M can be two-dimensional image data representing multiple two-dimensional models of learning objects generated by virtually projecting a three-dimensional model of the learning object, represented by the CAD data corresponding to the learning object, from multiple different directions onto virtual planes orthogonal to each of the multiple different directions. Therefore, the reference image data IMG_2M can be referred to as model data of the learning object. Furthermore, the reference image data IMG_2M can also be the same as the two-dimensional model data IMG_2M (refer to "(2-2) 2D Matching Processing"). The data generation unit 111 can perform matching processing on the image data IMG_2D#p1 using the learning object photographed in the two-dimensional image represented by the reference image data IMG_2M as a template. Alternatively, the data generation unit 111 may not generate the reference image data IMG_2M. The data generation unit 111 can also read the reference image data IMG_2M stored in the storage device 120 and use it as described later.

[0363] The data generation unit 111 can move, enlarge, reduce, and / or rotate the learning object captured in the two-dimensional image represented by the reference image data IMG_2M, so that the feature parts (e.g., at least one of feature points and edges) of the learning object captured in the two-dimensional image represented by the reference image data IMG_2M are close to (typically identical to) the feature parts of the learning object captured in the image represented by the image data IMG_2D#p1. That is, the data generation unit 111 can change the positional relationship between the coordinate system of the reference image data IMG_2M (e.g., the coordinate system of a CAD model) and the 2D shooting coordinate system based on the shooting device 21 that captures the learning object, so that the feature parts of the learning object captured in the two-dimensional image represented by the reference image data IMG_2M are close to (typically identical to) the feature parts of the learning object captured in the image represented by the image data IMG_2D#p1.

[0364] Furthermore, the data generation unit 111 can move, enlarge, reduce, and / or rotate the learning object captured in the two-dimensional image represented by the reference image data IMG_2M, so that the feature parts of the learning object captured in the two-dimensional image represented by the reference image data IMG_2M are close to the feature parts of the learning object captured in the image represented by the image data IMG_2D#p1. The user can move, enlarge, reduce, and / or rotate the learning object captured in the two-dimensional image represented by the reference image data IMG_2M displayed on the display, so that it is close to (typically identical to) the learning object captured in the image represented by the image data IMG_2D#p1 displayed on the display 150, which is the output device, via the input device 140. In this case, after the user moves, zooms in, zooms out, and / or rotates the learning object captured in the two-dimensional image represented by the reference image data IMG_2M to make it close to (typically identical to) the learning object captured in the image represented by the image data IMG_2D#p1, the data generation unit 111 further performs the matching process. In this case, the data generation unit 111 may repeatedly move, zoom in, zoom out, and / or rotate the learning object captured in the two-dimensional image represented by the reference image data IMG_2M in minute increments, and use a differential evolution method to make the feature parts of the learning object captured in the two-dimensional image represented by the reference image data IMG_2M close to (typically identical to) the feature parts of the learning object captured in the image represented by the image data IMG_2D#p1.

[0365] Furthermore, even if the user does not move, enlarge, reduce, and / or rotate the learning object captured in the two-dimensional image represented by the reference image data IMG_2M to make it close to (typically identical to) the learning object captured in the image represented by the image data IMG_2D#p1, the data generation unit 111 can repeatedly perform minor parallel movements, enlargements, reductions, and / or rotations on the learning object captured in the two-dimensional image represented by the reference image data IMG_2M, so that the feature parts of the learning object captured in the two-dimensional image represented by the reference image data IMG_2M are close to (typically identical to) the feature parts of the learning object captured in the image represented by the image data IMG_2D#p1 through differential evolution.

[0366] The result of the 2D matching process is that the data generation unit 111 can determine the positional relationship between the coordinate system of the reference image data IMG_2M and the 2D shooting coordinate system. Then, based on the positional relationship between the coordinate system of the reference image data IMG_2M and the 2D shooting coordinate system, the data generation unit 111 can calculate the position and pose of the learning object in the 2D shooting coordinate system according to the position and pose of the learning object in the coordinate system of the reference image data IMG_2M. The calculated position and pose of the learning object in the 2D shooting coordinate system can be the position and pose of the center of gravity of the learning object in the 2D shooting coordinate system. The data generation unit 111 can generate data representing the calculated position and pose of the learning object in the 2D shooting coordinate system (e.g., the position and pose of the center of gravity of the learning object in the 2D shooting coordinate system) as target position and pose data involved in the 2D matching process.

[0367] Here, the target position and pose data involved in the 2D matching process refers to position and pose data representing at least one of the position and pose of the learning object in the 2D shooting coordinate system: in the 2D matching process where the learning object recognition unit 112 uses the learning image data, at least one of the position and pose of the learning object captured in the image represented by the learning image data is close to (typically consistent with) the target.

[0368] Furthermore, the differential evolution method can make the feature parts of the learning object captured in the 2D image represented by the reference image data IMG_2M closely approximate (typically, identical to) the feature parts of the learning object captured in the image represented by the image data IMG_2D#p1. Therefore, in the 2D matching process, after determining the positional relationship between the coordinate system of the reference image data IMG_2M and the 2D shooting coordinate system through the differential evolution method, it is expected that the position and pose of the learning object in the 2D shooting coordinate system calculated based on the determined positional relationship will be infinitely close to the true position and pose.

[0369] Thus, in the 2D matching process performed by the data generation unit 111, the position and pose of the learning object in the 2D shooting coordinate system, which are infinitely close to the true value, can be calculated. That is, it can be said that the 2D matching process performed by the data generation unit 111 is a 2D matching process that can calculate at least one of the position and pose of the learning object in the 2D shooting coordinate system with high precision. Therefore, the 2D matching process performed by the data generation unit 111 can be called a high-precision matching process. Therefore, the 2D matching process performed by the data generation unit 111 will henceforth be appropriately referred to as "high-precision matching process".

[0370] Furthermore, in the 2D matching process, the differential evolution method is not limited to it; different methods (algorithms) can also be used. Additionally, the data generation unit 111 can calculate only one of the position and pose of the learning object in the 2D shooting coordinate system based on the positional relationship between the coordinate system of the determined reference image data IMG_2M and the 2D shooting coordinate system. That is, the data generation unit 111 can also calculate at least one of the position and pose of the learning object in the 2D shooting coordinate system based on the positional relationship between the coordinate system of the determined reference image data IMG_2M and the 2D shooting coordinate system.

[0371] High-precision matching processing can calculate at least one of the position and pose of the learning object in the 2D imaging coordinate system with high accuracy, thus it is suitable for calculating (generating) target position and pose data. On the other hand, high-precision matching processing requires a long time. Therefore, from the viewpoint of the robot's processing speed, it is impractical for the object recognition unit 115 to perform the same 2D matching processing as high-precision matching processing. Furthermore, the same applies to the learning object recognition unit 112, which performs the same 2D matching processing as the object recognition unit 115. Therefore, the 2D matching processing performed by the object recognition unit 115 and the learning object recognition unit 112 differs from high-precision matching processing in that its processing speed is faster than high-precision matching processing. The 2D matching processing performed by the object recognition unit 115 and the learning object recognition unit 112 also differs from high-precision matching processing in that its accuracy is lower.

[0372] As described above, the image represented by image data IMG_2D#p1, which serves as learning image data, simultaneously captures a learning object disposed on a platform ST and at least one marker M disposed on the platform ST. The data generation unit 111 can calculate at least one of the position and orientation of the platform ST in the 2D shooting coordinate system (e.g., the position and orientation of the center of the platform ST in the 2D shooting coordinate system) based on the at least one marker M detected from the image represented by image data IMG_2D#p1. Consequently, the data generation unit 111 can calculate at least one of the relative position and orientation of the platform ST and the learning object in the 2D shooting coordinate system based on at least one of the position and orientation of the learning object in the 2D shooting coordinate system and at least one of the position and orientation of the platform ST in the 2D shooting coordinate system.

[0373] Furthermore, the data generation unit 111 calculates at least one of the position and orientation of the stage ST based on at least one marker M detected from the image represented by the learning image data (e.g., image data IMG_2D#p1), which may be at least one of the position and orientation of the stage ST in the global coordinate system. In this case, the data generation unit 111 can use a transformation matrix for converting the position and orientation in the global coordinate system to the position and orientation in the 2D shooting coordinate system, to convert at least one of the position and orientation of the stage ST in the global coordinate system to at least one of the position and orientation of the stage ST in the 2D shooting coordinate system.

[0374] Next, if the data generation unit 111 calculates at least one of the position and posture of the learning object captured in the image represented by image data IMG_2D#p2, which is different from image data IMG_2D#p1, the following processing can be performed. Here, the relative position and posture of the platform ST and the learning object captured in the image represented by image data IMG_2D#p1 are the same as the relative position and posture of the platform ST and the learning object captured in the image represented by image data IMG_2D#p2. That is, image data IMG_2D#p1 and image data IMG_2D#p2 are image data generated by changing the positional relationship between the shooting unit 20 and the learning object without changing the relative position and posture of the platform ST and the learning object, so that the single monocular camera included in the shooting device 21 of the shooting unit 20 captures the learning object.

[0375] In this case, the data generation unit 111 can calculate at least one of the position and orientation of the stage ST in the 2D shooting coordinate system based on at least one marker M detected from the image represented by the image data IMG_2D#p2. The data generation unit 111 can also calculate at least one of the position and orientation of the learning object captured in the image represented by the image data IMG_2D#p2 in the 2D shooting coordinate system based on (i) at least one of the position and orientation of the stage ST in the 2D shooting coordinate system calculated based on the image data IMG_2D#p2, and (ii) at least one of the relative position and orientation of the stage ST and the learning object in the 2D shooting coordinate system calculated based on the image data IMG_2D#p1.

[0376] As an example, the data generation unit 111 can calculate a transformation matrix for converting at least one of the position and orientation of the platform ST in the 2D shooting coordinate system calculated based on image data IMG_2D#p2, and at least one of the position and orientation of the platform ST in the 2D shooting coordinate system calculated based on image data IMG_2D#p1, into at least one of the position and orientation of the platform ST in the 2D shooting coordinate system calculated based on image data IMG_2D#p2. The data generation unit 111 can use the calculated transformation matrix to convert at least one of the relative position and orientation of the platform ST and the learning object in the 2D shooting coordinate system calculated based on image data IMG_2D#p1 into at least one of the relative position and orientation of the platform ST and the learning object in the 2D shooting coordinate system captured in the image represented by image data IMG_2D#p2. The data generation unit 111 can calculate at least one of the positions and orientations of the platform ST in the 2D shooting coordinate system in the image represented by the image data IMG_2D#p2, based on at least one of the positions and orientations of the platform ST in the 2D shooting coordinate system calculated from the image data IMG_2D#p2, and at least one of the relative positions and orientations of the platform ST and the learning object in the 2D shooting coordinate system after the transformation.

[0377] The high-precision matching process can be performed using one of the multiple learning image datasets to calculate at least one of the positions and poses of the learning object captured in the image represented by the one learning image dataset in the 2D shooting coordinate system. The multiple learning image datasets are generated by changing the positional relationship between the shooting unit 20 and the learning object without changing the relative positions and poses of the platform ST and the learning object, so that a single monocular camera, which is the shooting device 21 included in the shooting unit 20, captures the learning object. At least one of the positions and poses of the learning object captured in the 2D shooting coordinate system in the images represented by the remaining learning image datasets can be calculated based on at least one of the relative positions and poses of the platform ST and the learning object captured in the 2D shooting coordinate system in the image represented by one learning image dataset, and at least one of the positions and poses of the platform ST captured in the 2D shooting coordinate system in the images represented by the remaining learning image datasets. That is, the high-precision matching process may not be performed on the remaining learning image datasets.

[0378] Here, the high-precision matching process requires a relatively long time. If the calculation of at least one of the position and pose of the learning object captured in the image represented by the remaining learning image data in the plurality of learning image data can be avoided by avoiding high-precision matching processing, then the time required to calculate at least one of the position and pose of the learning object captured in the image represented by the remaining learning image data in the 2D shooting coordinate system can be shortened.

[0379] In addition, in order to calculate at least one of the positions and poses of the learning object captured in the 2D shooting coordinate system in each of the multiple images represented by the multiple learning image data, high-precision matching processing using each of the multiple learning image data can also be performed.

[0380] As described above, the image data IMG_2D#p1, which serves as learning image data, can be generated by photographing the learning object using a single monocular camera, which is the photographing device 21. Therefore, the image data IMG_2D#p1 can be referred to as first learning image data and / or first monocular image data. Furthermore, as described above, the image data IMG_2D#p2, which serves as learning image data, can also be generated by photographing the learning object using a single monocular camera, which is the photographing device 21. Therefore, the image data IMG_2D#p2 can be referred to as second learning image data and / or second monocular image data.

[0381] That is, the learning image data can also be a concept that includes both first monocular image data and second monocular image data. Furthermore, the learning image data can also be a concept that includes other monocular image data besides the first and second monocular image data. The learning image data can also be a concept that includes only one of the first and second monocular image data. That is, the learning image data can also be a concept that includes at least one of the first and second monocular image data.

[0382] Since image data IMG_2D#p1 can be referred to as first learning image data and / or first monocular image data, the position of the learning object captured in the image represented by image data IMG_2D#p1 in the 2D shooting coordinate system, calculated by the data generation unit 111 through 2D matching processing using image data IMG_2D#p1 and reference image data IMG_2M, can be referred to as the first learning object position. Similarly, the pose of the learning object captured in the image represented by image data IMG_2D#p1 in the 2D shooting coordinate system, calculated by the data generation unit 111 through 2D matching processing using image data IMG_2D#p1 and reference image data IMG_2M, can be referred to as the first learning object pose.

[0383] Since the image data IMG_2D#p1 can be referred to as the first learning image data and / or the first monocular image data, the position of the stage ST in the 2D shooting coordinate system calculated by the data generation unit 111 based on at least one marker M set on the stage ST in the image represented by the image data IMG_2D#p1 can be referred to as the first stage position. Similarly, the pose of the stage ST in the 2D shooting coordinate system calculated by the data generation unit 111 based on at least one marker M set on the stage ST in the image represented by the image data IMG_2D#p1 can be referred to as the first stage pose.

[0384] Since image data IMG_2D#p1 can be referred to as second learning image data and / or second monocular image data, the position of the stage ST in the 2D shooting coordinate system calculated by the data generation unit 111 based on at least one marker M set on the stage ST in the image represented by image data IMG_2D#p2 can be referred to as the second stage position. Similarly, the pose of the stage ST in the 2D shooting coordinate system calculated by the data generation unit 111 based on at least one marker M set on the stage ST in the image represented by image data IMG_2D#p2 can be referred to as the second stage pose.

[0385] When the shooting unit 20 includes a single monocular camera as shooting device 21 and a stereo camera as shooting device 22, the data generation unit 111 can generate target position and pose data involved in 3D matching processing based on the position and pose data involved in 2D matching processing (for example, position and pose data representing at least one of the position and pose of the learning object captured in the image represented by image data IMG_2D#p1).

[0386] Here, the target position and pose data involved in the 3D matching process refers to position and pose data representing at least one of the position and pose of the learning object in the 3D shooting coordinate system: in the 3D matching process where the learning object recognition unit 112 uses the learning image data, at least one of the position and pose of the learning object captured in the image represented by the learning image data is close to (typically consistent with) the target.

[0387] Specifically, the data generation unit 111 can use a transformation matrix M to convert the position and pose in the 2D shooting coordinate system to the position and pose in the 3D shooting coordinate system. 23 This process transforms at least one of the position and pose of the learning object in the 2D shooting coordinate system, represented by the position and pose data involved in the 2D matching process, into at least one of the position and pose of the learning object in the 3D shooting coordinate system. If at least one of the position and pose of the learning object in the 2D shooting coordinate system is infinitely close to the true value, then the transformation matrix M is used. 23 The transformed learning object's position and pose in the 3D shooting coordinate system are at least one of infinitely close to the true value.

[0388] The data generation unit 111 can generate a representation using the transformation matrix M. 23 The transformed learning object's position and pose in the 3D shooting coordinate system, at least one of these, are used as the target position and pose data involved in the 3D matching process. Furthermore, the transformation matrix M... 23 It can be calculated using mathematical methods based on external parameters representing the positional relationship between the shooting device 21 and the shooting device 22.

[0389] Furthermore, if the imaging unit 20 does not include a single monocular camera as the imaging device 21, but includes a stereo camera as the imaging device 22, the data generation unit 111 can, for example, use image data (i.e., two-dimensional image data) to calculate at least one of the position and posture of the learning object. This image data (i.e., two-dimensional image data) is generated by photographing the learning object, which is not projected with patterned light from the projection device 23 included in the imaging unit 20, by having one of the two monocular cameras of the stereo camera as the imaging device 22 photograph it in a predetermined positional relationship. In this case, at least one of the position and posture of the learning object calculated by the data generation unit 111 can be at least one of the position and posture of the learning object in a 3D imaging coordinate system based on the imaging device 22. Furthermore, at least one of the position and posture of the learning object can be calculated using the same method as the high-precision matching process described above. The data generation unit 111 can generate data representing at least one of the position and posture of the learning object in the 3D imaging coordinate system as target position and posture data involved in the 3D matching process.

[0390] In this case, when generating the target 3D position data described later, it is not necessary to convert the position and pose of the learning object into the position and pose of the learning object in the 3D shooting coordinate system. That is, it is not necessary to change the viewpoint from the viewpoint of the single monocular camera of the shooting device 21 to the viewpoint of one of the two monocular cameras of the stereo camera of the shooting device 22.

[0391] (4-3-2) Generation and processing of three-dimensional position data

[0392] When the imaging unit 20 includes a single monocular camera as an imaging device 21 and a stereo camera as an imaging device 22, the data generation unit 111 can generate three-dimensional position data using image data generated by photographing the learning object using the single monocular camera as an imaging device 21. Here, an example of image data generated by photographing the learning object using the single monocular camera as an imaging device 21 is given, specifically the image data IMG_2D#p1.

[0393] The data generation unit 111 can perform viewpoint change processing on the image data IMG_2D#p1. This viewpoint change processing changes the viewpoint from that of a single monocular camera (as part of the imaging device 21) to that of one of the two monocular cameras (as part of the stereo camera) of the imaging device 22. The viewpoint change processing can utilize a transformation matrix M. 23 This will be done. Furthermore, since various existing objects can be applied in viewpoint change processing, detailed descriptions of them are omitted.

[0394] The data generation unit 111 can detect at least one marker M (for example, multiple markers M set on the stage ST) in the image represented by the image data IMG_2D#p1, which has undergone viewpoint change processing. The data generation unit 111 can calculate at least one of the three-dimensional positions and orientations of the detected at least one marker M (for example, each of the multiple markers M set on the stage ST) in the 3D shooting coordinate system. Furthermore, since the image represented by the image data IMG_2D#p1, which has undergone viewpoint change processing, is an image from the viewpoint of one of the two monocular cameras of the stereo camera of the shooting device 22, the calculated position and orientation of at least one marker M becomes at least one of the position and orientations in the 3D shooting coordinate system based on the shooting device 22.

[0395] The data generation unit 111 can calculate the three-dimensional positions of multiple points on the stage ST based on at least one of the positions and poses of at least one marker M (for example, multiple markers M set on the stage ST) in the 3D shooting coordinate system, and the constraint that the stage ST is a plane. Alternatively, the data generation unit 111 may not use the constraint that the stage ST is a plane.

[0396] Furthermore, the data generation unit 111 may, for example, use a transformation matrix M. 23 The system converts at least one of the position and orientation of the platform ST in the 2D shooting coordinate system from the image represented by image data IMG_2D#p1 (i.e., the image data before the viewpoint change processing is performed) into at least one of the position and orientation of the platform ST in the 3D shooting coordinate system. Furthermore, when the data generation unit 111 calculates the three-dimensional positions of multiple points of the platform ST, it can also use the converted position and orientation of the platform ST in the 3D shooting coordinate system. In this way, more accurate three-dimensional positions of the multiple points of the platform ST can be calculated.

[0397] The data generation unit 111 can also use the transformation matrix M 23 The image data IMG_2D#p1, calculated through high-precision matching, is used to convert at least one of the position and pose of the learning object in the 2D shooting coordinate system (refer to "(4-3-1) Generation Processing of Position and Pose Data") into at least one of the position and pose of the learning object in the 3D shooting coordinate system. Here, the position and pose of the learning object in the 2D shooting coordinate system calculated through high-precision matching is infinitely close to the true value (refer to "(4-3-1) Generation Processing of Position and Pose Data"). Therefore, it can be said that using the transformation matrix M... 23The transformed learning object's position and pose in the 3D shooting coordinate system are at least one of infinitely close to the true value.

[0398] The data generation unit 111 can modify the positional relationship between the learning object captured in the image represented by the viewpoint-change processed image data IMG_2D#p1 and its two-dimensional model, based on at least one of the transformed position and pose of the learning object in the 3D shooting coordinate system. This changes the two-dimensional model of the learning object captured in the two-dimensional image represented by the reference image data IMG_2M and the learning object captured in the image represented by the viewpoint-change processed image data IMG_2D#p1, making them close (typically identical). In other words, the data generation unit 111 can also align the two-dimensional model of the learning object.

[0399] The data generation unit 111 can also calculate the three-dimensional positions of multiple points on the learning object based on the alignment results of the two-dimensional model of the learning object. The data generation unit 111 can also generate data representing the calculated three-dimensional positions of multiple points on the platform ST and the calculated three-dimensional positions of multiple points on the learning object, as target three-dimensional position data. Furthermore, the target three-dimensional position data can be, for example, depth image data or point group data.

[0400] The data generation unit 111 can align the two-dimensional model of the learning object based on at least one of the transformed position and pose of the learning object in the 3D shooting coordinate system (i.e., based on at least one of the position and pose of the learning object in the 3D shooting coordinate system that is infinitely close to the true value). Therefore, the data generation unit 111 can align the two-dimensional model of the learning object with the learning object captured in the image represented by the image data IMG_2D#p1 after viewpoint change processing with high precision. Therefore, it is expected that the three-dimensional positions of multiple points of the learning object calculated based on the alignment result of the two-dimensional model of the learning object will be infinitely close to the true value.

[0401] The data generation unit 111 may also replace the method for generating the three-dimensional position data or, based thereon, generate the target three-dimensional position data by the method described below.

[0402] As a premise, let's assume that, under a predetermined positional relationship, a single monocular camera (which serves as the imaging device 21) included in the imaging unit 20 photographs a learning object without a projected pattern, thereby generating image data IMG_2D#p1. Alternatively, let's assume that, under the predetermined positional relationship, a stereo camera (which serves as the imaging device 22) included in the imaging unit 20 photographs a learning object without a projected pattern, thereby generating image data IMG_3D#p1. Furthermore, let's assume that, under the predetermined positional relationship, a stereo camera (which serves as the imaging device 22) included in the imaging unit 20 photographs a learning object with a projected pattern, thereby generating image data IMG_3D#p2.

[0403] The data generation unit 111 can perform high-precision matching processing using image data IMG_2D#p1 (refer to "(4-3-1) Generation and processing of position and pose data") to calculate at least one of the position and pose of the learning object represented by image data IMG_2D#p1 in the 2D shooting coordinate system.

[0404] The data generation unit 111 can detect at least one marker M captured in the image represented by one of the two image data contained in the image data IMG_3D#p1 (i.e., image data captured by the two monocular cameras of the stereo camera). Based on the detected at least one marker M, the data generation unit 111 can calculate at least one of the position and orientation of the stage ST in the 3D shooting coordinate system. Furthermore, one of the image data may, for example, correspond to image data IMG_2D#p1 that has undergone viewpoint change processing.

[0405] The data generation unit 111 can use the transformation matrix M 23 The calculated position and pose of the learning object in the 2D shooting coordinate system are transformed into at least one position and pose in the 3D shooting coordinate system. The transformed position and pose of the learning object in the 3D shooting coordinate system corresponds to at least one of the positions and poses of the learning object captured in the image represented by one of the image data in the 3D shooting coordinate system. The position and pose of the learning object in the 2D shooting coordinate system calculated through high-precision matching processing is infinitely close to the true value (refer to "(4-3-1) Generation Processing of Position and Pose Data"). Therefore, it can be said that using the transformation matrix M... 23 The transformed learning object's position and pose in the 3D shooting coordinate system are at least one of infinitely close to the true value.

[0406] The data generation unit 111 can establish correspondences between each part (e.g., each pixel) of the images represented by the two image data contained in the image data IMG_3D#p2 using well-known methods such as SGBM and SAD. As a result, the data generation unit 111 can calculate the three-dimensional positions of multiple points of the captured object (e.g., the learning object, or the learning object and the platform ST) in the images represented by the two image data contained in the image data IMG_3D#p2.

[0407] The data generation unit 111 can calculate the three-dimensional positions of the learning object and the multiple points of the platform ST based on at least one of the position and posture of the platform ST in the 3D shooting coordinate system, at least one of the position and posture of the learning object in the 3D shooting coordinate system, and the three-dimensional positions of multiple points of the object.

[0408] Specifically, the data generation unit 111 can modify the three-dimensional positions of multiple points of the object, such that at least one of the position and pose of the platform ST in the 3D shooting coordinate system calculated based on the three-dimensional positions of the multiple points of the object is close to (typically identical to) at least one of the position and pose of the platform ST in the 3D shooting coordinate system calculated based on one of the two image data contained in the image data IMG_3D#p1, and makes at least one of the position and pose of the learning object in the 3D shooting coordinate system calculated based on the three-dimensional positions of the multiple points of the object, and the position and pose calculated using the transformation matrix M 23 The transformed learning object has at least one of its position and pose in the 3D shooting coordinate system that is close (typically consistent), thereby, for example, calculating the three-dimensional positions of the learning object and multiple points of the stage ST.

[0409] The data generation unit 111 can also generate data representing the three-dimensional positions of multiple points of the learning object and the platform ST, as target three-dimensional position data. The data generation unit 111 can modify the three-dimensional positions of the multiple points of the object, such that at least one of the position and pose of the learning object in the 3D shooting coordinate system calculated based on the three-dimensional positions of the multiple points of the object, and the position calculated using the transformation matrix M... 23 The transformed learning object's position and pose in the 3D shooting coordinate system are close to (typically consistent with) the true value, and at least one of these (i.e., at least one of the learning object's position and pose in the 3D shooting coordinate system that is infinitely close to the true value) is used to calculate the individual 3D positions of multiple points of the learning object. Therefore, it can be expected that the calculated individual 3D positions of multiple points of the learning object are infinitely close to the true value.

[0410] Furthermore, if the shooting unit 20 does not include a stereo camera as a shooting device 22, but includes a single monocular camera as a shoo...

Claims

1. A computing device, comprising: A control unit, which outputs a control signal for controlling the photographing unit and a robot provided with the photographing unit; as well as a learning unit that generates a model for determining parameters of arithmetic processing in the control unit by learning using a result of imaging the learning target object by the imaging unit, The control unit outputs a first control signal, wherein the first control signal is used to drive the robot so that the imaging unit is in a predetermined positional relationship with respect to the learning object, and the imaging unit images the learning object in the predetermined positional relationship. The learning unit generates the model by learning using learning image data, wherein the learning image data is generated by causing the imaging unit to image the learning target object at the predetermined positional relationship under control based on the first control signal. The control unit performs the arithmetic processing using the parameters determined using the model and the processing target image data generated by the imaging unit imaging the processing target object having substantially the same shape as the learning target object, to calculate at least one of the position and posture of the processing target object.

2. The computing device according to claim 1, wherein: The predetermined positional relationship is a plurality of mutually different positional relationships of the imaging unit with respect to the learning target object. The first control signal is a signal for driving the robot to change to each of the plurality of positional relationships, and causing the imaging unit to image the learning object each time the robot changes to each of the plurality of positional relationships. The learning department The model is generated by learning using a plurality of learning image data generated by causing the imaging unit to image the learning target object in each of the plurality of positional relationships under control based on the first control signal.

3. The computing device according to claim 1 or 2, wherein: 251429 1PWCN;154051-CN-1543-PCT The control unit The parameters used in the operation processing are determined, and the operation processing is used to calculate at least one of the position and posture of the processing object using the processing object image data generated by photographing the processing object by the photographing unit and the model generated by learning using the learning image data.

4. The computing device according to any one of claims 1 to 3, wherein: The positional relationship includes a posture of the imaging unit relative to the learning target object.

5. The computing device according to claim 4, wherein: The posture is a posture around at least one of the first axis, the second axis and the third axis in a coordinate system of the learning object defined by a first axis along the optical axis of the optical system of the shooting unit, a second axis orthogonal to the first axis, and a third axis orthogonal to the first axis and the second axis.

6. The computing device according to any one of claims 1 to 5, wherein: The positional relationship includes the distance from the learning target object to the imaging unit.

7. The computing device according to claim 6, wherein: The distance is the distance from the learning object to the imaging unit on the first axis in the coordinate system of the learning object defined by a first axis along the optical axis of the optical system of the imaging unit, a second axis orthogonal to the first axis, and a third axis orthogonal to the first axis and the second axis.

8. The computing device according to any one of claims 1 to 7, wherein: The control unit receiving an input of a range in which the positional relationship is to be changed, and Based on the input range, the predetermined positional relationship is determined within the range.

9. The computing device according to any one of claims 1 to 8, wherein: The robot is also provided with a processing device for processing the processing object. The control unit Based on at least one of the calculated position and posture of the processing target object, a second control signal is outputted, the second control signal being used to drive the robot so as to process the processing target object using the processing device.

10. The computing device according to claim 9, wherein: 251429 1PWCN;154051-CN-1543-PCT The control unit controls the processing device. The second control signal is a signal used to control driving of the robot and processing of the processing target object by the processing device so that the processing target object is processed by the processing device.

11. The computing device according to any one of claims 2 to 10, wherein: The first learning image data among the plurality of learning image data is generated by causing the imaging unit to image the learning target object at a first relationship among the plurality of positional relationships. The second learning image data among the plurality of learning image data is generated by causing the imaging unit to image the learning object at a second relationship among the plurality of positional relationships that is different from the first relationship. The learning department calculating at least one of a first learning object position and a first learning object posture of the learning target object based on the first learning image data; generating first correct answer data as a parameter used in the calculation process under the first relationship using at least one of the first learning object position and the first learning object posture; calculating at least one of a second learning object position and a second learning object posture of the learning target object based on the second learning image data; generating second correct answer data as a parameter used in the calculation process under the second relationship using at least one of the second learned object position and the second learned object posture; The model is generated through learning using the first learning image data, the second learning image data, the first correct answer data, and the second correct answer data.

12. The computing device according to claim 11, wherein: The learning object is arranged on a stage provided with at least one mark, The first learning image data is generated by the imaging unit of the first relationship imaging at least one of the learning target object and the at least one mark. The second learning image data is generated by the imaging unit of the second relationship imaging at least one of the learning target object and the at least one mark.

13. The computing device according to claim 12, wherein: The at least one mark is a plurality of marks, The first learning image data is generated by the imaging unit of the first relationship imaging the learning target object and at least one of the plurality of markers together. The second learning image data is generated by the imaging unit of the second relationship imaging the learning target object together with at least one other of the plurality of markers.

14. The computing device according to claim 12, wherein: The at least one mark is a mark, The first learning image data and the second learning image data are generated by causing the imaging unit to image the learning target object and the one mark together in each of the first relationship and the second relationship.

15. The computing device according to any one of claims 12 to 14, wherein: The marker is an augmented reality marker.

16. The computing device according to any one of claims 11 to 15, wherein: The photographing unit includes a single-lens reflex camera, The processing target image data is generated by photographing the processing target object with the single-lens camera. The arithmetic processing includes two-dimensional matching processing using the processing target image data generated by the single-lens camera and two-dimensional model data of the processing target object. The parameters of the operation processing are the two-dimensional matching parameters of the two-dimensional matching processing, The first learning image data is first monocular image data generated by the monocular camera of the first relationship, the second learning image data is second monocular image data generated by the monocular camera of the second relationship, The control unit calculating at least one of the position and posture of the learning object by a two-dimensional matching process using the first monocular image data and the two-dimensional model data of the learning object under a first learning two-dimensional matching parameter, By using the second monocular image data and the two-dimensional model data of the learning object under the second learning two-dimensional matching parameters, at least one of the position and posture of the learning object 251429 1PWCN; 154051-CN-1543-PCT is calculated, The learning department The first correct answer data is calculated based on a first degree of consistency between at least one of the position and posture of the learning object calculated in the two-dimensional matching process under the first learning two-dimensional matching parameters and at least one of the first learning object position and the first learning object posture, The second correct answer data is calculated based on a second degree of consistency between at least one of the position and posture of the learning object calculated in the two-dimensional matching process under the second learning two-dimensional matching parameters and at least one of the second learning object position and the second learning object posture, The model is generated through learning using the first monocular image data, the second monocular image data, the first correct answer data, and the second correct answer data.

17. The computing device according to claim 16, wherein: The learning department The first two-dimensional matching parameter for learning is calculated as the first correct answer data as follows: at least one of the position and posture of the learning object calculated in the two-dimensional matching processing under the first two-dimensional matching parameter for learning is close to at least one of the first learning object position and the first learning object posture, The following second learning two-dimensional matching parameters are calculated as the second correct answer data: at least one of the position and posture of the learning object object calculated in the two-dimensional matching processing under the second learning two-dimensional matching parameters is close to at least one of the second learning object position and the second learning object posture.

18. The computing device according to claim 16 or 17, wherein: The learning department calculating the first correct answer data based on the first degree of consistency and the time required for the two-dimensional matching process under the first two-dimensional matching parameter for learning, The second correct answer data is calculated based on the second degree of consistency and the time required for the two-dimensional matching process under the second learning two-dimensional matching parameters.

19. The computing device according to claim 18, wherein: The learning department The following first learning two-dimensional matching parameters are calculated as the first correct answer data: at least one of the position and posture of the learning object object calculated in the two-dimensional matching processing under the first learning two-dimensional matching parameters is close to at least one of the first learning object position and the first learning object posture, and the time required for the two-dimensional matching processing under the first learning two-dimensional matching parameters is shortened, The following second learning two-dimensional matching parameters are calculated as the second correct answer data: at least one of the position and posture of the learning object object calculated in the two-dimensional matching processing under the second learning two-dimensional matching parameters is close to at least one of the second learning object position and the second learning object posture, and the time required for the two-dimensional matching processing under the second learning two-dimensional matching parameters is shortened.

20. The computing device according to any one of claims 12 to 19, wherein: The learning department using the first learning image data and the model data of the learning object, calculating at least one of the first learning object position and the first learning object posture of the learning object; calculating at least one of a first stage position and a first stage posture of the stage based on image data of at least one of the markers included in the first learning image data, using at least one of the first learning object position and the first learning object posture, and at least one of the first stage position and the first stage posture, to calculate at least one of the relative position and the relative posture of the learning object and the stage, calculating at least one of a second stage position and a second stage posture of the stage based on image data of at least one of the markers included in the second learning image data, At least one of the second learning object position and the second learning object posture is calculated using at least one of the relative position and the relative posture and at least one of the second stage position and the second stage posture.

21. The computing device according to any one of claims 11 to 19, wherein: The imaging unit includes a single-lens camera and a stereo camera including two single-lens cameras different from the single-lens camera. The first learning image data is first monocular image data generated by the monocular camera of the first relationship, the second learning image data is second monocular image data generated by the monocular camera of the second relationship, The processing target image data is generated by photographing the processing target object by the stereo camera. The arithmetic processing includes a position calculation process using the processing target image data generated by the stereo camera, The parameters of the calculation processing are the position calculation processing parameters of the position calculation processing, The third learning image data among the plurality of learning image data is first stereoscopic image data generated by photographing the learning target object by the stereoscopic camera of the first relationship, The fourth learning image data among the plurality of learning image data is second stereoscopic image data generated by photographing the learning target object with the stereoscopic camera of the second relationship, The control unit The position of the learning target object is calculated by performing a position calculation process using a first learning position calculation parameter using the first stereoscopic image data. The position of the learning target object is calculated by performing a position calculation process using a second learning position calculation parameter using the second stereoscopic image data. The learning department calculating the first correct answer data based on a third degree of consistency between the position of the learning object calculated in the position calculation process under the first learning position calculation parameter and the first learning object position; calculating the second correct answer data based on a fourth degree of consistency between the position of the learning object calculated in the position calculation process under the second learning position calculation parameters and the second learning object position; The model is learned using the first stereo image data, the second stereo image data, the first correct answer data, and the second correct answer data.

22. The computing device according to claim 21, wherein: The learning department The first learning position calculation parameter is calculated as the first correct answer data: the position of the learning object calculated in the position calculation process under the first learning position calculation parameter is close to the first learning object position, The second learning position calculation parameter is calculated as the second correct answer data, wherein the position of the learning object calculated in the position calculation process under the second learning position calculation parameter is close to the second learning object position.

23. The computing device according to claim 21 or 22, wherein: The learning department calculating the first correct answer data based on the third degree of consistency and the time required for the position calculation process under the first learning position calculation parameter, The second correct answer data is calculated based on the fourth degree of consistency and the time required for the position calculation process under the second learning position calculation parameters.

24. The computing device according to claim 23, wherein: The learning department The first learning position calculation parameter is calculated as the first correct answer data as follows: the position of the learning object object calculated in the position calculation process under the first learning position calculation parameter is close to the first learning object position, and the time required for the position calculation process under the first learning position calculation parameter is shortened, The following second learning position calculation parameters are calculated as the second correct answer data: the position of the learning object object calculated in the position calculation processing under the second learning position calculation parameters is close to the second learning object position, and the time required for the position calculation processing under the second learning position calculation parameters is shortened.

25. The computing device according to any one of claims 12 to 19 and 21 to 24, wherein: The imaging unit includes a single-lens camera and a stereo camera including two single-lens cameras different from the single-lens camera. The first learning image data is first monocular image data generated by the monocular camera of the first relationship, the second learning image data is second monocular image data generated by the monocular camera of the second relationship, The learning department converting the first learning image data into first converted image data representing an image of the learning target object photographed by the stereo camera, using the first converted image data and the model data of the learning object, calculating the first learning object position of the learning object, converting the second learning image data into second converted image data representing an image of the learning target object photographed by the stereo camera, The second learning object position of the learning object is calculated using the second converted image data and the model data of the learning object.

26. The computing device according to any one of claims 11 to 19 and 21 to 24, wherein: The imaging unit includes a single-lens camera and a stereo camera including two single-lens cameras different from the single-lens camera. The first learning image data is first monocular image data generated by the monocular camera of the first relationship, the second learning image data is second monocular image data generated by the monocular camera of the second relationship, The processing target image data is generated by photographing the processing target object by the stereo camera. The calculation processing includes a three-dimensional matching processing using the position data of the processing target object generated from the processing target image data generated by the stereo camera and the three-dimensional model data of the processing target object, The parameters of the operation processing are the three-dimensional matching parameters of the three-dimensional matching processing, The third learning image data among the plurality of learning image data is first stereoscopic image data generated by photographing the learning target object by the stereoscopic camera of the first relationship, The fourth learning image data among the plurality of learning image data is second stereoscopic image data generated by photographing the learning target object with the stereoscopic camera of the second relationship, The control unit calculating at least one of the position and posture of the learning object by performing a three-dimensional matching process using the position data of the learning object generated from the first stereoscopic image data and the three-dimensional model data of the learning object under the first learning three-dimensional matching parameters, calculating at least one of the position and posture of the learning object by performing a three-dimensional matching process under a second learning three-dimensional matching parameter using the position data of the learning object generated from the second stereoscopic image data and the three-dimensional model data of the learning object, The learning department The first correct answer data is calculated based on a fifth degree of consistency between at least one of the position and posture of the learning object calculated by the three-dimensional matching processing under the first learning three-dimensional matching parameters and at least one of the first learning object position and the first learning object posture, The second correct answer data is calculated based on a sixth degree of consistency between at least one of the position and posture of the learning object calculated by the three-dimensional matching processing under the second learning three-dimensional matching parameters and at least one of the second learning object position and the second learning object posture, The model is learned using the first stereo image data, the second stereo image data, the first correct answer data, and the second correct answer data.

27. The computing device according to claim 26, wherein: The learning department The first three-dimensional matching parameters for learning are calculated as the first correct answer data: at least one of the position and posture of the learning object calculated by the three-dimensional matching processing under the first three-dimensional matching parameters for learning is close to at least one of the first learning object position and the first learning object posture, The following second learning three-dimensional matching parameters are calculated as the second correct answer data: at least one of the position and posture of the learning object calculated by the three-dimensional matching processing under the second learning three-dimensional matching parameters is close to at least one of the second learning object position and the second learning object posture.

28. The computing device according to claim 26 or 27, wherein: The learning department calculating the first correct answer data based on the fifth degree of consistency and the time required for the three-dimensional matching process under the first learning three-dimensional matching parameters, The second correct answer data is calculated based on the sixth degree of consistency and the time required for the three-dimensional matching process under the second learning three-dimensional matching parameters.

29. The computing device according to claim 28, wherein: The learning department The following first learning three-dimensional matching parameters are calculated as the first correct answer data: at least one of the position and posture of the learning object calculated in the three-dimensional matching processing under the first learning three-dimensional matching parameters is close to at least one of the first learning object position and the first learning object posture, and the time required for the three-dimensional matching processing under the first learning three-dimensional matching parameters is shortened, The following second learning three-dimensional matching parameters are calculated as the second correct answer data: at least one of the position and posture of the learning object object calculated in the three-dimensional matching processing under the second learning three-dimensional matching parameters is close to at least one of the second learning object position and the second learning object posture of the learning object object, and the time required for the three-dimensional matching processing under the second learning three-dimensional matching parameters is shortened.

30. The computing device according to any one of claims 12 to 19, 21 to 24, and 26 to 28, wherein: The imaging unit includes a single-lens camera and a stereo camera including two single-lens cameras different from the single-lens camera. The first learning image data is first monocular image data generated by the monocular camera of the first relationship, the second learning image data is second monocular image data generated by the monocular camera of the second relationship, The learning department using the first learning image data and the model data of the learning object, calculating at least one of a first position and a first posture of the learning object in the coordinate system of the single-lens camera, calculating at least one of a first stage position and a first stage posture of the stage in a coordinate system of the single-lens camera based on image data of at least one of the markers included in the first learning image data, using at least one of the position and posture of the learning object in the coordinate system of the SLR camera and at least one of the first stage position and the first stage posture of the stage in the coordinate system of the SLR camera, to calculate at least one of the relative position and the relative posture of the learning object and the stage, calculating at least one of a second stage position and a second stage posture of the stage based on image data of at least one of the markers included in the second learning image data, using at least one of the relative position and the relative posture and at least one of the second stage position and the second stage posture, to calculate at least one of the second position and the second posture of the learning object in the coordinate system of the SLR camera, by converting at least one of the first position and the first posture of the learning object in the coordinate system of the monocular camera into the coordinate system of the stereo camera, calculating at least one of the first learning object position and the first learning object posture of the learning object, At least one of the second learning object position and the second learning object posture of the learning object is calculated by converting at least one of the second position and the second posture of the learning object in the coordinate system of the monocular camera into the coordinate system of the stereo camera.

31. A computing device according to any one of claims 22 to 30, wherein: A light projection device for projecting pattern light is provided in the robot, The processing target image data is generated by photographing the processing target object on which the pattern light is projected from the light projection device by the stereo camera. The first stereoscopic image data is data generated by photographing the learning object on which the pattern light is projected from the light projection device by the stereoscopic camera of the first relationship, The second stereoscopic image data is data generated by capturing an image of the learning target object on which the pattern light is projected from the light projection device by the stereoscopic camera of the second relationship.

32. The computing device according to any one of claims 1 to 31, wherein: The imaging unit includes a single-lens camera and a stereo camera including two single-lens cameras different from the single-lens camera.

33. A computing system comprising: A computing device according to any one of claims 1 to 32; as well as The shooting unit.

34. A robot system comprising: A computing device according to any one of claims 1 to 32; The shooting unit; as well as The robot.

35. The robot system according to claim 34, Also included is a processing device for processing the processing target object.

36. A computing method, comprising: Outputting a first control signal, the first control signal being used to drive a robot provided with a photographing unit so that the photographing unit is in a predetermined positional relationship with respect to a learning object, and the photographing unit photographs the learning object in the predetermined positional relationship; generating a model for determining parameters of a computational process by learning using learning image data, wherein the learning image data is generated by causing the imaging unit to image the learning target object at the predetermined positional relationship under control based on the first control signal; as well as The arithmetic processing is performed using the parameters determined using the model and the processing target image data generated by the imaging unit imaging a processing target object having substantially the same shape as the learning target object to calculate at least one of the position and posture of the processing target object.

37. A computer program causing a computer to execute the operation method according to claim 36.

38. A computing device comprising a recording medium recording the computer program according to claim 37, and capable of executing the computer program.

Citation Information

Patent Citations

  • Method and apparatus for calculating parameter

    JP2010188439A

  • Coordinate system calibration method and robot system

    JP2012091280A