Control device, control method, computer program, machine learning device, and machine learning method
The control device adjusts the position of robots and workpieces using a mobile machine to avoid interference, ensuring task completion and improving efficiency by resolving initial positional conflicts.
Patent Information
- Application Number
- PCT/JP2024/003340
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-01
- Publication Date
- 2025-08-07
AI Technical Summary
Robots fixed in a predetermined position may be unable to perform tasks on workpieces due to interference with environmental objects, leading to decreased efficiency.
A control device that includes a task determination unit to assess whether a robot can perform a task and a movement control unit to move the robot or the workpiece relative to each other, using a mobile machine to adjust their positions to avoid interference and ensure task completion.
Enhances task efficiency by effectively resolving positional conflicts between robots and environmental objects, allowing tasks to be performed even when initial positions are unsuitable.
Smart Images

Figure JP2024003340_07082025_PF_FP_ABST
Abstract
Description
Control device, control method, computer program, machine learning device, and machine learning method
[0001] The present disclosure relates to a control device, a control method, a computer program, a machine learning device, and a machine learning method for a robot or a mobile machine.
[0002] A system is known that determines whether a robot can perform a task on a workpiece (specifically, whether interference will occur between the robot and an environmental object) when the robot is to perform the task (for example, Patent Document 1).
[0003] JP 2019-28775 A
[0004] Conventionally, a robot fixed in a predetermined position may be unable to perform a task on a workpiece due to interference with an environmental object, placement at a singular point, etc. In such cases, the efficiency of the task may decrease.
[0005] In one aspect of the present disclosure, a control device for controlling a robot that performs a predetermined task on a workpiece and a mobile machine that moves the workpiece and the robot relatively includes a task determination unit that determines whether the robot, which is placed at a reference position with respect to the workpiece by the mobile machine, is capable of performing the task, and a movement control unit that operates the mobile machine to move the workpiece or the robot from the reference position when the task determination unit determines that the task cannot be performed. The task determination unit determines whether the robot, after being moved by the movement control unit, is capable of performing the task.
[0006] In another aspect of the present disclosure, a method for controlling a robot that performs a predetermined task on a workpiece and a mobile machine that moves the workpiece and the robot relative to each other includes a processor determining whether the robot, which has been placed at a reference position relative to the workpiece by the mobile machine, is capable of performing the task, and if it is determined that the task cannot be performed, operating the mobile machine to move the workpiece or the robot from the reference position, and determining whether the robot after movement is capable of performing the task.
[0007] In yet another aspect of the present disclosure, a machine learning device that learns the amount of movement of a robot that performs a predetermined task on a workpiece using a mobile machine that moves the workpiece relative to the robot includes a state observation unit that observes position data indicating the position of the robot relative to the workpiece and judgment data indicating whether the robot can perform the task on the workpiece at that position as state variables that represent the current state of the environment in which the robot performs the task, and a learning unit that uses the state variables to learn the amount of movement in relation to whether the task can be performed.
[0008] In yet another aspect of the present disclosure, a machine learning method for learning the amount of movement of a workpiece or a robot by a mobile machine that moves a robot performing a predetermined task on a workpiece relative to the workpiece observes position data indicating the position of the robot relative to the workpiece and judgment data indicating whether the robot can perform a task on the workpiece at that position as state variables representing the current state of the environment in which the robot performs the task, and uses the state variables to learn the amount of movement in association with whether the task can be performed.
[0009] FIG. 1 is a block diagram of a robot system according to an embodiment. FIG. 2 is a perspective view of a robot, a mobile machine, and a visual sensor according to an embodiment. FIG. 3 is a flowchart showing an example of an operation flow of the robot system of FIG. 1. FIG. 4 is a schematic diagram showing an example of the data structure of the position database created in step S2 in FIG. 3. FIG. 5 is a flowchart showing an example of step S5 in FIG. 3. FIG. 6 is an example of a change in the status of the position database shown in FIG. 4. FIG. 7 is a diagram showing the robot after movement in step S6 in FIG. 3. FIG. 8 is a block diagram of a robot system according to yet another embodiment. FIG. 9 is a flowchart showing an example of an operation flow of the robot system of FIG. 8. FIG. 10 is a block diagram of a robot system according to yet another embodiment. FIG. 11 is a flowchart showing an example of a learning flow executed in the robot system of FIG. 8.
[0010] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. In various embodiments described below, like elements will be designated by like reference numerals, and duplicated descriptions will be omitted. First, a robot system 10 according to one embodiment will be described with reference to FIGS. 1 and 2. The robot system 10 includes a robot 12, a mobile machine 14, a visual sensor 16, and a control device 50.
[0011] The robot 12 performs a predetermined task (such as workpiece handling, painting, or welding) on a workpiece W. As shown in FIG. 2 , in this embodiment, the robot 12 is a vertically articulated robot that includes a robot base 20, a rotating body 22, a lower arm 24, an upper arm 26, a wrist 28, and an end effector 30. The rotating body 22 is attached to the robot base 20 so as to be rotatable about a vertical axis. The lower arm 24 is attached to the rotating body 22 so as to have a base end that is rotatable about a horizontal axis. The upper arm 26 has a base end that is rotatably attached to a tip end of the lower arm 24. The wrist 28 is attached to a tip end of the upper arm 26 so as to have a base end that is rotatable about two axes that are perpendicular to each other.
[0012] The end effector 30 is, for example, a robot hand, a paint applicator, or a welding gun, and is detachably attached to the tip of the wrist 28. The end effector 30 performs work (work handling, painting, welding, etc.) on the workpiece W. The following describes a case in which the end effector 30 is a robot hand, and the robot 12 performs workpiece handling by gripping and removing workpieces W (not shown) stacked in bulk in a container 100 shown in FIG. 2 with the end effector 30.
[0013] The robot 12 is provided with a plurality of servo motors 32 (FIG. 1). These servo motors 32 rotate each of the moving components of the robot 12 (i.e., the rotating body 22, the lower arm 24, the upper arm 26, the wrist 28, and the end effector 30) around a drive axis. In this way, the robot 12 moves the end effector 30.
[0014] The moving machine 14 moves the workpiece W in the container 100 and the robot 12 relative to one another. In this embodiment, the moving machine 14 is a traveling device that moves the robot 12 along a linear traveling axis A. More specifically, the moving machine 14 has a rail portion 34, a slider 36, a timing belt 38, and a servo motor 40 ( FIG. 1 ). The rail portion 34 extends linearly along the traveling axis A. The slider 36 engages with the rail portion 34 so as to be slidable along the traveling axis A. By engaging with the rail portion 34, the slider 36 is guided to move along the traveling axis A. The robot base 20 of the robot 12 is fixed on the slider 36.
[0015] The timing belt 38 is rotatably mounted on the rail portion 34 and engages with the lower portion of the slider 36. The servo motor 40 generates power to rotate the timing belt 38. When the servo motor 40 rotates the timing belt 38, the slider 36 engaged with the timing belt 38 moves along the traveling axis A, and as a result, the robot 12 mounted on the slider 36 moves along the traveling axis A.
[0016] The visual sensor 16 is disposed above the container 100 so that the container 100 can be accommodated within its field of view. Specifically, the visual sensor 16 is a three-dimensional visual sensor or the like, and includes an image sensor (CCD, CMOS, etc.), an optical lens (collimator lens, focus lens, etc.) that guides an image of a subject to the image sensor, an image processor, etc.
[0017] The visual sensor 16 captures an image of the workpiece W in the container 100 and supplies the image data ID of the captured image of the workpiece W to the control device 50. The image data ID is, for example, three-dimensional point cloud data that represents the visual features (surfaces, edges, etc.) of the workpiece W as a three-dimensional point cloud, or distance image data that represents the visual features of the workpiece W as shades of color according to the distance from the visual sensor 16.
[0018] The control device 50 controls the operations of the robot 12, the mobile machine 14, and the visual sensor 16. As shown in Fig. 1, the control device 50 is a computer having a processor 52, a memory 54, and an I / O interface 56. The processor 52 has a CPU, a GPU, or the like, and is communicatively connected to the memory 54 and the I / O interface 56 via a bus 58.
[0019] The processor 52 communicates with the memory 54 and the I / O interface 56 and performs arithmetic processing to execute operations on the workpiece W. The memory 54 has RAM, ROM, or the like, and temporarily or permanently stores various data. The memory 54 may be a computer-readable recording medium such as a semiconductor memory, a magnetic recording medium, or an optical recording medium.
[0020] The I / O interface 56 has, for example, an Ethernet (registered trademark) port, a USB port, an optical fiber connector, or an HDMI (registered trademark) terminal, and communicates data with external devices via a wired or wireless connection under instructions from the processor 52. The servo motors 32 and 40 and the visual sensor 16 are connected to the I / O interface 56 so as to be able to communicate via a wired or wireless connection.
[0021] 2, a robot coordinate system C1 and a tool coordinate system C2 are set for the robot 12. The robot coordinate system C1 is a control coordinate system for automatically controlling the movable components of the robot 12. In this embodiment, the robot coordinate system C1 is fixed relative to the robot base 20 so that its origin is located at the center of the robot base 20 and its z-axis is parallel to (specifically, coincides with) the rotation axis of the rotating body 22 (i.e., the vertical direction). Therefore, when the mobile machine 14 moves the robot 12 along the traveling axis A, the robot coordinate system C1 also moves together with the robot 12.
[0022] The tool coordinate system C2 is a control coordinate system that defines the position of the end effector 30 in the robot coordinate system C1. In this embodiment, the tool coordinate system C2 is set with respect to the end effector 30 so that its origin (so-called TCP) is located at the working position of the end effector 30 (in this embodiment, the gripping position of the robot hand).
[0023] When moving the end effector 30, the processor 52 of the control device 50 sets a tool coordinate system C2 in the robot coordinate system C1 and generates commands (position commands, speed commands, torque commands, etc.) to each servo motor 32 of the robot 12 so as to place the end effector 30 at a position represented by the set tool coordinate system C2. In this manner, the processor 52 can position the end effector 30 at any position in the robot coordinate system C1. Note that in this disclosure, "position" may refer to both position and orientation.
[0024] On the other hand, a sensor coordinate system C3 is set for the visual sensor 16. The sensor coordinate system C3 is a control coordinate system that defines the position of the visual sensor 16 in the robot coordinate system C1 (i.e., the position and direction of the line of sight). In this embodiment, the sensor coordinate system C3 is set for the visual sensor 16 so that its origin is located at the center of the image sensor of the visual sensor 16 and its z-axis is parallel to (specifically, coincides with) the line of sight (or optical axis) of the visual sensor 16.
[0025] In addition, a mobile machine coordinate system C4 is set for the mobile machine 14. The mobile machine coordinate system C4 is a control coordinate system that defines the direction of the traveling axis A. In this embodiment, the mobile machine coordinate system C4 is fixed relative to the rail portion 34 so that its origin is located at an arbitrary position on the timing belt 38 and its y-axis is located parallel to the traveling axis A. Therefore, the position of the origin of the mobile machine coordinate system C4 remains unchanged even if the mobile machine 14 moves the robot 12.
[0026] The x-axis, y-axis, and z-axis of the robot coordinate system C1 may be parallel to the x-axis, y-axis, and z-axis of the movable machine coordinate system C4, respectively. In the following description, for ease of understanding, it is assumed that the positive x-axis, positive y-axis, and positive z-axis directions of the robot coordinate system C1 coincide with the positive x-axis, positive y-axis, and positive z-axis directions of the movable machine coordinate system C4, respectively.
[0027] The positional relationships between the robot coordinate system C1, tool coordinate system C2, sensor coordinate system C3, and mobile machine coordinate system C4 are known through calibration. Therefore, the coordinates of the robot coordinate system C1, tool coordinate system C2, sensor coordinate system C3, and mobile machine coordinate system C4 can be converted into each other via a known transformation matrix (e.g., a homogeneous transformation matrix). Furthermore, the relative relationship between the amount of movement δ by which the mobile machine 14 moves the robot 12 in the y-axis direction of the mobile machine coordinate system C4 and the amount of displacement of the position of the robot coordinate system C1 set for the robot 12 is also known through calibration. Even when the mobile machine 14 positions the robot 12 at any position on the traveling axis A, the positional relationship between the robot coordinate system C1 and the mobile machine coordinate system C4 is known.
[0028] Next, the operation of the robot system 10 will be described with reference to Figure 3. The processor 52 starts the flow of Figure 3 when it receives a work start command from an operator, a host controller, or the computer program PG1. In step S1, the processor 52 operates the visual sensor 16 to capture an image of the workpiece W in the container 100. The processor 52 obtains image data ID of the workpiece W from the visual sensor 16. In this image data ID, the visual features of the multiple workpieces W and the container 100 are represented as a three-dimensional point cloud (or a distance image), and each point (or each pixel of the distance image) constituting this three-dimensional point cloud is represented as a coordinate in a sensor coordinate system C3.
[0029] In step S2, the processor 52 acquires the position of the workpiece W. Specifically, the processor 52 detects the multiple workpieces W captured in the imaging data ID in the sensor coordinate system C3 by matching a workpiece model WM, which is a model of the workpiece W, with the three-dimensional point cloud (or distance image) captured in the imaging data ID. Then, the processor 52 calculates the coordinates Q3w of the multiple detected workpieces W in the sensor coordinate system C3. i are obtained respectively.
[0030] Then, the processor 52 calculates the coordinate Q3w in the sensor coordinate system C3. i is converted into the robot coordinate system C1, the coordinates Q1w of the multiple workpieces W detected from the image data ID in the robot coordinate system C1 are i Then, the processor 52 acquires the acquired coordinates Q1w i An example of the data structure of the location database 102 is shown in FIG.
[0031] 4 shows an example in which a total of six workpieces W are detected from the image data IDs captured in step S1. In the position database 102 shown in FIG. 4, column 104 indicates the rank i (i=1 to 6) of the workpiece W to be worked on, and column 106 indicates the coordinates Q1w of the detected workpiece W. i (x, y, z, w, p, r) Coordinates Q1w i Of these, the coordinates (x, y, z) indicate the position of the workpiece W in the robot coordinate system C1, and the coordinates (w, p, r) indicate the posture of the workpiece W in the robot coordinate system C1 (so-called yaw, pitch, roll).
[0032] On the other hand, column 108 indicates the status of the work on the work W. "Waiting for work" in FIG. 4 indicates that the work W with the rank i has not been worked on and work on this work W is scheduled to be performed. In this way, the processor 52 calculates the coordinates Q1w of the detected work W. i The position database 102 is created and stored in the memory 54 in association with the rank i and the status.
[0033] The processor 52 calculates the coordinates Q1w for the six detected workpieces W. iIn this case, the workpiece W located vertically uppermost (i.e., in the positive direction of the z axis of the robot coordinate system C1) among the six detected workpieces W is assigned the highest rank i=1.
[0034] Referring again to FIG. 3, in step S3, the processor 52 moves the robot 12 to the reference position P with respect to the container 100 (i.e., the workpiece W) by the mobile machine 14. 0 This reference position P 0 is predetermined by the operator as an arbitrary position on the traveling axis A. In this embodiment, the reference position P 0 is the coordinate y=y of the moving machine coordinate system C4 0 Hereinafter, the position of the robot base 20 (and the slider 36) shown in FIG. 2 on the traveling axis A will be referred to as the reference position P 0 (Y coordinate of the moving machine coordinate system C4 0 )
[0035] In step S4, the processor 52 sets the number of times n that it has been determined that the task cannot be performed in step S5, which will be described later, to n = 0. In step S5, the processor 52 executes a process of determining whether the robot 12 is capable of performing the task at this point in time. If step S5 is executed for the first time, the processor 52 sets the reference position P 0 It is determined whether the robot 12 placed in the location is capable of performing the task. This step S5 will be described with reference to FIG.
[0036] After the start of step S5, in step S11, the processor 52 determines whether the robot 12 will interfere with the environmental object SR in order to determine whether the task can be performed. Specifically, the processor 52 determines the coordinates Q1w of the workpiece W whose status is "waiting for task" and whose rank i is the highest in the position database 102 (FIG. 4) at this time. i (In the example shown in FIG. 4, the coordinate Q1w of the rank i=1 1 ) to obtain the
[0037] Then, the processor 52 calculates the coordinate Q1w of the acquired robot coordinate system C1. i The coordinate Q3w in the sensor coordinate system C3 corresponds to i and in the sensor coordinate system C3, the coordinate Q3w i The target position TP of the end effector 30 when gripping the workpiece W at the position is determined. This target position TP is expressed by the coordinate Q3t in the sensor coordinate system C3. i It is expressed as:
[0038] Here, the positional relationship between the workpiece W and the target position TP is predetermined by the operator, and positional relationship data Dr representing this positional relationship is stored in advance in the memory 54. The positional relationship data Dr may be expressed, for example, as coordinates Q5 of the tool coordinate system C2 in a workpiece coordinate system C5 (not shown) set for the workpiece W. The processor 52 calculates the coordinates Q3w i Based on the positional relationship data Dr (coordinate Q5), the target position TP is determined in the sensor coordinate system C3, and the coordinate Q3t of the target position TP is calculated. i Get.
[0039] The processor 52 then calculates the obtained coordinates Q3t i The processor 52 then uses an inverse kinematics algorithm to calculate the position of each moving component of the robot 12 (specifically, the rotation angle θ of each servo motor 32) for positioning the end effector 30 at the position represented by the set tool coordinate system C2. The processor 52 then determines whether each moving component (i.e., the rotating body 22, the lower arm 24, the upper arm 26, the wrist 22, and the end effector 30) placed at this position interferes with the environmental object SR in the sensor coordinate system C3.
[0040] The environmental object SR is, for example, the container 100 and the workpiece W (i.e., coordinates Q3w i and Q1w i For example, the processor 52 may use the coordinates Q3t iThe component models CM (i.e., the rotating torso model 22M, the lower arm model 24M, the upper arm model 26M, the wrist model 28M, and the end effector model 30M) that model each movable component of the robot 12 when the end effector 30 is placed on the sensor coordinate system C3 are simulated and placed on the sensor coordinate system C3. The processor 52 then detects interference between the three-dimensional point cloud of the environmental object SR and the movable component model CM in the sensor coordinate system C3.
[0041] Alternatively, the processor 52 may, instead of omitting the three-dimensional point cloud of the container 100 as the environmental object SR from the imaging data ID, place a container model 100M, which is a simplified model of the general shape of the container 100, in the sensor coordinate system C3. Then, the processor 52 may detect, in the sensor coordinate system C3, interference between the container model 100M and the movable component model CM, and interference between the three-dimensional point cloud of another workpiece W as the environmental object SR and the movable component model CM. Note that the processor 52 may also detect interference between at least one of the multiple component models CM (e.g., the end effector model 30M) and the three-dimensional point cloud of the environmental object SR (or the container model 100M).
[0042] Using the above method, the processor 52 determines whether the robot 12 will interfere with the environmental object SR. If interference between the robot 12 and the environmental object SR occurs, the processor 52 determines YES and proceeds to step S14. On the other hand, if the processor 52 determines NO, the processor 52 proceeds to step S12. The processor 52 may convert the three-dimensional point cloud of the container 100 and workpiece W captured in the image data ID from the sensor coordinate system C3 to the robot coordinate system C1, and detect interference between the environmental object SR (the three-dimensional point cloud or the container model 100M) and the robot 12 (the movable component model CM) in the robot coordinate system C1. In this case, the processor 52 simulates the movable component model CM (or the container model 100M) in the robot coordinate system C1.
[0043] In step S12, the processor 52 determines whether the robot 12 is positioned at a singular point SP to determine whether the task can be performed. The singular point SP corresponds to, for example, a position of the robot 12 where two of the drive axes that rotate the movable components of the robot 12 are aligned in a straight line.
[0044] In step S11, the processor 52 calculates the coordinate Q3t in the sensor coordinate system C3. i (Or, the coordinate Q1t of the target position TP in the robot coordinate system C1 i ) when the end effector model 30M is placed. If the singularity SP occurs in the component model CM, the processor 52 determines YES, and proceeds to step S14. On the other hand, if the processor 52 determines NO, the processor 52 proceeds to step S13.
[0045] In step S13, the processor 52 determines whether the robot 12 is outside the allowable operating range MR to determine whether the task can be performed. The allowable operating range MR includes, for example, a maximum reachable range MR1 that the robot 12 can reach with the end effector 30 and an allowable rotation range MR2 of each servo motor 32 of the robot 12.
[0046] The maximum reachable range MR1 may be defined as a range of a predetermined distance d (e.g., d = 3 m) from the origin of the robot coordinate system C1. The allowable rotation range MR2 may be defined as a range of rotation angles θ (e.g., -150°≦θ≦150°) of the servo motor 32. The allowable rotation range MR2 may be defined as a unique range that differs from one another for each of the multiple servo motors 32.
[0047] For example, the processor 52 calculates the coordinates Q3t of the target position TP of the end effector model 30M determined in step S11. i is converted from the sensor coordinate system C3 to the robot coordinate system C1, and the coordinates Q1t of the target position TP in the robot coordinate system C1 are calculated. i Then, the obtained coordinate Q1t iHowever, if the robot 12 is outside the maximum reachable range MR1 based on the origin of the robot coordinate system C1 at this time, it is determined that the robot 12 is outside the allowable operating range MR (i.e., YES).
[0048] The processor 52 also calculates the coordinate Q3t in the sensor coordinate system C3 calculated using the inverse kinematics algorithm in step S11. i The processor 52 acquires the rotation angle θ of each servo motor 32 when the end effector model 30M is placed on the robot 12. If at least one of the rotation angles θ of the plurality of servo motors 32 is outside the allowable rotation range MR2, the processor 52 determines that the robot 12 is outside the allowable operating range MR (i.e., YES). If the processor 52 determines YES, the process proceeds to step S14, but if the processor 52 determines NO, the process proceeds to step S7 in FIG. 3.
[0049] In this embodiment, the processor 52 executes steps S11 to S13 to determine whether the robot 12 can perform a task. Therefore, the processor 52 functions as a task determination unit 60 (FIG. 1) that determines whether a task can be performed. If a determination of YES is made in step S11, S12, or S13, it is determined that the task cannot be performed.
[0050] In step S14, the processor 52 increments the number of times n that the task is determined to be unexecutable by 1 (n=n+1). In step S15, the processor 52 increments the number of times n that the task is determined to be unexecutable by 1 (n=n+1). MAX (e.g., n MAX If processor 52 determines YES, it proceeds to step S8 in Fig. 3, whereas if processor 52 determines NO, it proceeds to step S6 in Fig. 3.
[0051] 3 again, in step S6, the processor 52 moves the workpiece W and the robot 12 relative to each other. Specifically, the processor 52 operates the mobile machine 14 to move the robot 12 in the y-axis direction of the mobile machine coordinate system C4 by a predetermined movement amount δ relative to the container 100 (i.e., the workpiece W). In this embodiment, the movement amount δ is predetermined as an arbitrary movement amount in the y-axis direction of the mobile machine coordinate system C4 (for example, δ = +500 [mm] or -500 [mm]).
[0052] If step S6 is executed for the first time, the processor 52 moves the robot 12 to the reference position P 0 (Coordinate y of the moving machine coordinate system C4 = y 0 As a result, the robot 12 moves from the target position P 1 (Coordinate y of the moving machine coordinate system C4 = y 0 The robot 12 is then placed at the position P 1 7 shows an example of the state in which the robot 12 is positioned at the target position. In this manner, in the present embodiment, the processor 52 functions as the movement control unit 62 (FIG. 1) that operates the mobile machine 14 to move the robot 12 when it is determined in step S5 that the task cannot be performed.
[0053] When the mobile machine 14 is operated to move the robot 12, the processor 52 calculates the coordinates Q1w stored in the position database 102 (FIG. 4) based on the amount of movement δ. i In this embodiment, as described above, the positive y-axis direction of the robot coordinate system C1 is set to coincide with the positive y-axis direction of the moving machine coordinate system C4.
[0054] Assume that in step S6, the robot 12 is moved by a movement amount δ=+500 [mm] in the y-axis direction of the moving machine coordinate system C4. In this case, the processor 52 calculates the coordinate Q1w stored in the position database 102. i In this way, the processor 52 corrects the coordinate Q1w in the position database 102 based on the amount of movement δ each time step S6 is executed. iCorrect the following.
[0055] The processor 52 determines the position of the workpiece W (specifically, the coordinates Q3w i or Q1w i ) in step S6, the direction in which the robot 12 is to be moved may be determined based on the coordinates Q1w of the workpiece W in the robot coordinate system C1 at this time. i If the coordinate y of the workpiece W is a positive value (in other words, if the detected workpiece W is located in the positive direction of the y axis of the moving machine coordinate system C4 relative to the robot 12), the processor 52 may set the sign of the movement amount δ to "+" (in other words, determine the movement direction of the robot 12 to be in the positive direction of the y axis of the moving machine coordinate system C4) (for example, δ = +500 [mm]).
[0056] On the other hand, at this point in time, the coordinate Q1w in the robot coordinate system C1 i If the coordinate y of the workpiece W is a negative value (in other words, if the detected workpiece W is located in the negative y-axis direction of the moving machine coordinate system C4 relative to the robot 12), the processor 52 may set the sign of the movement amount δ to "-" (in other words, determine the movement direction of the robot 12 to be in the negative y-axis direction of the moving machine coordinate system C4) (for example, δ = -500 [mm]).
[0057] After step S6, the processor 52 returns to step S5 and functions as the task determination unit 60 to execute steps S11 to S13 again to determine whether the robot 12 after the movement in the immediately preceding step S6 is capable of performing the task. In step S11 executed at this time, the processor 52 refers to the position database 102 at this time and calculates the coordinates Q1w corrected in the most recent step S6. i Get.
[0058] Then, the processor 52 calculates the corrected coordinate Q1w iand the positional relationship (i.e., the transformation matrix) between the robot coordinate system C1 and the sensor coordinate system C3 after the movement in step S6, the processor 52 determines whether or not there is interference with the environmental object SR. In this way, while the processor 52 determines in step S5 that the task cannot be performed (i.e., YES in step S11, S12, or S13) and the processor 52 determines NO in step S15, the processor 52 repeatedly executes the loop of steps S5 and S6.
[0059] Then, the processor 52 functions as the movement control unit 62 and causes the mobile machine 14 to move the robot 12 by the movement amount δ in the y-axis direction of the mobile machine coordinate system C4 each time step S6 is executed, thereby moving the robot 12 to the reference position P 0 From the position P after movement 1 , P 2 , P 3 , ... P m The position after the movement is P m (m=1, 2, 3, ...) is the coordinate y=y in the moving machine coordinate system C4 0 +mδ. In addition, the processor 52 calculates the coordinate Q1w in the position database 102 according to the amount of movement δ each time step S6 is executed. i The position database 102 is updated by correcting the above.
[0060] On the other hand, if it is determined in step S5 that the work can be performed (i.e., NO in steps S11, S12, and S13 in FIG. 5), in step S7, the processor 52 causes the robot 12 to perform the work. Specifically, the processor 52 determines the coordinates Q1w of the workpiece W whose status is "waiting for work" and whose rank i is the highest in the position database 102 (e.g., FIG. 4) at this time. i (For example, the coordinate Q1w of rank i=1 1 ) to obtain the
[0061] Then, the processor 52 calculates the obtained coordinate Q1w i Based on the positional relationship data Dr, a target position TP is determined in the robot coordinate system C1, and the coordinates Q1t of the target position TP are calculated. i Then, the processor 52 obtains the coordinate Q1t iThe robot 12 is operated so as to position the end effector 30 at the position represented by the set tool coordinate system C2. As a result, the end effector 30 is positioned at the workpiece W to be worked on.
[0062] The processor 52 then operates the end effector 30 to grip the workpiece W, and then operates the robot 12 to move the end effector 30 upward. In this manner, the robot 12 performs workpiece handling by gripping and removing the workpiece W from the container 100. In this manner, in this embodiment, the processor 52 functions as an operation command unit 64 ( FIG. 1 ) that causes the robot 12 to perform a task.
[0063] After executing step S7, the processor 52 changes the status of the row of rank i (e.g., i=1) in the position database 102 corresponding to the workpiece W for which the work has been successfully performed from "waiting for work" to "work successful," which indicates that the work has been appropriately successful. On the other hand, when the determination is YES in step S15 in Fig. 5, the processor 52 changes the status from "waiting for work" to "calculation failed," which indicates that the position of the robot 12 that can perform the work could not be calculated.
[0064] Such statuses are shown in Fig. 6. In the example shown in Fig. 6, the status of the item with rank i=1 is "Work Success", while the status of the item with rank i=2 is "Calculation Failed". In this way, in this embodiment, when the result of the determination in step S15 is YES (in other words, the number of times n that the task was determined to be impossible in step S5 is n=n MAX = 10), the operation on the workpiece W is canceled without executing step S7.
[0065] 3 again, in step S8, the processor 52 determines whether or not there is any work W that has not yet been worked on. Specifically, the processor 52 references the position database 102 at this time, and if there is any work W whose status is "waiting for work", the processor 52 determines YES and returns to step S3. On the other hand, if there is no work W whose status is "waiting for work", the processor 52 determines NO and proceeds to step S9.
[0066] 6, since the status of the workpieces W with ranks i=3 to 6 is "waiting for work", the processor 52 determines YES in step S8. Then, the processor 52 functions as the movement control unit 62 to execute step S3, and moves the robot 12 by the mobile machine 14 to the reference position P 0 Place it in.
[0067] If step S3 is executed after step S6 is executed m times, the processor 52 moves the robot 12 by the mobile machine 14 to the position P m (Coordinate y of the moving machine coordinate system C4 = y 0 +mδ) to the reference position P 0 (coordinate y = y 0 ) by the movement amount mδ. As a result, the robot 12 moves to the reference position P 0 Then, in the same manner as in step S6 described above, the processor 52 returns to the coordinate Q1w stored in the position database 102 at this point. i is corrected based on the amount of movement mδ (or the position database 102 initially created in step S2 is restored).
[0068] Thereafter, the processor 52 sequentially executes steps S4 to S8, and in step S5, functions as the operation determination unit 60 to determine the reference position P 0 In step S7, the processor 52 determines whether the robot 12 can perform the work on the next work W (in the example of FIG. 6, the work W with rank i=3). If it is determined that the work can be performed, the processor 52 functions as the operation command unit 64 and performs the work on the next work W.
[0069] In step S9, the processor 52 determines whether or not the work on all the workpieces W in the container 100 has been completed. If the processor 52 determines YES, it ends the flow of FIG. 3, whereas if the processor 52 determines NO, it returns to step S1 and sequentially executes steps S1 to S9. Note that the processor 52 may determine YES in step S9 when it receives a work completion command (or a shutdown command) from the operator, the upper controller, or the computer program PG1.
[0070] As described above, in this embodiment, the control device 50 has the functions of the work determination unit 60, the movement control unit 62, and the operation command unit 64, and controls the robot 12 and the mobile machine 14 to perform work on the workpiece W (in this embodiment, workpiece handling). In this control device 50, the work determination unit 60 determines whether the mobile machine 14 is at the reference position P 0 It is determined whether the robot 12 placed in the area is capable of performing the task (first step S5).
[0071] When the task determination unit 60 determines that the task cannot be performed (YES in step S11, S12, or S13), the movement control unit 62 operates the mobile machine 14 to move the robot 12 to the reference position P 0 Then, the task determining unit 60 determines whether the robot 12 after being moved by the movement control unit 62 is capable of performing the task (step S5 after step S6).
[0072] Here, the reference position P 0 7, the robot 12 placed at the reference position P may be unable to perform work on the workpiece W due to the above-mentioned interference, singularity, allowable operating range, etc. Even in such a case, the work may become possible by changing the positional relationship between the workpiece W (container 100) and the robot 12. According to this embodiment, the reference position P 0 If it is determined that the work cannot be performed, the robot 12 and the workpiece W are moved relatively by the mobile machine 14 to a position P after the movement as shown in FIG. mIn this way, it is possible to effectively resolve the state in which the robot 12 is unable to perform the task, thereby improving work efficiency.
[0073] In this embodiment, the task determination unit 60 determines whether the robot 12 will interfere with the workpiece W or the environmental object SR (step S11), whether the robot 12 will be positioned at a singular point SP (step S12), and whether the robot 12 will be outside the allowable operating range MR (specifically, the maximum reach range MR1 and the allowable rotation range MR2) (step S13) to determine whether the task can be performed. With this configuration, by determining whether the robot 12 will interfere with the environmental object SR, be positioned at the singular point SP, or move outside the allowable operating range MR, it is possible to effectively and accurately determine whether the task can be performed.
[0074] In this embodiment, when the task determination unit 60 determines that the robot 12 can perform the task after movement (when the determination is NO in steps S11 to S13 after step S6), the operation command unit 64 causes the robot 12 to perform the task (step S7). After the task is performed by the operation command unit 64, the movement control unit 62 operates the mobile machine 14 to move the robot 12 to the reference position P 0 (step S3).
[0075] Then, the operation determination unit 60 determines the reference position P 0 In step S5, it is determined whether the robot 12 can perform the work on the next work W (for example, the work W with the rank i=3 in FIG. 6). 0 , the subsequent steps S5 and S6 are performed at the reference position P 0 This can improve the efficiency of the calculation processes in steps S5 and S6.
[0076] In this embodiment, the mobile machine 14 has a traveling device that moves the robot 12 along the traveling axis A, and the movement control unit 62 operates the mobile machine 14 to move the robot 12 to a reference position P on the traveling axis A.0 According to this configuration, the control device 50 can control the position of the robot 12 on the traveling axis A with high precision by controlling the mobile machine 14 as a traveling device.
[0077] Next, other functions of the robot system 10 will be described with reference to Fig. 8. In this embodiment, the processor 52 executes the flow shown in Fig. 9. The flow shown in Fig. 9 differs from the flow in Fig. 3 in the following respects. That is, in the flow in Fig. 9, when a YES determination is made in step S8, step S10 is executed. Below, a case where step S10 is executed after steps S6 and S7 are executed will be described.
[0078] In step S10, the processor 52 moves the robot W to the reference position P based on the position of the unprocessed workpiece W. 0 Specifically, the processor 52 refers to the position database 102 at this time and determines whether the coordinates Q1w of the workpiece W whose status is "waiting for work" and whose rank i is the highest. i For example, in the example of FIG. 6, the processor 52 obtains the coordinate Q1w of the rank i=3. 3 Get.
[0079] Then, the processor 52 calculates the obtained coordinate Q1w 3 The position of the origin of the robot coordinate system C1 at this time and the reference position P 0 Based on this, the robot W is moved to the reference position P 0 As an example, the processor 52 determines whether the acquired coordinate Q1w of the robot coordinate system C1 should be returned to the 3 , and the coordinate Q4w of the moving machine coordinate system C4. 3 and convert the coordinate Q4w 3 The coordinate y = y 3 Get.
[0080] On the other hand, the processor 52 calculates the coordinate y=y of the origin of the robot coordinate system C1 at this time in the moving machine coordinate system C4. 1 and the reference position P 0 The coordinate y of the moving machine coordinate system C4 is y=y 0Then, the processor 52 obtains the coordinate y 3 and the coordinate y 1 Difference Δ 3_1 (= |y 3 -y 1 |) and the coordinate y 3 and the coordinate y 0 Difference Δ 3_0 (= |y 3 -y 0 |) and
[0081] The processor 52 calculates the distance Δ 3_1 is the distance Δ 3_0 If it is greater than (Δ 3_1 >Δ 3_0 ), the robot W is moved to the reference position P 0 Δ 3_1 >Δ 3_0 This means that the workpiece W to be worked on next is located closer to the reference position P than the robot 12 at this point in the y-axis direction of the moving machine coordinate system C4. 0 means close to.
[0082] In this way, the processor 52 determines the position (coordinates Q1w) of the unprocessed work W (for example, the work W with rank i=3). 3 , Q4w 3 ) based on which the robot W is moved to the reference position P 0 Therefore, the processor 52 determines whether the robot W should be returned to the reference position P 0 The function of the return determining unit 66 (FIG. 8) is to determine whether or not the process should be returned to the normal state.
[0083] If the determination in step S10 is YES, the processor 52 returns to step S3, and functions as the movement control unit 62 to operate the mobile machine 14 and move the robot 12 to the reference position P 0 On the other hand, if the determination in step S10 is NO, the processor 52 returns to step S4. Therefore, in this case, the processor 52 functions as the movement control unit 62 to maintain the mobile machine 14 in a stationary state, thereby returning the robot W to the position P after the movement in the most recently executed step S6. m Then, the processor 52 executes step S5 and keeps the position Pm In this step, it is determined whether the robot 12 is capable of performing the work on the next workpiece W.
[0084] As described above, in this embodiment, when the task determination unit 60 determines that the robot 12 after the movement in step S6 is capable of performing the task, the operation command unit 64 causes the robot 12 to perform the task (step S7 after step S6). Then, after the task is performed by the operation command unit 64, the task determination unit 60 determines the position P m In step S10, it is determined whether the robot 12 can perform the work on the next workpiece W (step S5 after determining NO in step S10). 0 , without returning to the position P after the movement in step S6. m This will enable you to determine whether or not the work can be done.
[0085] In this embodiment, after the operation command unit 64 executes the work, the return determination unit 66 determines the position of the unprocessed workpiece W (coordinates Q1w 3 , Q4w 3 ), the robot 12 is moved to the reference position P 0 Then, the movement control unit 62 determines whether or not the reference position P 0 If it is determined that the robot 12 should be returned to the reference position P (YES in step S10), the mobile machine 14 is operated to return the robot 12 to the reference position P 0 (step S3).
[0086] On the other hand, the movement control unit 62 determines the reference position P 0 If it is determined that the robot 12 should not be returned to the position P after the movement (NO in step S10), m According to this configuration, the reference position P 0 or to the position P m This allows you to choose whether to keep the work in progress or not, thereby reducing the work cycle time.
[0087] 3 or 9 in accordance with a computer program PG1 stored in the memory 54. The functions of the work determination unit 60, the movement control unit 62, the operation command unit 64, and the return determination unit 66 executed by the processor 52 may be functional modules realized by the computer program PG1.
[0088] It should be noted that various modifications can be made to the flow shown in Figure 3, Figure 5, or Figure 9. For example, the order of steps S11, S12, and S13 in Figure 5 may be changed, and the processor 52 may execute steps S13 → S11 → S12 in that order. Also, step S11, S12, or S13 may be omitted from the flow in Figure 5. That is, the processor 52 may execute at least one of steps S11, S12, and S13 in order to determine whether or not the task can be executed in step S5.
[0089] 5, steps S14 and S15 may be omitted, and the processor 52 may proceed to step S6 when determining YES in step S11, S12, or S13. The processor 52 may also execute step S6 and the subsequent step S5 in parallel in the flow of Fig. 3 or 9. That is, in this case, the processor 52 performs calculations for the determination of step S5 while moving the robot 12 in step S6.
[0090] 3 or 9, the processor 52 may omit step S7 and execute the determination of whether or not the work can be performed in step S5 without actually performing the work on the workpiece W. This configuration allows the operator to verify whether or not the work can be performed before actually performing the work. In other words, in this case, the operation command unit 64 can be omitted from the control device 50.
[0091] 3 or 9, the processor 52 actually moves the actual robot 12 by the mobile machine 14 in step S6. However, the present invention is not limited to this. The processor 52 may function as the movement control unit 62 and execute step S6' of simulating the movement of the robot 12 in a virtual space defined by the mobile machine coordinate system C4, instead of step S6 of moving the actual robot 12 in real space.
[0092] Then, the processor 52 may execute step S5 after step S6' and function as the task determination unit 60 to determine whether the robot 12 after movement in the virtual space is capable of performing the task. In this way, even when the robot 12 is moved simulated in the moving machine coordinate system C4 without moving the actual robot 12, the determination in step S5 can be performed by calculation.
[0093] If the processor 52 determines as a result of step S5 that the task can be performed, it may function as the movement control unit 62 to execute step S6 in which the actual robot 12 is moved in real space by the movement amount δ, and may function as the operation command unit 64 to execute step S7. In this case, the processor 52 may execute steps S6 and S7 in parallel, thereby reducing the task cycle time.
[0094] Next, with reference to Figure 10, further functions of the robot system 10 will be described. In this embodiment, the function of a machine learning device 70 that learns the movement amount δ for moving the robot 12 by the mobile machine 14 is implemented in the control device 50. The machine learning device 70 includes a state observation unit 72, a learning unit 74, and a decision-making unit 76. The state observation unit 72 observes position data Dp indicating the position P of the robot with respect to the workpiece W and determination data Dd indicating whether the robot 12 can perform work on the workpiece W at the position P, as state variables SV that represent the current state of the environment in which the robot 12 works.
[0095] The learning unit 74 uses the state variable SV to learn the movement amount δ in association with whether or not the task can be performed. The learning unit 74 repeatedly executes learning based on the state variable SV obtained by repeatedly attempting the task according to an arbitrary learning algorithm collectively known as machine learning. The decision-making unit 76 outputs a command value Cδ for the movement amount δ to be commanded to the mobile machine 14 when attempting the task, based on the learning results by the learning unit 74.
[0096] In this embodiment, the processor 52 functions as a machine learning device 70 (i.e., a state observation unit 72, a learning unit 74, and a decision-making unit 76), and proceeds with learning of the movement amount δ by executing the learning flow shown in Fig. 11. In the learning flow shown in Fig. 11, the processor 52 randomly selects a movement amount δ, moves the robot 12 by the movement amount δ using the mobile machine 14 (i.e., executes step S6), and determines the position P after the movement. m and then determining whether the robot 12 is capable of performing the task (i.e., executing step S5). Then, each time the processor 52 attempts this task, it functions as the state observing unit 72 and collects the state variables SV.
[0097] The learning flow shown in Figure 11 shows an example of executing reinforcement learning as an example of a learning algorithm. However, the learning algorithm is not particularly limited, and any learning algorithm may be adopted, such as supervised learning, unsupervised learning, or a neural network. The reinforcement learning algorithm is known as Q-learning, which is a method of learning a function EQ(s, a) that represents the value of an action when action a is selected in state s, using the state s of an agent and actions a that the agent can select in state s as independent variables.
[0098] The optimal solution is to select an action a that maximizes the value function EQ in state s. Q-learning is started when the correlation between state s and action a is unknown, and trial and error is repeated to select various actions a in any state s, thereby iteratively updating the value function EQ and approaching the optimal solution. Here, when the environment (i.e., state s) changes as a result of selecting action a in state s, a reward r (i.e., weighting of action a) corresponding to the change is obtained, and learning is guided to select an action a that obtains a higher reward r, thereby making it possible to approach the optimal solution with the value function EQ in a relatively short time.
[0099] The update equation for the value function EQ can generally be expressed as the following equation (1).
[0100]
[0101] In formula (1), s t and a t are the state and action at time t, respectively, and action a t The state is s t+1 It changes to r t+1 is the state s t From t+1 The term maxQ means the Q when action a is taken that is thought to have the maximum value Q at time t+1. α and γ are the learning coefficient and discount rate, respectively, and are set arbitrarily between 0<α≦1 and 0<γ≦1.
[0102] When Q-learning is performed, the state variable SV observed by the state observation unit 72 corresponds to the state s in the update equation, and the action indicating how far the robot 12 should be moved by the mobile machine 14 when the robot 12 in its current state performs a task (i.e., the movement amount δ) corresponds to the action a in the update equation. Furthermore, the reward R to be calculated corresponds to the reward r in the update equation. The learning unit 74 repeatedly updates the function EQ, which represents the value of the movement amount δ when the robot 12 in its current state performs a task, by Q-learning using the reward R.
[0103] In step S21, the processor 52 functions as the decision-making unit 76, references the value function EQ at this time, and randomly selects a movement amount δ as an action to be performed in the current state indicated by the state variables SV (position data Dp and judgment data Dd) observed by the state observing unit 72. The processor 52 outputs a command value Cδ for the selected movement amount δ to the mobile machine 14 (specifically, the servo motor 40), and operates the mobile machine 14 in accordance with the command value Cδ, thereby moving the robot 12 by the movement amount δ.
[0104] In step S22, the processor 52 functions as the state observation unit 72 and acquires a state variable SV of the current state. Specifically, the processor 52 acquires, as the state variable SV, position data Dp indicating the position P of the robot relative to the workpiece W. For example, the processor 52 acquires, as the position data Dp, the coordinates Q3w of the workpiece W to be worked on. i and Q1w i Furthermore, the processor 52 acquires, as the position data Dp, the positions P of the robot 12 before and after the robot 12 is moved by the movement amount δ selected in the most recent step S21 (i.e., the reference position P 0 , or position P m ) is acquired as the coordinate y of the moving machine coordinate system C4.
[0105] Furthermore, when the processor 52 has moved the robot 12 by the movement amount δ selected in the most recent step S21, the processor 52 executes step S5 described above to determine whether the task can be performed. If the processor 52 determines that the task can be performed (i.e., NO in steps S11 to S13 in FIG. 5), the processor 52 acquires, as the state variable SV, determination data Dd1 indicating that the task can be performed. Conversely, if the processor 52 determines that the task cannot be performed (i.e., YES in step S11, S12, or S13 in FIG. 5), the processor 52 acquires, as the state variable SV, determination data Dd2 indicating that the task cannot be performed.
[0106] In step S23, the processor 52 determines whether the work was possible. Specifically, if the processor 52 obtained determination data Dd1 indicating that the work was possible in the most recent step S22, the processor 52 determines it as YES. Conversely, if the processor 52 obtained determination data Dd2 indicating that the work was not possible in the most recent step S22, the processor 52 determines it as NO. If the processor 52 determines it as YES, the processor 52 proceeds to step S24, and if the processor 52 determines it as NO, the processor 52 proceeds to step S25.
[0107] In step S24, the processor 52 functions as the learning unit 74, calculates a positive reward +R, and applies the calculated positive reward +R to the update equation for the function EQ. On the other hand, if the determination in step S23 is NO, in step S25, the processor 52 calculates a negative reward −R and applies it to the update equation for the function EQ. Note that in step S25, instead of giving a negative reward −R, the processor 52 may apply a reward R=0 to the update equation for the function EQ. In this way, in this embodiment, the processor 52 functions as a reward calculation unit 78 ( FIG. 10 ) that calculates a reward related to whether or not the task can be completed. By giving a reward R according to whether or not the task can be completed in this way, learning of the movement amount δ is guided in the direction of selecting an action that enables the task to be completed.
[0108] In step S26, processor 52 functions as learning unit 74 to update value function EQ. As such, in this embodiment, processor 52 functions as function update unit 80 ( FIG. 10 ) that updates function EQ, which represents the value of movement amount δ. By repeating the learning cycle of steps S21 to S26, processor 52 iteratively updates value function EQ, and as a result, can progress learning of movement amount δ associated with the feasibility of work.
[0109] After obtaining the value function EQ as the learning result, the processor 52 executes the flow of FIG. 3 or FIG. 9, and determines the movement amount δ using the value function EQ in step S6. Specifically, in step S6, the processor 52 determines the coordinates Q3w of the workpiece W obtained in the most recent step S2. i or Q1w iand the data of the position P of the robot 12 at this time (reference position P 0 or position P m The value function EQ then outputs a movement amount δ corresponding to the input state variable SV. The processor 52 outputs a command value Cδ for the movement amount δ output by the value function EQ to the servo motor 40 of the mobile machine 14, and causes the mobile machine 14 to move the robot 12 by the movement amount δ.
[0110] As described above, in the machine learning device 70, the state observing unit 72 obtains the position data Dp (coordinates Q3 of the workpiece W) indicating the position P of the robot 12 relative to the workpiece W. i and Q1w i , the position P of the robot 12 0 or P m coordinate y) and the judgment data Dd (Dd1 or Dd2) indicating whether the robot 12 can perform work on the workpiece W at the position P are observed as state variables SV (step S22).
[0111] The learning unit 74 then uses the state variable SV to learn the movement amount δ in association with whether the task can be performed (steps S24 to S26). With this configuration, when moving the robot 12 in step S6 in the flow of Fig. 3 or 9, for example, the movement amount δ suitable for enabling the task to be performed can be automatically learned. This makes it possible to avoid the task becoming impossible to perform after step S6, thereby improving the efficiency of the task.
[0112] Furthermore, in this embodiment, the learning unit 74 has a reward calculation unit 78 that calculates a reward R related to whether or not the task can be completed, and a function update unit 80 that updates a function EQ that represents the value of the movement amount δ using the reward R. With this configuration, learning of the movement amount δ is guided in the direction of selecting an action that makes it possible to perform the task, and therefore, learning of the movement amount δ can be efficiently advanced.
[0113] Furthermore, in this embodiment, the decision-making unit 76 outputs a command value Cδ for the amount of movement δ to be commanded to the mobile machine 14 based on the learning results of the learning unit 74. The state observing unit 72 then observes the judgment data Dd obtained when the mobile machine 14 is operated in accordance with the command value Cδ as the state variable SV for the next learning cycle. With this configuration, the amount of movement δ can be automatically changed and trials of the learning task can be repeated, making it possible to automate the learning cycle.
[0114] The processor 52 may function as a reward calculation unit 78 to calculate a different reward R depending on the time T required for the task or the load τ applied to the robot 12 during the task. When performing such machine learning, the processor 52 selects the movement amount δ, moves the robot 12 (step S6), and calculates the position P after the movement. m In addition to determining whether the work can be performed (step S5), if it is determined that the work can be performed, the work is attempted on the workpiece W (i.e., step S7).
[0115] In step S22 described above, the processor 52 functions as the state observing unit 72 and acquires the time T required for the task as the state variable SV. For example, during the trial of the task, the processor 52 acquires the time T1 required for the movement of the robot 12 (step S6) and the task on the workpiece W (step S7). Alternatively, the processor 52 may acquire the time T2 required for a series of tasks including the selection of the movement amount δ, the movement of the robot 12 (step S6), the determination of whether the task can be completed (step S5), and the task on the workpiece W (step S7).
[0116] Then, the processor 52 determines whether the acquired time T (T1 or T2) is equal to or smaller than a predetermined threshold T th If T≦T th If T>T, the processor 52 calculates the positive reward +R in step S24 described above and applies it to the update formula for the function EQ. thIf so, in step S25 described above, processor 52 calculates a negative reward −R and applies it to the update equation for function EQ. By providing a reward R according to the time T required for the task in this way, learning of the movement amount δ is guided in the direction of selecting an action that reduces the time T.
[0117] Furthermore, in step S22 described above, the processor 52 functions as the state observation unit 72 and acquires, as the state variable SV, the load τ applied to the robot 12 when attempting to perform work on the workpiece W (step S7). For example, the processor 52 acquires, as the load τ, a load torque τ1 applied to each servo motor 32 of the robot 12 during execution of step S7. Alternatively, the processor 52 may acquire, as the load τ, a force τ2 applied to a component of the robot 12 (e.g., the robot base 20).
[0118] Then, the processor 52 determines whether the acquired load τ (load torque τ1 or force τ2) is greater than or equal to a predetermined threshold τ th If τ≦τ, then it is determined whether τ is equal to or smaller than τ. th If τ>τ, the processor 52 calculates the positive reward +R in step S24 described above and applies it to the update formula for the function EQ. th If so, the processor 52 calculates a negative reward −R in step S25 described above and applies it to the update formula for the function EQ. By providing a reward R according to the load τ in this way, the learning of the movement amount δ is guided in the direction of selecting an action that reduces the load τ.
[0119] In the present embodiment, the case where the functions of the machine learning device 70 are implemented in the control device 50 has been described. However, this is not limiting, and the functions of the machine learning device 70 can be implemented in any computer other than the control device 50, such as a higher-level controller or a PC.
[0120] In the above embodiment, the positive x-axis direction, positive y-axis direction, and positive z-axis direction of the robot coordinate system C1 are respectively identical to the positive x-axis direction, positive y-axis direction, and positive z-axis direction of the moving machine coordinate system C4. However, the present invention is not limited to this, and the robot coordinate system C1 and the moving machine coordinate system C4 may be set to have any positional relationship.
[0121] In the above embodiment, the robot coordinate system C1 is fixed to the robot base 20 and moves together with the robot 12 in the movable machine coordinate system C4. However, this is not limiting, and the robot coordinate system C1 may be fixed in the movable machine coordinate system C4. The processor 52 may also execute the above steps S2 to S8 and S10 based on the movable machine coordinate system C4. Alternatively, a world coordinate system C6 (not shown) may be set that defines the three-dimensional space of the work cell, and the processor 52 may execute the above steps S2 to S8 and S10 based on the world coordinate system C6.
[0122] In the above embodiment, the robot 12 is moved by the mobile machine 14. However, the present invention is not limited to this. The mobile machine 14 may move the container 100 (i.e., the workpiece W) to the reference position P 0 It should be understood that the concepts of the present disclosure described with reference to Figures 3, 9 and 11 are equally applicable to moving the workpiece W by the mobile machine 14.
[0123] Furthermore, the visual sensor 16 may be fixed at a fixed point above the container 100, or may be provided on a movable component of the robot 12 (e.g., the wrist 28 or the end effector 30) and moved by the robot 12. Furthermore, the robot 12 is not limited to a vertical articulated type, and may be any other type of robot, such as a horizontal articulated type or a parallel link type. Furthermore, the mobile machine 14 is not limited to a traveling device, and may be, for example, a work table device having a first ball screw mechanism that moves the robot 12 or the container 100 (workpiece W) in the x-axis direction of the mobile machine coordinate system C4 and a second ball screw mechanism that moves it in the y-axis direction, or any other mechanism.
[0124] Although the present disclosure has been described in detail above, the present disclosure is not limited to the individual embodiments described above. Various additions, substitutions, modifications, partial deletions, etc. are possible in these embodiments without departing from the gist of the present disclosure or the spirit of the present disclosure derived from the content of the claims and their equivalents. These embodiments can also be implemented in combination. For example, in the above-described embodiments, the order of each operation and the order of each process are shown as examples and are not limited to these. The same applies when numerical values or mathematical expressions are used in the description of the above-described embodiments.
[0125] The present disclosure describes the following aspects: (Aspect 1) A control device 50 controls a robot 12 that performs a predetermined task on a workpiece W and a mobile machine 14 that moves the workpiece W and the robot 12 relative to each other, and the mobile machine 14 moves the workpiece W to a reference position P 0 a work determination unit (60) that determines whether the robot (12) placed at the reference position P is capable of performing the work; and, when the work determination unit (60) determines that the work is not capable of performing the work, the mobile machine (14) is operated to move the workpiece (W) or the robot (12) to the reference position P. 0 and a movement control unit 62 that moves the robot 12 from the reference position P, and the work determination unit 60 determines whether the robot 12 is capable of performing the work after being moved by the movement control unit 62. (Aspect 2) The control device 50 according to Aspect 1, wherein the work determination unit 60 determines whether the robot 12 will interfere with the environmental object SR, determines whether the robot 12 is located at a singular point, or determines whether the robot 12 is outside the allowable operating range MR, in order to determine whether the work can be performed. (Aspect 3) The control device 50 further includes an operation command unit 64 that causes the robot 12 to perform the work when the work determination unit 60 determines that the robot 12 is capable of performing the work after being moved, and the movement control unit 62 operates the mobile machine 14 to move the workpiece W or the robot 12 to the reference position P after the work is performed by the operation command unit 64. 0 , and the operation determination unit 60 returns the reference position P 0The control device 50 according to aspect 1 or 2 determines whether the robot 12 can perform the work on the next workpiece W at the next workpiece W. (Aspect 4) The control device 50 further includes an operation command unit 64 that causes the robot 12 to perform the work when the operation determination unit 60 determines that the robot 12 can perform the work after the movement, and the operation determination unit 60 determines whether the robot 12 can perform the work on the next workpiece W at the next workpiece W at the next workpiece W. m The control device 50 according to the first or second aspect determines whether the robot 12 can perform the work on the next workpiece W. (Aspect 5) After the work is performed by the operation command unit 64, the control device 50 moves the workpiece W or the robot 12 to the reference position P based on the position of the unworked workpiece W. 0 The movement control unit 62 further includes a return determination unit 66 that determines whether or not the movement control unit 62 should return to the reference position P 0 If it is determined that the workpiece W or the robot 12 should be returned to the reference position P, the mobile machine 14 is operated to return the workpiece W or the robot 12 to the reference position P. 0 , while the return determination unit 66 returns the reference position P 0 If it is determined that the workpiece W or the robot 12 should not be returned to the position P m The control device 50 according to aspect 3 or 4, wherein the mobile machine 14 has a traveling device that moves the robot 12 along the traveling axis A, and the movement control unit 62 operates the mobile machine 14 to move the robot 12 to a reference position P on the traveling axis A. 0 A control device (50) according to any one of aspects (1) to (5) above, wherein the control device (50) controls the robot (12) to perform a predetermined task on the workpiece (W) and the mobile machine (14) to move the workpiece (W) relative to the robot (12), the control device (50) according to any one of aspects (1) to (5) above, the control device (50) controls the mobile machine (14) to move the workpiece (W) relative to the reference position (P 0 If it is determined that the robot 12 cannot perform the work, the mobile machine 14 is operated to move the workpiece W or the robot 12 to the reference position P. 0and determining whether the robot 12 is capable of performing the task after the movement. (Aspect 8) A computer program PG1 causing a processor 52 to execute the method described in Aspect 7. (Aspect 9) A machine learning device 70 for learning a movement amount δ for moving a robot 12 performing a predetermined task on a workpiece W or the robot 12 by a mobile machine 14 that moves the robot 12 relative to the workpiece W, the machine learning device 70 including: a state observation unit 72 that observes position data Dp indicating a position P of the robot 12 relative to the workpiece W and determination data Dd indicating whether the robot 12 is capable of performing the task on the workpiece W at the position P as state variables SV representing the current state of the environment in which the robot 12 performs the task; and a learning unit 74 that learns the movement amount δ by using the state variables SV to associate the movement amount δ with whether the task can be performed. (Aspect 10) The machine learning device 70 according to Aspect 9, wherein the learning unit 74 includes a reward calculation unit 78 that calculates a reward R related to whether the task can be performed, and a function update unit 80 that updates a function EQ representing the value of the movement amount δ using the reward R. (Aspect 11) The machine learning device 70 according to aspect 10, wherein the reward calculation unit 78 calculates a reward R that varies depending on the time T required for the task or the load τ applied to the robot 12 during the task. (Aspect 12) The machine learning device 70 according to any of aspects 9 to 11, further comprising a decision making unit 76 that outputs a command value Cδ of a movement amount δ to be commanded to the mobile machine 14 based on a learning result by the learning unit 74, and wherein the state observation unit 72 observes, as a state variable SV in the next learning cycle, determination data Dd obtained when the mobile machine 14 is operated in accordance with the command value Cδ. (Mode 13) A machine learning method for learning the movement amount δ for moving a workpiece W or a robot 12 by a mobile machine 14 that moves a robot 12 that performs a specified task on the workpiece W relative to the workpiece W, wherein position data Dp indicating the position P of the robot 12 relative to the workpiece W and judgment data Dd indicating whether the robot 12 is able to perform the task on the workpiece W at the position P are observed as state variables SV that represent the current state of the environment in which the robot 12 performs the task, and the state variables SV are used to learn the movement amount δ in association with whether the task can be performed.
[0126] REFERENCE SIGNS LIST 10 Robot system 12 Robot 14 Mobile machine 16 Visual sensor 50 Control device 52 Processor 60 Work determination unit 62 Movement control unit 64 Operation command unit 66 Return determination unit 70 Machine learning device 72 State observation unit 74 Learning unit 76 Decision-making unit 78 Reward calculation unit 80 Function update unit
Claims
1. A control device for controlling a robot that performs a specified task on a workpiece and a mobile machine that moves the workpiece and the robot relative to one another, comprising: a task determination unit that determines whether the robot, which has been placed at a reference position relative to the workpiece by the mobile machine, is capable of performing the task; and a movement control unit that operates the mobile machine to move the workpiece or the robot from the reference position when the task determination unit determines that the task cannot be performed, wherein the task determination unit determines whether the robot, after movement by the movement control unit, is capable of performing the task.
2. The control device according to claim 1, wherein the task determination unit determines whether the robot will interfere with an environmental object, determines whether the robot will be positioned at a singular point, or determines whether the robot will be outside of an allowable operating range in order to determine whether the task can be performed.
3. A control device as described in claim 1, further comprising an operation command unit that causes the robot to perform the task when the task determination unit determines that the robot is capable of performing the task after the movement, wherein the movement control unit operates the mobile machine to return the workpiece or the robot to the reference position after the task is performed by the task determination unit, and the task determination unit determines whether the robot is capable of performing the task on the next workpiece at the reference position.
4. The control device described in claim 1, further comprising an operation command unit that causes the robot to perform the task when the task determination unit determines that the robot is capable of performing the task after the movement, and the operation determination unit determines whether the robot is capable of performing the task on the next workpiece at the position after the movement after the task is performed by the task determination unit.
5. A control device as described in claim 3 or 4, further comprising a return judgment unit that judges whether or not the workpiece or the robot should be returned to the reference position based on the position of the unworked workpiece after the operation command unit has performed the work, and the movement control unit operates the mobile machine to return the workpiece or the robot to the reference position when the return judgment unit judges that the workpiece or the robot should be returned to the reference position, but leaves the workpiece or the robot in the position after the movement when the return judgment unit judges that the workpiece or the robot should not be returned to the reference position.
6. The control device according to claim 1, wherein the mobile machine has a traveling device that moves the robot along a traveling axis, and the movement control unit operates the mobile machine to move the robot from the reference position on the traveling axis.
7. A method for controlling a robot that performs a specified task on a workpiece and a mobile machine that moves the workpiece and the robot relatively, wherein a processor determines whether the robot, which has been placed at a reference position relative to the workpiece by the mobile machine, is capable of performing the task, and if it is determined that the task cannot be performed, operates the mobile machine to move the workpiece or the robot from the reference position, and determines whether the robot after the movement is capable of performing the task.
8. A computer program product causing said processor to perform the method of claim 7.
9. A machine learning device that learns the amount of movement of a workpiece or a robot by a mobile machine that moves a robot that performs a specified task on a workpiece relative to the workpiece, the machine learning device comprising: a state observation unit that observes position data that indicates the position of the robot relative to the workpiece and judgment data that indicates whether the robot can perform the task on the workpiece at that position as state variables that represent the current state of the environment in which the robot performs the task; and a learning unit that uses the state variables to learn the amount of movement in association with whether the task can be performed.
10. The machine learning device of claim 9, wherein the learning unit has: a reward calculation unit that calculates a reward related to whether the task is possible or not; and a function update unit that uses the reward to update a function that represents the value of the movement amount.
11. The machine learning device according to claim 10, wherein the reward calculation unit calculates different rewards depending on the time required to perform the task or the load placed on the robot during the task.
12. The machine learning device of claim 9, further comprising a decision-making unit that outputs a command value for the movement amount to be commanded to the mobile machine based on the learning results by the learning unit, and the state observation unit observes the judgment data when the mobile machine is operated in accordance with the command value as the state variable in the next learning cycle.
13. A machine learning method for learning the amount of movement of a workpiece or a robot by a mobile machine that moves a robot that performs a specified task on the workpiece relative to the workpiece, the machine learning method observing position data indicating the position of the robot relative to the workpiece and judgment data indicating whether the robot can perform the task on the workpiece at that position as state variables representing the current state of the environment in which the robot performs the task, and using the state variables to learn the amount of movement in association with whether the task can be performed.
Citation Information
Patent Citations
Robot system and control method of robot system
JP2020196058A
Conveyance robot and control method
JP2023106692A