Control device and control method

The control device employs a deep reinforcement learning model to determine optimal gripping points for robotic arms, addressing inefficiencies in workpiece handling by minimizing center of gravity deviation and enhancing gripping precision.

JP7848769B2Active Publication Date: 2026-04-21TOYOTA JIDOSHA KK
View PDF 9 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
TOYOTA JIDOSHA KK
Filing Date
2023-07-27
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing control devices for robotic arms struggle to accurately determine a gripping point for workpieces, leading to inefficient or unsuccessful picking operations.

Method used

A control device utilizing a deep reinforcement learning model to learn and adjust the gripping point of a workpiece based on the deviation from its center of gravity, incorporating joint information of the robotic arm, and employing a path generation unit to derive optimal joint configurations for precise gripping.

Benefits of technology

Enables accurate and efficient gripping of various workpieces by minimizing deviation from the center of gravity, reducing computational burden, and allowing for seamless handling of differently shaped or packaged items.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007848769000001
    Figure 0007848769000001
  • Figure 0007848769000002
    Figure 0007848769000002
  • Figure 0007848769000003
    Figure 0007848769000003
Patent Text Reader

Abstract

To provide a control device that can derive a grasping point of each work-piece at which each work-piece can be appropriately grasped.SOLUTION: A control device (5) according to one embodiment described in the disclosure comprises a reinforcement learning part (5b) that derives a gravity center point (G) of a work-piece (W) on the basis of detected information by a force sensor (3) provided in a robot arm (2), and learns a grasping point (P) of the work-piece (W) at which deviations of the grasping point from the gravity center point (G) of the work-piece (W) becomes minimum, using a deep reinforcement learning model in which deviations of the grasping point (P) of the work-piece (W) from the gravity center point (G) of the work-piece (W) and information on joints of a robot arm (2) are inputted and the grasping point (P) of the work-piece (W) is outputted. The reinforcement learning part (5b) re-learns the grasping point (P) of the work-piece (W), using the deep reinforcement learning model.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a control device and a control method, and more particularly, to a control device and a control method for a robotic arm that picks up and moves a workpiece.

Background Art

[0002] For example, the control device of Patent Document 1 calculates the center of gravity of a workpiece when the workpiece is picked up by a robotic arm based on detection information of a force sensor provided on the robotic arm, and determines the type of the workpiece by referring to the calculated center of gravity.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] The applicant of the present application has found the following problems. Although the control device of Patent Document 1 can discriminate the type of a workpiece, it has a problem that it cannot derive a gripping point that can pick up the workpiece well for each workpiece. }]

[0005] In view of such problems, the present disclosure provides a control device and a control method that can derive a gripping point that can pick up each workpiece well.

Means for Solving the Problems

[0006] A control device according to an aspect of the present disclosure is a control device for a robotic arm that picks up and moves a workpiece, A reinforcement learning unit learns the gripping point of the workpiece that minimizes the amount of deviation from the center of gravity of the workpiece when the robot arm picks up the workpiece, using a deep reinforcement learning model that takes the amount of deviation between the center of gravity of the workpiece and the gripping point of the workpiece when the robot arm picks up the workpiece and the joint information of the robot arm when the robot arm picks up the workpiece as inputs and the gripping point of the workpiece as output. A path generation unit that derives joint information of the robot arm so as to pick up the workpiece at the gripping point of the workpiece learned by the deep reinforcement learning model, A control unit that controls the robot arm based on the derived joint information of the robot arm, Equipped with, The reinforcement learning unit uses the deep reinforcement learning model to control the robot arm based on the derived joint information of the robot arm, and uses the amount of deviation between the center of gravity of the workpiece and the gripping point of the workpiece when the robot arm picks up the workpiece, along with the derived joint information of the robot arm, to relearn the gripping point of the workpiece.

[0007] The control device described above acquires information regarding the type of workpiece to be picked up by the robot arm, the order in which the workpieces are picked up, and information regarding the types of workpieces that have been learned. Based on the information regarding the type of workpiece to be picked up by the robot arm, the order in which the workpieces are picked up, and the types of workpieces that have been learned, the control device includes a determination unit that determines whether or not the workpiece to be picked up this time has been learned. In the present case, if the workpiece to be picked up has already been learned, it is preferable to omit the learning process using the deep reinforcement learning model, and if the workpiece to be picked up has not already been learned, it is preferable to perform the learning process using the deep reinforcement learning model.

[0008] In the control device described above, it is preferable that the joint information of the robot arm includes, in addition to the angle information of each joint of the robot arm, at least one of the angular velocity information, angular acceleration information, or torque information of the motor that drives the joint of the robot arm.

[0009] A control method relating to one aspect of this disclosure is a control method for a robot arm that picks up and moves a workpiece, A step of determining the center of gravity of the workpiece when the robot arm picks up the workpiece based on detection information from a force sensor provided on the robot arm, A deep reinforcement learning model takes the amount of deviation between the center of gravity of the workpiece and the gripping point of the workpiece when the robot arm picks up the workpiece as input, and the joint information of the robot arm when the robot arm picks up the workpiece as output, to learn the gripping point of the workpiece that minimizes the amount of deviation from the center of gravity of the workpiece. The process of deriving joint information of the robot arm so that the workpiece is picked up at the gripping point of the workpiece learned by the deep reinforcement learning model, A step of controlling the robot arm based on the joint information of the robot arm derived above, The process involves using the deep reinforcement learning model to control the robot arm based on the joint information of the robot arm derived above, to relearn the gripping point of the workpiece, using the amount of deviation between the center of gravity of the workpiece and the gripping point of the workpiece when the robot arm picks up the workpiece, and the derived joint information of the robot arm. It is equipped with.

[0010] The control method described above includes a step of acquiring information regarding the type of workpiece to be picked up by the robot arm, the order in which the workpieces are picked up, and information regarding the types of workpieces that have been learned, and determining whether the workpiece to be picked up this time has been learned based on the information regarding the type of workpiece to be picked up by the robot arm, the order in which the workpieces are picked up, and the types of workpieces that have been learned. In the present case, if the workpiece to be picked up has already been learned, it is preferable to omit the learning process using the deep reinforcement learning model, and if the workpiece to be picked up has not already been learned, it is preferable to perform the learning process using the deep reinforcement learning model. [Effects of the Invention]

[0011] According to this disclosure, a control device and control method can be realized that can determine a gripping point that allows for good gripping of each workpiece. [Brief explanation of the drawing]

[0012] [Figure 1] This is a block diagram showing the configuration of the robot arm system according to Embodiment 1. [Figure 2] This figure shows how the robot arm of the robot arm system of Embodiment 1 grips a workpiece. [Figure 3] This is a flowchart illustrating the flow of controlling the robot arm of the robot arm system of Embodiment 1. [Figure 4] This diagram shows the workpiece, picked up by the robot arm, stored in a storage box. [Figure 5] This is a block diagram showing the configuration of the robot arm system of Embodiment 2. [Figure 6] This is a flowchart illustrating the flow of controlling the robot arm of the robot arm system of Embodiment 2. [Modes for carrying out the invention]

[0013] The following describes specific embodiments applying this disclosure with reference to the drawings. However, this disclosure is not limited to the following embodiments. Also, the following description and drawings have been simplified as appropriate.

[0014] <Embodiment 1> First, the configuration of the robot arm system according to this embodiment will be described. FIG. 1 is a block diagram showing the configuration of the robot arm system according to this embodiment. FIG. 2 is a diagram showing a state in which the robot arm of the robot arm system according to this embodiment grips a workpiece.

[0015] As shown in FIG. 1, the robot arm system 1 according to this embodiment includes a robot arm 2, a force sensor 3, a camera 4, a control device 5, and a database (DB) 6. These robot arm 2, force sensor 3, camera 4, and DB 6 are communicably connected to the control device 5 by wire or wirelessly.

[0016] The robot arm 2 is an articulated robot arm similar to a general robot arm, and as shown in FIG. 2, includes a hand portion 2a and an arm portion 2b. The hand portion 2a can be configured by, for example, a two-finger hand. The arm portion 2b can be configured by, for example, a six-axis robot. Each joint portion of these hand portion 2a and arm portion 2b operates based on the driving force of the motor 2c.

[0017] As shown in FIG. 2, the force sensor 3 is disposed between the hand portion 2a and the arm portion 2b of the robot arm 2. That is, the hand portion 2a is connected to the arm portion 2b via the force sensor 3. The camera 4 can be configured by, for example, an RGBD camera.

[0018] The camera 4, for example, photographs the workpiece W before being gripped by the robot arm 2. The camera 4 is preferably provided on the robot arm 2, but may be disposed at a location where the workpiece W before being gripped by the robot arm 2 can be photographed.

[0019] As shown in FIG. 1, the control device 5 includes a shape recognition unit 5a, a reinforcement learning unit 5b, a path generation unit 5c, and a control unit 5d. The shape recognition unit 5a recognizes the shape of the workpiece W based on the image information photographed by the camera 4. A general method can be used as a method for recognizing the shape of the workpiece W.

[0020] As will be described in detail later, the reinforcement learning unit 5b derives the center of gravity G of the workpiece W when the robot arm 2 picks it up, based on the detection information from the force sensor 3, and learns the gripping point P of the workpiece W that minimizes the amount of deviation from the center of gravity G of the workpiece W using a deep reinforcement learning model.

[0021] The deep reinforcement learning model takes the amount of displacement between the center of gravity G of the workpiece W and the gripping point P of the workpiece W when the robot arm 2 picks up the workpiece W as input, and the joint information of the robot arm 2 when the robot arm 2 picks up the workpiece W as input, and outputs the gripping point P of the workpiece W. The center of gravity G, gripping point P, and displacement amount may be represented, for example, by three-dimensional coordinate information with a predetermined position as the origin.

[0022] The path generation unit 5c generates a path for the robot arm 2 to approach the gripping point P of the workpiece W, which is derived based on the shape of the recognized workpiece W, or the gripping point P of the workpiece W learned by the deep reinforcement learning model, and pick up the workpiece W. The unit also derives joint information for the robot arm 2, such as the angles of each joint, to realize the generated path.

[0023] Furthermore, the path generation unit 5c generates a path for moving the workpiece W from its gripping point P, which has been learned by a deep reinforcement learning model, to a predetermined location, and derives joint information for the robot arm 2, such as the angles of each joint, in order to realize the generated path.

[0024] The control unit 5d controls each motor 2c of the robot arm 2 based on the angle information of each joint of the robot arm 2 that it has derived. At this time, the control unit 5d may control the motor 2c based on the detection information of the encoder 2d provided on the motor 2c. DB6 stores the deep reinforcement learning model and result information such as whether the gripping of the workpiece W was successful or not for each gripping point P of the workpiece W.

[0025] Next, the control flow of the robot arm 2 of the robot arm system 1 of this embodiment will be explained. Figure 3 is a flowchart illustrating the control flow of the robot arm of the robot arm system of this embodiment.

[0026] Here, Figure 4 shows the state in which the workpiece W to be picked up by the robot arm 2 is stored in the storage box 11, and in the following description, we will assume that the workpiece W is being picked up. In the example of Figure 4, a plate-shaped workpiece W, which has an eccentric shape in side view as shown in Figure 2, is stored inside the storage box 11 and is propped up against a partition 11a provided in the storage box 11.

[0027] In this state, first, camera 4 photographs the workpiece W from above. Then, the shape recognition unit 5a of the control device 5 recognizes the shape of the workpiece W based on the image information of the workpiece W captured by camera 4 (S1). At this time, it is preferable for the shape recognition unit 5a to recognize the shape of the workpiece W as viewed from above. In the example in Figure 4, the shape recognition unit 5a recognizes that the top view of the workpiece W is a thin plate shape.

[0028] Next, the path generation unit 5c of the control device 5 derives the gripping point P of the workpiece W based on the shape of the recognized workpiece W (S2). For example, if the top view of the workpiece W is a thin plate shape, the path generation unit 5c sets the gripping point P to the center of the width dimension of the workpiece W (left-right direction in Figure 2), as shown in Figure 2. However, the path generation unit 5c may, for example, set the gripping point P to approximately the center of the top view of the workpiece W based on the shape of the recognized workpiece W.

[0029] Next, the path generation unit 5c of the control device 5 generates a path for the robot arm 2 to approach the gripping point P of the workpiece W and pick up the workpiece W, and derives joint information for the robot arm 2, such as the angles of each joint, in order to realize the generated path.

[0030] Then, the control unit 5d of the control device 5 controls each motor 2c of the robot arm 2 based on the angle information of each joint of the robot arm 2 that it has derived, to pick up the workpiece W (S3). Here, the workpiece W can be picked up by, for example, pulling the workpiece W up vertically to a predetermined height.

[0031] Next, the reinforcement learning unit 5b of the control device 5 derives the center of gravity G of the workpiece W when the robot arm 2 picks up the workpiece W, based on the detection information of the force sensor 3, and acquires the amount of deviation between the center of gravity G of the workpiece W and the gripping point P of the workpiece W when the robot arm 2 picks up the workpiece W, as well as the joint information of the robot arm 2 when the robot arm 2 picks up the workpiece W (S4).

[0032] Here, a general method can be used to determine the center of gravity G of the workpiece W using the force sensor 3. Furthermore, the amount of displacement between the center of gravity G of the workpiece W and the gripping point P of the workpiece W should be the amount of displacement in the horizontal direction.

[0033] Furthermore, the joint information of the robot arm 2 may include, in addition to the angle information of each joint when the robot arm 2 grasps the workpiece W in order to pick up the workpiece W, at least one of the angular velocity information, angular acceleration information, and torque information of the motor 2c.

[0034] Next, the reinforcement learning unit 5b of the control device 5 receives the amount of displacement between the center of gravity G of the workpiece W and the gripping point P of the workpiece W, which was obtained in step S4, and the joint information of the robot arm 2 when the robot arm 2 picked up the workpiece W, and learns the gripping point P of the workpiece W that minimizes the amount of displacement relative to the center of gravity G of the workpiece W (S5).

[0035] In this case, if the gripping point P of the workpiece W is in a localized area, it is conceivable that the robot arm 2 may not be able to grip the gripping point P of the workpiece W due to factors such as the shape of the workpiece W or interference with the storage box 11. Therefore, the reinforcement learning unit 5b of the control device 5 should distribute the gripping points P of the workpiece W within a predetermined range (in the example in Figure 2, a predetermined range in the width direction of the workpiece W with respect to the center of gravity G of the workpiece W).

[0036] Next, the path generation unit 5c of the control device 5 generates a path for the robot arm 2 to approach the learned gripping point P of the workpiece W and pick up the workpiece W, and derives joint information for the robot arm 2, such as the angles of each joint, to realize the generated path.

[0037] Then, the control unit 5d of the control device 5 controls each motor 2c of the robot arm 2 based on the angle information of each joint of the robot arm 2 that it has derived, in order to pick up the workpiece W (S6).

[0038] Next, the reinforcement learning unit 5b of the control device 5 acquires the center of gravity G of the workpiece W when the robot arm 2 picks up the workpiece W by grasping the gripping point P of the workpiece W that it has learned, based on the detection information of the force sensor 3. It also acquires the amount of deviation between the center of gravity G of the workpiece W and the gripping point P of the workpiece W when the robot arm 2 picks up the workpiece W, and the joint information of the robot arm 2 when the robot arm 2 picks up the workpiece W (S7).

[0039] Next, the reinforcement learning unit 5b of the control device 5 inputs the amount of deviation between the center of gravity G of the workpiece W and the gripping point P of the workpiece W, which was acquired in step S7, and the joint information of the robot arm 2 when the robot arm 2 picked up the workpiece W, and relearns the gripping point P of the workpiece W that minimizes the amount of deviation relative to the center of gravity G of the workpiece W (S8).

[0040] In this case, if the reinforcement learning unit 5b distributes the gripping points P of the workpiece W within a predetermined range relative to the centroid G of the workpiece W during step S5, it is advisable to input the result information of whether or not the gripping of the workpiece W was successful into the deep reinforcement learning model and retrain the gripping points P of the workpiece W that minimize the amount of displacement relative to the centroid G of the workpiece W.

[0041] This allows the workpiece W to be securely gripped at the gripping point P of the workpiece W, where the amount of displacement relative to the center of gravity G of the workpiece W is minimized. Therefore, even if the workpiece W is large, for example, the moment generated when the workpiece W is gripped can be reduced, and the workpiece W can be picked up smoothly.

[0042] Next, the path generation unit 5c of the control device 5 generates a path for the robot arm 2 to approach the gripping point P of the relearned workpiece W, pick up the workpiece W, and move it to a predetermined location. It then derives joint information for the robot arm 2, such as the angles of each joint, to realize the generated path.

[0043] Then, the control unit 5d of the control device 5 controls each motor 2c of the robot arm 2 based on the angle information of each joint of the robot arm 2 that it has derived, in order to pick up the workpiece W and move it to a predetermined location (S9).

[0044] Thus, the control device 5 and control method of this embodiment learn the gripping point P of the workpiece W that minimizes the amount of displacement relative to the centroid G of the workpiece W using a deep reinforcement learning model. Therefore, the control device 5 and control method of this embodiment can easily derive the gripping point P that allows for good gripping of each workpiece W.

[0045] Furthermore, there is no need for the operator to input the gripping point P of each workpiece W, and compared to the control device in Patent Document 1, it is possible to easily increase the number of workpieces W that can be gripped, and workpieces W can be picked up regardless of their shape or packaging.

[0046] <Embodiment 2> Figure 5 is a block diagram showing the configuration of the robot arm system of this embodiment. The robot arm system 21 of this embodiment is substantially the same as the robot arm system 1 of Embodiment 1, as shown in Figure 5.

[0047] Therefore, although we will omit redundant explanations, as shown in Figure 5, the control device 22 is equipped with a determination unit 22a, which determines whether the workpiece W to be picked up this time has already been learned, and based on the determination result, it determines whether or not to learn the gripping point P of the workpiece W.

[0048] DB23 stores information such as the type of workpiece W to be picked up by the robot arm 2 and the order in which the workpieces W are picked up, information about the types of workpieces W that have been learned, information about the gripping points P of the workpieces W that have been learned, and information about the joints of the workpieces W that have been learned.

[0049] Figure 6 is a flowchart illustrating the control flow of the robot arm of the robot arm system of this embodiment. The control flow of the robot arm 2 of the robot arm system 21 of this embodiment is substantially the same as the control flow of the robot arm 2 of the robot arm system 1 of Embodiment 1, as shown in Figure 6.

[0050] Therefore, although redundant explanations will be omitted, before step S1, the determination unit 22a of the control device 22 reads from DB23 information such as the type of workpiece W to be picked up by the robot arm 2 and the order in which the workpieces W are picked up, as well as information on the types of workpiece W that have been learned. Based on this information, the determination unit 22a determines whether the workpiece W to be picked up this time has been learned (S11).

[0051] If the workpiece W to be picked up is not already learned (NO in S11), steps S1 to S9 are executed to control the robot arm 2 to pick up and move the workpiece W. In other words, the control device 22 learns the gripping point of the workpiece W to be picked up this time.

[0052] On the other hand, if the workpiece W to be picked up this time has already been learned (YES in S11), the control unit 5d of the control device 22 reads the joint information of the workpiece W from DB23, and based on the read joint information of the workpiece W, controls each of the motors 2c of the robot arm 2 to attempt to grasp the workpiece W, and determines whether or not the grasping of the workpiece W was successful (S12).

[0053] If gripping of workpiece W fails (NO in S12), steps S1 to S9 are executed, and the robot arm 2 is controlled to pick up and move the workpiece W. In this case, even if workpiece W has been learned, the reason for the failure to grip it may be that the shape of the workpiece W has changed.

[0054] On the other hand, if the gripping of the workpiece W is successful (YES in S12), the robot arm 2 controls each of its motors 2c based on the read joint information of the workpiece W to continue moving the workpiece W (S13). In other words, the control device 22 omits learning the gripping point of the workpiece W to be picked up this time.

[0055] Thus, the control device 22 and control method of this embodiment also learn the gripping point P of the workpiece W that minimizes the amount of displacement relative to the centroid G of the workpiece W using a deep reinforcement learning model. Therefore, the control device 22 and control method of this embodiment can easily derive the gripping point P that can properly grip each workpiece W.

[0056] In particular, the control device 22 and control method of this embodiment determine whether the workpiece W to be picked up has already been learned, and based on the determination result, the learning by the deep reinforcement learning model is omitted. Therefore, it is not necessary to learn the gripping point P of the workpiece W each time a workpiece W is picked up, and the computational burden on the control device 22 can be reduced.

[0057] In this embodiment, the determination unit 22a of the control device 22 reads from DB23 information such as the type of workpiece W to be picked up by the robot arm 2 and the order in which the workpieces W are picked up, as well as information on the types of workpiece W that have been learned. Based on this information, it determines whether the workpiece W to be picked up this time has been learned or not. Alternatively, the determination may be made based on the shape of the workpiece W recognized from image information of the workpiece W to be picked up this time.

[0058] Furthermore, although step S12 is performed in this embodiment, step S12 may be omitted.

[0059] <Other Embodiments> In the embodiments described above, the disclosure was explained as a hardware configuration, but the disclosure is not limited thereto. The disclosure can also be implemented by having a CPU (Central Processing Unit) execute a program for each process.

[0060] Here, a program includes a set of instructions (or software code) that, when loaded into a computer, cause the computer to perform one or more functions. A program may be stored on a non-temporary computer-readable medium or a physical storage medium. Examples, but not limited to, include random-access memory (RAM), read-only memory (ROM), flash memory, solid-state drive (SSD) or other memory technologies, CD-ROM, digital versatile disc (DVD), Blu-ray® disc or other optical disc storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage devices. A program may be transmitted over a temporary computer-readable medium or a communication medium. Examples, but not limited to, include temporary computer-readable medium or a communication medium that includes electrically, optically, acoustically or otherwise propagating signals.

[0061] This disclosure is not limited to the embodiments described above, and may be modified as appropriate without departing from the spirit of the invention. [Explanation of symbols]

[0062] 1. Robotic Arm System 2 robot arm, 2a hand unit, 2b arm unit, 2c motor, 2d encoder 3. Force sensor 4 cameras 5 Control device, 5a Shape recognition unit, 5b Reinforcement learning unit, 5c Path generation unit, 5d Control unit 6 Databases 11 Storage Boxes 21 Robot Arm System 22 Control device, 22a Determination unit 23 Databases Center of gravity of G Work P is the gripping point of the workpiece. Double job

Claims

1. A control device for a robot arm that picks up and moves a workpiece, A reinforcement learning unit learns the gripping point of the workpiece that minimizes the amount of deviation from the center of gravity of the workpiece when the robot arm picks up the workpiece, using a deep reinforcement learning model that takes the amount of deviation between the center of gravity of the workpiece and the gripping point of the workpiece when the robot arm picks up the workpiece and the joint information of the robot arm when the robot arm picks up the workpiece as inputs and the gripping point of the workpiece as output. A path generation unit that derives joint information of the robot arm so as to pick up the workpiece at the gripping point of the workpiece learned by the deep reinforcement learning model, A control unit that controls the robot arm based on the derived joint information of the robot arm, Equipped with, The reinforcement learning unit is a control device that, using the deep reinforcement learning model, controls the robot arm based on the joint information of the robot arm derived to pick up the workpiece, and uses the amount of deviation between the center of gravity of the workpiece and the gripping point of the workpiece, along with the derived joint information of the robot arm, to relearn the gripping point of the workpiece.

2. The system includes a determination unit that acquires information regarding the type of workpiece to be picked up by the robot arm, the order in which the workpieces are picked up, and the types of workpieces that have been learned, and determines whether the workpiece to be picked up this time has been learned or not based on the information regarding the type of workpiece to be picked up by the robot arm, the order in which the workpieces are picked up, and the types of workpieces that have been learned. The control device according to claim 1, wherein if the workpiece to be picked up this time has already been learned, the learning by the deep reinforcement learning model is omitted, and if the workpiece to be picked up this time has not been learned, the learning by the deep reinforcement learning model is performed.

3. The control device according to claim 1 or 2, wherein, in addition to the angle information of each joint of the robot arm, the joint information of the robot arm includes at least one of angular velocity information, angular acceleration information, or torque information of a motor that drives the joint of the robot arm.

4. A method for controlling a robot arm that picks up and moves a workpiece, A step of determining the center of gravity of the workpiece when the robot arm picks up the workpiece based on detection information from a force sensor provided on the robot arm, A deep reinforcement learning model takes the amount of deviation between the center of gravity of the workpiece and the gripping point of the workpiece when the robot arm picks up the workpiece as input, and the joint information of the robot arm when the robot arm picks up the workpiece as output, to learn the gripping point of the workpiece that minimizes the amount of deviation from the center of gravity of the workpiece. The process of deriving joint information of the robot arm so that the workpiece is picked up at the gripping point of the workpiece learned by the deep reinforcement learning model, A step of controlling the robot arm based on the joint information of the robot arm derived above, The process involves using the deep reinforcement learning model to control the robot arm based on the joint information of the robot arm derived above, to relearn the gripping point of the workpiece, using the amount of deviation between the center of gravity of the workpiece and the gripping point of the workpiece when the robot arm picks up the workpiece, and the derived joint information of the robot arm. A control method comprising:

5. The system includes a step of acquiring information regarding the type of workpiece to be picked up by the robot arm, the order in which the workpieces are picked up, and the types of workpieces that have been learned, and determining whether the workpiece to be picked up this time has been learned based on the information regarding the type of workpiece to be picked up by the robot arm, the order in which the workpieces are picked up, and the types of workpieces that have been learned, The control method according to claim 4, wherein if the workpiece to be picked up this time has already been learned, the learning by the deep reinforcement learning model is omitted, and if the workpiece to be picked up this time has not been learned, the learning by the deep reinforcement learning model is performed.

Citation Information

Patent Citations

  • Control device of robot carrying work

    JP2014210311A

  • Machine learning device for learning taking-out operation of workpiece, robot system, and machine learning method

    JP2017064910A

  • Handling system and controller

    JP2018118343A

  • Controller and machine learning apparatus

    JP2019034836A

  • Holding device, control method, control program, control system and end effector

    JP2021171885A