Robot control system and picking method

JP7916837B2Active Publication Date: 2026-09-08TOYOTA JIDOSHA KK
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2023107283
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-06-29
Publication Date
2026-09-08
Estimated Expiration
2043-06-29

AI Technical Summary

Benefits of technology

【0007】 本開示によれば、ピッキングロボットにおいて、制御負荷を低減させつつロボットアームを精確に制御できる。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007916837000001
    Figure 0007916837000001
  • Figure 0007916837000002
    Figure 0007916837000002
  • Figure 0007916837000003
    Figure 0007916837000003
Patent Text Reader

Abstract

To provide a robot control system, which controls a robot arm accurately while reducing control loads in a picking robot.SOLUTION: A robot control system inputs data which include deviation-information showing amounts of deviations between a position (a first position) of a picking mechanism at the time of off-line control of a robot arm, which is based on an off-line control command value generated by off-line teaching and a position of the picking mechanism at the time of on-line control of the robot arm, the off-line control command value and information on characteristics of the robot arm, and executes deep reinforcement machine learning so that the amounts of deviations become minimum, so as to generate a learning control command value at the time of on-line control. The robot control system inputs, as portion of input data, updated deviation-information which is information showing amounts of deviations between the first position and a position (a third position) of the picking mechanism at the time of on-line control based on the learning control command value, and executes the deep reinforcement machine learning so that the amounts of deviations between the first position and the third position become minimum, so as to update the learning control command value.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a robot control system and a picking method.

Background Art

[0002] Patent Document 1 discloses a technique in which a robot hand for gripping an article is mounted on a robot arm, and a control value is generated by offline teaching through simulation for an arm robot that performs picking. In the technique described in Patent Document 1, the robot arm is actually controlled online based on the generated control value, and correction control is performed based on the actual position of the robot arm detected by a sensor.

Prior Art Literature

Patent Literature

[0003]

Patent Document 1

Summary of Invention

Problem to be Solved by Invention

[0004] However, in the technique described in Patent Document 1, it is necessary to perform correction control each time based on the actual position of the robot arm detected by a sensor, which increases the control load. Therefore, for picking robots, development of a technique that enables accurate control of the robot arm while reducing the control load is desired.

Means for Solving the Problem

[0005] The robot control system according to this disclosure, for a picking robot that picks items with a robot arm having a picking mechanism, includes an acquisition unit that acquires information indicating the amount of deviation between a first position, which is the position of the picking mechanism when the robot arm is controlled offline, and a second position, which is the position of the picking mechanism when the robot arm is actually controlled online based on the offline control command value, which is a control command value that controls the robot arm generated by offline teaching through simulation, and the offline control command value, the deviation information, and the operating characteristics of the robot arm. The system includes a generation unit that generates a learning control command value, which is a control command value for actually controlling the robot arm online, by performing deep reinforcement machine learning on data including characteristic information, which is input data, so as to minimize the amount of deviation between the first position and the second position. The acquisition unit acquires updated deviation information, which is information indicating the amount of deviation between the first position and a third position, which is the position of the picking mechanism when the robot arm is actually controlled online based on the learning control command value. The generation unit inputs the updated deviation information as part of the input data and updates the learning control command value by performing deep reinforcement machine learning so as to minimize the amount of deviation between the first position and the third position.

[0006] The picking method relating to this disclosure is a picking method in which the picking robot is controlled by the robot control system and the items are picked. [Effects of the Invention]

[0007] According to this disclosure, in a picking robot, the robot arm can be precisely controlled while reducing the control load. [Brief explanation of the drawing]

[0008] [Figure 1] This is a block diagram showing an example configuration of a robot control system according to an embodiment. [Figure 2] This figure illustrates the overview of the processing in the robot control system shown in Figure 1. [Figure 3] Figure 2 illustrates the discrepancy in the position of the picking mechanism between the offline picking robot and the actual picking robot. [Figure 4] Figure 1 is a flowchart illustrating an example of processing in the robot control system. [Figure 5] This is a flowchart illustrating another processing example in the robot control system shown in Figure 1. [Figure 6] This is a flowchart illustrating further examples of processing in the robot control system shown in Figure 1. [Modes for carrying out the invention]

[0009] The present invention will be described below through embodiments, but the claims are not limited to the following embodiments. Furthermore, not all of the configurations described in the embodiments are necessarily essential for solving the problem.

[0010] (Embodiment) First, an example of the robot control system according to this embodiment (hereinafter referred to as "this system") will be described using Figures 1 to 3. Figure 1 is a block diagram showing an example configuration of this system. Figure 2 is a diagram illustrating the overview of the processing in this system shown in Figure 1. Figure 3 is a diagram illustrating the positional difference of the picking mechanism between the offline picking robot and the actual picking robot in Figure 2.

[0011] As shown in Figure 1, the system 1 includes, for example, a picking robot 10 that picks up items with a robotic arm, and an information processing device 20. Hereinafter, the picking robot will be simply referred to as the robot, the robotic arm as the arm, and the items as workpieces. The robot 10 includes, for example, a control unit 11, a drive control unit 12, a drive unit 13, a sensor group 14, an input unit 15, and a storage unit 16. As illustrated in Figure 2, the robot 10 is a robot that picks up workpieces W placed on a table M or the ground with an arm 150. The arm 150 includes a picking mechanism 151 such as a gripping mechanism, and also includes multiple joint mechanisms that constitute one or more joint sections. Each joint mechanism is movable by rotation, extension, etc. Of course, the gripping mechanism can also include joint mechanisms. The robot 10 can pick up workpieces W with the picking mechanism 151 and place the workpieces W in any location.

[0012] Figure 2 shows an example of the appearance of a robot 10 in which the arm 150 rotates on 6 axes, but its shape, the number, position, shape, and type of joint mechanisms, and the shape and type of the picking mechanism 151 are not restricted. The picking mechanism 151 can be any gripping mechanism that grips the workpiece W, as shown in the example below. However, if the workpiece is metal, the picking mechanism 151 may be a mechanism that attracts the workpiece using the magnetic force of a permanent magnet or electromagnetic force, in which case the surface on which the picking mechanism 151 attracts the workpiece may be a flat surface.

[0013] The control unit 11 controls the drive control unit 12, the sensor group 14, the input unit 15, and the memory unit 16. The control unit 11 can be implemented by a computer consisting of, for example, a processor such as a CPU (Central Processing Unit), working memory, and a non-volatile storage device. The control program executed by the processor is stored in this storage device, and the processor can perform the functions of the control unit 11 by reading the control program from the working memory and executing it. This storage device can also utilize the memory unit 16. Of course, the control unit 11 may also be configured as a dedicated control circuit.

[0014] In the following explanation, we will use an example in which the control unit 11, input unit 15, and storage unit 16 are built into the main body of the robot 10, as shown in Figure 1. However, the control unit 11, input unit 15, and storage unit 16 may be provided in a controller such as a computer located at a distance from the main body of the robot 10. In that case, if the main body of the robot 10 and the controller can communicate via wired or wireless connection, the robot 10 can be controlled from the controller.

[0015] The drive control unit 12 controls the drive unit 13 according to the control unit 11. The drive unit 13 can be composed of actuators such as motors that drive the joint mechanism to perform actions such as relatively rotating multiple parts connected via the joint mechanism or extending and retracting parts. Figures 2 and 3 show an example in which the robot can rotate on the pivot axes of the six joint mechanisms, as indicated by arrows in the direction of rotation, and the picking mechanism 151 can grip the workpiece W with its gripping parts 152 and 153. One or more drive units 13 are provided for each joint mechanism, but for the sake of simplicity, an example in which only one drive unit 13 is provided for each joint mechanism is given. If the robot 10 is provided with mechanisms other than the joint mechanism, the drive unit 13 is also provided to drive those mechanisms.

[0016] The sensor group 14 consists of multiple sensors provided in each drive unit 13 or each joint mechanism. The sensor group 14 includes, for example, a sensor 14a such as a camera provided in or near the picking mechanism 151. The sensor 14a is provided to recognize the workpiece W.

[0017] The input unit 15 is the part that receives information from the information processing device 20 or the like. The storage unit 16 can be a storage device that can read and write information from the control unit 11, and can store the learning model 17, which will be described later. The control unit 11 may include a program for executing learning using the learning model 17 as part of the control program described above.

[0018] The information processing apparatus 20 includes, for example, an input unit 22, an output unit 23, a storage unit 24, and a control unit 21 that controls these components. The control unit 21 can be implemented by, for example, a computer including a processor, a working memory, and a non-volatile storage device. A control program executed by the processor is stored in this storage device, and the functions of the control unit 21 can be achieved when the processor reads the control program into the working memory and executes it. The storage unit 24 may also be used as this storage device.

[0019] The input unit 22 receives information necessary for simulation by a simulation program 25 described below from an external device, and stores the information in the storage unit 24 in a readable state during simulation. Hereinafter, the simulation program 25 is simply referred to as the program 25. The output unit 23 outputs result information indicating a simulation result to the robot 10 via a network. Of course, the result information may be output to a portable recording medium and input to the robot 10 via the portable recording medium. The storage unit 24 can store the program 25 and the like. The control unit 21 may preferably include, as part of the aforementioned control program, a program for executing simulation by the program 25.

[0020] The present system 1 uses deep reinforcement learning to suppress positional deviation of the picking mechanism 151 in the physical robot 10 after offline teaching. This will be specifically described with reference to FIG. 2 and FIG. 3. Hereinafter, models corresponding to the robot 10, the arm 150, and the picking mechanism 151 in the program 25 are described as the robot 25a, the arm 250, and the picking mechanism 251, respectively. Similarly, models corresponding to the gripping parts 152, 153, the sensor 14a, the table M, and the workpiece W in the program 25 are described as the gripping parts 252, 253, the sensor 254, the table Ms, and the workpiece Ws, respectively. The robot 25a is a model on the program 25 before manufacturing of the robot 10, but it may also be a model after manufacturing of the robot 10. Note that, in FIG. 2, for convenience, the learning model 17 and the control unit 11 are drawn outside the robot 10, but in actuality they are as shown in FIG. 1.

[0021] First, the control unit 21 of the information processing device 20 executes the program 25 for the robot 25a, thereby performing simulation by offline teaching, and generates an offline control command value that is a control command value for controlling the arm 250. The offline control command value is generated as an optimal control command value. It can be said that the control unit 21 includes a generation unit that generates the offline control command value. The offline control command value preferably includes control command values for driving a plurality of driving units 13 for driving each joint mechanism. The offline control command value actually refers to an offline control command value for specifying a target position where a certain work Ws exists and performing control to move toward the target position. The control unit 21 preferably generates a plurality of offline control command values in advance respectively for a plurality of conceivable target positions, that is, respectively for a plurality of works Ws installed at different positions.

[0022] Further, the control unit 21 also outputs a first position according to the program 25. The first position is the position of the picking mechanism 251 when the arm 250 is controlled offline based on the offline control command value, and is, for example, coordinates indicated by the midpoint H2 of the gripping parts 252 and 253 in FIG. 3. The position of the picking mechanism 251 is preferably, for example, a picking position when the arm 250 is controlled to pick the work Ws. More specifically, the position of the picking mechanism 251 may be a position before starting gripping when the distance or positional relationship with the work Ws becomes a predetermined distance or positional relationship, or may be a position at which the work Ws is gripped. The control unit 21 instructs the output unit 23 to output the offline control value as result information and the first position to the robot 10.

[0023] An example of how result information is calculated using program 25 is given. By reading program 25, the control unit 21 configures, for example, the following object recognition processing unit 211, gripping point search unit 212, and path generation unit 213. In other words, program 25 can include programs to enable the control unit 21 to implement the functions of each unit 211 to 213. The object recognition processing unit 211 processes the recognition of the workpiece Ws by the sensor 254, and if the workpiece Ws is recognized, it passes the position of the workpiece Ws to the gripping point search unit 212. Based on the received position of the workpiece Ws and the current position of the picking mechanism 251, the gripping point, which is the gripping position on the workpiece Ws, is searched for, and the searched gripping point is passed to the path generation unit 213. The gripping position on the workpiece Ws may be defined, for example, as the center of the position that contacts the picking mechanism 251 when gripped. The path generation unit 213 generates a movement path for the picking mechanism 251 from the gripping point and passes it to the gripping point search unit 212. The gripping point search unit 212 and the path generation unit 213 exchange information as described above and repeatedly generate movement paths to the searched gripping points to determine the optimal gripping point and movement path. The optimal gripping point and movement path are determined, for example, as a gripping point and movement path that can grip the workpiece Ws in the fastest possible time.

[0024] The path generation unit 213 generates offline control command values ​​indicating the drive angle, velocity, acceleration, and torque for multiple drive units corresponding to each joint mechanism of the arm 250, for moving the picking mechanism 251 to the determined gripping point along the determined movement path. Note that the drive angle for these drive units refers to the joint angle, and the drive velocity, acceleration, and torque refer to the velocity, acceleration, and torque that can move the joint, respectively. The same applies to the drive angle, etc., of the drive unit 13 described later. The offline control command values ​​may also be time-series control command values ​​from the current position to the gripping point. The path generation unit 213 instructs the output unit 23 to output the determined optimal gripping point and optimal offline control command value as the first position and offline control value, respectively, for each workpiece Ws position, i.e., each target position, to the robot 10. The path generation unit 213 should be instructed to output the result information including the first position and the offline control value, as well as the target position from which they were calculated, and the position of the picking mechanism 251 before it was driven, which is the current position of the picking mechanism 251.

[0025] An example configuration of the control unit 11 will be described. The control unit 11 includes, for example, an object recognition processing unit 211, a gripping point search unit 212, and a path generation unit 213, which perform processing corresponding to the object recognition processing unit 211, the gripping point search unit 212, and the path generation unit 213, respectively. The functions of each unit 111 to 113 can be realized, for example, by including a program that executes the processing of each unit 111 to 113 in the control program executed by the processor.

[0026] The object recognition processing unit 111 recognizes the workpiece W with the sensor 14a. The gripping point search unit 112 searches for a gripping point, which is the gripping position on the workpiece W, based on the position of the recognized workpiece W and the current position of the picking mechanism 151. The path generation unit 113 generates a movement path for the picking mechanism 151 from the searched gripping point. The gripping point search unit 112 and the path generation unit 113 exchange information and repeatedly generate movement paths to the searched gripping point, determining the optimal gripping point and movement path, for example, as the gripping point and movement path that can grip the workpiece W in the fastest possible time. The path generation unit 113 generates control command values ​​indicating the drive angle, speed, acceleration, and torque for each of the drive units 13 corresponding to each joint mechanism of the arm 150, in order to move the picking mechanism 151 to the determined gripping point along the determined movement path. The generated control command values, which are online control command values, may also be time-series control command values ​​from the current position to the gripping point. The path generation unit 113 sends an instruction to the drive control unit 12 to control the drive unit 13 of the arm 150 based on the online control command value generated for the workpiece W.

[0027] The control unit 11 of the robot 10 has this configuration and also acquires information indicating the amount of deviation between the first position and the second position. In other words, the robot 10 is equipped with an acquisition unit that acquires deviation information. Furthermore, the deviation information may be acquired as information indicating the distance or the amount of deviation indicating the distance and direction of the deviation between the first position and the second position, but it may also be acquired as information indicating the first position and information indicating the second position.

[0028] The first position is received by the input unit 15 as part of the result information. The control unit 11 may temporarily store the result information received by the input unit 15 from the information processing device 20 in the storage unit 16. The second position is the position of the picking mechanism 151 when the control unit 11 actually controls the arm 150 online based on the offline control command value received by the input unit 15 from the information processing device 20. The second position may be obtained as a sensor value detected by the sensor group 14. The second position is the position of the picking mechanism 151 of the robot 10, which corresponds to the picking mechanism 251 of the robot 25a, and is the coordinate shown by the midpoint H1 of the gripping parts 152 and 153 in Figure 3, for example. However, the midpoint H1 may deviate from the midpoint H2. However, in this embodiment, this deviation can be reduced by deep reinforcement machine learning, which will be described later.

[0029] This section describes an example of control based on offline control command values, as well as an example of detecting a second position. The object recognition processing unit 111 processes the recognition of the workpiece W by the sensor 14a. The gripping point search unit 112 searches for the gripping point, which is the gripping position on the workpiece W, based on the recognized position of the workpiece W and the current position of the picking mechanism 151, and passes the searched gripping point and the current position of the picking mechanism 151 to the path generation unit 113. The path generation unit 113 determines the optimal gripping point and movement path while exchanging information with the gripping point search unit 112, and searches for offline control command values ​​corresponding to the gripping point and the current position of the picking mechanism 151 from the offline control command values ​​received from the information processing device 20 and stored in the storage unit 16. Based on the offline control command values ​​found, the path generation unit 113 sends an instruction to the drive control unit 12 to control the drive unit 13 of the arm 150. The second position is detected by the sensor group 14 as the position of the picking mechanism 151 after the drive unit 13 is driven in accordance with this instruction. Similarly, for the positions of other target workpieces W, i.e., other target positions, the drive unit 13 is controlled using offline control command values ​​selected based on the current position of the picking mechanism 151, and the second position is detected by the sensor group 14. The control unit 11 acquires the second positions for multiple target positions from the sensor group 14 in this manner.

[0030] Next, the control unit 11 inputs the data, including the offline control command values, deviation information, and characteristic information described later, acquired as described above, into the learning model 17 as input data. From this input data, the control unit 11 performs deep reinforcement machine learning in the learning model 17 so as to minimize the amount of deviation between the first position and the second position.

[0031] The control unit 11 generates control command values ​​to actually control the arm 150 online using this deep reinforcement machine learning. These generated control command values ​​are called learned control command values. The learning model 17 can be a model that performs reinforcement learning using DQN (Deep Q-Network), but any deep reinforcement learning model is acceptable, regardless of its algorithm or number of layers. The learned control command values ​​may include control command values ​​to drive multiple drive units 13 for driving each joint mechanism. The control unit 11 can be said to have a generation unit that generates such learned control command values ​​using the learning model 17. The learned control command values ​​may also be generated for multiple target positions using the same approach as the offline control command values. In other words, the control unit 11 may generate multiple learned control command values ​​for each of multiple workpieces W with different installation positions.

[0032] The characteristic information input to the learning model 17 as part of the input data is information indicating the operational characteristics of the arm 150. The characteristic information can be, for example, information indicating the motor characteristics of the motor provided as a drive unit 13 on a pivot axis, the backlash of the gears included in the joint mechanism or drive unit 13, and the rigidity of the arm 150. If the robot 10 is equipped with actuators other than motors, it is preferable to input information indicating the actuator characteristics instead of the motor characteristics as characteristic information. The rigidity of the arm 150 may be, for example, the rigidity of the entire arm 150, or the rigidity of each arm member that makes up the arm 150.

[0033] Furthermore, the characteristic information should at least partially consist of information indicating the current operating characteristics. Specifically, at least a portion of the characteristic information should be sensor values ​​detected when the arm 150 is actually controlled online by one or more sensors equipped on the arm 150 that detect operating characteristics. The one or more sensors mentioned above are included in the sensor group 14. In this way, by including the sensor values, which are the result of detecting characteristic information by the sensor group 14, in the input data, it is possible to generate learned control command values ​​that reflect the current characteristics of the arm 150, taking into account changes in the robot 10 over time.

[0034] Furthermore, the input data may include sensor values ​​detected when the arm 150 is actually controlled online by one or more sensors equipped on the arm 150 that detect the operating state. The one or more sensors mentioned above are included in the sensor group 14. The sensor group 14 detects information indicating the angle, velocity, acceleration, and torque of the joint mechanism as, for example, the operating state. In this way, by including the sensor values, which are the result of detecting the operating state by the sensor group 14, in the input data, it is possible to generate a learning control command value that reflects the operating state.

[0035] Furthermore, the control unit 11 acquires updated shift information, which is information indicating the amount of deviation between the first position and the third position. The third position is the position of the picking mechanism 151 when the arm 150 is actually controlled online based on the learned control command value. The control unit 11 is equipped with an acquisition unit that acquires such updated shift information for input into the learning model 17. The updated shift information may be acquired as information indicating the result of calculating the amount of deviation, which is the distance or distance and direction of the deviation between the first position and the third position, but since the first information has already been input into the learning model 17, the updated shift information may be acquired as information indicating the third position.

[0036] The detection process for the third position is basically performed using the same procedure as the detection process for the second position. However, the third position is detected as the position of the picking mechanism 151 by the sensor group 14 after the drive unit 13 is controlled based on a learned control command value instead of an offline control command value.

[0037] Furthermore, the control unit 11 further inputs the update misalignment information as part of the input data and updates the learning control command value by performing deep reinforcement machine learning so that the amount of misalignment between the first position and the third position is minimized. For this retraining, the learning control command value obtained during training may also be input as part of the input data. The control unit 11 can be said to include a generation unit that generates such updated learning control command values ​​using the learning model 17.

[0038] Next, the robot control method according to this embodiment will be described using Figure 4. Figure 4 is a flowchart illustrating an example of processing in the system shown in Figure 1. This robot control method includes, for example, a first generation step, a first acquisition step, a second generation step, a second acquisition step, and an update step. The first generation step is a step performed by the information processing device 20, and the other steps are steps performed by the robot 10.

[0039] The first generation step is to generate offline control command values ​​for controlling the arm 250 of the robot 25a by offline teaching through simulation. Specifically, the control unit 21 of the information processing device 20 performs offline teaching through simulation on the robot 25a equipped with the arm 250 (step S1). Through this teaching, the control unit 21 generates offline control command values ​​indicating the angle, velocity, acceleration, and torque for controlling the drive unit of the arm 250, and transmits them to the robot 10, which then acquires these offline control command values ​​(step S2). For example, the robot 10 acquires an offline control command value Ia for one drive unit of the arm 250 for input to the learning model 17, and similarly acquires offline control command values ​​for other drive units.

[0040] The first acquisition step involves acquiring deviation information indicating the amount of deviation between the first position and the second position. Specifically, the control unit 11 acquires deviation information indicating the amount of deviation between when it operates itself as an actual machine using the offline control command value and when it operates offline (step S3). For example, the control unit 11 acquires the first position Ib for a certain drive unit of the arm 250 for input to the learning model 17, and similarly acquires the first position for the other drive units. Furthermore, the control unit 11 acquires the second position of the drive unit 13 corresponding to a certain drive unit of the arm 250 when the drive unit 13 is controlled using the offline control command value, using the sensor group 14, and similarly acquires the second position for the other drive units 13.

[0041] Furthermore, in step S3, the control unit 11 uses the sensor group 14 to detect the angle, velocity, acceleration, and torque of each drive unit 13 corresponding to each joint mechanism when the drive unit 13 is controlled by the offline control command value, and acquires joint information which is the result of these detections.

[0042] The second generation step uses data including offline control command values, deviation information, and characteristic information as input data to perform deep reinforcement machine learning so that the amount of deviation between the first position and the second position is minimized. This learning generates a learned control command value that actually controls the arm 150 online. Specifically, the control unit 11 acquires characteristic information from the sensor group 14, such as the motor characteristics of the motors of each rotation axis, the backlash of the gears, and the rigidity of each arm member of the arm 150 (step S4). Characteristic information and joint information may be acquired simultaneously. Next, the control unit 11 inputs all the acquired data into the learning model 17, which is a deep reinforcement learning model, and performs machine learning (step S5). This machine learning learns the amount of deviation and its cause, and generates a learned control command value that minimizes the amount of deviation.

[0043] The second acquisition step involves acquiring updated deviation information that shows the amount of deviation between the first position and the third position. Specifically, the control unit 11 controls the robot 10, including the arm 150, based on the learning control command value generated in step S5 (step S6), and acquires updated deviation information showing the amount of deviation from the offline position and information for each joint when the robot 10 operates under that control (step S7). For example, the control unit 11 acquires the third position Id as updated deviation information for input to the learning model 17. The control unit 11 may also further store the learning control command value Ic for one of the drive units 13 of the arm 150 in the storage unit 16 for input to the learning model 17, and similarly store the learning control command values ​​for other drive units 13. Information for each joint is acquired in the same manner as in step S3.

[0044] The update step updates the learning control command value by inputting the update misalignment information as part of the input data and performing deep reinforcement machine learning to minimize the misalignment between the first position and the third position. Specifically, the control unit 11 inputs all data, including the update misalignment information and joint information acquired in step S7, into the learning model 17 and performs retraining (step S8). All data subject to retraining may include data already input in step S5. Through this machine learning, the amount of misalignment and its cause are learned again, and an updated learning control command value Oa is generated that minimizes the update misalignment. Then, the control unit 11 uses the updated learning control command value Oa generated in step S8, which corresponds to the optimal gripping point for the target workpiece W and the current position of the picking mechanism 151, as the online control command value to control and operate the robot 10 including the arm 150 (step S9), and terminates the process.

[0045] As described above, according to this embodiment, the discrepancy between the results obtained during offline teaching and the results obtained on the actual robot can be suppressed in the robot 10, that is, the arm 150 can be controlled accurately, thereby reducing the amount of manual adjustment work required at the installation location of the robot 10. Furthermore, according to this embodiment, since the drive control of the drive unit 13 is based on data using the updated learned control command values, there is no need to perform dynamic correction control for the amount of discrepancy, as in rule-based drive control, thus reducing the control load on the robot 10.

[0046] Furthermore, by controlling the robot 10 using the robot control method described above and employing a picking method that picks the workpiece W, that is, by picking the workpiece W with the robot 10 as described above, the workpiece W can be accurately picked even if there is deterioration due to aging or other factors.

[0047] As a comparative example, it is conceivable to use online teaching only, without using offline teaching. However, in this comparative example, teaching can only be performed after a real robot like robot 10 is completed and the operating environment of the robot, including the target workpiece W and the arrangement of each workpiece W, is in place. Therefore, in this comparative example, online teaching is required every time the operating environment changes, resulting in a longer lead time, increased man-hours, and higher costs. Furthermore, with the technology of the comparative example, online teaching is required every time the condition of the robot's hardware changes due to aging or other factors, for example, every time characteristic information such as the motor characteristics of each axis, gear backlash, and arm rigidity changes. In contrast, this embodiment can resolve the problems described in these comparative examples.

[0048] It should be noted that the present invention is not limited to the embodiments described above, and can be modified as appropriate without departing from the spirit of the invention. For example, if the influence of the actuators and gear backlash of each pivot axis is small, and the stiffness of each arm member constituting the arm 150 is the main cause of the misalignment, the following processing may be performed. That is, the characteristic information input to the learning model 17 may include, for example, information indicating the stiffness of the arm 150, but may not include information indicating the motor characteristics or gear backlash. In this case, machine learning will learn that the stiffness of each arm member is the cause of the misalignment, and updated learning control command values ​​that take into account the stiffness of each arm member will be generated so as to minimize the misalignment. On the other hand, if each arm member constituting the arm 150 is highly rigid, but the motor and gears are inexpensive and there is a large variation in motor characteristics and gear backlash, the following processing may be performed. That is, the characteristic information input to the learning model 17 may include, for example, information indicating the motor characteristics or gear backlash, but may not include information indicating the stiffness of the arm 150. In this case, machine learning will learn that motor characteristics and gear backlash are the causes of the deviation, and updated learning control command values ​​will be generated that take into account motor attributes and gear backlash to minimize the deviation.

[0049] Furthermore, relearning is not limited to a one-time process. Figures 5 and 6 are flowcharts illustrating other processing examples in the system 1 shown in Figure 1. As shown in Figure 5, after processing in step S9, the control unit 11 determines whether the arm 150 has been operated for a predetermined period or a predetermined number of times (step S10). If the result is YES, it returns to step S4 and performs relearning based on the current characteristic information. This allows the learning control command value to be updated each time the predetermined period or number of operations are performed. Also, as shown in Figure 6, after processing in step S9, the control unit 11 determines whether a picking error has occurred (step S11). If the result is YES, it returns to step S4 and performs relearning based on the current characteristic information. This allows the learning control command value to be updated when a picking error occurs. The occurrence of a picking error can be detected by sensors such as a camera, which detect that the workpiece W could not be picked, or that the picking mechanism 151 has come into contact with a workpiece W other than the intended picking target. Alternatively, the system may incorporate the judgments of both steps S10 and S11, so that it returns to step S4 when either event occurs. [Explanation of symbols]

[0050] 1 Robot control system, 10 Picking robot, 11, 21 Control unit, 12 Drive control unit, 13 Drive unit, 14 Sensor group, 15, 22 Input unit, 16, 24 Memory unit, 17 Learning model, 23 Output unit, 25 Simulation program

Claims

1. A picking robot that picks items with a robot arm having a picking mechanism, and an acquisition unit that acquires deviation information indicating the amount of deviation between a first position, which is the position of the picking mechanism when the robot arm is controlled offline, and a second position, which is the position of the picking mechanism when the robot arm is actually controlled online based on the offline control command value, which is a control command value that controls the robot arm generated by offline teaching through simulation, and A generation unit generates a learned control command value, which is a control command value for actually controlling the robot arm online, by performing deep reinforcement machine learning on data including the offline control command value, the deviation information, and characteristic information indicating the operating characteristics of the robot arm as input data, so as to minimize the amount of deviation between the first position and the second position. Equipped with, At least a portion of the characteristic information is a sensor value detected by a sensor provided on the robot arm to detect the motion characteristics when the robot arm is actually controlled online. The acquisition unit acquires updated deviation information, which is information indicating the amount of deviation between the first position and the third position, which is the position of the picking mechanism when the robot arm is actually controlled online based on the learned control command value. The generation unit inputs the update misalignment information as part of the input data and updates the learning control command value by performing deep reinforcement machine learning so that the amount of misalignment between the first position and the third position is minimized. Robot control system.

2. The generating unit is If the stiffness of each of the multiple arm members constituting the robot arm is the main cause of the misalignment, the characteristic information should include information indicating the stiffness, but not information indicating the motor characteristics of the motor included in the drive unit that drives the joint mechanism of the robot arm, and information indicating the backlash of the gears included in the joint mechanism or the drive unit. If there is a large variation in the motor characteristics of the motor or the backlash of the gear, the characteristic information will include information indicating the motor characteristics of the motor and information indicating the backlash of the gear, but will not include information indicating the rigidity. The robot control system according to claim 1.

3. The aforementioned input data includes sensor values ​​detected by a sensor that detects the operating state of the robot arm when the robot arm is actually controlled online. The robot control system according to claim 1 or 2.

4. The generation unit updates the learned control command value when a picking error occurs, or each time an operation is performed for a predetermined period or a predetermined number of times. The robot control system according to claim 1 or 2.

5. A picking method for picking an article by controlling the picking robot with the robot control system described in claim 1 or 2.

Citation Information

Patent Citations

  • Assembly robot

    JP1997319420A

  • Machine learning device for performing learning by use of simulation result, mechanical system, manufacturing system and machine learning method

    JP2017185577A

  • Control system and control method

    JP2021000678A

  • Operation control program, operation control method, and operation control device

    JP2022077228A

  • Warehouse System

    JP6905147B2