Route determination system and method

The path determination system optimizes robot arm movements using a camera, sensor, and deep reinforcement learning to minimize joint angle changes, ensuring safe and intimidation-free operations near operators.

JP7848759B2Active Publication Date: 2026-04-21TOYOTA JIDOSHA KK
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
TOYOTA JIDOSHA KK
Filing Date
2023-06-27
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Picking robots near operators can perform large movements that intimidate or intimidate surrounding individuals, even without physical contact, due to unpredictable joint angle changes.

Method used

A path determination system using a camera, sensor, and deep reinforcement learning model to minimize joint angle changes and ensure safe, intimidation-free movements by optimizing robot arm paths.

Benefits of technology

The system suppresses large, intimidating movements of robot arms by generating paths that reduce joint angle changes and maintain safe distances from individuals and obstacles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007848759000001
    Figure 0007848759000001
  • Figure 0007848759000002
    Figure 0007848759000002
  • Figure 0007848759000003
    Figure 0007848759000003
Patent Text Reader

Abstract

To provide a route determination system and a method which suppresses large movements of a robot arm through machine learning.SOLUTION: A route determination system 1 comprises: a camera 17 which is mounted on a robot arm 10 to image an object; a sensor which is mounted on a joint part of the robot arm to acquire angle data of the joint part; a robot controller 100 which controls behavior of the robot arm 10; and a deep enhancement learning model 20 which optimizes the behavior of the robot arm. The robot controller recognizes the object based on an image imaged by the camera, finds a gripping point of the recognized object, and creates a route where a gripping part at a tip of the robot arm is moved to the found gripping point. The deep enhancement learning model uses the angle of the prescribed joint part of the robot arm moving along the created route as input to determine the route of the robot arm to minimize variation of the angle of the prescribe joint part as much as possible.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to a path determination system and method.

Background Art

[0002] Patent Document 1 discloses a path generation device that generates a plurality of gripping poses for gripping the work of a picking robot and generates an operation path of the picking robot based on the gripping poses that satisfy predetermined conditions.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] When operating a picking robot near an operator, the picking robot may perform large movements that the operator cannot anticipate. Even if it does not come into contact with the operator or surrounding objects, such large movements may give a sense of intimidation.

[0005] This disclosure has been made to solve such problems, and an object thereof is to provide a path determination system and a path determination method that suppress large movements of a robot arm.

Means for Solving the Problems

[0006] A path determination system according to one aspect of this disclosure is a camera mounted on a robot arm for photographing an object, and a sensor mounted on a joint portion of the robot arm for acquiring angle data of the joint portion, and a robot controller for controlling the operation of the robot arm, and A deep reinforcement learning model that optimizes the operation of the robot arm, Equipped with, The aforementioned robot controller is The camera captures an image and recognizes an object. Search for the gripping point of the recognized object, A path is generated to move the gripping part at the tip of the robot arm to the searched gripping point. The deep reinforcement learning model is configured to take the angles of predetermined joints of the robot arm as input when moving along the generated path, learn to minimize the change in the angles of the predetermined joints, and determine the path of the robot arm based on the learned deep reinforcement learning model.

[0007] A method for determining a route according to one aspect of this disclosure is: A camera mounted on a robotic arm to photograph objects, A sensor mounted on the joint of the robot arm, which acquires angle data of the joint, A robot controller that controls the movement of the robot arm, A deep reinforcement learning model that optimizes the operation of the robot arm, A route determination method using a route determination system equipped with, The camera recognizes an object from the image it captures, Search for the gripping point of the recognized object, A path is generated to move the gripping part at the tip of the robot arm to the searched gripping point. The paths of the robot arm are determined by inputting the angles of predetermined joints of the robot arm as it moves along the generated path, and based on the deep reinforcement learning model, which has been trained by deep reinforcement learning to minimize the amount of change in the angles of the predetermined joints. [Effects of the Invention]

[0008] This disclosure provides a path determination system and method that suppresses large movements of a robot arm. [Brief explanation of the drawing]

[0009] [Figure 1] This is a diagram illustrating an example of a robotic arm configuration. [Figure 2] This is a conceptual diagram showing the configuration of the route determination system. [Figure 3] This diagram illustrates the posture of the robot arm at each stage of a pick-and-place operation. [Figure 4] This is a flowchart illustrating a route determination method for several embodiments. [Figure 5] This is a flowchart illustrating a route determination method according to another embodiment. [Figure 6] This diagram illustrates an example of a hardware configuration for a route determination system. [Modes for carrying out the invention]

[0010] Specific embodiments to which the present invention is applied will be described in detail below with reference to the drawings. However, the present invention is not limited to the following embodiments. Also, for clarity of explanation, the following description and drawings have been simplified as appropriate.

[0011] The robot arm 10 shown in Figure 1 is also called a multi-joint robot and comprises a rotatable base 150, multiple links 130 and joints 13a to 13f that rotatably connect each link 130, and a hand 15 for gripping a workpiece. In this example, it is a 6-axis robot. The hand 15 for picking up the workpiece, which is the object to be gripped, is connected to the tip of the robot arm 10. The hand 15 has at least two gripping parts 16 (also called claws). By gripping an appropriate part of the workpiece with at least two gripping parts 16 of the hand 15, the robot arm 10 can lift the workpiece and transport it to any desired location. The robot arm is sometimes called a picking robot. Note that the configuration of the robot arm is not limited to that shown in the figure, and various robot arms can be used.

[0012] Each joint part is provided with a rotation sensor such as an encoder for detecting the rotation information of each joint part, an actuator such as a servo motor for driving each joint part, and a force sensor for detecting the operating force of each joint part. The force sensor is, for example, a torque sensor for detecting the torque of each joint part. Each joint part is provided with a reduction mechanism or the like.

[0013] The robot arm 10 is installed in a warehouse, a factory, or the like, and recognizes a corresponding part from a large number of parts constituting, for example, a vehicle or the like from a captured image captured by a sensor 17 such as a camera or LiDAR (Light Detection And Ranging), picks up the part or the like, and places the part at an appropriate location (for example, inside a box or the like). The sensor 17 can also acquire three-dimensional point cloud data. The sensor 17 is provided, for example, near the hand 15 so as to be able to capture the gripping object. The robot arm 10 supports a part of the work of an operator working in the surroundings. Sensors 19 are provided at appropriate locations (for example, each joint part) of the robot arm 10 so that the picking and placing operations do not contact surrounding people, carts, obstacles, or the like. The sensors 19 can include a human sensor, LiDAR, a distance sensor, or the like. When the distance between the sensor 19 and surrounding people, carts, obstacles, or the like becomes less than a threshold value, an alarm can be issued or an emergency stop process can be performed.

[0014] The robot controller 100 of the robot arm 10 is composed of a processor, a memory, etc., processes the captured image, recognizes the object in the image, automatically determines the path to the automatically recognized object and the path to the destination, and automatically controls the pick-and-place operation. At this time, when the robot arm 10 performs a large pick-and-place operation, it may give a sense of intimidation to the people working around. As a result of intensive studies, the inventor discovered that when the amount of change in the rotation angle of the joint parts excluding the wrist increases, it will give a sense of intimidation to the people working around. Since there are a large number of parts constituting the vehicle, a method for automatically determining the optimal path to an appropriate location (for example, each box) for each type of part is required.

[0015] FIG. 2 is a conceptual diagram showing the configuration of the path determination system. The path determination system 1 includes the robot arm 10, the deep reinforcement learning model 20, and the robot controller 100 shown in FIG. 1. Each component is communicably connected to each other via a wired or wireless network.

[0016] The robot controller 100 controls the operation of the robot arm 10. The robot controller 100 controls the actuator of each joint part based on, for example, the rotation information (for example, rotation angle, etc.) from the rotation sensor of each joint part and the operating force (for example, speed, acceleration, etc.) from the force sensor, thereby performing feedback control of the robot arm. Also, the robot controller 100 can grip the workpiece by bringing at least two gripping parts 16, 16 of the hand 15 close to each other. The maximum distance between at least two gripping parts 16, 16 of the hand 15 is arbitrarily determined. In some embodiments, the hand 15 may have three or more gripping parts (claws).

[0017] As shown in Figure 2, the robot controller 100 includes an object recognition unit 101, a gripping point search unit 102, and a path generation unit 103. The object recognition unit 101 uses known object recognition techniques to recognize one part from among many parts that make up a vehicle or the like from an image captured by a sensor 17 such as a camera. The gripping point search unit 102 searches for a gripping point to be grasped on the recognized part. For example, in the case of a rectangular part, the gripping point search unit 102 may determine the center of the rectangle of the part as the gripping point. The path generation unit 103 generates one or more paths from the hand to the gripping point and determines one of them. After that, the path generation unit 103 also generates a path to an appropriate location (for example, a box for each part) for each recognized part. The robot controller 100 transmits control signals such as the rotation angle, velocity, and acceleration of each joint to the robot arm 10 in order to move the robot arm 10 along the determined path.

[0018] The deep reinforcement learning model 20 can be used to optimize the movement of a robot arm. The deep reinforcement learning model 20 has an input layer 201, a hidden layer 202, and an output layer 203. The input layer 201 of the deep reinforcement learning model 20 is input with acquired rotation angle data. In some embodiments, the rotation angles (and changes) of the rotation axes (1 to 5 axes), excluding the 6 axes of the wrist (i.e., the rotation axes of the end of the robot arm), may be input. In particular, in other embodiments, the rotation angles (and changes) of intermediate rotation axes such as the 3rd and 4th axes, excluding the 6 axes of the wrist and the 1st axis of the base, may be input. This is because these intermediate rotation axes have a significant impact on large pick-and-place movements. The input layer 201 is also called the rotation angle data acquisition unit. In this example, a Deep Q-Network is used as the deep reinforcement learning model, but it is not limited to this, and various neural network models or reinforcement learning models that are understandable to those skilled in the art can be used. Between each layer, there are synapses (not shown) connecting each neuron, and a set of weight parameters consisting of the weights of each synapse can be adjusted by machine learning.

[0019] The deep reinforcement learning model is trained to generate a path that minimizes the change in joint angles in order to reduce the pick-and-place motion as much as possible. For example, in each of the three actions shown in Figure 3—(a) picking, (b) action along the path, and (c) action during placement—the rotation angle of each joint is obtained, and the deep reinforcement learning model is trained to generate a path that minimizes the change in joint angles. For example, by performing deep reinforcement learning using the change in joint angles when a first action path is executed as input information, a second action path with an even smaller change in the joint angles can be output. In other embodiments, deep reinforcement learning may be performed to generate a path that minimizes the change in joint angles for the entire process from (a) to (c).

[0020] Furthermore, in some embodiments, the deep reinforcement learning model is subjected to deep reinforcement learning to generate paths such that the distance to people, moving objects, obstacles, etc., acquired by sensors 19 mounted on the robot arm 10 is greater than or equal to a threshold. For example, the threshold distance between the robot arm 10 and a person can be set to a sufficiently long distance so as not to intimidate people working in the vicinity.

[0021] In other embodiments, if the sensor 19 detects that no worker is present near the robot arm 10, any path may be output regardless of the amount of change in the angle of a predetermined joint. On the other hand, if the sensor 19 detects that a worker is present near the robot arm 10, the system may be configured to determine a path for the robot arm that minimizes the amount of change in the angle of a predetermined joint and ensures that the distance to the detected worker is greater than or equal to a threshold.

[0022] Figure 4 is a flowchart illustrating the route determination method. Step S41: Data (e.g., angle, velocity, acceleration) is acquired from each joint of a 6-axis picking robot equipped with sensors that detect surrounding obstacles (e.g., motion sensors, distance sensors). The acquired data (e.g., angle, velocity, acceleration) for each joint is input to the input layer of a deep reinforcement learning model (e.g., Deep QNetwork) (Step S42). Deep reinforcement learning is performed to minimize the change in angle of each joint (e.g., axes 2-5, preferably axes 3 and 4, especially axis 4), except for the wrist area near the hand (e.g., axis 6, 13f in Figure 1), which is less likely to lead to large movements that may intimidate a person, and to avoid contact with obstacles (e.g., a person) (Step S43). In some embodiments, the deep reinforcement learning model may be trained so that the distance between the robot arm (e.g., joint) and the obstacle is greater than a threshold. Based on the deep reinforcement learning model thus trained, a path that does not intimidate a person can be determined. The robot controller assigns the determined path to the picking robot (step S44). It acquires joint data for each movement along the assigned path and uses this data to retrain the deep reinforcement learning model (step S45). Based on the results, it gives instructions to the picking robot (step S46).

[0023] Figure 5 is a flowchart illustrating a route determination method according to another embodiment. Step S51: Data (e.g., angle, velocity, acceleration) is acquired from each joint of a 6-axis picking robot equipped with sensors to detect surrounding obstacles (e.g., motion sensors, distance sensors). The acquired data for each joint (e.g., angle, velocity, acceleration) is input to the input layer of a deep reinforcement learning model (e.g., Deep QNetwork) (Step S52). Deep reinforcement learning is performed to minimize the change in angle of each joint (e.g., axes 2-5, preferably axes 3 and 4, especially axis 4), except for the wrist area near the hand (e.g., axis 6, 13f in Figure 1), which is less likely to lead to large movements that may intimidate people, and to avoid contact with obstacles (e.g., people) (Step S53). In some embodiments, the deep reinforcement learning model may be trained so that the distance between the robot arm (e.g., joint) and the obstacle is greater than a threshold. Alternatively, the weights for axes 3 and 4 may be given greater weight than the other axes during training. In some cases, better learning results can be obtained by inputting velocity and acceleration in addition to angle as input data.

[0024] Here, sensors mounted on the picking robot that detect surrounding obstacles (e.g., motion sensors, distance sensors) determine whether or not there are people nearby (step S54). If the sensors determine that a person is near the picking robot (YES in step S54), the path determination system 1 uses a deep reinforcement learning model to determine a path that will not intimidate the person and assigns it to the picking robot (step S55). The deep reinforcement learning model acquires information on each joint moving along the assigned path (e.g., angle, velocity, acceleration) and retrains based on the acquired information (step S56). The path determination system 1 gives instructions to the picking robot based on the results of the retraining (step S57).

[0025] On the other hand, if the sensor determines that no person is near the picking robot (NO in step S54), the path determination system 1 refrains from selecting a path determined by the deep reinforcement learning model described above, and instead provides the robot with optimal path information (step S59). Here, optimal path information for the robot may be, for example, the path that minimizes the time required to execute the task. In other words, it may involve performing large pick-and-place operations.

[0026] As described above, the path determination system according to this embodiment can, in particular, allow a robot to perform pick-and-place operations along a path that does not intimidate a person, rather than a path that is optimal for the robot, especially when a person and a robot need to work in close proximity.

[0027] Figure 6 is a block diagram showing an example configuration of a routing system. Referring to Figure 6, the routing system (and robot controller, etc.) includes a network interface 1201, a processor 1202, and memory 1203. The network interface 1201 is used to communicate with other network node devices that constitute the communication system. The network interface 1201 may also be used for wireless communication. For example, the network interface 1201 may be used for wireless LAN communication as defined in the IEEE 802.11 series, or for mobile communication as defined in 3GPP (registered trademark) (3rd Generation Partnership Project). Alternatively, the network interface 1201 may include, for example, a network interface card (NIC) compliant with the IEEE 802.3 series. Various sensors such as the aforementioned camera sensor 17, human presence sensor 19, joint rotation sensors, and force sensors are connected to the network interface 1201.

[0028] The processor 1202 reads and executes software (computer programs) from the memory 1203 to perform the processing of the route determination system described using a flowchart or sequence in the above embodiment. The processor 1202 may be, for example, a microprocessor, an MPU (Micro Processing Unit), a CPU (Central Processing Unit), or a GPU (Graphics Processing Unit). The processor 1202 may include multiple processors.

[0029] Memory 1203 is composed of a combination of volatile and non-volatile memory. Memory 1203 may also include storage located away from the processor 1202. In this case, the processor 1202 may access memory 1203 via an I / O interface not shown.

[0030] In the example shown in Figure 6, memory 1203 is used to store a group of software modules. The processor 1202 can read these software modules from memory 1203 and execute them to perform the routing system processing described in the above embodiment.

[0031] As illustrated with Figures 4 and 5, each processor executes one or more programs containing a set of instructions for causing the computer to perform the algorithms illustrated in the diagrams.

[0032] In the examples described above, the program includes a set of instructions (or software code) that, when loaded into a computer, cause the computer to perform one or more of the functions described in the embodiments. The program may be stored on a non-temporary computer-readable medium or a physical storage medium. Examples, but not limited to, include random-access memory (RAM), read-only memory (ROM), flash memory, solid-state drive (SSD) or other memory technologies, CD-ROM, digital versatile disc (DVD), Blu-ray® disc or other optical disc storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage devices. The program may be transmitted over a temporary computer-readable medium or a communication medium. Examples, but not limited to, include temporary computer-readable medium or a communication medium that includes electrically, optically, acoustically or otherwise propagating signals.

[0033] It should be noted that the present invention is not limited to the embodiments described above, and can be modified as appropriate without departing from the spirit of the invention. [Explanation of Symbols]

[0034] 1. Route determination system 10 Robot Arms 13a~f Joint area (1 axis~6 axes) 15 Hand 16 Gripping part (claw) 17. Sensor (camera) 19 sensors 20 Deep Reinforcement Learning Models 100 Robot Controllers 101 Object recognition section 102 Grasping point search unit 103 Route generation unit 130 links 150 base 201 Input Layer 202 Hidden Layer 203 Output Layer

Claims

1. A camera mounted on a robotic arm to photograph objects, A sensor mounted on the joint of the robot arm, which acquires angle data of the joint, The robot arm and a sensor that detects the distance to obstacles, including workers, A robot controller that controls the movement of the robot arm, A deep reinforcement learning model that optimizes the operation of the robot arm, Equipped with, The aforementioned robot controller is The camera captures an image and recognizes an object. Search for the gripping point of the recognized object, A path is generated to move the gripping part at the tip of the robot arm to the searched gripping point. The deep reinforcement learning model is configured to take the angles of predetermined joints of the robot arm as input when moving along the generated path, to minimize the amount of change in the angles of the predetermined joints, and to ensure that the distance between the robot arm and obstacles, including the worker, as detected by the sensor, is greater than or equal to a threshold, and to determine the path of the robot arm based on the learned deep reinforcement learning model. If the sensor detects that no worker is present near the robot arm, it refrains from outputting the path determined based on the deep reinforcement learning model, and instead outputs an arbitrary path regardless of the amount of change in the angle of a predetermined joint. A path determination system configured to determine the path of the robot arm based on a deep reinforcement learning model that has been trained to minimize the change in the angle of a predetermined joint and to ensure that the distance to the detected worker is greater than or equal to a threshold, when the sensor detects the presence of a worker in the vicinity of the robot arm.

2. The path determination system according to claim 1, wherein the predetermined joint is one or more joints selected from joints other than the joint located at the tip of the robot arm.

3. A camera mounted on a robotic arm to photograph objects, A sensor mounted on the joint of the robot arm, which acquires angle data of the joint, The robot arm and a sensor that detects the distance to obstacles, including workers, A robot controller that controls the movement of the robot arm, A deep reinforcement learning model that optimizes the operation of the robot arm, A route determination method using a route determination system equipped with, The camera captures an image and recognizes an object. Search for the gripping point of the recognized object, A path is generated to move the gripping part at the tip of the robot arm to the searched gripping point. The angles of predetermined joints of the robot arm moving along the generated path are input, and the path of the robot arm is determined based on the deep reinforcement learning model learned by deep reinforcement learning so as to minimize the amount of change in the angles of the predetermined joints, and so that the distance between the robot arm detected by the sensor and the obstacles including the worker is greater than or equal to a threshold. If the sensor detects that no worker is present near the robot arm, it refrains from outputting the path determined based on the deep reinforcement learning model, and instead outputs an arbitrary path regardless of the amount of change in the angle of a predetermined joint. A path determination method for determining the path of a robot arm when the presence of a worker is detected in the vicinity of the robot arm by the aforementioned sensor, based on a deep reinforcement learning model that has been trained to minimize the change in the angle of a predetermined joint and to ensure that the distance to the detected worker is greater than or equal to a threshold.

Citation Information

Patent Citations

  • Machine learning device, machine learning method, and machine learning program

    JP2018202550A

  • Robot system

    JP2020089944A

  • Control device, control method and control program

    JP2021035714A

  • Route generating device and route generating program of robot arm

    JP2021062416A

  • Robot control system, robot control method and program

    JP2022076572A