Learning device, handling system, learning method, program, and storage medium

A three-stage learning process for handling robots, combining simulation and real-world training, addresses high learning costs by using pre-trained policies and sensor information to enhance grasping efficiency and reduce training time.

JP2026057206APending Publication Date: 2026-04-02KK TOSHIBA
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-20
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

Existing handling robots require high learning costs for grasping operations due to the need for extensive data collection and training, especially when the gripping part is replaced or object characteristics change.

Method used

A three-stage learning process involving simulation, real-world environment training, and model refinement is employed, utilizing pre-trained policies and sensor information to reduce training costs and improve grasping efficiency.

Benefits of technology

The method significantly reduces the learning cost and time required for handling robots to grasp objects by leveraging simulation-based policy learning, real-world adaptation, and image-based model training, enhancing the success rate of grasping operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026057206000001_ABST
    Figure 2026057206000001_ABST
Patent Text Reader

Abstract

The present invention provides learning devices, handling systems, learning methods, programs, and storage media that can reduce learning costs. [Solution] The learning device according to the embodiment performs a first learning, a second learning, and a third learning. In the first learning, a first policy for determining the gripping motion of a robot arm having a gripping part is learned in a simulation environment. In the second learning, a second policy for determining the gripping motion of a robot arm is learned in a real environment. In the third learning, a model is learned that outputs gripping information for gripping an object in response to an image input. In the second learning, the learning device learns the second policy using the output from the learned first policy and sensor information acquired by the sensor during the gripping motion. In the third learning, the learning device learns the model using a first image that reflects the real first environment and gripping information output from the second policy learned for the first environment as training data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to a learning device, a handling system, a learning method, a program, and a storage medium.

Background Art

[0002] There is a handling robot that transports or picks up an object. For handling robots, a technology that can reduce the learning cost is required.

Prior Art Documents

Patent Documents

[0003]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] The problem to be solved by the embodiments of the present invention is to provide a learning device, a handling system, a learning method, a program, and a storage medium capable of reducing the learning cost.

Means for Solving the Problems

[0005] The learning device according to the embodiment performs a first learning, a second learning, and a third learning. In the first learning, a first policy for determining the gripping motion of a robot arm having a gripping part is learned in a simulation environment. In the second learning, a second policy for determining the gripping motion of a robot arm is learned in a real environment. In the third learning, a model is learned that outputs gripping information for gripping an object in response to an image input. In the second learning, the learning device learns the second policy using the output from the learned first policy and sensor information acquired by the sensor during the gripping motion. In the third learning, the learning device learns the model using a first image that reflects the real first environment and gripping information output from the second policy learned for the first environment as training data. [Brief explanation of the drawing]

[0006] [Figure 1] Figure 1 is a perspective view showing an example of a handling system according to an embodiment. [Figure 2] Figure 2 is a schematic diagram illustrating a learning method using a learning device according to an embodiment. [Figure 3] Figure 3 is a flowchart showing the processing during the first learning stage. [Figure 4] Figure 4 is a schematic diagram showing the data flow during the first learning stage. [Figure 5] Figure 5 is a flowchart showing the processing in the second learning stage. [Figure 6] Figure 6 is a schematic diagram showing the data flow in the second learning stage. [Figure 7] Figure 7 is a flowchart showing the processing in the third learning stage. [Figure 8] Figure 8 is a flowchart showing the handling method using a pre-trained model. [Figure 9] Figure 9 is a flowchart illustrating a method for processing sensor information. [Figure 10] Figure 10 is a schematic diagram illustrating the processing method shown in Figure 9. [Figure 11] Figure 11 is a schematic diagram representing the hardware configuration. [Modes for carrying out the invention]

[0007] The embodiments of the present invention will be described below with reference to the drawings. The drawings are schematic or conceptual, and the relationship between the thickness and width of each part, the ratio of the sizes of the parts, etc., are not necessarily the same as those of reality. Furthermore, even when representing the same part, the dimensions and ratios may be represented differently in the drawings. In this specification and each drawing, elements similar to those already described are denoted by the same reference numerals, and detailed explanations are omitted as appropriate.

[0008] In logistics operations, handling robots are used to transport or pick objects. Handling robots include a multi-jointed robotic arm, with a gripping section at the end of the arm. The gripping section can grasp objects by suction or clamping.

[0009] When a robotic arm grasps an object, the necessary grasping information is calculated. For example, object recognition, calculation of the grasping point, and generation of a plan are performed. In recent years, machine learning models have been used to calculate this grasping information in a shorter time. By using these models, grasping information can be acquired in a shorter time, and the robotic arm's movement can be started earlier. This allows for more efficient operation of the handling robot.

[0010] On the other hand, training a model requires a vast amount of data and training time. Furthermore, if the gripping part is replaced or the characteristics of the object being gripped change, the model needs to be updated. "Update" involves either retraining the model or replacing it. Updating the model requires data and time again for training.

[0011] Here, the data required for learning and the time required for learning are referred to as the "learning cost". As described above, when obtaining grasping information using a model, a large amount of learning cost is required in advance. And learning cost also occurs when the model is updated. An embodiment of the present invention aims to provide a technique capable of reducing this learning cost.

[0012] FIG. 1 is a perspective view showing an example of a handling system according to an embodiment. The handling system 1 shown in FIG. 1 executes handling of an object using a learned model. Specifically, the handling system 1 includes a handling robot 10, a sensor 20, and a processing device 30.

[0013] The handling robot 10 includes a robot arm 11 and a base 12. The robot arm 11 includes a plurality of links 11a and a plurality of rotating shafts 11b. The links 11a are connected to each other by the rotating shafts 11b. In the illustrated example, the robot arm 11 is a vertically articulated type. The robot arm 11 may be a horizontally articulated type.

[0014] By operating each rotating shaft 11b, the position and orientation (angle) of the tip of the robot arm 11 change. The tip of the robot arm 11 preferably has six degrees of freedom. The gripping part 15 is attached to the tip of the robot arm 11. In the illustrated example, the gripping part 15 includes a suction mechanism 16 and a clamping mechanism 17.

[0015] The suction mechanism 16 grips an object by suction. The suction mechanism 16 includes one or more suction pads 16a. In a state where the suction pad 16a contacts the object, the inside of the suction pad 16a is depressurized by a decompression device (not shown). Thereby, the object is adsorbed to the suction pad 16a. The number of suction pads 16a may be less or more than that in the illustrated example.

[0016] The clamping mechanism 17 grips an object by clamping. The clamping mechanism 17 includes a plurality of rod-shaped support portions 17a. The object is sandwiched and gripped by the plurality of support portions 17a. The clamping mechanism 17 may include more support portions 17a than the illustrated example. The support portion 17a may be configured in a finger shape including one or more joints.

[0017] The gripping portion 15 further includes a switching mechanism 18. The suction mechanism 16 and the clamping mechanism 17 are connected to the switching mechanism 18. The switching mechanism 18 rotates the suction mechanism 16 and the clamping mechanism 17. By rotating the suction mechanism 16 and the clamping mechanism 17, the mechanism used for gripping the object can be switched.

[0018] Not limited to the illustrated example, the gripping portion 15 may include only one of the suction mechanism 16 and the clamping mechanism 17. In that case, the switching mechanism 18 is unnecessary.

[0019] The robot arm 11 is further provided with a sensor 13. The sensor 13 can detect one or more selected from the group consisting of the load applied to the gripping portion 15, the torque applied to the gripping portion 15, the acceleration of the gripping portion 15, and the angular velocity of the gripping portion 15. For example, the sensor 13 includes one or more selected from the group consisting of a force sensor, an acceleration sensor, and an angular velocity sensor.

[0020] Near the handling robot 10, two containers C1 and C2 are placed. The handling robot 10 grips an object O housed in the container C1 and transports it to the container C2.

[0021] The sensor 20 is provided to detect the internal state of the container C1. For example, the sensor 20 includes one or more selected from an image sensor and a depth sensor. The sensor 20 may be fixed above the container C1 or attached to the handling robot 10.

[0022] The processing unit 30 acquires the image obtained by the sensor 20. The processing unit 30 also refers to the first model M1. The first model M1 outputs gripping information for gripping an object in response to the image input. The "gripping information" includes, for example, the gripping point. The gripping point represents the position and orientation (angle) of the gripping unit 15 when gripping an object. The gripping information may further include information indicating the type of gripping unit 15. If there are multiple objects in container C1 and the multiple objects are transported sequentially, the gripping information may include the gripping point and the order of transport for each object.

[0023] The processing unit 30 inputs the image to the first model M1 and acquires grasping information for grasping the object in the image. The robot arm 11 operates according to the grasping information and grasps the object.

[0024] The first model M1 is pre-trained by the learning device 40. The handling system 1 may include the learning device 40. The processing unit 30 may also have the function of the learning device 40.

[0025] Figure 2 is a schematic diagram illustrating a learning method using a learning device according to an embodiment. The learning device according to the embodiment performs the learning method shown in Figure 1. The learning method includes a first learning step (step S10), a second learning step (step S20), and a third learning step (step S30).

[0026] In the first learning step (step S10), the learning device teaches the first policy P1. The first policy P1 is a rule for determining the grasping motion of the robot arm. The first policy P1 is learned in a computer-based simulation environment. The first policy P1 is learned in a way that improves the success rate of the robot arm's grasping.

[0027] In the second learning step (step S20), the learning device trains the second policy P2. The second policy P2 is a set of rules for determining the grasping motion of the robot arm. The second policy P2 is trained in a real environment using an actual robot arm. Information about the robot arm, sensor information acquired by sensors, and the output from the first policy P1 are used for training the second policy P2. The second policy P2 is trained to improve the success rate of the robot arm's grasping.

[0028] In the third learning step (step S30), the learning device trains the first model M1. The first model M1 outputs grasping information for grasping an object in response to image input. The first model M1 is trained using images (first images) that depict the real environment (first environment), outputs from the second policy P2, etc. The first model M1 is trained to improve the success rate of grasping by the robot arm.

[0029] In the first training stage, a simulation environment is used. Training using a simulation environment is easier than training using a real environment, thus reducing training costs. In the second training stage, a real environment is used. In this stage, using the pre-trained first policy can speed up the training of the second policy, further reducing training costs. In the third training stage, the output of the pre-trained second policy is used for training. By preparing high-quality training data using the second policy, the training of the first model M1 can be accelerated, further reducing training costs.

[0030] The following sections will explain specific examples of each learning method. Here, we will explain an example where reinforcement learning is used for learning.

[0031] Figure 3 is a flowchart showing the processing in the first learning stage. Figure 4 is a schematic diagram showing the data flow in the first learning stage. In the first learning stage, the state of the robot arm (agent) and the simulation environment are initialized on the simulator (step S11). Next, the simulation environment is created (step S12). The simulation environment is created and saved by the user. For example, a physics engine such as Bullet can be used for the simulator. In addition, a robot visualization tool such as rviz can be used for the simulation. By using a visualization tool, the robot's sensor information, the robot's state, and a three-dimensional model of the environment can be displayed. The simulation environment is created to mimic the environment of a real handling robot 10.

[0032] Reinforcement learning is performed using the generated simulation environment. Specifically, the learning device acquires first information i1 and second information i2 (shown in Figure 4) from the simulation environment (step S13).

[0033] The first piece of information i1 includes one or more selected from a group consisting of information indicating the state of the robot arm 11, information indicating the state of the area around the robot arm 11, and information indicating the action of the robot arm 11. For example, "state of the robot arm" may include whether or not it is grasping an object, the position and orientation of the gripping part 15, and the position and orientation of the robot arm 11. For example, the position and orientation of the robot arm 11 and the position and orientation of the gripping part 15 can be represented by a combination of rotation angles of each axis (motor). "State of the area around the robot arm" may include how many objects are in the container, the arrangement of objects in the container, and whether or not there are partitions in the container. "Action of the robot arm" may include picking, knocking over objects, etc.

[0034] The second piece of information i2 includes one or more selected from the group consisting of information indicating the characteristics of the object O being grasped, information indicating the characteristics of the grasping unit 15 (robot hand), sensor information, and image information of the object O. Sensor information is acquired by sensor 13. Image information is acquired by sensor 20.

[0035] In the example shown in Figure 4, multiple first policies P1a to P1d are learned. First policy P1a relates to the operation of the robot arm 11 when grasping an object by suction. First policy P1a outputs grasping information for the robot arm 11 to grasp the object. If the output of first policy P1a is adopted in the plan, the robot arm 11 attempts to grasp the object according to the grasping information output from first policy P1a.

[0036] The first policy P1b relates to the operation of the robot arm 11 when gripping an object by pinching. The first policy P1b outputs gripping information for the robot arm 11 to grip the object. If the output of the first policy P1b is adopted in the plan, the robot arm 11 attempts to grip the object according to the gripping information output from the first policy P1b.

[0037] The first strategy P1c relates to the operation of the robot arm 11 before grasping an object. As an example, several thin objects are placed upright in a container C1. When attempting to grasp any one of the objects by suction, the contact area between the gripping part 15 and the object O is small, making suction-based grasping difficult. In such cases, it is effective to bring the gripping part 15 into contact with the object and cause it to tip over (collapse). When the object tips over, a larger surface of the object faces upward. By grasping this larger surface, the success rate of grasping increases. If the output of the first strategy P1c is adopted in the plan, the robot arm 11 will attempt to collapse the object according to the operation information output from the first strategy P1c.

[0038] The first strategy P1d relates to the operation of the robot arm 11 after an attempt to grasp an object. As an example, the gripping unit 15 grasps a cylindrical or circular object. When the gripping unit 15 comes into contact with an object, the object may rotate. When the object rotates, the contact state between the gripping unit 15 and the object changes. For example, if the contact area between the gripping unit 15 and the object decreases, the success rate of grasping decreases. When the contact state changes, or when grasping fails, the success rate of grasping can be increased by temporarily releasing the gripping unit 15 from the object and then making contact again. If the output of the first strategy P1d is adopted in the plan, the robot arm 11 will attempt to re-contact the object with the gripping unit 15 according to the operation information output from the first strategy P1d.

[0039] The first strategies P1a to P1d receive the first information i1 and the second information i2 as inputs. The first strategies P1a to P1d output motion information for the robot arm 11 in response to the input of this information.

[0040] In the illustrated example, the second information i2 is dimensionally compressed by encoders e1a to e1d (step S14). The first information i1 and the dimensionally compressed second information i2 are input to each of the first policies P1a to P1d. The second information i2 includes information that is not important for determining the action. Encoding abstracts the information of the second information i2, increasing the versatility of the first policies P1a to P1d. Also, since the first information i1 contains little or no unnecessary information, the first information i1 is input to the first policies P1a to P1d without going through the encoder.

[0041] Encoders e1a to e1d may be the same or different from each other. Preferably, encoders e1a to e1d are different from each other so that dimensionality reduction suitable for the first policy P1a to P1d is achieved. For example, variational autoencoders (VAEs) and convolutional autoencoders (CAEs) can be used for encoders e1a to e1d. Generative adversarial networks (GANs) may be combined with VAEs.

[0042] When the first information i1 and the second information i2 are input to the first policy P1a to P1d, the first policy P1a to P1d outputs operation information (step S15). For example, the first policy P1a to P1d outputs the position and orientation of the gripping part 15 when performing each operation.

[0043] The learning device 40 adopts any of the operation information output from the first policies P1a to P1d and decides on an action based on that operation information (step S16). For example, the learning device 40 generates a plan based on the operation information, the first information, and the second information. The plan includes the object to be transported, the position and orientation of the gripping part 15 when gripping the object, the intermediate positions, the position and orientation of the gripping part 15 when releasing the object, the gripping force by the gripping part 15, the method of gripping the object, the transport speed, etc. When the object is gripped by suction, the gripping force is expressed as pressure (vacuum). When the object is gripped by clamping, the gripping force is expressed as the motor current.

[0044] The learning device 40 operates the robot arm 11 in the simulation environment according to the action determined in step S16. If the intended result is obtained, the learning device 40 returns a reward to the first policy that output the adopted action information (step S17). For example, if the action is determined based on the output from the suction or gripping policy, and suction or gripping is successful, the suction or gripping policy is rewarded. If the action is determined based on the output from the dislodgement policy, and dislodgement is successful, or grasping after the dislodgement is successful, the dislodgement policy is rewarded. If the action is determined based on the output from the re-contact policy, and grasping is successful after re-contact, the re-contact policy is rewarded. The first policies P1a to P1d are learned in such a way that the reward is maximized.

[0045] After learning is performed in the simulation environment generated in step S12, the learning device 40 determines whether to terminate the learning (step S18). For example, if the cumulative reward or average reward obtained by the agent within a certain period exceeds a preset threshold, the learning device 40 terminates the first learning. Alternatively, if learning has been performed in a preset number of simulation environments, the learning device 40 terminates the first learning.

[0046] If learning continues, the learning device 40 acquires the next simulation environment and uses that simulation environment to execute steps S12 to S17 again.

[0047] Through the above process, one or more first policies are learned. The learning device 40 stores the learned first policies.

[0048] Figure 5 is a flowchart showing the processing in the second learning stage. Figure 6 is a schematic diagram showing the data flow in the second learning stage. As shown in Figure 5, in the second learning stage, third information i3 and sensor information i4 (shown in Figure 6) are acquired (step S21). The third information i3, like the second information i2, includes one or more selected from the group consisting of information indicating the characteristics of the object O to be grasped, information indicating the characteristics of the gripping unit 15 (robot hand), sensor information, and image information of the object O. The sensor information i4 includes information acquired by the sensors of the actual handling system 1. The sensor information i4 includes the load applied to the gripping unit 15, the torque applied to the gripping unit 15, the acceleration of the gripping unit 15, the angular velocity of the gripping unit 15, the rotation angle of the motor included in the robot arm 11, the rotation speed of the motor, contact information between the gripping unit 15 and the object, or an image of the handling robot 10. The sensor information i4 may also include multiple consecutive images (videos).

[0049] The learning device 40 processes the sensor information i4 (step S22). For example, the processing removes noise contained in the sensor information i4, or the sensor information i4 is abstracted for learning.

[0050] The learning device 40 inputs the third information i3 and sensor information i4 to encoders e2a and e2b, respectively, to perform dimensionality reduction (step S23). The learning device 40 also inputs state information i5 of the actual robot arm 11 to the first policy (step S24). The state information i5 includes one or more selected from the group consisting of information indicating the state of the robot arm 11 and information indicating the state of the surroundings of the robot arm 11.

[0051] Using the dimensionally compressed third information i3, the dimensionally compressed sensor information i4, the processed sensor information i4, and the output from the first policy, the learning device 40 performs imitation learning on the second policy (step S25). The output of the first policy may be obtained from the output layer of the first policy. Alternatively, the latent vector in the intermediate layer between the input layer and the output layer of the first policy may be extracted as the output of the first policy. The output of the first policy may be distilled and used for learning the second policy. In imitation learning, the output of the second policy when the third information i3 and sensor information i4 are input is learned to imitate the output of the first policy. In the example shown in Figure 6, outputs are obtained from multiple first policies. These outputs may be averaged or weighted averaged.

[0052] Instead of the example shown in Figure 6, the sensor information i4 may be processed and dimensionally compressed. In this case, the second policy is learned by imitation using the dimensionally compressed third information i3, the processed and dimensionally compressed sensor information i4, and the output from the first policy.

[0053] The learned second policy P2 outputs motion information of the robot arm 11 (step S26). The learning device 40 determines the action of the robot arm 11 based on the output of the second policy P2 (step S27). The learning device 40 operates the robot arm 11 in the real environment according to the action determined in step S26. If the intended result is obtained, the learning device 40 returns a reward to the second policy (step S28). The second policy P2 is learned to maximize the reward.

[0054] The learning device 40 determines whether to terminate the learning process (step S29). For example, if the cumulative reward or average reward obtained by the agent within a certain period exceeds a preset threshold, the learning device 40 terminates the second learning process. Alternatively, if a preset number of imitation learning sessions have been performed, the learning device 40 terminates the second learning process.

[0055] If learning continues, the next real-world environment is prepared. In that real-world environment, the learning device 40 executes steps S21 to S28 again.

[0056] Through the above process, the second policy is learned. The learning device 40 saves the learned second policy.

[0057] Figure 7 is a flowchart showing the processing in the third learning stage. As shown in Figure 7, in the third learning stage, third information and sensor information from the real environment are acquired (step S31). The learning device 40 compresses the dimensions of this information (step S32) and inputs it to the second policy. The learning device 40 acquires the grasping information output from the second policy (step S33). The learning device 40 also acquires images of the real environment acquired by the sensor 20 (step S34). The learning device 40 sets the images in the input layer and the grasping information from the second policy in the output layer to perform supervised learning on the model (step S35).

[0058] The learning device 40 determines whether to terminate the learning process (step S36). For example, if the loss of the learned model is less than a preset threshold, the learning device 40 terminates the third learning process. Alternatively, if a preset number of learning cycles have been performed, the learning device 40 terminates the third learning process. If learning continues, the next real-world environment is prepared. In that real-world environment, the learning device 40 executes steps S31 to S35 again.

[0059] Through the above process, a model for outputting grasping information is trained. After the model training is complete, grasping information is acquired using that model.

[0060] Figure 8 is a flowchart showing the handling method using a pre-trained model. After the learning is complete, the processing unit 30 uses the learned first model M1 to cause the handling robot 10 to grasp an object. Specifically, as shown in Figure 8, the processing unit 30 acquires the image obtained by the sensor 20 (step S41).

[0061] The processing unit 30 inputs an image to the classifier D and obtains the gripping method output from the classifier D (step S42). The classifier D outputs a determination result indicating whether to use suction or pinching gripping method, according to the image input. The classifier D is pre-trained. For example, the classifier D includes a neural network. Preferably, the classifier D includes a convolutional neural network (CNN).

[0062] The processing unit 30 performs object segmentation and recognition using the image (step S43). Segmentation and recognition are performed by a trained recognition model. For example, the recognition model includes a neural network. Preferably, the recognition model M includes a CNN. The recognition model M outputs an image showing the recognition result.

[0063] The processing unit 30 inputs the image output from the recognition model M to the trained first model M1 and acquires the grasping information output from the first model M1 (step S44).

[0064] Furthermore, the processing unit 30 acquires third information and sensor information (step S45). Based on this information, the processing unit 30 generates various plans. Plan generation includes the generation of a gripping plan (step S46), a motion plan (step S47), a task plan (step S48), and a release plan (step S49). The gripping plan includes the position and orientation of the gripping unit 15 when gripping an object, the gripping method, the gripping force, etc. The motion plan includes the movement of the gripping unit 15 when gripping an object, the movement of the object being gripped, etc. The task plan includes the transit points of the gripping unit 15 from gripping the object to releasing the object. The release plan includes the position and orientation of the gripping unit 15 when releasing the object.

[0065] The advantages of the embodiment will be explained. In embodiments of the present invention, first learning, second learning, and third learning are performed. In the first learning, a first policy for determining the grasping motion of the robot arm 11 is learned in a simulation environment. Learning in a simulation environment allows the first policy to be learned in a shorter time compared to learning in a real environment. In the second learning, the second policy is learned in a real environment. In this learning, sensor information acquired by sensors during the grasping motion and the output from the learned first policy are used. The first policy is sufficiently learned in the simulation environment. Therefore, the output from the first policy can be used as high-quality learning data. By using the output from the first policy for learning, the time required to learn the second policy can be shortened. Furthermore, by using sensor information from the real environment for learning, the time required to learn the second policy can be further shortened. In the third learning, a model that outputs grasping information based on images is learned. The second policy is sufficiently learned in a real environment. Therefore, the output from the second policy can be used as high-quality learning data. By using the output from the second policy for training, the time required to train the model can be reduced.

[0066] According to this embodiment, the cost required to train a model for obtaining grasping information can be reduced.

[0067] In learning the second strategy, sensor information may be used in any way. Here, we describe a specific example where the sensor information includes force information in time series data. The force information is at least one selected from the group consisting of load, torque, acceleration, and angular velocity.

[0068] Figure 9 is a flowchart illustrating a method for processing sensor information. Figure 9 shows a specific example of step S22 in Figure 5. First, the learning device 40 performs a frequency conversion on the force information (step S22a). The Fast Fourier Transform (FFT) can be used for the frequency conversion. By performing the frequency conversion, noise that is unnecessary for learning is removed.

[0069] Next, the learning device 40 performs patterning on the time series data (step S22b). Through patterning, the time series data is divided into multiple intervals, and further, it is determined which operation is being performed in each interval of the time series data. The k-means method can be used for patterning.

[0070] The learning device 40 corrects the patterned time series data using a pre-prepared theoretical model (step S22c). The correction includes one or more methods selected from the group consisting of comparison with a threshold, filtering, and integration with a stiffness matrix. For example, weak noise contained in the time series data can be removed by comparison with a threshold or filtering. Specific moments contained in the time series data can be emphasized by integration with a stiffness matrix. In the correction, the threshold, filter, or stiffness matrix may be prepared for each object.

[0071] The learning device 40 selects the planning hierarchy to which the processed time-series data will be input (step S22d). Various processes are performed before the final movement of the robot arm 11 is determined. These include, for example, object recognition, task plan calculation, motion plan calculation, and grasping plan calculation. Step S22d selects which of these hierarchies the time-series data will be input to. Subsequently, at the selected planning level, the time-series data is used to learn the second policy.

[0072] As a concrete example, when generating a certain action, a gripping plan, motion plan, task plan, and release plan are generated, as shown in Figure 8. In the generation of the motion plan and release plan, the position and force over a relatively short period are determined. In the generation of the gripping plan and task plan, the actions over a relatively long period are determined. Actions over a long period include, for example, performing a release motion followed by a gripping motion. In step S22d, it is selected whether to use the time-series data from when a certain action was performed to generate a plan for a relatively short period or a plan for a relatively long period. For example, by specifying the control period for the time-series data, it is possible to select whether to use the time-series data to generate a plan for a short period or a plan for a long period. As an example, if the time-series data is used to generate a plan for a short period, the control period is set to 1 millisecond. If the time-series data is used to generate a plan for a long period, the control period is set to 10 milliseconds.

[0073] Figure 10 is a schematic diagram illustrating the processing method shown in Figure 9. First, time-series data TD1 of force information is acquired from sensor 13. In time-series data TD1, the horizontal axis represents time t, and the vertical axis represents the value v detected by the sensor. Time-series data TD1 is frequency-converted to obtain time-series data TD2. By patterning time-series data TD2, each timing in time-series data TD2 is classified as being in the course of one of the operations.

[0074] For example, if a gripping operation fails, past force feedback information from successful gripping operations is referenced. The learning device 40 trains a second policy so that the force feedback information from failed operations approaches that of successful operations. This improves the success rate of gripping using the output from the second policy.

[0075] For example, an object is grasped by suction. Even if the vacuum level is sufficiently high, there are cases where the object falls during the upward movement after being grasped. According to this embodiment, force information and posture information are acquired at the start of the upward movement, and if the object appears likely to fall, the lifting operation is stopped. The grasping method is then changed from suction to clamping. This makes it possible to avoid the object falling.

[0076] As another example, a cylindrical or spherical object is grasped by suction. If there are errors in distance measurement by the sensor 20 or errors in position calculation in the plan, the gripping part 15 may shift position when the object rolls while grasping it. According to this embodiment, when the position of the gripping part 15 shifts, the shift can be detected and the gripping part 15 can be moved in the opposite direction to the shift. This can increase the success rate of grasping cylindrical or spherical objects.

[0077] In the second learning phase, information (such as force feedback) from the gripping motion of the actual handling robot 10 is continuously input to the second policy P2. According to the second policy P2, the robot switches between actions such as suction, pinching, and re-contact depending on the level of reward in any given state. Training data can be obtained by acquiring images and gripping information output from the second policy P2 at each timing during the operation of the handling robot 10. The first model M1 is trained using these images and gripping information. As a result, the first model M1 is trained to output appropriate gripping information at each timing during the operation of the handling robot 10.

[0078] After training with the first model M1, images acquired during the operation of the handling robot 10 are input to the first model M1 in real time. This allows appropriate gripping information at each timing to be obtained from the first model M1. By reflecting this gripping information in the plan, it becomes possible to stop the lifting operation and change the gripping method as described above. Alternatively, it becomes possible to move the gripping part 15 in the opposite direction to the displacement.

[0079] Figure 11 is a schematic diagram representing the hardware configuration. As the processing unit 30 or learning device 40, for example, the computer 90 shown in Figure 11 is used. The computer 90 includes a processing circuit 91, ROM 92, RAM 93, storage device 94, input interface 95, output interface 96, and communication interface 97.

[0080] ROM92 stores programs that control the operation of computer 90. ROM92 contains the programs necessary for computer 90 to perform each of the processes described above. RAM93 functions as a memory area where the programs stored in ROM92 are loaded.

[0081] The processing circuit 91 includes an arithmetic processing unit such as a CPU or GPU. The processing circuit 91 uses RAM 93 as work memory and executes a program stored in at least one of ROM 92 or storage device 94. During program execution, the processing circuit 91 controls each component via the system bus 98 and performs various processes.

[0082] The memory device 94 stores data necessary for program execution and data obtained through program execution.

[0083] The input interface (I / F) 95 can connect the computer 90 and the input device 95a. The input I / F 95 is, for example, a serial bus interface such as USB. The processing circuit 91 can read various data from the input device 95a via the input I / F 95.

[0084] The output interface (I / F) 96 can connect the computer 90 and the output device 96a. The output I / F 96 is a video output interface such as a Digital Visual Interface (DVI) or a High-Definition Multimedia Interface (HDMI®). The processing circuit 91 can transmit data to the output device 96a via the output I / F 96 and display an image on the output device 96a.

[0085] The communication interface (I / F) 97 can connect the computer 90 to a server 97a located outside the computer 90. The communication I / F 97 is, for example, a network card such as a LAN card. The processing circuit 91 can read various data from the server 97a via the communication I / F 97.

[0086] The storage device 94 includes one or more selected from Hard Disk Drives (HDDs) and Solid State Drives (SSDs). The input device 95a includes one or more selected from a mouse, keyboard, microphone (voice input), and touchpad. The output device 96a includes one or more selected from a monitor, projector, printer, and speaker. Devices that have the functions of both input device 95a and output device 96a, such as a touch panel, may also be used.

[0087] Each process performed by the processing unit 30 or the learning device 40 may be implemented by a single computer 90, or by the cooperation of multiple computers 90. A single computer 90 may have the functions of both the processing unit 30 and the learning device 40.

[0088] The processing of the various data described above may be recorded as a program that can be executed by a computer on a magnetic disk (flexible disk and hard disk, etc.), an optical disk (CD-ROM, CD-R, CD-RW, DVD-ROM, DVD±R, DVD±RW, etc.), a semiconductor memory, or another non-transitory computer-readable storage medium.

[0089] For example, data on a recording medium is read by a computer (or embedded system). The recording format (storage format) on the recording medium is arbitrary. For example, a computer reads a program from the recording medium and causes the CPU to execute instructions based on this program. The acquisition (or reading) of the program by the computer may be done via a network.

[0090] According to the embodiments described above, a learning device, a handling system, a learning method, a program, and a storage medium are provided that can reduce the learning cost.

[0091] In this specification, "or" indicates that "at least one" of the items listed in the text may be adopted.

[0092] Although several embodiments of the present invention have been illustrated above, these embodiments are presented as examples only and are not intended to limit the scope of the invention. These novel embodiments can be implemented in various other forms, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. These embodiments and their variations are included in the scope and spirit of the invention, as well as in the claims of the invention and its equivalents. Furthermore, the embodiments described above can be implemented in combination with each other. [Explanation of symbols]

[0093] 1: Handling system, 10: Handling robot, 11: Robot arm, 11a: Link, 11b: Rotating axis, 12: Base, 13: Sensor, 15: Gripping part, 16: Suction mechanism, 16a: Suction pad, 17: Clamping mechanism, 17a: Support part, 18: Switching mechanism, 20: Sensor, 30: Processing unit, 40: Learning device, 90: Computer, C1, C2: Container, D: Classifier, M: Recognition model, M1: First model, O: Object, P1, P1a~P1d: First policy, P2: Second policy, TD1: Time series data, TD2: Time series data, e1a~e1d, e2a, e2b: Encoder, i1: First information, i2: Second information, i3: Third information, i4: Sensor information, i5: State information

Claims

1. A first learning process involves learning a first strategy for determining the gripping motion of a robot arm having a gripping part in a simulation environment, and The second learning process involves training a second strategy for determining the grasping motion of the robot arm in a real-world environment, The third learning process involves training a model that outputs grasping information for grasping an object in response to an image input, A learning device that performs the following: In the second learning process, the second policy is trained using the output from the first policy that has been learned and the sensor information acquired by the sensor during the gripping operation. A learning device that, in the third learning stage, trains the model using a first image that reflects the real first environment and grasping information output from the second policy that has been trained for the first environment as training data.

2. In the first learning process, multiple first policies are trained, One of the aforementioned first measures relates to the operation of the robot arm when grasping an object, Another of the plurality of first measures is the learning device according to claim 1, relating to the movement of the robot arm before grasping an object.

3. The learning device according to claim 2, wherein in the second learning, the second policy is learned using the outputs of each of the plurality of first policies.

4. In the first learning process, the first policy is input with first information and second information. The first information includes one or more selected from the group consisting of information indicating the state of the robot arm, information indicating the state of the area around the robot arm, and information indicating the behavior of the robot arm. The learning device according to claim 1, wherein the second information includes one or more selected from the group consisting of information indicating the characteristics of the object to be grasped, information indicating the characteristics of the grasping part, sensor information acquired by a sensor provided on the robot arm, and image information of the object to be grasped.

5. The learning device according to claim 4, wherein the second information is dimensionally compressed by an encoder and input to the first policy.

6. The learning device according to any one of claims 1 to 5, wherein in the second learning, the sensor information includes one or more selected from the group consisting of load on the gripping part, acceleration of the gripping part, and torque on the gripping part.

7. The second measure outputs a gripping point that represents the position and orientation of the gripping part, The learning device according to any one of claims 1 to 6, wherein in the third learning, the model is learned using the gripping points output by the second policy as training data.

8. A learning device according to any one of claims 1 to 7, A handling robot including the aforementioned robotic arm, A handling system equipped with [this feature].

9. On the computer, A first learning process involves learning a first strategy for determining the gripping motion of a robot arm having a gripping part in a simulation environment, and The second learning process involves training a second strategy for determining the grasping motion of the robot arm in a real-world environment, The third learning process involves training a model that outputs grasping information for grasping an object in response to an image input, Make it run, In the second learning process, the second policy is trained using the output from the first policy that has been learned and the sensor information acquired by the sensor during the gripping operation. A learning method in which, in the third learning stage, the model is trained using a first image that reflects the real first environment and grasping information output from the second policy that has been trained for the first environment as training data.

10. A program that causes a computer to execute the learning method described in claim 9.

11. A storage medium storing the program described in claim 10.