Object grasping method, object grasping program, object grasping device, learning method, learning program, and learning device

The machine learning device autonomously learns optimal workpiece retrieval operations for robots, addressing the need for human intervention in conventional systems by using a state quantity observation unit and learning unit to enhance performance in handling randomly placed workpieces.

JP7856691B2Active Publication Date: 2026-05-11FANUC LTD +1
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
FANUC LTD
Filing Date
2024-03-12
Publication Date
2026-05-11

AI Technical Summary

Technical Problem

Conventional robot systems require human intervention to pre-program the extraction of workpieces from three-dimensional images and robot operations, leading to decreased success rates when settings are inappropriate, necessitating repetitive refinement to optimize performance.

Method used

A machine learning device that observes and learns the optimal operation of a robot using a hand unit to pick up randomly placed workpieces, incorporating a state quantity observation unit, operation result acquisition unit, and learning unit to calculate rewards and update value functions based on success or failure, enabling autonomous learning without human intervention.

Benefits of technology

The machine learning device enables the robot to learn optimal workpiece retrieval operations without human intervention, improving success rates by adapting to varying workpiece configurations and environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007856691000003
    Figure 0007856691000003
  • Figure 0007856691000004
    Figure 0007856691000004
  • Figure 0007856691000005
    Figure 0007856691000005
Patent Text Reader

Abstract

To provide a machine-learning device, a robot system, and a machine-learning method which can select optimal motion of a robot which is extracting a work-piece, without intervention by a person.SOLUTION: A machine-learning device comprises: a state quantity observing part 21 that observes output data from a three-dimensional measuring instrument 15 that measures at least a three-dimensional map for each of work-pieces 12 placed in a random manner including a bulky manner; a motion-result obtaining part 26 that obtains a result of extraction motion of a robot 14 that extracts the work-pieces using a hand part 13; and a learning part 22 that receives output from the state quantity observing part and the motion-result obtaining part to learn the extraction motion of the work-pieces. The state quantity observing part further observes output data from a coordinate calculating part that calculates a three-dimension position for each of the work-pieces on the basis of output from the three-dimensional measuring instrument. The learning part comprises a reward calculating part 23 that calculates a reward on the basis of a result of determination of a success or a failure of extraction of the work-piece which is output from the motion-result obtaining part, and a value function updating part 24 that has a value function for determining a value of the work-piece extraction motion and updates the value function in accordance with the reward.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a machine learning device, a robot system, and a machine learning method for learning how to retrieve workpieces that are placed haphazardly, including in a loosely stacked state. [Background technology]

[0002] For some time, robot systems have been known that, for example, grasp and transport workpieces that are loosely stacked in a basket-like box using the robot's hand (see, for example, Patent Documents 1 and 2). In such robot systems, for example, a three-dimensional measuring instrument installed above the basket-like box is used to acquire positional information of multiple workpieces, and the workpieces are picked up one by one by the robot's hand based on that positional information. [Prior art documents] [Patent Documents]

[0003] [Patent Document 1] Patent No. 5642738 [Patent Document 2] Patent No. 5670397 [Overview of the project] [Problems that the invention aims to solve]

[0004] However, in the conventional robot systems described above, it is necessary to pre-program, for example, how to extract a workpiece from distance images of multiple workpieces measured by a three-dimensional measuring instrument, and which workpiece to extract from which position. Furthermore, it is necessary to pre-program how the robot's hand will operate when extracting a workpiece. Specifically, for example, it may be necessary for a human to teach the robot the workpiece extraction action using a teaching pendant.

[0005] Therefore, if the settings for extracting the workpiece from distance images of multiple workpieces are not appropriate, or if the robot's motion program is not properly created, the success rate of the robot picking up and transporting the workpiece will decrease. Furthermore, to improve this success rate, it is necessary for humans to repeatedly experiment and refine the workpiece detection settings and the robot's motion program while searching for the optimal robot operation.

[0006] Therefore, in view of the circumstances described above, the object of the present invention is to provide a machine learning device, a robot system, and a machine learning method that can learn the optimal operation of a robot when picking up workpieces that are placed haphazardly, including in a stacked state, without human intervention. [Means for solving the problem]

[0007] According to a first configuration example of a first embodiment of the present invention, a machine learning device is provided for learning the operation of a robot that uses a hand unit to pick up workpieces from a plurality of randomly placed workpieces, including a stacked state, and comprising: a state quantity observation unit that observes output data of a three-dimensional measuring instrument that measures at least a three-dimensional map for each workpiece; an operation result acquisition unit that acquires the result of the robot's pick-up operation of the workpiece using the hand unit; and a learning unit that receives the output from the state quantity observation unit and the operation result acquisition unit and learns the pick-up operation of the workpiece, wherein the state quantity observation unit also observes output data of a coordinate calculation unit that calculates the three-dimensional position of each workpiece based on the output of the three-dimensional measuring instrument, and the learning unit comprises a reward calculation unit that calculates a reward based on the success or failure of the pick-up of the workpiece, which is the output of the operation result acquisition unit, and a value function update unit that has a value function that defines the value of the pick-up operation of the workpiece and updates the value function according to the reward. According to a second configuration example of the first embodiment of the present invention, a machine learning device is provided for learning the operation of a robot that picks up workpieces from a plurality of randomly placed workpieces, including a stacked state, using a hand unit, comprising: a state quantity observation unit that observes output data of a three-dimensional measuring instrument that measures at least a three-dimensional map for each workpiece; an operation result acquisition unit that acquires the result of the robot's picking up operation of the workpieces using the hand unit; and a learning unit that receives the output from the state quantity observation unit and the output from the operation result acquisition unit and learns the picking up operation of the workpieces, wherein the state quantity observation unit also observes output data of a coordinate calculation unit that calculates the three-dimensional position of each workpiece based on the output of the three-dimensional measuring instrument, and the learning unit has a learning model for learning the picking up operation of the workpieces, and comprises an error calculation unit that calculates an error based on the success or failure of the picking up of the workpieces, which is the output of the operation result acquisition unit, and the learning model, and a learning model update unit that updates the learning model according to the error. Preferably, the machine learning device further includes a decision-making unit that refers to the output from the learning unit and determines command data to instruct the robot to perform the workpiece removal operation.

[0008] According to a second embodiment of the present invention, a machine learning device is provided for learning the operation of a robot that uses a hand unit to pick up workpieces from a plurality of randomly placed workpieces, including a stacked state, the device comprising: a state quantity observation unit that observes state quantities of the robot, including output data from a three-dimensional measuring instrument that measures a three-dimensional map of each workpiece; an operation result acquisition unit that acquires the results of the robot's pick-up operation in which the hand unit picks up the workpieces; and a learning unit that receives the output from the state quantity observation unit and the output from the operation result acquisition unit, and learns manipulated variables, including the measurement parameters of the three-dimensional measuring instrument, in relation to the state quantities of the robot and the results of the pick-up operation. Preferably, the machine learning device further comprises a decision-making unit that determines the measurement parameters of the three-dimensional measuring instrument by referring to the manipulated variables learned by the learning unit.

[0009] The state quantity observation unit can also observe the state quantities of the robot, including output data from a coordinate calculation unit that calculates the three-dimensional position of each workpiece based on the output of the three-dimensional measuring instrument. The coordinate calculation unit may further calculate the orientation of each workpiece and output the calculated three-dimensional position and orientation data for each workpiece. The operation result acquisition unit can utilize the output data of the three-dimensional measuring instrument. The machine learning device further includes a pre-processing unit that processes the output data of the three-dimensional measuring instrument before inputting it to the state quantity observation unit, and it is preferable that the state quantity observation unit receives the output data from the pre-processing unit as the state quantities of the robot. The pre-processing unit can ensure that the orientation and height of each workpiece in the output data of the three-dimensional measuring instrument are consistent. The operation result acquisition unit can acquire at least one of the following: the success or failure of workpiece removal, the damage status of the workpiece, and the degree of achievement when the removed workpiece is handed over to a subsequent process.

[0010] The learning unit may include a reward calculation unit that calculates a reward based on the output of the operation result acquisition unit, and a value function update unit that has a value function that determines the value of the workpiece retrieval operation and updates the value function according to the reward. The learning unit may also include an error calculation unit that has a learning model that learns the workpiece retrieval operation, calculates an error based on the output of the operation result acquisition unit and the output of the learning model, and a learning model update unit that updates the learning model according to the error. The machine learning device preferably has a neural network.

[0011] According to a third embodiment of the present invention, a robot system is provided comprising a machine learning device for learning the operation of a robot that uses a hand unit to pick up workpieces from a plurality of randomly placed workpieces, including a stacked state, the device comprising: a state quantity observation unit that observes state quantities of the robot, including output data from a three-dimensional measuring instrument that measures a three-dimensional map for each workpiece; an operation result acquisition unit that acquires the result of a pick-up operation of the robot in which the hand unit picks up the workpieces; and a learning unit that receives the output from the state quantity observation unit and the output from the operation result acquisition unit and learns an operation quantity, including command data that commands the robot to perform the pick-up operation of the workpieces, in relation to the state quantities of the robot and the result of the pick-up operation, the device comprising a robot, a three-dimensional measuring instrument, and control devices that control the robot and the three-dimensional measuring instrument, respectively.

[0012] According to a fourth embodiment of the present invention, a robot system is provided comprising a machine learning device for learning the operation of a robot that uses a hand unit to pick up workpieces from a plurality of randomly placed workpieces, including a stacked state, the device comprising: a state quantity observation unit that observes state quantities of the robot, including output data from a three-dimensional measuring instrument that measures a three-dimensional map for each workpiece; an operation result acquisition unit that acquires the result of the robot's pick-up operation in which the hand unit picks up the workpieces; and a learning unit that receives the output from the state quantity observation unit and the output from the operation result acquisition unit and learns manipulated variables, including measurement parameters of the three-dimensional measuring instrument, in relation to the state quantities of the robot and the result of the pick-up operation, the device comprising a robot, a three-dimensional measuring instrument, and control devices that control the robot and the three-dimensional measuring instrument, respectively.

[0013] The robot system comprises a plurality of robots, and each robot is provided with a machine learning device. It is preferable that the plurality of machine learning devices provided on the plurality of robots share or exchange data with each other via a communication medium. The machine learning devices may reside on a cloud server.

[0014] A fifth embodiment of the present invention provides a machine learning method for learning the operation of a robot that picks up workpieces from a plurality of randomly placed workpieces, including a loosely stacked state, using a hand unit, the method comprising: observing the state quantities of the robot, including output data from a three-dimensional measuring instrument that measures a three-dimensional map of each workpiece; obtaining the results of the robot's pick-up operation in which the hand unit picks up the workpieces; receiving the output from the state quantity observation unit and the output from the operation result acquisition unit; and learning an operation quantity, including command data that commands the robot to perform the pick-up operation of the workpieces, in relation to the state quantities of the robot and the results of the pick-up operation. [Effects of the Invention]

[0015] According to the machine learning device, robot system, and machine learning method according to the present invention, there is an effect that the optimal operation of the robot when taking out randomly placed workpieces including the stacked state can be learned without human intervention.

Brief Description of the Drawings

[0016] [Figure 1] FIG. 1 is a block diagram showing a conceptual configuration of a robot system according to an embodiment of the present invention. [Figure 2] FIG. 2 is a diagram schematically showing a model of a neuron. [Figure 3] FIG. 3 is a diagram schematically showing a three-layer neural network configured by combining the neurons shown in FIG. 2. [Figure 4] FIG. 4 is a flowchart showing an example of the operation of the machine learning device shown in FIG. 1. [Figure 5] FIG. 5 is a block diagram showing a conceptual configuration of a robot system according to another embodiment of the present invention. [Figure 6] FIG. 6 is a diagram for explaining an example of the processing of the preprocessing unit in the robot system shown in FIG. 5. [Figure 7] FIG. 7 is a block diagram showing a modification example of the robot system shown in FIG. 1.

Embodiments for Carrying Out the Invention

[0017] Hereinafter, embodiments of the machine learning device, robot system, and machine learning method according to the present invention will be described in detail with reference to the accompanying drawings. Here, in each drawing, the same members are denoted by the same reference numerals. Also, it is assumed that those having the same reference numerals in different drawings are components having the same function. Note that, for ease of understanding, the scales of these drawings are appropriately changed.

[0018] Figure 1 is a block diagram showing the conceptual configuration of a robot system according to one embodiment of the present invention. The robot system 10 of this embodiment includes a robot 14 to which a hand portion 13 for gripping workpieces 12 piled loosely in a cage-like box 11 is attached, a three-dimensional measuring instrument 15 for measuring a three-dimensional map of the surface of the workpieces 12, a control device 16 for controlling the robot 14 and the three-dimensional measuring instrument 15 respectively, a coordinate calculation unit 19, and a machine learning device 20.

[0019] Here, the machine learning device 20 comprises a state quantity observation unit 21, an operation result acquisition unit 26, a learning unit 22, and a decision-making unit 25. As will be described in detail later, the machine learning device 20 learns and outputs manipulated quantities such as command data that instructs the robot 14 to perform the workpiece removal operation, or measurement parameters of the three-dimensional measuring instrument 15.

[0020] The robot 14 is, for example, a 6-axis articulated robot, and the drive axes of the robot 14 and the hand unit 13 are controlled by the control device 16. The robot 14 is also used to pick up workpieces 12 one by one from a box 11 placed in a predetermined location and move them sequentially to a designated location, for example, a conveyor or workbench (not shown).

[0021] Incidentally, when removing the loosely stacked workpieces 12 from the box 11, the hand unit 13 or the workpiece 12 may collide with or come into contact with the wall of the box 11. Alternatively, the hand unit 13 or the workpiece 12 may get caught on another workpiece 12. In such cases, a function to detect the force acting on the hand unit 13 is necessary to immediately avoid overloading the robot 14. For this reason, a 6-axis force sensor 17 is provided between the tip of the arm of the robot 14 and the hand unit 13. Furthermore, the robot system 10 of this embodiment also has a function to estimate the force acting on the hand unit 13 based on the current value of the motor (not shown) that drives the drive axis of each joint of the robot 14.

[0022] Furthermore, since the force sensor 17 can detect the force acting on the hand unit 13, it can also determine whether the hand unit 13 is actually gripping the workpiece 12. In other words, when the hand unit 13 grips the workpiece 12, the weight of the workpiece 12 acts on the hand unit 13. Therefore, after performing the workpiece removal operation, if the detected value of the force sensor 17 exceeds a predetermined threshold, it can be determined that the hand unit 13 is gripping the workpiece 12. The determination of whether the hand unit 13 is gripping the workpiece 12 can also be made, for example, by capturing data from a camera used in the three-dimensional measuring instrument 15 or by using the output of a photoelectric sensor (not shown) attached to the hand unit 13. Alternatively, the determination may be made based on the data from the pressure gauge of the suction-type hand described later.

[0023] Here, the hand portion 13 may have various forms as long as it is capable of holding the workpiece 12. For example, the hand portion 13 may be in a form that grips the workpiece 12 by opening and closing two or more claw portions, or it may be equipped with an electromagnet or negative pressure generating device that generates an attractive force on the workpiece 12. In other words, although the hand portion 13 is depicted in Figure 1 as gripping the workpiece with two claw portions, it goes without saying that it is not limited to this.

[0024] The three-dimensional measuring instrument 15 is positioned above multiple workpieces 12 by a support 18 in order to measure multiple workpieces 12. As the three-dimensional measuring instrument 15, for example, a three-dimensional vision sensor can be used to acquire three-dimensional position information by processing image data of the workpieces 12 captured by two cameras (not shown). Specifically, a three-dimensional map (the positions of the surfaces of multiple randomly stacked workpieces 12) is measured by applying methods such as triangulation, light section method, time-of-flight method, depth from defocus method, or a combination thereof.

[0025] The coordinate calculation unit 19 takes the three-dimensional map obtained from the three-dimensional measuring instrument 15 as input and calculates (measures) the positions of the surfaces of the multiple randomly stacked workpieces 12. That is, by using the output of the three-dimensional measuring instrument 15, it is possible to obtain three-dimensional position data (x,y,z), or three-dimensional position data (x,y,z) and orientation data (w,p,r) for each workpiece 12. Here, the state quantity observation unit 21 receives both the three-dimensional map from the three-dimensional measuring instrument 15 and the position data (or orientation data) from the coordinate calculation unit 19 to observe the state quantities of the robot 14. However, for example, it is also possible to observe the state quantities of the robot 14 by receiving only the three-dimensional map from the three-dimensional measuring instrument 15. Furthermore, as will be explained later with reference to Figure 5, it is also possible to add a preprocessing unit 50, which processes (preprocesses) the three-dimensional map from the three-dimensional measuring instrument 15 before inputting it to the state quantity observation unit 21.

[0026] The correlation position between the robot 14 and the three-dimensional measuring instrument 15 is assumed to be determined in advance by calibration. In addition, a laser distance measuring instrument can be used instead of a three-dimensional vision sensor in the three-dimensional measuring instrument 15 of the present invention. That is, the three-dimensional position data and orientation (x, y, z, w, p, r) of multiple randomly stacked workpieces 12 may be obtained by measuring the distance from the position where the three-dimensional measuring instrument 15 is installed to the surface of each workpiece 12 by laser scanning, or by using various sensors such as a monocular camera and a tactile sensor.

[0027] In other words, in this invention, any three-dimensional measuring instrument 15 that employs any three-dimensional measurement method can be used, as long as it can acquire data (x, y, z, w, p, r) for each workpiece 12. Furthermore, the manner in which the three-dimensional measuring instrument 15 is installed is not particularly limited; for example, it may be fixed to the floor or wall, or it may be attached to the arm of a robot 14.

[0028] The three-dimensional measuring instrument 15 acquires a three-dimensional map of multiple workpieces 12 randomly stacked in the box 11 based on a command from the control device 16. The coordinate calculation unit 19 then acquires (calculates) the three-dimensional position (orientation) data of the multiple workpieces 12 based on this three-dimensional map, and outputs this data to the control device 16 and the state quantity observation unit 21 and operation result acquisition unit 26 of the machine learning device 20, which will be described later. In particular, the coordinate calculation unit 19 estimates, for example, the boundary between one workpiece 12 and another workpiece 12, and the boundary between a workpiece 12 and the box 11, based on the image data of the multiple workpieces 12 that have been photographed, and acquires the three-dimensional position data for each workpiece 12.

[0029] The three-dimensional position data for each workpiece 12 refers to data obtained, for example, by estimating the location and achievable position of each workpiece 12 from the positions of multiple points on the surface of multiple workpieces 12 that are stacked randomly. Of course, the three-dimensional position data for each workpiece 12 may also include data on the orientation of the workpiece 12.

[0030] Furthermore, the acquisition of three-dimensional position and orientation data for each workpiece 12 in the coordinate calculation unit 19 may also involve the use of machine learning techniques. For example, it is possible to apply object recognition and angle estimation from input images or laser rangefinders using supervised learning techniques, as described later.

[0031] Then, when the three-dimensional position data for each workpiece 12 is input from the three-dimensional measuring instrument 15 to the control device 16 via the coordinate calculation unit 19, the control device 16 controls the operation of the hand unit 13 to pick up a certain workpiece 12 from the box 11. At this time, the motors (not shown) of each axis of the hand unit 13 and the robot 14 are driven based on command values ​​(operated variables) corresponding to the optimal position, orientation, and picking direction of the hand unit 13 obtained by the machine learning device 20, which will be described later.

[0032] Furthermore, the machine learning device 20 can learn the variables of the shooting conditions of the camera used in the three-dimensional measuring instrument 15 (measurement parameters of the three-dimensional measuring instrument 15: for example, the exposure time adjusted during shooting using a light meter, the illuminance of the lighting system that illuminates the object to be photographed, etc.), and control the three-dimensional measuring instrument 15 via the control device 16 based on the learned measurement parameter manipulation amounts. Here, the position and posture estimation condition variables that the three-dimensional measuring instrument 15 uses to estimate the existing position and posture and the holdable position and posture of each workpiece 12 from the positions of multiple workpieces 12 measured may be included in the output data of the three-dimensional measuring instrument 15 as described above.

[0033] Furthermore, as mentioned above, the output data from the three-dimensional measuring instrument 15 can be pre-processed by the pre-processing unit 50, etc., which will be described in detail later with reference to Figure 5, and the processed data (image data) can be provided to the state quantity observation unit 21. The operation result acquisition unit 26 can, for example, acquire the result of the robot 14's hand unit 13 picking up the workpiece 12 from the output data from the three-dimensional measuring instrument 15 (output data from the coordinate calculation unit 19). In addition, it goes without saying that operation results such as the degree of achievement when the picked-up workpiece 12 is passed to the next process, and whether there are any changes in the state of the picked-up workpiece 12 such as damage, can also be acquired via other means (for example, a camera or sensor installed in the next process). In summary, the state quantity observation unit 21 and the operation result acquisition unit 26 are functional blocks, and it is of course possible to consider them as a single block that achieves both functions.

[0034] Next, we will describe in detail the machine learning device 20 shown in Figure 1. The machine learning device 20 has the function of extracting useful rules, knowledge representations, and judgment criteria from a set of data input to the device through analysis, outputting the judgment results, and performing knowledge learning (machine learning). There are various machine learning methods, but they can be broadly divided into, for example, "supervised learning," "unsupervised learning," and "reinforcement learning." Furthermore, in order to realize these methods, there is a method called "deep learning," which learns the extraction of features themselves. These machine learning processes (machine learning device 20) can be performed using general-purpose computers or processors, but processing can be done at a faster speed by applying GPGPU (General-Purpose computing on Graphics Processing Units) or large-scale PC clusters.

[0035] First, supervised learning involves feeding a large amount of data pairs of inputs and results (labels) to a machine learning device 20, learning the features in those datasets, and inductively acquiring a model that estimates results from inputs, i.e., the relationship between them. When this supervised learning is applied to this embodiment, it can be used, for example, in the part that estimates the workpiece position from sensor input, or in the part that estimates the success probability of acquiring a candidate workpiece. For example, it can be implemented using algorithms such as neural networks, which will be described later.

[0036] Unsupervised learning is a method in which a learning system learns the distribution of input data by feeding it a large amount of data, and then learns how to compress, classify, and shape the input data without providing corresponding supervised output data. For example, it can be used to cluster similar features in these datasets. Using these results, it is possible to predict the output by assigning outputs that optimize based on certain criteria.

[0037] Furthermore, there is a problem setting called semi-supervised learning, which is an intermediate between unsupervised and supervised learning. This corresponds, for example, to cases where only some input and output data sets exist, and the rest are input-only data. In this embodiment, learning can be performed efficiently by using data that can be obtained without actually moving the robot (such as image data or simulation data) in unsupervised learning.

[0038] Next, I will explain reinforcement learning. First, let's consider the problem setting for reinforcement learning as follows. The robot observes the state of the environment and decides on its actions. The environment changes according to certain rules, and furthermore, our own actions can also cause changes in the environment. Each time you take action, you receive a reward signal. The goal is to maximize the total (discount) rewards received over time. Learning begins with a state of having no knowledge, or only partial knowledge, of the consequences of an action. In other words, a robot can only obtain data on the results after actually taking action. This means that it needs to explore the optimal action through trial and error. • By using pre-trained data (such as the supervised learning or inverse reinforcement learning methods mentioned earlier) as the initial state, similar to how human actions are mimicked, learning can be started from a favorable starting point.

[0039] Here, reinforcement learning is not just about judgment and classification, but also about learning actions, learning appropriate actions based on the interaction of those actions with the environment, that is, learning how to maximize future rewards. In this embodiment, this means that, for example, it is possible to acquire actions that have an impact on the future, such as making it easier to obtain work 12 in the future by clearing the pile of work 12. The following explanation will continue with the example of Q-learning, but it is not limited to Q-learning.

[0040] Q-learning is a method of learning the value Q(s, a) for selecting an action a under a certain environmental state s. That is, when in a certain state s, the action a with the highest value of Q(s, a) should be selected as the optimal action. However, initially, for the combination of state s and action a, the correct value of Q(s, a) is completely unknown. Therefore, the agent (the acting entity) selects various actions a under a certain state s, and a reward is given for the action a at that time. Thereby, the agent learns to select better actions, that is, to learn the correct value of Q(s, a).

[0041] Furthermore, since it is desired to maximize the total reward obtained over time as a result of the action, ultimately, we aim to make Q(s, a) = E[Σ(γ t )r t . Here, E[] represents the expected value, t is the time, γ is a parameter called the discount rate described later, r t is the reward at time t, and Σ is the sum over time t. The expected value in this formula is taken when the state changes according to the optimal action, and since that is unknown, learning will be done while exploring. Such an update formula for the value Q(s, a) can be represented, for example, by the following formula (1).

[0042] [Number] In the above formula (1), s t represents the state of the environment at time t, and a t represents the action at time t. By the action a t , the state changes to s t+1 . r t+1 represents the reward obtained due to that state change. Also, the term with max is the product of γ and the Q-value when the action a with the highest Q-value known at that time is selected under the state s t+1 . Here, γ is a parameter where 0 < γ ≤ 1 and is called the discount rate. Also, α is the learning coefficient and is in the range 0 < α ≤ 1.

[0043] The formula (1) described above is for the trial at As a result, the reward r t+1 Based on this, state s t Action a in t Evaluation value Q(s t ,a t This represents how to update the evaluation value Q(s) of action a in state s. t ,a t ) is better than reward r t+1 Q(s) is the evaluation value of the best action max a in the following state based on action a. t+1 ,max a t+1 If the sum of ) is greater, then Q(s t ,a t If Q(s) is increased, and conversely if it is decreased, t ,a t This indicates that the value of a certain action in a given state is brought closer to the value of the best action in the next state resulting from that action, and the immediate reward that comes as a result.

[0044] Here, there are two ways to represent Q(s,a) on a computer: one is to store its value in a table for all state-action pairs (s,a), and the other is to prepare a function that approximates Q(s,a). In the latter method, equation (1) mentioned above can be realized by adjusting the parameters of the approximation function using methods such as stochastic gradient descent. A neural network, as described later, can be used as the approximation function.

[0045] Furthermore, neural networks can be used as learning models for supervised learning and unsupervised learning, or as approximation algorithms for value functions in reinforcement learning. Figure 2 is a schematic diagram of a neuron model, and Figure 3 is a schematic diagram of a three-layer neural network constructed by combining the neurons shown in Figure 2. In other words, a neural network is composed of a computing unit and memory that mimics a neuron model, such as the one shown in Figure 2.

[0046] As shown in Figure 2, a neuron outputs an output (result) y for multiple inputs x (in Figure 2, inputs x1 to x3 as an example). Each input x (x1, x2, x3) is multiplied by a weight w (w1, w2, w3) corresponding to that input x. As a result, the neuron outputs a result y expressed by the following equation (2). Note that the inputs x, result y, and weights w are all vectors. Also, in equation (2) below, θ is the bias, and f k This is the activation function.

number

[0047] Referring to Figure 3, we will explain the three-layer neural network constructed by combining the neurons shown in Figure 2. As shown in Figure 3, multiple inputs x (here, for example, inputs x1 to x3) are input from the left side of the neural network, and results y (here, for example, results y1 to input y3) are output from the right side. Specifically, inputs x1, x2, and x3 are input to each of the three neurons N11 to N13, each multiplied by the corresponding weight. These weights multiplied by these inputs are collectively denoted as W1.

[0048] Neurons N11 to N13 output z11 to z13, respectively. In Figure 3, these z11 to z13 are collectively denoted as feature vector Z1 and can be considered as a vector in which the features of the input vector are extracted. This feature vector Z1 is the feature vector between weights W1 and W2. For each of the two neurons N21 and N22, z11 to z13 are input multiplied by their corresponding weights. The weights multiplied by these feature vectors are collectively denoted as W2.

[0049] Neurons N21 and N22 output z21 and z22, respectively. In Figure 3, these z21 and z22 are collectively labeled as feature vector Z2. This feature vector Z2 is the feature vector between weights W2 and W3. For each of the three neurons N31 to N33, z21 and z22 are input multiplied by their corresponding weights. These weights multiplied by these feature vectors are collectively labeled as W3.

[0050] Finally, neurons N31 to N33 output results y1 to y3, respectively. Neural networks operate in two modes: learning mode and value prediction mode. For example, in learning mode, weights W are learned using a training dataset, and these parameters are used in prediction mode to make decisions about the robot's actions. Although we have written "prediction" for convenience, it goes without saying that a variety of tasks such as detection, classification, and inference are possible.

[0051] Here, it's possible to either immediately learn from data obtained by actually moving the robot in prediction mode and reflect that in the next action (online learning), or to perform batch learning using a pre-collected dataset and then continuously use those parameters in detection mode (batch learning). Alternatively, an intermediate approach is possible, where learning mode is interspersed after a certain amount of data has been accumulated.

[0052] Furthermore, weights W1 to W3 can be learned using backpropagation. Note that error information enters from the right and flows to the left. Backpropagation is a method that adjusts (learns) the weights of each neuron to minimize the difference between the output y when input x is received and the true output y (teacher).

[0053] Such neural networks can have more than three layers, and even more layers can be added (this is called deep learning). Furthermore, it is possible to automatically acquire a computational unit that performs stepwise feature extraction from the input and regresses the results, using only the training data.

[0054] Therefore, the machine learning device 20 of this embodiment, in order to perform the above-mentioned Q-learning, includes a state quantity observation unit 21, an operation result acquisition unit 26, a learning unit 22, and a decision-making unit 25, as shown in Figure 1. However, as mentioned above, the machine learning method applied to the present invention is not limited to Q-learning. That is, various methods that can be used in a machine learning device, such as "supervised learning," "unsupervised learning," "semi-supervised learning," and "reinforcement learning," can be applied. In addition, these machine learning methods (machine learning device 20) may use a general-purpose computer or processor, but processing can be done at a faster speed by applying GPGPU or a large-scale PC cluster.

[0055] In other words, according to this embodiment, a machine learning device for learning the operation of a robot 14 that picks up workpieces 12 from a plurality of randomly placed workpieces 12, including a loosely stacked state, using a hand unit 13, comprises: a state quantity observation unit 21 that observes state quantities of the robot 14, including output data from a three-dimensional measuring instrument 15 that measures the three-dimensional position (x,y,z) or the three-dimensional position and orientation (x,y,z,w,p,r) of each workpiece 12; an operation result acquisition unit 26 that acquires the results of the pick-up operation of the robot 14 that picks up the workpieces 12 using the hand unit 13; and a learning unit 22 that receives the output from the state quantity observation unit 21 and the output from the operation result acquisition unit 26, and learns operation quantities, including command data that commands the robot 14 to pick up the workpieces 12, in relation to the state quantities of the robot 14 and the results of the pick-up operation.

[0056] The state variables observed by the state variable observation unit 21 may include, for example, state variables that set the position, orientation, and retrieval direction of the hand unit 13 when a workpiece 12 is retrieved from the box 11. The learned manipulated variables may also include, for example, command values ​​such as torque, speed, and rotational position that are given from the control device 16 to each drive axis of the robot 14 and the hand unit 13 when the workpiece 12 is retrieved from the box 11.

[0057] Then, when the learning unit 22 picks up one of the multiple randomly stacked workpieces 12, it learns the above state variables in relation to the result of the workpiece 12 picking operation (output of the operation result acquisition unit 26). In other words, the control device 16 randomly sets the output data of the three-dimensional measuring instrument 15 (coordinate calculation unit 19) and the command data of the hand unit 13, respectively, or sets them intentionally based on predetermined rules, and the hand unit 13 performs the workpiece 12 picking operation. Here, the predetermined rules include, for example, picking up the workpieces 12 in order from the highest height (z) direction among the multiple randomly stacked workpieces 12. As a result, the output data of the three-dimensional measuring instrument 15 and the command data of the hand unit 13 correspond to the act of picking up a certain workpiece. Then, success or failure occurs in picking up the workpiece 12, and each time such success or failure occurs, the learning unit 22 evaluates the state variables composed of the output data of the three-dimensional measuring instrument 15 and the command data of the hand unit 13.

[0058] Furthermore, the learning unit 22 stores the output data of the three-dimensional measuring instrument 15 and the command data of the hand unit 13 when the workpiece 12 is removed, along with an evaluation of the result of the workpiece removal operation. Examples of failures include cases where the hand unit 13 fails to grasp the workpiece 12, or where, even if the workpiece 12 is grasped, the workpiece 12 collides with or comes into contact with the wall of the box 11. The success or failure of such workpiece removal is determined based on the detection value of the force sensor 17 and the image data captured by the three-dimensional measuring instrument. Here, the machine learning device 20 can also perform learning using, for example, a portion of the command data of the hand unit 13 output from the control device 16.

[0059] In this embodiment, the learning unit 22 preferably includes a reward calculation unit 23 and a value function update unit 24. For example, the reward calculation unit 23 calculates a reward, such as a score, based on the success or failure of retrieving the workpiece 12 due to the state variables mentioned above. The reward is set to be higher for successful retrieval of the workpiece 12 and lower for unsuccessful retrieval of the workpiece 12. Alternatively, the reward may be calculated based on the number of times the workpiece 12 is successfully retrieved within a predetermined time. Furthermore, when calculating this reward, the reward may be calculated according to each stage of retrieval of the workpiece 12, such as successful gripping by the hand unit 13, successful transport by the hand unit 13, and successful placement of the workpiece 12.

[0060] The value function update unit 24 has a value function that determines the value of the workpiece retrieval operation, and updates the value function according to the reward described above. The update formula for value Q(s,a) described above is used to update this value function. Furthermore, it is preferable to create an action value table when this update is performed. The action value table referred to here is a record that associates the output data of the three-dimensional measuring instrument 15 and the command data of the hand unit 13 when the workpiece 12 is retrieved with the value function (i.e., evaluation value) updated according to the result of retrieving the workpiece 12 at that time.

[0061] Furthermore, it is possible to use a function approximated using the aforementioned neural network as this action value table, which is particularly effective when the amount of information about state s is enormous, such as in image data. Also, the above value function is not limited to one type. For example, a value function that evaluates the success or failure of gripping the workpiece 12 by the hand unit 13, or a value function that evaluates the time (cycle time) required to grip and transport the workpiece 12 by the hand unit 13 can be considered.

[0062] Furthermore, as the value function described above, a value function that evaluates the interference between the box 11 and the hand unit 13 or the workpiece 12 during workpiece removal may be used. In order to calculate the reward used to update this value function, it is preferable for the state quantity observation unit 21 to observe the force acting on the hand unit 13, for example, the value detected by the force sensor 17. When the amount of change in the force detected by the force sensor 17 exceeds a predetermined threshold, it can be estimated that the above interference has occurred, and in that case, it is preferable to set the reward to a negative value, for example, so that the value defined by the value function becomes lower.

[0063] Furthermore, according to this embodiment, it is also possible to learn the measurement parameters of the three-dimensional measuring instrument 15 as manipulated variables. That is, according to this embodiment, a machine learning device for learning the operation of a robot 14 that picks up workpieces 12 from a plurality of randomly placed workpieces 12, including a stacked state, using a hand unit 13, comprises: a state quantity observation unit 21 that observes the state quantities of the robot 14, including the output data of a three-dimensional measuring instrument 15 that measures the three-dimensional position (x,y,z) or the three-dimensional position and orientation (x,y,z,w,p,r) of each workpiece 12; an operation result acquisition unit 26 that acquires the results of the pick-up operation of the robot 14 that picks up the workpieces 12 using the hand unit 13; and a learning unit 22 that receives the output from the state quantity observation unit 21 and the output from the operation result acquisition unit 26, and learns manipulated variables, including the measurement parameters of the three-dimensional measuring instrument 15, in relation to the state quantities of the robot 14 and the results of the pick-up operation.

[0064] Furthermore, the robot system 10 of this embodiment may be equipped with an automatic hand replacement device (not shown) that replaces the hand unit 13 attached to the robot 14 with a hand unit 13 of a different form. In that case, the value function update unit 24 may have the above value function for each hand unit 13 of a different form, and update the value function of the replaced hand unit 13 according to the reward. As a result, the optimal operation of the hand unit 13 can be learned for each of the multiple hands 13 of different forms, making it possible to have the automatic hand replacement device select the hand unit 13 with a higher value function.

[0065] Next, the decision-making unit 25 preferably refers to the action value table created as described above and selects the output data of the three-dimensional measuring instrument 15 and the command data of the hand unit 13 that correspond to the highest evaluation value. After that, the decision-making unit 25 outputs the optimal data of the selected hand unit 13 and three-dimensional measuring instrument 15 to the control device 16.

[0066] The control device 16 then uses the optimal data from the hand unit 13 and the three-dimensional measuring instrument 15 output by the learning unit 22 to control the three-dimensional measuring instrument 15 and the robot 14, respectively, to retrieve the workpiece 12. For example, it is preferable that the control device 16 operates the drive axes of the hand unit 13 and the robot 14 based on state variables that set the optimal position, orientation, and retrieval direction of the hand unit 13, respectively, obtained by the learning unit 22.

[0067] In the above-described embodiment, the robot system 10 is equipped with one machine learning device 20 for each robot 14, as shown in Figure 1. However, in the present invention, the number of robots 14 and machine learning devices 20 is not limited to one. For example, the robot system 10 may be equipped with multiple robots 14, and one or more machine learning devices 20 may be provided corresponding to each robot 14. Furthermore, it is preferable that the robot system 10 shares or exchanges the optimal state variables of the three-dimensional measuring instrument 15 and the hand unit 13 acquired by the machine learning device 20 of each robot 14 via a communication medium such as a network. This allows the optimal operation results acquired by the machine learning device 20 of another robot 14 to be used for the operation of that robot 14, even if the operating rate of one robot 14 is lower than that of another robot 14. In addition, the time required for learning can be shortened by sharing learning models among multiple robots, or by sharing the manipulated variables including the measurement parameters of the three-dimensional measuring instrument 15, the state variables of the robot 14, and the results of the retrieval operation.

[0068] Furthermore, the machine learning device 20 may be located inside or outside the robot 14. Alternatively, the machine learning device 20 may be located inside the control device 16, or it may be located on a cloud server (not shown).

[0069] Furthermore, if the robot system 10 includes multiple robots 14, it is possible to have another robot 14's hand unit perform the task of picking up the workpiece 12 while one robot 14 is transporting the workpiece 12 grasped by its hand unit 13. The value function update unit 24 can also update the value function during the time it takes for the robot 14 to switch which robot is responsible for picking up the workpiece 12. In addition, the machine learning device 20 can have state variables for multiple hand models, perform pick-up simulations with multiple hand models during the workpiece 12 pick-up operation, and learn the state variables of the multiple hand models in relation to the results of the workpiece 12 pick-up operation based on the results of these simulations.

[0070] In the machine learning device 20 described above, the output data from the three-dimensional measuring instrument 15 when acquiring three-dimensional map data for each workpiece 12 is transmitted from the three-dimensional measuring instrument 15 to the state quantity observation unit 21. Since such transmitted data may contain abnormal data, the machine learning device 20 can be equipped with an abnormal data filtering function, that is, a function that allows the user to select whether or not to input data from the three-dimensional measuring instrument 15 to the state quantity observation unit 21. This enables the learning unit 22 of the machine learning device 20 to efficiently learn the optimal operation of the hand unit 13 by the three-dimensional measuring instrument 15 and the robot 14.

[0071] Furthermore, in the machine learning device 20 described above, the control device 16 receives output data from the learning unit 22, but it is not guaranteed that the output data from the learning unit 22 does not contain abnormal data. Therefore, the control device 16 may be provided with an abnormal data filtering function, that is, a function that allows it to select whether or not to output data from the learning unit 22 to the control device 16. This enables the control device 16 to allow the robot 14 to perform optimal movements of the hand unit 13 more safely.

[0072] Furthermore, the aforementioned anomalous data can be detected by the following procedure: estimate the probability distribution of the input data, use the probability distribution to derive the probability of a new input occurring, and if the probability of occurrence is below a certain level, it is considered anomalous data that deviates significantly from typical behavior.

[0073] Next, an example of the operation of the machine learning device 20 provided in the robot system 10 of this embodiment will be described. Figure 4 is a flowchart showing an example of the operation of the machine learning device shown in Figure 1. As shown in Figure 4, in the machine learning device 20 shown in Figure 1, when the learning operation (learning process) starts, the three-dimensional measuring instrument 15 performs three-dimensional measurement and outputs the results (step S11 in Figure 4). That is, in step S11, for example, a three-dimensional map (output data of the three-dimensional measuring instrument 15) of each randomly placed workpiece 12, including a loosely stacked state, is acquired and output to the state quantity observation unit 21, and the coordinate calculation unit 19 receives the three-dimensional map of each workpiece 12, calculates the three-dimensional position (x, y, z) of each workpiece 12, and outputs it to the state quantity observation unit 21, the operation result acquisition unit 26, and the control device 16. Here, the coordinate calculation unit 19 may also calculate and output the orientation (w, p, r) of each workpiece 12 from the output of the three-dimensional measuring instrument 15.

[0074] As explained with reference to Figure 5, the output (three-dimensional map) of the three-dimensional measuring instrument 15 may be input to the state quantity observation unit 21 via a pre-processing unit 50 that processes the output before it is input to the state quantity observation unit 21. Also, as explained with reference to Figure 7, only the output of the three-dimensional measuring instrument 15 may be input to the state quantity observation unit 21, or only the output of the three-dimensional measuring instrument 15 may be input to the state quantity observation unit 21 via the pre-processing unit 50. Thus, the implementation and output of the three-dimensional measurement in step S11 can include a variety of things.

[0075] Specifically, in the case of Figure 1, the state quantity observation unit 21 observes state quantities (output data from the three-dimensional measuring instrument 15) such as a three-dimensional map of each workpiece 12 from the three-dimensional measuring instrument 15, and the three-dimensional position (x, y, z) and orientation (w, p, r) of each workpiece 12 from the coordinate calculation unit 19. The operation result acquisition unit 26 acquires the results of the robot 14's pick-up operation, in which the hand unit 13 picks up the workpiece 12, based on the output data from the three-dimensional measuring instrument 15 (output data from the coordinate calculation unit 19). In addition to the output data from the three-dimensional measuring instrument, the operation result acquisition unit 26 can also acquire results of the pick-up operation such as the degree of achievement when the picked-up workpiece 12 is handed over to the next process, or damage to the picked-up workpiece 12.

[0076] Furthermore, for example, the machine learning device 20 determines the optimal operation based on the output data of the three-dimensional measuring instrument 15 (step S12 in Figure 4), and the control device 16 outputs command data (operation amount) for the hand unit 13 (robot 14) to perform the workpiece retrieval operation (step S13 in Figure 4). The result of the workpiece retrieval is then acquired by the operation result acquisition unit 26 described above (step S14 in Figure 4).

[0077] Next, the success or failure of retrieving the workpiece 12 is determined based on the output from the operation result acquisition unit 26 (step S15 in Figure 4). If the retrieval of workpiece 12 is successful, a positive reward is set (step S16 in Figure 4). If the retrieval of workpiece 12 fails, a negative reward is set (step S17 in Figure 4). Then, the action value table (value function) is updated (step S18 in Figure 4).

[0078] Here, the success or failure of removing the workpiece 12 can be determined, for example, based on the output data of the three-dimensional measuring instrument 15 after the workpiece 12 removal operation. Furthermore, the success or failure of removing the workpiece 12 is not limited to evaluating the success or failure of removing the workpiece 12, but may also be evaluated based on, for example, the degree of achievement when the removed workpiece 12 is handed over to the next process, whether there are any changes in the condition of the removed workpiece 12 such as damage, or the time (cycle time) and energy (electrical energy) required to grasp and transport the workpiece 12 with the hand unit 13.

[0079] The reward value is calculated based on the success or failure of retrieving work 12 by the reward calculation unit 23, and the action value table is updated by the value function update unit 24. Specifically, when the learning unit 22 successfully retrieves work 12, it sets a positive reward in the reward in the aforementioned value Q(s,a) update formula (S16), and when it fails to retrieve work 12, it sets a negative reward in the reward in the update formula (S17). The learning unit 22 then updates the aforementioned action value table each time work 12 is retrieved (S18). By repeating the above steps S11 to S18, the learning unit 22 continues to update (learn) the action value table.

[0080] In the above, the data input to the state quantity observation unit 21 is not limited to the output data of the three-dimensional measuring instrument 15, but may also include, for example, data from the output of other sensors, and it is also possible to use a portion of the command data from the control device 16. In this way, the control device 16 uses the command data (operated variables) output from the machine learning device 20 to cause the robot 14 to perform the workpiece removal operation. As mentioned above, the learning by the machine learning device 20 is not limited to the workpiece removal operation, but may also include, for example, the measurement parameters of the three-dimensional measuring instrument 15.

[0081] As described above, the robot system 10 equipped with the machine learning device 20 of this embodiment can learn the operation of the robot 14 to pick up workpieces 12 from multiple workpieces 12 that are placed haphazardly, including in a loosely stacked state, using the hand unit 13. As a result, the robot system 10 can learn to select the optimal operation for the robot 14 to pick up loosely stacked workpieces 12 without human intervention.

[0082] Figure 5 is a block diagram showing the conceptual configuration of a robot system according to another embodiment of the present invention, and shows a robot system to which supervised learning is applied. As is clear from comparing Figure 5 with Figure 1 described above, the robot system 10' to which supervised learning is applied shown in Figure 5 is further equipped with a result (labeled) data recording unit 40 compared to the robot system 10 to which Q-learning (reinforcement learning) is applied shown in Figure 1. Furthermore, the robot system 10' shown in Figure 5 is equipped with a pre-processing unit 50 for pre-processing the output data of the three-dimensional measuring instrument 15. It goes without saying that the pre-processing unit 50 can also be provided in, for example, the robot system 10 shown in Figure 1.

[0083] As shown in Figure 5, the machine learning device 30 in the robot system 10' to which supervised learning is applied comprises a state quantity observation unit 31, an operation result acquisition unit 36, a learning unit 32, and a decision-making unit 35. The learning unit 32 includes an error calculation unit 33 and a learning model update unit 34. In this embodiment of the robot system 10', the machine learning device 30 learns and outputs manipulated quantities such as command data that instructs the robot 14 to perform a workpiece 12 removal operation, or measurement parameters of the three-dimensional measuring instrument 15.

[0084] In other words, in the robot system 10' to which supervised learning is applied as shown in Figure 5, the error calculation unit 33 and the learning model update unit 34 correspond to the reward calculation unit 23 and the value function update unit 24 in the robot system 10 to which Q-learning is applied as shown in Figure 1, respectively. The configurations of other components, such as the three-dimensional measuring instrument 15, the control device 16, and the robot 14, are the same as those described in Figure 1 above, and their explanation is omitted.

[0085] The error calculation unit 33 calculates the error between the result (label) output from the operation result acquisition unit 36 ​​and the output of the learning model implemented in the learning unit. Here, the result (labeled) data recording unit 40 can, for example, if the shape of the workpiece 12 and the processing by the robot 14 are the same, store the result (labeled) data obtained up to the day before the predetermined day on which the robot 14 is to perform the work, and on that predetermined day, provide the result (labeled) data stored in the result (labeled) data recording unit 40 to the error calculation unit 33. Alternatively, it is also possible to provide data obtained from simulations performed outside the robot system 10', or result (labeled) data from another robot system, to the error calculation unit 33 of that robot system 10' via a memory card or communication line. Furthermore, the result (labeled) data recording unit 40 can be configured with non-volatile memory such as flash memory, and the result (labeled) data recording unit (non-volatile memory) 40 can be built into the learning unit 32, allowing the result (labeled) data held in the result (labeled) data recording unit 40 to be used directly by the learning unit 32.

[0086] Figure 6 is a diagram illustrating an example of the preprocessing in the robot system shown in Figure 5. Figure 6(a) shows an example of the three-dimensional position (orientation) data of multiple workpieces 12 stacked loosely in a box 11, i.e., the output data of the three-dimensional measuring instrument 15. Figures 6(b) to 6(d) show examples of image data after preprocessing has been performed on the workpieces 121 to 123 in Figure 6(a).

[0087] Here, the workpiece 12 (121-123) is assumed to be a cylindrical metal part, and the hand (13) is assumed to be a suction pad that uses negative pressure to suck up the longitudinal center of the cylindrical workpiece 12, rather than gripping the workpiece with two claws. Therefore, if the position of the longitudinal center of the workpiece 12 is known, the workpiece 12 can be removed by moving the suction pad (13) to that position and sucking it up. The numerical values ​​in Figures 6(a) to 6(d) are expressed in [mm] and indicate the x, y, and z directions, respectively. The z direction corresponds to the height (depth) direction of the image data captured by a three-dimensional measuring instrument 15 (for example, having two cameras) located above a box 11 in which multiple workpieces 12 are stacked.

[0088] As is clear from the comparison between Figure 6(a) and Figures 6(b) to 6(d), one example of the processing performed by the preprocessing unit 50 in the robot system 10' shown in Figure 5 is to rotate the workpiece 12 of interest (for example, three workpieces 121 to 123) from the output data (three-dimensional image) of the three-dimensional measuring instrument 15, and process it so that the height of the center becomes '0'.

[0089] In other words, the output data of the three-dimensional measuring instrument 15 includes, for example, information on the three-dimensional position (x, y, z) and orientation (w, p, r) of the longitudinal center portion of each workpiece 12. At this time, as shown in Figures 6(b), 6(c), and 6(d), the three workpieces of interest 121, 122, and 123 are rotated by -r and z is subtracted from each to bring them all under the same conditions. By performing such preprocessing, it is possible to reduce the load on the machine learning device 30.

[0090] Here, the three-dimensional image shown in Figure 6(a) is not the output data of the three-dimensional measuring instrument 15 itself, but rather, for example, an image obtained by a program that defines the order in which the workpieces 12 are picked up, with a lower threshold for selection. This processing itself can also be performed by the pre-processing unit 50. Needless to say, the processing performed by the pre-processing unit 50 can vary in various ways depending on various conditions, including the shape of the workpiece 12 and the type of hand 13.

[0091] Thus, the output data of the three-dimensional measuring instrument 15 (three-dimensional map for each workpiece 12) processed by the preprocessing unit 50 before input to the state quantity observation unit 31 is input to the state quantity observation unit 31. Referring again to Figure 5, the error calculation unit 33, which receives the results (labels) output from the operation result acquisition unit 36, considers, for example, that if the output of the neural network shown in Figure 3 as the learning model is y, there is an error of -log(y) if the workpiece 12 extraction operation was successful, and an error of -log(1-y) if it failed, and processes with the goal of minimizing this error. The input to the neural network shown in Figure 3 may be, for example, image data of the workpieces 121 to 123 of interest after preprocessing as shown in Figures 6(b) to 6(d), as well as data of the three-dimensional position and orientation (x, y, z, w, p, r) for each of the workpieces 121 to 123 of interest.

[0092] Figure 7 is a block diagram showing a modified version of the robot system shown in Figure 1. As is clear from comparing Figure 7 with Figure 1, in the modified version of the robot system 10 shown in Figure 7, the coordinate calculation unit 19 is removed, and the state quantity observation unit 21 receives only a three-dimensional map from the three-dimensional measuring instrument 15 to observe the state quantities of the robot 14. It goes without saying that the control device 16 can be configured to correspond to the coordinate calculation unit 19. Furthermore, the configuration shown in Figure 7 can also be applied to the robot system 10' to which supervised learning, as explained with reference to Figure 5, is applied. That is, in the robot system 10' shown in Figure 5, it is also possible to remove the preprocessing unit 50 and have the state quantity observation unit 31 receive only a three-dimensional map from the three-dimensional measuring instrument 15 to observe the state quantities of the robot 14. Thus, each of the embodiments described above can be modified and transformed in various ways.

[0093] As detailed above, this embodiment makes it possible to provide a machine learning device, robot system, and machine learning method that can learn the optimal robot operation for picking up randomly placed workpieces, including those in a loosely stacked state, without human intervention. It should be noted that the machine learning devices 20 and 30 in this invention are not limited to those applying reinforcement learning (e.g., Q-learning) or supervised learning, and various machine learning algorithms can be applied.

[0094] While embodiments have been described above, all examples and conditions described herein are provided for the purpose of aiding the understanding of the concept of the invention as applied to the invention and the technology, and are not intended to limit the scope of the invention. Nor are such descriptions in the specification intended to indicate the advantages or disadvantages of the invention. Although embodiments of the invention have been described in detail, it should be understood that various changes, substitutions, and modifications can be made without departing from the spirit and scope of the invention. [Explanation of symbols]

[0095] 10,10' Robot System 11 boxes 12 Work 13 Hand section 14 Robots 15. Three-dimensional measuring instruments 16 Control device 17 Force Sensor 18 Support part 19 Coordinate calculation section 20,30 Machine Learning Devices 21,31 State Variable Observation Unit 22,32 Learning Department 23 Remuneration Calculation Department 24 Value function update section 25,35 Decision making department 26,36 Operation result acquisition part 33 Error calculation section 34. Learning Model Update Unit 40 Data recording unit with result (label) 50 Pre-processing section

Claims

1. The steps include: acquiring image data in which at least one processor includes at least either first image information of a plurality of objects or second image information of a plurality of objects obtained by processing the first image information; The steps include: the at least one processor inputting the image data of the plurality of objects into a neural network to obtain information for grasping one of the plurality of objects; The process includes the step of the at least one processor controlling a gripping unit to grip the object using information for gripping the object, The aforementioned neural network was trained based on data obtained at least from object grasping simulations. The information for grasping the object includes at least one of the following: robot control information, position information of the grasping part, posture information of the grasping part, information regarding the retrieval direction of the grasping part, position information of the object, information regarding the success probability of grasping the object, or control information of a measuring instrument. How to grasp objects.

2. The aforementioned neural network was trained using rewards calculated based on object grasping information. The method for gripping an object according to claim 1.

3. The information on gripping the object includes at least one of the following: whether or not the object was gripped, the number of times the object was successfully gripped, the time required to grip and transport the object, the force applied to the gripping part, the degree of achievement in the subsequent process after gripping the object, the state of the object, or the energy required to grip and transport the object. The method for gripping an object according to claim 2.

4. The aforementioned neural network is a value function in reinforcement learning, The method for gripping an object according to any one of claims 1 to 3.

5. Based on the type of gripping part used to grip the object, the value function is selected from a plurality of value functions. The method for gripping an object according to claim 4.

6. The aforementioned neural network is trained to minimize the error calculated based on the data obtained from the simulation and the output of the neural network. The method for gripping an object according to claim 1.

7. The neural network outputs at least one of the following: location information of the object or information regarding the probability of successfully grasping the object. The method for gripping an object according to claim 6.

8. The system further comprises the step of determining whether the information for grasping the object is abnormal or not. The method for gripping an object according to any one of claims 1 to 7.

9. The first image information includes distance information from the measuring instrument to the surfaces of the plurality of objects. The method for gripping an object according to any one of claims 1 to 8.

10. The second image information includes distance information from the measuring instrument to the surfaces of the plurality of objects. The method for gripping an object according to any one of claims 1 to 9.

11. The measuring instrument is mounted on the arm of a robot that grasps the object. The method for gripping an object according to claim 9 or 10.

12. An object gripping program for causing at least one computer to execute the object gripping method according to any one of claims 1 to 11.

13. At least one memory, It comprises at least one processor, The at least one processor performs the object gripping method according to any one of claims 1 to 11. Object grasping device.

14. The steps include: acquiring image data in which at least one processor includes at least either first image information of a plurality of objects or second image information of a plurality of objects obtained by processing the first image information; The steps include: the at least one processor inputting the image data of the plurality of objects into a neural network, causing the neural network to output information for grasping one of the plurality of objects; The steps include: the at least one processor calculating an error based on the output of the neural network and the label; The steps include: the at least one processor updating the neural network according to the error; Equipped with, The aforementioned labels are data obtained from an object grasping simulation. The information for grasping the object includes at least one of the following: location information of the object, information regarding the probability of successfully grasping the object, information indicating whether or not the object was grasped, or information indicating the result of grasping the object. Learning methods.

15. The first image information includes distance information from the measuring instrument to the surfaces of the plurality of objects. The learning method according to claim 14.

16. The second image information includes distance information from the measuring instrument to the surfaces of the plurality of objects. The learning method according to claim 14 or 15.

17. The measuring instrument is mounted on the arm of a robot that grasps the object. The learning method according to claim 15 or 16.

18. A learning program for causing at least one computer to execute the learning method described in any one of claims 14 to 17.

19. At least one memory, It comprises at least one processor, The at least one processor performs the learning method according to any one of claims 14 to 18. Learning device.