System, server and robot
A machine learning device enhances robot systems by autonomously learning optimal workpiece picking operations, addressing the need for human intervention in conventional systems and improving success rates.
Patent Information
- Application Number
- JP2025102531
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2015-07-31
- Filing Date
- 2025-06-18
- Publication Date
- 2025-09-02
AI Technical Summary
Conventional robot systems require human intervention to set up and program the operation for picking up workpieces from a disorderly pile, leading to decreased success rates if settings or programs are inappropriate.
A machine learning device that observes and learns the optimal operation of a robot to pick up workpieces using a hand unit, incorporating a state quantity observation unit, operation result acquisition unit, and a learning unit that calculates rewards and updates value functions based on the success of the picking operation.
Enables the robot to learn optimal behavior for picking up disorderly placed workpieces without human intervention, improving success rates and reducing the need for trial and error.
Smart Images

Figure 2025128375000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a machine learning device, a robot system, and a machine learning method that learn the operation of picking up workpieces that are placed in a disorderly manner, including in a random pile. [Background technology]
[0002] For example, robot systems have been known in the past that use a robot hand to grasp and transport workpieces randomly stacked in a basket-like box (see, for example, Patent Documents 1 and 2). In such robot systems, position information of multiple workpieces is obtained using, for example, a three-dimensional measuring device installed above the basket-like box, and the workpieces are picked up one by one by the robot hand based on the position information. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Patent No. 5642738 [Patent Document 2] Patent No. 5670397 Summary of the Invention [Problem to be solved by the invention]
[0004] However, in the conventional robot system described above, it is necessary to set in advance, for example, how to extract the workpiece to be picked from a range image of multiple workpieces measured by a three-dimensional measuring device, and the position of the workpiece to be picked. Furthermore, it is also necessary to program in advance how to operate the robot's hand when picking up a workpiece. Specifically, for example, a human must use a teaching pendant to teach the robot how to pick up the workpiece.
[0005] Therefore, if the settings for extracting the workpiece to be picked from the range images of multiple workpieces are inappropriate, or if the robot's operation program is not created appropriately, the success rate of the robot picking up and transporting the workpiece will decrease.In addition, to increase the success rate, humans need to improve the workpiece detection settings and the robot's operation program while searching for the optimal robot operation through trial and error.
[0006] Therefore, in view of the above-described situation, an object of the present invention is to provide a machine learning device, a robot system, and a machine learning method that can learn the optimal operation of a robot when picking up workpieces that are placed in a disorderly manner, including in a loose pile, without human intervention. [Means for solving the problem]
[0007] According to a first configuration example of the first embodiment of the present invention, there is provided a machine learning device that learns the behavior of a robot that uses a hand unit to pick up a plurality of workpieces that are placed in a disorderly manner, including in a loosely piled state, the machine learning device comprising: a state quantity observation unit that observes output data from a three-dimensional measuring device that measures at least a three-dimensional map of each of the workpieces; an operation result acquisition unit that acquires results of the picking operation of the robot that picks up the workpieces using the hand unit; and a learning unit that receives output from the state quantity observation unit and output from the operation result acquisition unit and learns the picking operation of the workpieces, wherein the state quantity observation unit further observes output data from a coordinate calculation unit that calculates the three-dimensional position of each of the workpieces based on the output of the three-dimensional measuring device, and the learning unit comprises: a reward calculation unit that calculates a reward based on the determination result of whether the picking of the workpieces was successful, which is the output of the operation result acquisition unit; and a value function update unit that has a value function that determines the value of the picking operation of the workpieces, and updates the value function in accordance with the reward. According to a second configuration example of the first embodiment of the present invention, there is provided a machine learning device that learns the operation of a robot that uses a hand unit to pick up a workpiece from a plurality of workpieces that are placed in a disorderly manner, including a loosely piled state, the machine learning device comprising: a state quantity observation unit that observes output data from a three-dimensional measuring device that measures at least a three-dimensional map of each of the workpieces; an operation result acquisition unit that acquires the results of the picking operation of the robot that picks up the workpieces using the hand unit; and a learning unit that receives output from the state quantity observation unit and output from the operation result acquisition unit and learns the picking operation of the workpieces, wherein the state quantity observation unit further observes output data from a coordinate calculation unit that calculates the three-dimensional position of each of the workpieces based on the output of the three-dimensional measuring device, and the learning unit has a learning model that learns the picking operation of the workpieces, and comprises: an error calculation unit that calculates an error based on the learning model, which is the output of the operation result acquisition unit, and a learning model update unit that updates the learning model in accordance with the error. It is preferable that the machine learning device further includes a decision-making unit that refers to the output from the learning unit and determines command data to instruct the robot to perform an operation to pick up the workpiece.
[0008] According to a second embodiment of the present invention, there is provided a machine learning device that learns the operation of a robot that uses a hand unit to pick up a workpiece from a plurality of workpieces that are placed in a disorderly manner, including a loose pile, the machine learning device including: a state quantity observation unit that observes state quantities of the robot including output data from a three-dimensional measuring device that measures a three-dimensional map of each of the workpieces; an operation result acquisition unit that acquires a result of the picking operation of the robot that picks up the workpiece using the hand unit; and a learning unit that receives outputs from the state quantity observation unit and the operation result acquisition unit, and learns operation quantities including measurement parameters of the three-dimensional measuring device by associating the operation quantities of the robot and the result of the picking operation. Preferably, the machine learning device further includes a decision-making unit that determines the measurement parameters of the three-dimensional measuring device by referring to the operation quantities learned by the learning unit.
[0009] The state quantity observation unit can further observe state quantities of the robot, including output data from a coordinate calculation unit that calculates the three-dimensional position of each workpiece based on the output of the three-dimensional measuring device. The coordinate calculation unit may further calculate the orientation of each workpiece and output the calculated three-dimensional position and orientation data for each workpiece. The operation result acquisition unit can use the output data of the three-dimensional measuring device. The machine learning device further includes a preprocessing unit that processes the output data of the three-dimensional measuring device before inputting it to the state quantity observation unit, and the state quantity observation unit preferably receives the output data of the preprocessing unit as the state quantity of the robot. The preprocessing unit can align the orientation and height of each workpiece in the output data of the three-dimensional measuring device to a constant value. The operation result acquisition unit can acquire at least one of the success or failure of the workpiece removal, the damage state of the workpiece, and the degree of achievement when transferring the removed workpiece to a subsequent process.
[0010] The learning unit may include a reward calculation unit that calculates a reward based on the output of the operation result acquisition unit, and a value function update unit that has a value function that determines the value of the workpiece pick-up operation and updates the value function in accordance with the reward. The learning unit may also include a learning model that learns the workpiece pick-up operation, an error calculation unit that calculates an error based on the output of the operation result acquisition unit and the output of the learning model, and a learning model update unit that updates the learning model in accordance with the error. The machine learning device preferably includes a neural network.
[0011] According to a third embodiment of the present invention, there is provided a robot system equipped with a machine learning device that learns the operation of a robot that uses a hand unit to pick up a workpiece from a plurality of workpieces that are placed in a disorderly manner, including a loosely piled state, the machine learning device comprising: a state quantity observation unit that observes state quantities of the robot including output data from a three-dimensional measuring device that measures a three-dimensional map of each of the workpieces; an operation result acquisition unit that acquires the result of the picking operation of the robot that picks up the workpiece using the hand unit; and a learning unit that receives output from the state quantity observation unit and output from the operation result acquisition unit, and learns operation quantities that include command data that instruct the robot to perform the picking operation of the workpiece by associating them with the state quantities of the robot and the result of the picking operation, the robot system comprising: the robot; the three-dimensional measuring device; and a control device that controls each of the robot and the three-dimensional measuring device.
[0012] According to a fourth embodiment of the present invention, there is provided a robot system equipped with a machine learning device that learns the operation of a robot that uses a hand unit to pick up a workpiece from a plurality of workpieces that are placed in a disorderly manner, including a loosely piled state, the machine learning device comprising: a state quantity observation unit that observes state quantities of the robot including output data from a three-dimensional measuring device that measures a three-dimensional map of each of the workpieces; an operation result acquisition unit that acquires the result of the picking operation of the robot that picks up the workpiece using the hand unit; and a learning unit that receives output from the state quantity observation unit and output from the operation result acquisition unit, and learns operation quantities including measurement parameters of the three-dimensional measuring device by associating them with the state quantities of the robot and the result of the picking operation, the robot system comprising: the robot; the three-dimensional measuring device; and a control device that controls each of the robot and the three-dimensional measuring device.
[0013] Preferably, the robot system includes a plurality of the robots, each of the robots is provided with a machine learning device, and the machine learning devices provided on the robots share or exchange data with each other via a communication medium. The machine learning devices may reside on a cloud server.
[0014] According to a fifth embodiment of the present invention, there is provided a machine learning method for learning the operation of a robot that uses a hand unit to pick up a workpiece from a plurality of workpieces that are placed in a disorderly manner, including a loosely piled state, the machine learning method comprising: observing state quantities of the robot including output data from a three-dimensional measuring device that measures a three-dimensional map of each of the workpieces; acquiring the results of the robot's picking operation that picks up the workpiece using the hand unit; receiving output from the state quantity observation unit and output from the operation result acquisition unit; and learning an operation quantity that includes command data that instructs the robot to perform the picking operation of the workpiece by associating it with the state quantities of the robot and the result of the picking operation. [Effects of the Invention]
[0015] The machine learning device, robot system, and machine learning method according to the present invention have the effect of being able to learn the optimal robot behavior when picking up workpieces that are placed in a disorderly manner, including those that are piled loose, without human intervention. [Brief explanation of the drawings]
[0016] [Figure 1] FIG. 1 is a block diagram showing a conceptual configuration of a robot system according to an embodiment of the present invention. [Figure 2] FIG. 2 is a diagram showing a schematic diagram of a neuron model. [Figure 3] FIG. 3 is a diagram showing a schematic diagram of a three-layer neural network configured by combining the neurons shown in FIG. [Figure 4] FIG. 4 is a flowchart illustrating an example of the operation of the machine learning device illustrated in FIG. [Figure 5] FIG. 5 is a block diagram showing a conceptual configuration of a robot system according to another embodiment of the present invention. [Figure 6] FIG. 6 is a diagram for explaining an example of processing by the pre-processing unit in the robot system shown in FIG. [Figure 7] FIG. 7 is a block diagram showing a modification of the robot system shown in FIG. DETAILED DESCRIPTION OF THE INVENTION
[0017] Hereinafter, embodiments of a machine learning device, a robot system, and a machine learning method according to the present invention will be described in detail with reference to the accompanying drawings. In each drawing, the same components are designated by the same reference symbols. Also, components designated by the same reference symbols in different drawings have the same functions. The scales of these drawings have been appropriately changed to facilitate understanding.
[0018] 1 is a block diagram showing the conceptual configuration of a robot system according to one embodiment of the present invention. The robot system 10 of this embodiment includes a robot 14 equipped with a hand unit 13 for grasping workpieces 12 randomly stacked in a basket-like box 11, a three-dimensional measuring device 15 for measuring a three-dimensional map of the surface of the workpieces 12, a control device 16 for controlling the robot 14 and the three-dimensional measuring device 15, a coordinate calculation unit 19, and a machine learning device 20.
[0019] Here, the machine learning device 20 includes a state quantity observing unit 21, an operation result acquiring unit 26, a learning unit 22, and a decision making unit 25. As will be described in detail later, the machine learning device 20 learns and outputs operation quantities such as command data for instructing the robot 14 to perform an operation to pick up the workpiece 12 or measurement parameters of the three-dimensional measuring device 15.
[0020] The robot 14 is, for example, a six-axis articulated robot, and the drive axes of the robot 14 and the hand unit 13 are controlled by a control device 16. The robot 14 is used to take out the workpieces 12 one by one from a box 11 installed at a predetermined position and move them sequentially to a specified location, for example, a conveyor or a workbench (not shown).
[0021] Incidentally, when randomly stacked workpieces 12 are removed from the box 11, the hand unit 13 or the workpiece 12 may collide with or come into contact with the wall of the box 11. Alternatively, the hand unit 13 or the workpiece 12 may get caught on another workpiece 12. In such a case, a function for detecting the force acting on the hand unit 13 is required so that an overload on the robot 14 can be immediately avoided. For this reason, a six-axis force sensor 17 is provided between the tip of the arm of the robot 14 and the hand unit 13. The robot system 10 of this embodiment also has a function for estimating the force acting on the hand unit 13 based on the current value of a motor (not shown) that drives the drive shaft of each joint of the robot 14.
[0022] Furthermore, because the force sensor 17 can detect the force acting on the hand unit 13, it can also determine whether the hand unit 13 is actually gripping the workpiece 12. In other words, when the hand unit 13 grips the workpiece 12, the weight of the workpiece 12 acts on the hand unit 13, and therefore, if the detection value of the force sensor 17 exceeds a predetermined threshold value after the workpiece 12 is removed, it can be determined that the hand unit 13 is gripping the workpiece 12. Whether the hand unit 13 is gripping the workpiece 12 can also be determined, for example, from photographic data taken by a camera used in the three-dimensional measuring device 15 or the output of a photoelectric sensor (not shown) attached to the hand unit 13. It may also be determined based on data from a pressure gauge of the suction hand, which will be described later.
[0023] Here, the hand unit 13 may have various forms as long as it is capable of holding the workpiece 12. For example, the hand unit 13 may be of a form that grips the workpiece 12 by opening and closing two or more claws, or may be of a form that includes an electromagnet or a negative pressure generator that generates an attractive force on the workpiece 12. That is, although the hand unit 13 is depicted in FIG. 1 as gripping the workpiece with two claws, it goes without saying that the present invention is not limited to this.
[0024] The three-dimensional measuring device 15 is provided at a predetermined position above the plurality of workpieces 12 by a support part 18 in order to measure the plurality of workpieces 12. As the three-dimensional measuring device 15, for example, a three-dimensional visual sensor that acquires three-dimensional position information by image processing of image data of the workpieces 12 photographed by two cameras (not shown) can be used. Specifically, a three-dimensional map (the positions of the surfaces of the plurality of workpieces 12 stacked in bulk) is measured by applying a triangulation method, a light section method, a time-of-flight method, a depth from defocus method, or a method that combines these methods.
[0025] The coordinate calculation unit 19 receives the three-dimensional map obtained by the three-dimensional measuring device 15 as input and calculates (measures) the surface positions of the multiple workpieces 12 stacked in a random order. That is, using the output of the three-dimensional measuring device 15, it is possible to obtain three-dimensional position data (x, y, z) for each workpiece 12, or the three-dimensional position data (x, y, z) and orientation data (w, p, r) for each workpiece 12. Here, the state quantity observation unit 21 receives both the three-dimensional map from the three-dimensional measuring device 15 and the position data (orientation data) from the coordinate calculation unit 19 to observe the state quantities of the robot 14. However, it is also possible to receive only the three-dimensional map from the three-dimensional measuring device 15 to observe the state quantities of the robot 14. Furthermore, as will be described later with reference to FIG. 5, a preprocessing unit 50 may be added, and the preprocessing unit 50 may process (preprocess) the three-dimensional map from the three-dimensional measuring device 15 before inputting it to the state quantity observation unit 21.
[0026] It is assumed that the relative position between the robot 14 and the three-dimensional measuring device 15 has been determined in advance by calibration. Also, a laser distance measuring device can be used for the three-dimensional measuring device 15 of the present invention instead of a three-dimensional visual sensor. That is, the three-dimensional position data and orientation (x, y, z, w, p, r) of multiple randomly piled workpieces 12 can be obtained by measuring the distance from the position where the three-dimensional measuring device 15 is installed to the surface of each workpiece 12 by laser scanning, or by using various sensors such as a monocular camera or a tactile sensor.
[0027] That is, in the present invention, any three-dimensional measuring device 15 that applies any three-dimensional measuring method can be used as long as it can acquire data (x, y, z, w, p, r) of each workpiece 12. Furthermore, the manner in which the three-dimensional measuring device 15 is installed is not particularly limited, and it may be fixed to the floor or wall, or attached to the arm of the robot 14, for example.
[0028] The three-dimensional measuring device 15 acquires a three-dimensional map of the multiple workpieces 12 randomly piled in the box 11 in response to a command from the control device 16, and the coordinate calculation unit 19 acquires (calculates) data on the three-dimensional positions (postures) of the multiple workpieces 12 based on the three-dimensional map, and outputs the data to the control device 16 and a state quantity observation unit 21 and an operation result acquisition unit 26 of the machine learning device 20, which will be described later. In particular, the coordinate calculation unit 19 estimates, for example, the boundaries between one workpiece 12 and another workpiece 12 and the boundaries between the workpieces 12 and the box 11 based on image data of the multiple photographed workpieces 12, and acquires three-dimensional position data for each workpiece 12.
[0029] The three-dimensional position data for each workpiece 12 refers to data acquired by, for example, estimating the position where each workpiece 12 exists or the position where it can be held from the positions of multiple points on the surface of multiple randomly piled workpieces 12. Of course, the three-dimensional position data for each workpiece 12 may also include data on the posture of the workpiece 12.
[0030] Furthermore, the acquisition of the three-dimensional position and orientation data for each workpiece 12 in the coordinate calculation unit 19 may involve the use of machine learning techniques. For example, it is also possible to apply object recognition and angle estimation from an input image or a laser distance measuring device using techniques such as supervised learning, which will be described later.
[0031] Then, when the three-dimensional position data for each workpiece 12 is input from the three-dimensional measuring device 15 to the control device 16 via the coordinate calculation unit 19, the control device 16 controls the operation of the hand unit 13 to pick up a certain workpiece 12 from the box 11. At this time, motors (not shown) for each axis of the hand unit 13 and the robot 14 are driven based on command values (operation amounts) corresponding to the optimal position, posture, and pick-up direction of the hand unit 13 obtained by the machine learning device 20, which will be described later.
[0032] The machine learning device 20 can also learn variables of the shooting conditions of the camera used in the three-dimensional measuring device 15 (measurement parameters of the three-dimensional measuring device 15: for example, the exposure time adjusted during shooting using an exposure meter, the illuminance of the lighting system that illuminates the object to be photographed, etc.), and control the three-dimensional measuring device 15 via the control device 16 based on the learned measurement parameter manipulation amounts. Here, the variables of the position / posture estimation conditions used by the three-dimensional measuring device 15 to estimate the existing position / posture and the maintainable position / posture of each workpiece 12 from the measured positions of multiple workpieces 12 may be included in the output data of the above-mentioned three-dimensional measuring device 15.
[0033] Furthermore, as mentioned above, the output data from the three-dimensional measuring device 15 can be pre-processed by a pre-processing unit 50, which will be described in detail later with reference to FIG. 5, and the processed data (image data) can be provided to the state quantity observing unit 21. Note that the operation result acquiring unit 26 can acquire, for example, the result of picking up the workpiece 12 by the hand unit 13 of the robot 14 from the output data from the three-dimensional measuring device 15 (output data from the coordinate calculation unit 19). However, it goes without saying that the operation result, such as the degree of accomplishment when the picked up workpiece 12 is delivered to a subsequent process, and whether or not there has been any change in the state of the picked up workpiece 12, such as damage, can also be acquired via other means (for example, a camera or sensor provided in the subsequent process). In the above, the state quantity observing unit 21 and the operation result acquiring unit 26 are functional blocks, and it goes without saying that the functions of both can be achieved by a single block.
[0034] Next, the machine learning device 20 shown in FIG. 1 will be described in detail. The machine learning device 20 analyzes and extracts useful rules, knowledge expressions, and judgment criteria from a set of data input to the device, outputs the judgment results, and has the function of learning knowledge (machine learning). There are various machine learning techniques, which can be broadly categorized into, for example, "supervised learning," "unsupervised learning," and "reinforcement learning." Furthermore, to realize these techniques, there is a technique called "deep learning," which learns to extract features themselves. Note that these machine learning techniques (machine learning device 20) may use a general-purpose computer or processor, but faster processing is possible by applying a GPGPU (General-Purpose Computing on Graphics Processing Units) or a large-scale PC cluster.
[0035] First, supervised learning is a method of providing a large amount of data sets of certain inputs and results (labels) to the machine learning device 20, learning the features of those data sets, and inductively acquiring a model that estimates results from inputs, i.e., the relationships between them. When this supervised learning is applied to this embodiment, it can be used, for example, in the part that estimates the work position from sensor input, or in the part that estimates the probability of successful acquisition of a candidate work. For example, it can be realized using an algorithm such as a neural network, which will be described later.
[0036] Unsupervised learning is a technique in which a learning device is provided with a large amount of input data only, and learns the distribution of the input data. This is done by using a device that compresses, classifies, and formats the input data without providing the corresponding supervised output data. For example, the features in these data sets can be clustered together into similar sets. Using this result, it is possible to set some kind of criterion and assign outputs that optimize it, thereby realizing output prediction.
[0037] Note that there is also something called semi-supervised learning, which is an intermediate problem setting between unsupervised learning and supervised learning, and corresponds to a case where, for example, only some of the data have pairs of input and output, and the rest are input-only data. In this embodiment, learning can be performed efficiently by using data that can be obtained without actually moving the robot (image data, simulation data, etc.) in unsupervised learning.
[0038] Next, we will explain reinforcement learning. First, consider the following problem setting for reinforcement learning. The robot observes the state of the environment and decides its actions. The environment changes according to certain rules, and our own actions can also bring about changes in the environment. Every time you take action, you receive a reward signal. What we want to maximize is the total (discounted) reward over the future. Learning begins with complete or partial knowledge of the consequences of actions. In other words, the robot can only obtain data on the consequences of actions once it has actually taken action. This means that the robot must search for optimal actions through trial and error. - To mimic human behavior, it is possible to start learning from a good starting point by using pre-trained conditions (using methods such as supervised learning and inverse reinforcement learning) as the initial state.
[0039] Here, reinforcement learning refers not only to judgment and classification, but also to learning behaviors, thereby learning appropriate behaviors based on the interactions between the behaviors and the environment, i.e., learning methods for maximizing future rewards. In this embodiment, this means that it is possible to acquire behaviors that will have an impact on the future, such as knocking down a pile of workpieces 12 to make it easier to pick up the workpieces 12 in the future. Below, we will continue to explain the case of Q-learning as an example, but the present invention is not limited to Q-learning.
[0040] Q-learning is a method for learning the value Q(s, a) of selecting action a under a certain environmental state s. In other words, in a certain state s, the action a with the highest value Q(s, a) is selected as the optimal action. However, initially, the correct value Q(s, a) for the combination of state s and action a is completely unknown. Therefore, the agent (subject of action) selects various actions a under a certain state s, and is given a reward for each action a. In this way, the agent learns to select better actions, i.e., the correct value Q(s, a).
[0041] Furthermore, we want to maximize the total rewards we will receive in the future as a result of our actions, so ultimately we have Q(s,a)=E[Σ(γ t )r t ] where E[] represents the expected value, t is the time, γ is a parameter called the discount rate, which will be described later, and r t is the reward at time t, and Σ is the sum at time t. The expected value in this equation is taken when the state changes according to the optimal action, and since this is not known, it is learned through exploration. The update equation for such value Q(s, a) can be expressed, for example, as the following equation (1).
[0042]
number
[0043] The above equation (1) ist As a result, the reward returned is r t+1 Based on the state s t Actions in a t The evaluation value Q(s t ,a t ) is a method for updating the evaluation value Q(s t ,a t ) than reward r t+1 and the evaluation value Q(s t+1 ,max a t+1 ) is greater than the sum of Q(s t ,a t ) is increased, and conversely, if it is small, Q(s t ,a t ) is small. In other words, the value of an action in a state is made to approach the value of the immediate reward that results from that action and the best action in the next state.
[0044] Here, there are two ways to represent Q(s, a) on a computer: one is to store the values for all state-action pairs (s, a) as a table, and the other is to prepare a function that approximates Q(s, a). In the latter method, the above-mentioned formula (1) can be realized by adjusting the parameters of the approximation function using a method such as stochastic gradient descent. Note that a neural network, which will be described later, can be used as the approximation function.
[0045] Furthermore, neural networks can be used as learning models for supervised learning, unsupervised learning, or as approximation algorithms for value functions in reinforcement learning. Figure 2 is a diagram that schematically illustrates a neuron model, and Figure 3 is a diagram that schematically illustrates a three-layer neural network configured by combining the neurons shown in Figure 2. That is, a neural network is configured, for example, with a computing device and memory that mimic the neuron model shown in Figure 2.
[0046] As shown in FIG. 2, a neuron outputs an output (result) y for a plurality of inputs x (in FIG. 2, inputs x1 to x3 are used as an example). Each input x (x1, x2, x3) is multiplied by a weight w (w1, w2, w3) corresponding to this input x. As a result, the neuron outputs a result y expressed by the following equation (2). Note that the input x, result y, and weight w are all vectors. In the following equation (2), θ is a bias, and f k is the activation function.
number
[0047] With reference to Fig. 3, a three-layered neural network constructed by combining the neurons shown in Fig. 2 will be described. As shown in Fig. 3, multiple inputs x (here, as an example, inputs x1 to x3) are input from the left side of the neural network, and results y (here, as an example, results y1 to y3) are output from the right side. Specifically, inputs x1, x2, and x3 are multiplied by corresponding weights and input to each of three neurons N11 to N13. The weights multiplied by these inputs are collectively denoted as W1.
[0048] Neurons N11 to N13 output z11 to z13, respectively. In FIG. 3, z11 to z13 are collectively referred to as feature vector Z1, and can be regarded as a vector obtained by extracting the feature quantity of an input vector. This feature vector Z1 is a feature vector between weights W1 and W2. z11 to z13 are multiplied by the corresponding weights and input to two neurons N21 and N22. The weights multiplied by these feature vectors are collectively referred to as W2.
[0049] Neurons N21 and N22 output z21 and z22, respectively. In FIG. 3, z21 and z22 are collectively labeled as feature vector Z2. This feature vector Z2 is a feature vector between weights W2 and W3. z21 and z22 are multiplied by the corresponding weights and input to each of three neurons N31 to N33. The weights multiplied by these feature vectors are collectively labeled as W3.
[0050] Finally, neurons N31 to N33 output results y1 to y3, respectively. Neural networks operate in a learning mode and a value prediction mode. For example, in learning mode, a weight W is learned using a training data set, and in prediction mode, this parameter is used to determine the robot's behavior. Although we have written prediction for convenience, it goes without saying that a variety of tasks, such as detection, classification, and inference, are possible.
[0051] Here, you can actually operate the robot in prediction mode and immediately learn the data obtained and reflect it in the next action (online learning), or you can perform collective learning using a group of data collected in advance and then continue to use those parameters in detection mode (batch learning).Alternatively, you can do something in between, where you switch to learning mode every time a certain amount of data is accumulated.
[0052] The weights W1 to W3 can be learned using backpropagation. Error information enters from the right side and flows to the left side. Backpropagation is a technique for adjusting (learning) each weight for each neuron so as to reduce the difference between the output y when an input x is input and the true output y (teacher).
[0053] Such neural networks can have three or more layers (known as deep learning). It is also possible to automatically acquire a computing device that performs step-by-step feature extraction of inputs and regresses the results from training data alone.
[0054] Therefore, in order to perform the above-mentioned Q-learning, the machine learning device 20 of this embodiment includes a state quantity observing unit 21, an operation result acquiring unit 26, a learning unit 22, and a decision-making unit 25, as shown in FIG. 1. However, as mentioned above, the machine learning method applied to the present invention is not limited to Q-learning. In other words, various methods that can be used in machine learning devices, such as "supervised learning," "unsupervised learning," "semi-supervised learning," and "reinforcement learning," can be applied. Note that these machine learning methods (machine learning device 20) may use a general-purpose computer or processor, but faster processing is possible by applying a GPGPU, a large-scale PC cluster, or the like.
[0055] That is, according to this embodiment, the machine learning device learns the operation of a robot 14 that uses a hand unit 13 to pick up workpieces 12 from a plurality of workpieces 12 that are placed in a disorderly manner, including a loosely piled state, and includes a state quantity observation unit 21 that observes the state quantities of the robot 14, including output data from a three-dimensional measuring device 15 that measures the three-dimensional position (x, y, z) or the three-dimensional position and orientation (x, y, z, w, p, r) of each workpiece 12, an operation result acquisition unit 26 that acquires the results of the picking operation of the robot 14 that picks up the workpieces 12 using the hand unit 13, and a learning unit 22 that receives the output from the state quantity observation unit 21 and the output from the operation result acquisition unit 26, and learns the operation quantities, including command data that instructs the robot 14 to pick up the workpieces 12, by associating them with the state quantities of the robot 14 and the results of the picking operation.
[0056] The state quantities observed by the state quantity observing unit 21 may include, for example, state variables that respectively set the position, posture, and pick-up direction of the hand unit 13 when a certain workpiece 12 is picked up from the box 11. The operation quantities to be learned may also include, for example, command values such as torque, speed, and rotational position that are given from the control device 16 to each drive axis of the robot 14 and the hand unit 13 when the workpiece 12 is picked up from the box 11.
[0057] When picking up one of the multiple workpieces 12 stacked randomly, the learning unit 22 learns the state variables by associating them with the result of the picking operation of the workpiece 12 (the output of the operation result acquisition unit 26). That is, the control device 16 randomly sets the output data of the three-dimensional measuring device 15 (coordinate calculation unit 19) and the command data of the hand unit 13, or sets them intentionally based on a predetermined rule, and the hand unit 13 performs the picking operation of the workpiece 12. Here, the predetermined rule may be, for example, that, of the multiple workpieces 12 stacked randomly, the workpieces tallest in the height (z) direction are picked up in order. As a result, the output data of the three-dimensional measuring device 15 and the command data of the hand unit 13 correspond to the action of picking up a certain workpiece. Then, success and failure in picking up the workpieces 12 occur, and each time such success and failure occur, the learning unit 22 evaluates the state variables composed of the output data of the three-dimensional measuring device 15 and the command data of the hand unit 13.
[0058] The learning unit 22 also stores the output data of the three-dimensional measuring device 15 and the command data of the hand unit 13 when removing the workpiece 12, in association with an evaluation of the result of the operation to remove the workpiece 12. Examples of failures include when the hand unit 13 is unable to grasp the workpiece 12, or when the hand unit 13 is able to grasp the workpiece 12 but collides with or comes into contact with the wall of the box 11. The success or failure of such removal of the workpiece 12 is determined based on the detection value of the force sensor 17 and the photographic data captured by the three-dimensional measuring device. Here, the machine learning device 20 can also perform learning using, for example, part of the command data of the hand unit 13 output from the control device 16.
[0059] Here, the learning unit 22 of this embodiment preferably includes a reward calculation unit 23 and a value function update unit 24. For example, the reward calculation unit 23 calculates a reward, for example, a score, based on the success or failure of picking up the workpiece 12 due to the above-mentioned state variables. A higher reward is given for successful picking up of the workpiece 12, and a lower reward is given for unsuccessful picking up of the workpiece 12. The reward may also be calculated based on the number of times picking up of the workpiece 12 is successful within a predetermined time. Furthermore, when calculating this reward, the reward may be calculated according to each stage of picking up the workpiece 12, such as successful grasping by the hand unit 13, successful transport by the hand unit 13, successful placing of the workpiece 12, etc.
[0060] The value function update unit 24 has a value function that determines the value of the action of picking up the workpiece 12, and updates the value function in accordance with the above-mentioned reward. The update formula for the value Q(s, a) as described above is used to update this value function. Furthermore, it is preferable to create an action value table during this update. The action value table here refers to a record of the output data of the three-dimensional measuring device 15 and the command data of the hand unit 13 when the workpiece 12 is picked up, and the value function (i.e., the evaluation value) updated in accordance with the picking result of the workpiece 12 at that time, in association with each other.
[0061] It is also possible to use a function approximated by the aforementioned neural network as this action value table, which is particularly effective when the amount of information in the state s is enormous, such as in image data. The above value function is not limited to one type. For example, a value function that evaluates whether the hand unit 13 has successfully grasped the workpiece 12, or a value function that evaluates the time (cycle time) required for the hand unit 13 to grasp and transport the workpiece 12, can be considered.
[0062] Furthermore, a value function that evaluates interference between box 11 and hand unit 13 or workpiece 12 when picking up a workpiece may be used as the above-mentioned value function. In order to calculate the reward used to update this value function, state quantity observing unit 21 preferably observes the force acting on hand unit 13, for example, the value detected by force sensor 17. If the amount of change in force detected by force sensor 17 exceeds a predetermined threshold, it can be assumed that the above-mentioned interference has occurred, and therefore it is preferable to set the reward in this case to, for example, a negative value so that the value determined by the value function becomes lower.
[0063] Furthermore, according to this embodiment, it is also possible to learn the measurement parameters of the three-dimensional measuring device 15 as operation variables. That is, according to this embodiment, the machine learning device learns the operation of the robot 14 that uses the hand unit 13 to pick up workpieces 12 from a plurality of workpieces 12 that are placed in a disorderly manner, including a randomly piled state, and includes a state quantity observing unit 21 that observes state quantities of the robot 14 including output data from the three-dimensional measuring device 15 that measures the three-dimensional position (x, y, z) or the three-dimensional position and orientation (x, y, z, w, p, r) of each workpiece 12, an operation result acquiring unit 26 that acquires the result of the picking operation of the robot 14 that picks up the workpieces 12 using the hand unit 13, and a learning unit 22 that receives the output from the state quantity observing unit 21 and the output from the operation result acquiring unit 26, and learns the operation variables including the measurement parameters of the three-dimensional measuring device 15 by associating them with the state quantities of the robot 14 and the result of the picking operation.
[0064] Furthermore, the robot system 10 of this embodiment may be provided with an automatic hand exchange device (not shown) that exchanges the hand 13 attached to the robot 14 for a hand 13 of a different configuration. In this case, the value function update unit 24 preferably has the above-mentioned value function for each hand 13 of a different configuration, and updates the value function of the exchanged hand 13 in accordance with the reward. This makes it possible to learn the optimal operation of the hand 13 for each of a plurality of hands 13 of different configurations, and allows the automatic hand exchange device to select a hand 13 with a higher value function.
[0065] Next, the decision-making unit 25 preferably refers to the action value table created as described above, for example, and selects the output data of the three-dimensional measuring device 15 and the command data of the hand unit 13 that correspond to the highest evaluation value. Thereafter, the decision-making unit 25 outputs the selected optimal data for the hand unit 13 and three-dimensional measuring device 15 to the control device 16.
[0066] Then, the control device 16 uses the optimal data of the hand unit 13 and the three-dimensional measuring device 15 output by the learning unit 22 to control the three-dimensional measuring device 15 and the robot 14, respectively, to pick up the workpiece 12. For example, the control device 16 preferably operates each drive axis of the hand unit 13 and the robot 14 based on state variables that respectively set the optimal position, posture, and pick-up direction of the hand unit 13 obtained by the learning unit 22.
[0067] As shown in FIG. 1 , the robot system 10 of the above-described embodiment includes one machine learning device 20 for one robot 14. However, in the present invention, the number of each of the robots 14 and the machine learning devices 20 is not limited to one. For example, the robot system 10 may include multiple robots 14, with one or more machine learning devices 20 provided for each robot 14. The robot system 10 preferably shares or mutually exchanges the optimal state variables of the three-dimensional measuring device 15 and the hand unit 13 acquired by the machine learning devices 20 of each robot 14 via a communication medium such as a network. This allows the optimal operation results acquired by the machine learning device 20 of one robot 14 to be used for the operation of the other robot 14, even if the operation rate of one robot 14 is lower than that of another robot 14. Furthermore, by sharing a learning model among multiple robots, or by sharing the operation variables including the measurement parameters of the three-dimensional measuring device 15, the state variables of the robots 14, and the results of the pick-up operation, the time required for learning can be shortened.
[0068] Furthermore, the machine learning device 20 may be located within the robot 14 or external to the robot 14. Alternatively, the machine learning device 20 may be located within the control device 16 or on a cloud server (not shown).
[0069] Furthermore, when the robot system 10 includes multiple robots 14, it is possible for one robot 14 to carry a workpiece 12 grasped by its hand unit 13 while another robot 14 has its hand unit perform the task of picking up the workpiece 12. The value function update unit 24 can then use the time while the robot 14 picking up the workpiece 12 is switched to another robot 14 to update the value function. Furthermore, the machine learning device 20 can have state variables of multiple hand models, perform pick-up simulations with the multiple hand models during the pick-up operation of the workpiece 12, and learn the state variables of the multiple hand models in association with the results of the pick-up operation of the workpiece 12 according to the results of the pick-up simulation.
[0070] In the above-described machine learning device 20, the output data of the three-dimensional measuring device 15 when three-dimensional map data for each workpiece 12 is acquired is transmitted from the three-dimensional measuring device 15 to the state quantity observing unit 21. Since such transmitted data does not necessarily include abnormal data, the machine learning device 20 can be provided with a filtering function for abnormal data, that is, a function that allows selection of whether or not to input data from the three-dimensional measuring device 15 to the state quantity observing unit 21. This allows the learning unit 22 of the machine learning device 20 to efficiently learn optimal operations of the three-dimensional measuring device 15 and the hand unit 13 of the robot 14.
[0071] Furthermore, in the above-described machine learning device 20, the control device 16 receives output data from the learning unit 22, but the output data from the learning unit 22 is not necessarily free from abnormal data, so the control device 16 may be provided with a filtering function for abnormal data, i.e., a function that allows the control device 16 to select whether or not to output data from the learning unit 22 to the control device 16. This enables the control device 16 to cause the robot 14 to perform optimal operations of the hand unit 13 more safely.
[0072] The above-mentioned anomalous data can be detected by the following procedure: Estimate the probability distribution of input data, derive the occurrence probability of a new input using the probability distribution, and if the occurrence probability is below a certain level, regard the data as anomalous because it deviates significantly from typical behavior.
[0073] Next, an example of the operation of the machine learning device 20 included in the robot system 10 of this embodiment will be described. FIG. 4 is a flowchart showing an example of the operation of the machine learning device shown in FIG. 1. As shown in FIG. 4, when the learning operation (learning process) starts in the machine learning device 20 shown in FIG. 1, the three-dimensional measuring device 15 performs three-dimensional measurement and outputs the results (step S11 in FIG. 4). That is, in step S11, for example, a three-dimensional map (output data of the three-dimensional measuring device 15) of each of the workpieces 12 placed in a disorderly manner, including a state where the workpieces are piled up, is acquired and output to the state quantity observing unit 21. At the same time, the coordinate calculating unit 19 receives the three-dimensional map of each of the workpieces 12, calculates the three-dimensional position (x, y, z) of each of the workpieces 12, and outputs it to the state quantity observing unit 21, the operation result acquiring unit 26, and the control device 16. Here, the coordinate calculating unit 19 may calculate and output the orientation (w, p, r) of each of the workpieces 12 from the output of the three-dimensional measuring device 15.
[0074] 5, the output (three-dimensional map) of three-dimensional measuring device 15 may be input to state quantity observing unit 21 via a pre-processing unit 50 that processes the output before being input to state quantity observing unit 21. Also, as will be described with reference to FIG. 7, only the output of three-dimensional measuring device 15 may be input to state quantity observing unit 21, or only the output of three-dimensional measuring device 15 may be input to state quantity observing unit 21 via pre-processing unit 50. In this way, the implementation and output of three-dimensional measurement in step S11 can include a variety of things.
[0075] 1, the state quantity observation unit 21 observes the three-dimensional map for each workpiece 12 from the three-dimensional measuring device 15, and state quantities (output data of the three-dimensional measuring device 15) such as the three-dimensional position (x, y, z) and orientation (w, p, r) for each workpiece 12 from the coordinate calculation unit 19. The operation result acquisition unit 26 acquires the result of the pick-up operation of the robot 14, which picks up the workpiece 12 using the hand unit 13, based on the output data of the three-dimensional measuring device 15 (output data of the coordinate calculation unit 19). In addition to the output data of the three-dimensional measuring device, the operation result acquisition unit 26 can also acquire the results of the pick-up operation, such as the degree of success when the picked workpiece 12 is handed over to a subsequent process and any damage to the picked workpiece 12.
[0076] Furthermore, for example, the machine learning device 20 determines an optimal operation based on the output data of the three-dimensional measuring device 15 (step S12 in FIG. 4), and the control device 16 outputs command data (operation amount) for the hand unit 13 (robot 14) to perform the operation of picking up the workpiece 12 (step S13 in FIG. 4). Then, the result of picking up the workpiece is acquired by the operation result acquisition unit 26 described above (step S14 in FIG. 4).
[0077] Next, the success or failure of the removal of the work 12 is determined based on the output from the operation result acquisition unit 26 (step S15 in FIG. 4). If the removal of the work 12 is successful, a positive reward is set (step S16 in FIG. 4). If the removal of the work 12 is unsuccessful, a negative reward is set (step S17 in FIG. 4). Then, the action value table (value function) is updated (step S18 in FIG. 4).
[0078] Here, the success or failure of the removal of the workpiece 12 can be determined based on, for example, output data of the three-dimensional measuring device 15 after the removal operation of the workpiece 12. Furthermore, the success or failure of the removal of the workpiece 12 is not limited to an evaluation of the success or failure of the removal of the workpiece 12, but may be an evaluation of, for example, the degree of success when the removed workpiece 12 is handed over to a subsequent process, whether there is any change in the condition of the removed workpiece 12 such as damage, or the time (cycle time) and energy (amount of electricity) required for the hand unit 13 to grasp and transport the workpiece 12.
[0079] The calculation of the reward value based on the success or failure of the extraction of the work 12 is performed by the reward calculation unit 23, and the update of the action value table is performed by the value function update unit 24. That is, when the extraction of the work 12 is successful, the learning unit 22 sets a positive reward to the reward in the update formula for the value Q(s, a) described above (S16), and when the extraction of the work 12 is unsuccessful, the learning unit 22 sets a negative reward to the reward in the update formula (S17). Then, the learning unit 22 updates the action value table described above (S18) every time the work 12 is extracted. By repeating the above steps S11 to S18, the learning unit 22 continues (learns) updating the action value table.
[0080] In the above, the data input to the state quantity observing unit 21 is not limited to the output data of the three-dimensional measuring device 15, but may include, for example, data such as the output of other sensors, and it is also possible to use part of the command data from the control device 16. In this way, the control device 16 causes the robot 14 to perform the operation of picking up the workpiece 12 using the command data (operation amount) output from the machine learning device 20. Note that, as mentioned above, learning by the machine learning device 20 is not limited to the operation of picking up the workpiece 12, but may also be, for example, measurement parameters of the three-dimensional measuring device 15.
[0081] As described above, the robot system 10 equipped with the machine learning device 20 of this embodiment can learn the behavior of the robot 14 that uses the hand unit 13 to pick up workpieces 12 from a plurality of workpieces 12 that are placed in a disorderly manner, including in a randomly piled state. This enables the robot system 10 to learn to select the optimal behavior of the robot 14 that picks up workpieces 12 that are randomly piled, without human intervention.
[0082] FIG. 5 is a block diagram showing the conceptual configuration of a robot system according to another embodiment of the present invention, which illustrates a robot system that employs supervised learning. As is apparent from a comparison of FIG. 5 with FIG. 1, the robot system 10′ employing supervised learning shown in FIG. 5 further includes a result (labeled) data recording unit 40 in addition to the robot system 10 employing Q-learning (reinforcement learning) shown in FIG. 1. The robot system 10′ shown in FIG. 5 further includes a preprocessing unit 50 that preprocesses output data from the three-dimensional measuring device 15. It goes without saying that the preprocessing unit 50 can also be provided in the robot system 10 shown in FIG. 1, for example.
[0083] 5, the machine learning device 30 in the robot system 10′ to which supervised learning is applied includes a state quantity observing unit 31, an operation result acquiring unit 36, a learning unit 32, and a decision-making unit 35. The learning unit 32 includes an error calculating unit 33 and a learning model updating unit 34. Note that in the robot system 10′ of this embodiment as well, the machine learning device 30 learns and outputs operation quantities such as command data for instructing the robot 14 to perform an operation to pick up the workpiece 12 or measurement parameters of the three-dimensional measuring device 15.
[0084] That is, in robot system 10' to which supervised learning is applied shown in Fig. 5, error calculation unit 33 and learning model update unit 34 respectively correspond to reward calculation unit 23 and value function update unit 24 in robot system 10 to which Q-learning is applied shown in Fig. 1. Note that other components, such as three-dimensional measuring device 15, control device 16, and robot 14, are the same as those shown in Fig. 1 above, and therefore description thereof will be omitted.
[0085] The error calculation unit 33 calculates the error between the result (label) output from the operation result acquisition unit 36 and the output of the learning model implemented in the learning unit. Here, the result (label) attached data recording unit 40 can store the result (label) attached data obtained up to the day before the predetermined day on which the robot 14 is to perform the work, for example, when the shape of the workpiece 12 and the processing by the robot 14 are the same, and provide the result (label) attached data stored in the result (label) attached data recording unit 40 to the error calculation unit 33 on that predetermined day. Alternatively, data obtained by a simulation or the like performed outside the robot system 10' or result (label) attached data of another robot system can be provided to the error calculation unit 33 of that robot system 10' via a memory card or a communication line. Furthermore, the result (labeled) data recording unit 40 can be configured using non-volatile memory such as flash memory, and the result (labeled) data recording unit (non-volatile memory) 40 can be built into the learning unit 32, and the result (labeled) data stored in the result (labeled) data recording unit 40 can be used directly in the learning unit 32.
[0086] Figure 6 is a diagram for explaining an example of processing by the pre-processing unit in the robot system shown in Figure 5, where Figure 6(a) shows an example of three-dimensional position (posture) data of multiple workpieces 12 randomly piled in a box 11, i.e., output data of a three-dimensional measuring device 15, and Figures 6(b) to 6(d) show examples of image data after pre-processing has been performed on workpieces 121 to 123 in Figure 6(a).
[0087] Here, the workpieces 12 (121-123) are assumed to be cylindrical metal parts, and the hand (13) is assumed to be, for example, a suction pad that uses negative pressure to suck the longitudinal center portion of the cylindrical workpiece 12 rather than gripping the workpiece with two claws. Therefore, for example, if the position of the longitudinal center portion of the workpiece 12 is known, the suction pad (13) can be moved to that position to suck the workpiece 12, thereby removing the workpiece 12. The numerical values in Figures 6(a) to 6(d) are expressed in [mm] and indicate the x, y, and z directions, respectively. The z direction corresponds to the height (depth) direction of image data captured by a three-dimensional measuring device 15 (e.g., having two cameras) installed above the box 11 in which multiple workpieces 12 are piled.
[0088] As is clear from a comparison of Figure 6(a) with Figures 6(b) to 6(d), an example of processing by the pre-processing unit 50 in the robot system 10' shown in Figure 5 is to rotate the workpiece 12 of interest (e.g., three workpieces 121 to 123) from the output data (three-dimensional image) of the three-dimensional measuring device 15 and process it so that the height of the center becomes "0".
[0089] That is, the output data of the three-dimensional measuring device 15 includes, for example, information on the three-dimensional position (x, y, z) and orientation (w, p, r) of the longitudinal center portion of each workpiece 12. At this time, as shown in Figures 6(b), 6(c), and 6(d), the three workpieces 121, 122, and 123 of interest are each rotated by -r and have z subtracted from them to align them all under the same conditions. By performing such preprocessing, it is possible to reduce the load on the machine learning device 30.
[0090] 6(a) is not the output data of the three-dimensional measuring device 15 itself, but is, for example, an image obtained by lowering the threshold value for selection from images obtained by a program that defines the pick-up order of the workpieces 12 that has been implemented in the past, and this processing itself can also be performed by the pre-processing unit 50. It goes without saying that the processing by the pre-processing unit 50 can vary in various ways depending on various conditions, including the shape of the workpieces 12 and the type of hand 13.
[0091] In this way, the output data (three-dimensional map for each workpiece 12) of the three-dimensional measuring device 15, which has been processed by the preprocessing unit 50 before being input to the state quantity observing unit 31, is input to the state quantity observing unit 31. Referring again to FIG. 5, the error calculation unit 33, which receives the results (labels) output from the operation result acquisition unit 36, assumes that when the output of the neural network shown in FIG. 3 as a learning model is y, there is an error of -log(y) if the actual removal operation of the workpiece 12 is successful, and an error of -log(1-y) if the operation is unsuccessful, and performs processing with the goal of minimizing this error. Note that, as input to the neural network shown in FIG. 3, for example, image data of the target workpieces 121-123 after preprocessing as shown in FIGS. 6(b) to 6(d) and data on the three-dimensional position and orientation (x, y, z, w, p, r) of each of the target workpieces 121-123 are provided.
[0092] FIG. 7 is a block diagram showing a modification of the robot system shown in FIG. 1. As is clear from a comparison of FIG. 7 with FIG. 1, in the modification of the robot system 10 shown in FIG. 7, the coordinate calculation unit 19 is eliminated, and the state quantity observation unit 21 receives only the three-dimensional map from the three-dimensional measuring device 15 and observes the state quantities of the robot 14. It goes without saying that a configuration corresponding to the coordinate calculation unit 19 can be provided in the control device 16. The configuration shown in FIG. 7 can also be applied to the robot system 10′ to which supervised learning is applied, as described with reference to FIG. 5. That is, in the robot system 10′ shown in FIG. 5, the preprocessing unit 50 can be eliminated, and the state quantity observation unit 31 can receive only the three-dimensional map from the three-dimensional measuring device 15 and observe the state quantities of the robot 14. As such, various changes and modifications can be made to the above-described embodiments.
[0093] As described above in detail, this embodiment makes it possible to provide a machine learning device, a robot system, and a machine learning method that can learn the optimal robot operation when picking up workpieces that are placed in a disorderly manner, including in a loose pile, without human intervention. Note that the machine learning devices 20 and 30 of the present invention are not limited to those that apply reinforcement learning (e.g., Q-learning) or supervised learning, and various machine learning algorithms can be applied.
[0094] Although the embodiments have been described above, all examples and conditions described herein are described for the purpose of helping to understand the concept of the invention as applied to the invention and technology, and the particularly described examples and conditions are not intended to limit the scope of the invention. Furthermore, such descriptions in the specification are not intended to show the advantages and disadvantages of the invention. Although the embodiments of the invention have been described in detail, it should be understood that various changes, substitutions, and modifications can be made without departing from the spirit and scope of the invention. [Explanation of symbols]
[0095] 10,10' Robot System 11 boxes 12 Work 13 Hand section 14. Robot 15 Three-dimensional measuring instrument 16 Control device 17 Force Sensor 18 Support part 19 Coordinate calculation section 20,30 Machine learning devices 21,31 State Observation Unit 22,32 Learning Department 23 Remuneration Calculation Department 24 Value function update section 25,35 Decision making department 26,36 Operation result acquisition part 33 Error calculation section 34 Learning model update unit 40 Data recording section with results (label) 50 Pretreatment section
Claims
1. A machine learning device that learns the operation of a robot (14) that uses a hand unit (13) to pick up a workpiece (12) from a plurality of workpieces (12) that are randomly placed, including a randomly piled state, a state quantity observation unit (21, 31) that observes output data of a three-dimensional measuring device (15) that measures at least a three-dimensional map for each of the workpieces (12); an operation result acquisition unit (26, 36) that acquires a result of a pick-up operation of the robot (14) that picks up the workpiece (12) by the hand unit (13); a learning unit (22, 32) that receives an output from the state quantity observing unit (21, 31) and an output from the operation result acquiring unit (26, 36) and learns the picking operation of the work (12), The state quantity observation unit (21, 31) further observes output data of a coordinate calculation unit (19) that calculates the three-dimensional position of each of the workpieces (12) based on the output of the three-dimensional measuring device (15), The learning unit (22) a reward calculation unit (23) that calculates a reward based on the result of the determination of success or failure of the workpiece pick-up, which is the output of the operation result acquisition unit (26); a value function update unit (24) that has a value function that determines the value of the pick-up operation of the work (12) and updates the value function according to the reward, A machine learning device characterized by:
2. A machine learning device that learns the operation of a robot (14) that uses a hand unit (13) to pick up a workpiece (12) from a plurality of workpieces (12) that are randomly placed, including a randomly piled state, a state quantity observation unit (21, 31) that observes output data of a three-dimensional measuring device (15) that measures at least a three-dimensional map for each of the workpieces (12); an operation result acquisition unit (26, 36) that acquires a result of a pick-up operation of the robot (14) that picks up the workpiece (12) by the hand unit (13); a learning unit (22, 32) that receives an output from the state quantity observing unit (21, 31) and an output from the operation result acquiring unit (26, 36) and learns the picking operation of the work (12), The state quantity observation unit (21, 31) further observes output data of a coordinate calculation unit (19) that calculates the three-dimensional position of each of the workpieces (12) based on the output of the three-dimensional measuring device (15), The learning unit (32) has a learning model that learns the pick-up operation of the workpiece (12), an error calculation unit (33) that calculates an error based on the determination result of whether the workpiece has been picked up, which is the output of the operation result acquisition unit (26), and the learning model; a learning model update unit (34) that updates the learning model in accordance with the error, A machine learning device characterized by:
3. moreover, a decision-making unit (25, 35) that refers to the output from the learning unit (22, 32) and determines command data for instructing the robot (14) to perform an operation of picking up the workpiece (12); The machine learning device according to claim 1 or 2.
4. A machine learning device that learns the operation of a robot (14) that uses a hand unit (13) to pick up a workpiece (12) from a plurality of workpieces (12) that are randomly placed, including a randomly piled state, a state quantity observation unit (21, 31) that observes output data of a three-dimensional measuring device (15) that measures at least a three-dimensional map for each of the workpieces (12) and measurement parameters of the three-dimensional measuring device (15); an operation result acquisition unit (26, 36) that acquires a result of a pick-up operation of the robot (14) that picks up the workpiece (12) by the hand unit (13); a learning unit (22, 32) that receives an output from the state quantity observing unit (21, 31) and an output from the operation result acquiring unit (26, 36) and learns the picking operation of the work (12), The state quantity observation unit (21, 31) further observes output data of a coordinate calculation unit (19) that calculates the three-dimensional position of each of the workpieces (12) based on the output of the three-dimensional measuring device (15), The learning unit (22) a reward calculation unit (23) that calculates a reward based on the result of the determination of success or failure of the workpiece pick-up, which is the output of the operation result acquisition unit (26); a value function update unit (24) that has a value function that determines the value of the pick-up operation of the work (12) and updates the value function according to the reward, A machine learning device characterized by:
5. A machine learning device that learns the operation of a robot (14) that uses a hand unit (13) to pick up a workpiece (12) from a plurality of workpieces (12) that are randomly placed, including a randomly piled state, a state quantity observation unit (21, 31) that observes output data of a three-dimensional measuring device (15) that measures at least a three-dimensional map for each of the workpieces (12) and measurement parameters of the three-dimensional measuring device (15); an operation result acquisition unit (26, 36) that acquires a result of a pick-up operation of the robot (14) that picks up the workpiece (12) by the hand unit (13); a learning unit (22, 32) that receives an output from the state quantity observing unit (21, 31) and an output from the operation result acquiring unit (26, 36) and learns the picking operation of the work (12), The state quantity observation unit (21, 31) further observes output data of a coordinate calculation unit (19) that calculates the three-dimensional position of each of the workpieces (12) based on the output of the three-dimensional measuring device (15), The learning unit (32) has a learning model that learns the pick-up operation of the workpiece (12), an error calculation unit (33) that calculates an error based on the determination result of whether the workpiece has been picked up, which is the output of the operation result acquisition unit (26), and the learning model; a learning model update unit (34) that updates the learning model in accordance with the error, A machine learning device characterized by:
6. moreover, a decision-making unit (25, 35) that refers to the output from the learning unit (22, 32) and determines command data for instructing the robot (14) to perform an operation to pick up the workpiece (12) and the measurement parameters of the three-dimensional measuring device (15).
6. The machine learning device according to claim 4 or claim 5.
7. The coordinate calculation unit (19) further Calculating the posture of each of the workpieces (12) and outputting the calculated three-dimensional position and posture data of each of the workpieces (12). The machine learning device according to any one of claims 1 to 6.
8. The operation result acquisition unit (26, 36) uses output data from the three-dimensional measuring device (15). The machine learning device according to any one of claims 1 to 7.
9. moreover, a pre-processing unit (50) that processes output data from the three-dimensional measuring device (15) before inputting the data to the state quantity observing unit (21, 31); The state quantity observation unit (21, 31) receives output data from a preprocessing unit (50). The machine learning device according to any one of claims 1 to 8.
10. The pre-processing unit (50) aligns the direction and height of each of the workpieces (12) in the three-dimensional position data (x, y, z) and posture data (w, p, r) output from the three-dimensional measuring device (15). The machine learning device according to claim 9 .
11. The operation result acquisition unit (26, 36) At least one of the following is acquired: whether the workpiece (12) was successfully removed; the state of damage to the workpiece (12); and the time and energy required to grip and transport the workpiece (12) by the hand unit (13), or an evaluation of the time and energy. The machine learning device according to any one of claims 1 to 10.
12. The machine learning device includes: The machine learning device according to any one of claims 1 to 11, comprising a neural network.
13. A robot system (10, 10') including the machine learning device (20, 30) according to any one of claims 1 to 12, The robot (14); The three-dimensional measuring device (15), and a control device (16) that controls the robot (14) and the three-dimensional measuring device (15), A robot system characterized by:
14. The robot system (10, 10') comprises a plurality of the robots (14), the machine learning devices (20, 30) are provided for each of the robots (14); The machine learning devices (20) provided on the robots (14) are configured to share or exchange data with each other via a communication medium. The robot system according to claim 13 .
15. The machine learning device (20, 30) exists on a cloud server. The robot system according to claim 14 .
16. A machine learning method for learning the operation of a robot (14) that uses a hand unit (13) to pick up a workpiece (12) from a plurality of workpieces (12) that are randomly placed, including a state where the workpieces are piled up randomly, comprising: Observing output data from a three-dimensional measuring device (15) that measures at least a three-dimensional map for each of the workpieces (12); Obtaining the result of the pick-up operation of the robot (14) that picks up the workpiece (12) by the hand unit (13); receiving at least a three-dimensional map for each of the observed workpieces (12) and the acquired results of the picking operation of the robot (14) to learn the picking operation of the workpieces (12); The observation of the output data of the three-dimensional measuring device (15) further includes observing three-dimensional position data calculated based on the output of the three-dimensional measuring device (15) for each of the workpieces (12); The learning of the pick-up operation is calculating a reward based on the result of the determination of whether the picking up of the workpiece was successful or not, which is the result of the picking up operation of the robot (14); A value function that determines the value of the removal operation of the work (12) is updated according to the reward. A machine learning method characterized by:
17. A machine learning method for learning the operation of a robot (14) that uses a hand unit (13) to pick up a workpiece (12) from a plurality of workpieces (12) that are randomly placed, including a state where the workpieces are piled up randomly, comprising: Observing output data from a three-dimensional measuring device (15) that measures at least a three-dimensional map for each of the workpieces (12); Obtaining the result of the pick-up operation of the robot (14) that picks up the workpiece (12) by the hand unit (13); receiving at least a three-dimensional map for each of the observed workpieces (12) and the acquired results of the picking operation of the robot (14) to learn the picking operation of the workpieces (12); The observation of the output data of the three-dimensional measuring device (15) further includes observing three-dimensional position data calculated based on the output of the three-dimensional measuring device (15) for each of the workpieces (12); The learning of the pick-up operation is calculating an error based on the acquired result of the picking operation of the robot (14) and a learning model that learns the picking operation of the workpiece (12); updating the learning model according to the error; A machine learning method characterized by:
18. A machine learning method for learning the operation of a robot (14) that uses a hand unit (13) to pick up a workpiece (12) from a plurality of workpieces (12) that are randomly placed, including a state where the workpieces are piled up randomly, comprising: Observing output data of a three-dimensional measuring device (15) that measures at least a three-dimensional map for each of the workpieces (12) and measurement parameters of the three-dimensional measuring device (15); Obtaining the result of the pick-up operation of the robot (14) that picks up the workpiece (12) by the hand unit (13); receiving at least a three-dimensional map for each of the observed workpieces (12) and the acquired results of the picking operation of the robot (14) to learn the picking operation of the workpieces (12); The observation of the output data of the three-dimensional measuring device (15) further includes observing three-dimensional position data calculated based on the output of the three-dimensional measuring device (15) for each of the workpieces (12); The learning of the pick-up operation is calculating a reward based on the result of the determination of whether the picking up of the workpiece was successful or not, which is the result of the picking up operation of the robot (14); A value function that determines the value of the removal operation of the work (12) is updated according to the reward. A machine learning method characterized by:
19. A machine learning method for learning the operation of a robot (14) that uses a hand unit (13) to pick up a workpiece (12) from a plurality of workpieces (12) that are randomly placed, including a state where the workpieces are piled up randomly, comprising: Observing output data of a three-dimensional measuring device (15) that measures at least a three-dimensional map for each of the workpieces (12) and measurement parameters of the three-dimensional measuring device (15); Obtaining the result of the pick-up operation of the robot (14) that picks up the workpiece (12) by the hand unit (13); receiving at least a three-dimensional map for each of the observed workpieces (12) and the acquired results of the picking operation of the robot (14) to learn the picking operation of the workpieces (12); The observation of the output data of the three-dimensional measuring device (15) further includes observing three-dimensional position data calculated based on the output of the three-dimensional measuring device (15) for each of the workpieces (12); The learning of the pick-up operation is calculating an error based on the acquired result of the picking operation of the robot (14) and a learning model that learns the picking operation of the workpiece (12); updating the learning model according to the error; A machine learning method characterized by:
Citation Information
Patent Citations
Hydraulic pressure shock absorbing device for vehicles
JP1981042738A
Method of and apparatus for advancing cylinder groups shallowly covered with soil
JP1981070397A