Machine learning device, numerical control device, machine tool and machine learning method
The machine learning apparatus addresses the challenge of workpiece transfer failure by learning and adapting position commands for the driving device, reducing the offset of the workpiece transfer position and enhancing transfer reliability.
Patent Information
- Application Number
- DE112018007687
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2018-07-06
- Publication Date
- 2025-05-22
- Estimated Expiration
- 2038-07-06
AI Technical Summary
Existing technologies, such as those described in Patent Literature 1, fail to reduce the probability of workpiece transfer failure due to a fixed function expressing correlation, which does not adapt effectively to varying workpiece shapes and sizes.
A machine learning apparatus that learns a position command for a driving device moving a first chuck, utilizing a state observation unit and a learning unit. The learning unit includes a reward calculation unit and a function update unit to adjust the position command based on feedback data, reducing the offset of the workpiece transfer position between chucks.
The machine learning apparatus significantly reduces the probability of workpiece transfer failure by adapting the position command in real-time, improving the accuracy and reliability of the transfer process.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Area
[0001] The present invention relates to a machine learning apparatus that learns operations of transferring a workpiece, a numerical control apparatus, a machine tool, and a machine learning method. background
[0002] In a machine tool such as a lathe, during a workpiece transfer operation between a chuck that grips the workpiece and transfers the workpiece and a chuck that grips the workpiece to receive the workpiece, a loader that transfers the workpiece moves the workpiece gripped by the transferring chuck to a workpiece transfer position. In some cases, the workpiece cannot be moved to the center position of a chuck portion of the chuck that receives the workpiece because, for example, the workpiece is an elongated workpiece, the workpiece is curved, or the chuck fails to grip the workpiece.Because, as described above, when transferring a workpiece, the transfer may fail if the transfer position is offset from a suitable position, a technology for reducing or preventing offset of a transfer position is desired.
[0003] A loader control device described in Patent Literature 1 generates a function of the correlation between an offset value of a workpiece transfer position and a motor torque of a servo motor that drives a loader, estimates an offset value of the transfer position based on the correlation and a measured motor torque, and corrects the transfer position based on the estimated offset value. Citation listPatent literature
[0004] Patent literature 1: JP 2002 - 187 040 A
[0005] US 2002 / 0 091 461 A1 discloses a conventional control device for a loader which adjusts a control command for a motor based on a measured torque of the motor.
[0006] DE 10 2016 009 030 A1 discloses a robot system for picking up workpieces in which machine learning is used.
[0007] KOBER, Jens; BAGNELL, J. Andrew; PETERS, Jan. Reinforcement learning in robotics: A survey. Int. Journal of Robotics Research (IJRR), Vol. 32, Issue 11. September 2013, pp. 1238-1274. URL: doi.org / 10.1177 / 0278364913495721 reveals applications of reinforcement learning in the field of robotics. SummaryTechnical problem
[0008] However, in the above-mentioned Patent Literature 1, because the function expressing the correlation is a fixed function, the probability of failure of transferring a workpiece does not decrease even after repetitions of the workpiece transfer operation.
[0009] The present invention has been made in view of the above, and an object thereof is to provide a machine learning apparatus capable of reducing the probability of failure of transferring a workpiece by learning operations of transferring a workpiece. Solution to the problem
[0010] The technical problem is solved by the invention according to the subject matter of the independent claims. Advantageous developments of the invention are defined in the dependent claims.
[0011] One aspect of the present invention relates to a machine learning device that learns a position command to a driving device that moves a first chuck when transferring a workpiece between the first chuck that grips and transfers the workpiece and a second chuck that grips to receive the workpiece. The machine learning device includes: a state observation unit for observing, as a state variable, the position command to the driving device and feedback data from the driving device; and a learning unit for learning the position command that reduces an offset of a transfer position of the workpiece between the first chuck and the second chuck, depending on a data set generated based on the state variable.The feedback data includes a feedback position indicating a position of a conveying device for conveying the workpiece. The learning unit includes: a reward calculation unit for calculating a reward based on a difference between a position specified by the position command to the driving device and the feedback position; and a function updating unit for updating, based on the reward, a function for determining the position command.
[0012] Another aspect of the present invention relates to a machine learning method for learning a position command to a driver device that moves a first chuck when a workpiece is transferred between the first chuck that grips and transfers the workpiece and a second chuck that grips to receive the workpiece.The machine learning method includes: a state observation step of observing, as a state variable, the position command to the driving device and feedback data from the driving device; and a learning step of learning the position command that reduces an offset of a transfer position of the workpiece between the first chuck and the second chuck, depending on a data set generated based on the state variable, wherein the feedback data includes a feedback position indicating a position of a conveying device for conveying the workpiece.The learning step includes: a reward calculation step of calculating a reward based on a difference between a position specified by the position command to the driving device and the feedback position; and a function updating step of updating, based on the reward, a function for determining the position command. Advantageous effects of the invention
[0013] The invention according to the subject matters of the independent claims reduces the probability of failure of the transfer of a workpiece by learning operations of transferring a workpiece. Short description of drawings Fig. 1 is a diagram showing a configuration of a machining system according to an embodiment. Fig. 2 is a diagram showing a configuration of a control system including a numerical control apparatus according to the embodiment. Fig. 3 is a flowchart showing operation procedures of the machining system according to the embodiment. Fig. 4 is a diagram for explaining a first learning example performed by a machine learning apparatus according to the embodiment. Fig. 5 is a diagram for explaining a second learning example performed by the machine learning apparatus according to the embodiment. Fig. 6 is a diagram for explaining relative positions of a lathe chuck of a machine tool according to the embodiment and a workpiece. Fig. 7 is a diagram showing an example of a hardware configuration of the numerical control apparatus according to the embodiment. Description of embodiments
[0014] A machine learning apparatus, a numerical control apparatus, a machine tool, and a machine learning method according to embodiments of the present invention will be described in detail below with reference to the drawings. Note that the present invention is not limited to the embodiments. Embodiment
[0015] Fig. 1 is a diagram showing a configuration of a machining system according to an embodiment. Fig. 1 shows a machining system 1 viewed in the vertical direction. In the present embodiment, a case will be described where the vertical direction is a Y-axis direction, and the horizontal directions, which are the directions of movement of a workpiece 40, are an X-axis direction and a Z-axis direction.
[0016] The machining system 1 includes a machine tool 2 that machines a workpiece 40, and a control system 3 that controls the operation of the machine tool 2. Examples of the machine tool 2 include a lathe and a machining center. A case in which the machine tool 2 is a lathe will be described below.
[0017] The machine tool 2 includes a rotary unit 35, a loader chuck 32, which is a first chuck, a lathe chuck 31, which is a second chuck, and a loader 36, which is a conveying device for conveying a workpiece 40. The operation of the loader 36 is controlled by the control system 3. The loader chuck 32 is connected to the loader 36 and moves together with the loader 36. The loader chuck 32 is capable of gripping a workpiece 40, which is an object to be machined. Examples of the loader chuck 32 include a three-jaw chuck and a collet chuck. When machining of a workpiece 40 is started, the loader chuck 32 transfers the workpiece 40 to the lathe chuck 31, and after completion of the machining of the workpiece 40, the loader chuck 32 receives the workpiece 40 from the lathe chuck 31.
[0018] The rotary unit 35 rotates around the Z-axis, which is a main axis, as a rotary axis. The lathe chuck 31 is connected to the rotary unit 35 and rotates together with the rotary unit 35. The lathe chuck 31 is capable of gripping a workpiece 40. Examples of the lathe chuck 31 include a three-jaw chuck and a collet. When the workpiece 40 is machined, the rotary unit 35 rotates with the lathe chuck 31 gripping the workpiece 40 to rotate the workpiece 40. An example of the rotary unit 35 is a lathe device.
[0019] When the rotary unit 35 is loaded with a workpiece 40, the machine tool 2 grips one end of the workpiece 40 with the loader chuck 32. In this state, the loader 36 moves along the negative X-axis direction and stops at a position opposite the lathe chuck 31 (s0). Then, the loader 36 moves along the negative Z-axis direction. In this way, the loader 36 moves the workpiece 40 to a position where the lathe chuck 31 can grip the workpiece 40 (s1). The position of the loader 36 where the lathe chuck 31 can grip the workpiece 40 is a desired transfer position. The machine tool 2 starts a closing operation of the lathe chuck 31 using an auxiliary function, such as M codes (s2), and waits until the closing operation of the lathe chuck 31 is completed (s3). By closing the lathe chuck 31, the lathe chuck 31 grips the other end of the workpiece 40.
[0020] Thereafter, the machine tool 2 starts an opening operation of the loader chuck 32 (s4) and waits until the opening operation of the loader chuck 32 is completed (s5). The loader 36 then moves along the positive Z-axis direction. In this way, the loader 36 retracts in the direction away from the workpiece 40 (s6) and further moves along the positive X-axis direction.
[0021] When the machining of the workpiece 40 is completed, the machine tool 2 unloads the workpiece 40 from the rotary unit 35 through processes in the reverse order of the above-described processes s0 to s6. In this case, each of the processes s0 to s6 is a process in the reverse direction. Therefore, in each of the processes s0 to s6, the movement directions of the loader 36 during loading and unloading are opposite to each other. During unloading, the closing operation of the loader chuck 32 and the opening operation of the lathe chuck 31 are performed.
[0022] Specifically, during unloading, the loader 36 moves along the negative X-axis direction and further moves toward the workpiece 40. The loader chuck 32 then starts a closing operation. When the closing operation is completed, the lathe chuck 31 starts an opening operation. When the opening operation is completed, the loader 36 retracts in the direction away from the workpiece 40 and further moves along the positive X-axis direction.
[0023] Because the processes for loading a workpiece 40 into the rotary unit 35 and the processes for unloading a workpiece 40 from the rotary unit 35 are similar processes, processes for loading a workpiece 40 into the rotary unit 35 are explained below.
[0024] The mesh tool 2 has some characteristics specific to machines that grip, convey, and the like a workpiece 40. Therefore, a relationship caused by the machine-specific characteristics exists between a position command to the loader chuck 32 and the actual position of the loader chuck 32 in the X-axis direction. Therefore, even if a suitable position command is provided for use in loading a workpiece 40, the position of the workpiece 40 may be offset in the X-axis direction when the loader chuck 32 transfers the workpiece 40 to the lathe chuck 31, and the workpiece 40 may hit one end 30 of the lathe chuck 31 on the positive side in the Z-axis direction. In this case, the loader chuck 32 cannot transfer the workpiece 40 to the lathe chuck 31.
[0025] In the present embodiment, a numerical control device (NC device) 10, which will be described later and is included in the control system 3, learns the position of the loader chuck 32 in the X-axis direction when the loader chuck 32 transfers a workpiece 40 to the lathe chuck 31. Therefore, the numerical control device 10 reduces the probability of failure to transfer a workpiece 40 by learning position commands to the loader chuck 32. When the posture and shape of a workpiece 40 gripped by the loader chuck 32 are the same every time the workpiece 40 is gripped, the position of the workpiece 40 in the X-axis direction corresponds to the position of the loader chuck 32 in the X-axis direction. Therefore, in the present embodiment, an offset value of the loader chuck 32 in the X-axis direction and an offset value of a workpiece 40 in the X-axis direction are used synonymously.
[0026] Next, a configuration of the numerical control device 10 that controls the operation of the machine tool 2 will be described. Fig. Figure 2 is a diagram showing a configuration of the control system including the numerical control device according to the embodiment. The control system 3 includes the numerical control device 10, a driver unit 37, and a servo motor 38.
[0027] The numerical control unit 10 is a computer that controls the position of the loader 36 by sending a position command 53 to the driver unit 37. The position command 53 that the numerical control unit 10 sends to the driver unit 37 is a command indicating the position of the loader 36 and includes a position command in the X-axis direction and a position command in the Z-axis direction. The numerical control unit 10 controls the transfer of a workpiece 40 from the loader chuck 32 to the lathe chuck 31, and if the transfer fails, changes the position command 53 in the X-axis direction to the loader chuck 32 to control the transfer again.The numerical control device 10 learns an appropriate position command 53 in the X-axis direction to the loader chuck 32 based on a position command 53 in the X-axis direction to the loader chuck 32 and based on whether the result of a transfer using the position command 53 in the X-axis direction was failure or success.
[0028] The drive unit 37 is a drive device for moving the loader 36 by driving the servo motor 38. The drive unit 37 calculates a current value to be transmitted to the servo motor 38 based on the position command 53 from the numerical control device 10. The drive unit 37 drives the servo motor 38 by transmitting a current associated with the position command 53 to the servo motor 38. The drive unit 37 transmits a feedback (FB) current 55 to the numerical control device 10, which is data indicating the current to be transmitted to the servo motor 38. The FB current 55 is an example of feedback data from the drive unit 37 to the numerical control device 10.
[0029] When information indicating the rotational speed of the servo motor 38 is transmitted from an encoder 39, the driver unit 37 calculates an FB position 54, which is data indicating the current position of the workpiece 40, based on the rotational speed and transmits the FB position 54 to the numerical control device 10. The FB position 54 is an example of the feedback data from the driver unit 37 to the numerical control device 10.
[0030] The servomotor 38 is connected to the loader 36 and moves the loader 36 depending on the current from the drive unit 37. The servomotor 38 includes a servomotor for moving the loader 36 in the X-axis direction and a servomotor for moving the loader 36 in the Z-axis direction. The sensor 39, which detects the rotational speed of the servomotor 38, is attached to each servomotor 38 for the X-axis direction and for the Z-axis direction. The sensor 39 transmits to the drive unit 37 information indicating the detected rotational speed.
[0031] The numerical control device 10 includes a machining control program storage unit 11, an analysis unit 12, a control unit 13, a storage unit 14, and a machine learning device 20. The machining control program storage unit 11 stores machining control programs used for machining a workpiece 40. A machining control program includes a loading command 61 for loading the rotary unit 35 with a workpiece 40, a machining command for machining the workpiece 40, and an unloading command for unloading the workpiece 40 from the rotary unit 35. Fig. 2 shows a loading command 61. Among these commands, the loading command 61 and the unloading command are dedicated commands for transferring a workpiece 40. The loading command 61 is transmitted to the analysis unit 12 in the form of a G-code 51 for positioning the loader 36.
[0032] The analysis unit 12 analyzes the machining control programs. The analysis unit 12 determines whether an analyzed command is a dedicated command or not, and if the analyzed command is a dedicated command, for example, the loading command 61, generates transfer position information 52 based on the G code 51, indicating the position at which a workpiece 40 is to be transferred. Specifically, the analysis unit 12 generates the transfer position information 52 based on a command for positioning the loader 36 included in the G code 51. The transfer position information 52 is information about the position at which the workpiece 40 is transferred between the loader chuck 32 and the lathe chuck 31. Specifically, the transfer position information 52 is an end point of the loader 36.At the first transfer of a workpiece 40, information used for the transfer operation is set by arguments of the dedicated command.
[0033] The dedicated command includes the following command arguments (A1) to (A5): (A1) the end point of the loader 36, which is the transfer position information 52 of the workpiece 40; (A2) a reference current value A, which is a reference for determining whether or not the workpiece 40 strikes during transfer; (A3) a retraction value Lz of the loader 36 when it has been determined that the workpiece 40 is struck; (A4) the direction and a movement value Lx in which and with which the loader 36 is moved during learning; and (A5) a maximum movement distance Lmax in one direction during learning.
[0034] The transfer position information 52 of the aforementioned (A1) includes an X coordinate and a Z coordinate. The reference current value A of (A2) is a value for determining whether the transfer is in an abnormal state or not, and is compared with the FB current 55, which is a current sent to the servo motor 38. A case in which an FB current 55 greater than the reference current value A is sent to the servo motor 38 is such an abnormal state in which the workpiece 40 hits the lathe chuck 31 at a position different from the transfer position of the transfer by the lathe chuck 31. An example of the position at which the workpiece 40 hits something other than the transfer position of the transfer by the lathe chuck 31 is the end 30 of the lathe chuck 31 mentioned above.The retraction value Lz of (A3) is a distance by which the workpiece 40 is moved back along the Z-axis direction when the workpiece 40 has hit the end 30 and the movement of the workpiece 40 stopped.
[0035] When a workpiece 40 has hit the end 30 and stopped moving, the machine learning device 20 learns the position command 53. The movement value Lx of (A4) is a distance by which the workpiece 40 is moved in the X-axis direction during learning. The workpiece 40 is moved by the movement value Lx in the X-axis direction and then moved to the lathe chuck 31 in the Z-axis direction. The workpiece 40 is moved by the movement value Lx each time until the position to which the workpiece 40 was moved for transfer is within an allowable range. Furthermore, the movement distance Lmax of (A5) is a distance limit by which the workpiece 40 is moved in the X-axis direction during learning. Therefore, even during learning, the workpiece 40 is not moved further than the movement distance Lmax.
[0036] The analysis unit 12 sends the handover position information 52 of (A1) and the retraction value Lz of (A3) to the control unit 13. The analysis unit 12 also sends the values of the command arguments of (A2), (A4), and (A5) described above to the machine learning device 20. Note that the analysis unit 12 is not limited to the case where the values of the command arguments are obtained from the dedicated command, and may obtain values associated with the command arguments from parameters. In this case, the values associated with the command arguments are stored as parameters in the storage unit 14.
[0037] The control unit 13 generates a position command 53 based on the transfer position information 52 sent from the analysis unit 12 or based on an action 58 provided by the machine learning device 20. The action 58 is the next position command 53 in the X-axis direction. The control unit 13 sends the position command 53 to the driver unit 37 and the machine learning device 20. After receiving a notification from the machine learning device 20 indicating that the position to which the workpiece 40 has been moved is within the allowable range, the control unit 13 controls the further operation of transferring the workpiece 40.
[0038] Upon receiving a notification from the machine learning device 20 indicating that the position to which the workpiece 40 has been moved is outside the allowable range, the control unit 13 moves the workpiece 40 back by the retraction value Lz in the Z-axis direction and generates a position command 53 depending on the action 58.
[0039] The machine learning device 20 includes a state observation unit 25 and a learning unit 21. The state observation unit 25 receives the reference current value A of (A2) of the command arguments from the analysis unit 12, and the learning unit 21 receives the movement value Lx of (A4) and the movement distance Lmax of (A5) of the command arguments from the analysis unit 12.
[0040] The state observation unit 25 receives the position commands 53 in the X-axis direction and in the Z-axis direction from the control unit 13, the FB positions 54 in the X-axis direction and in the Z-axis direction from the driver unit 37, and the FB currents 55 in the X-axis direction and in the Z-axis direction from the driver unit 37.
[0041] To determine whether the position to which the workpiece 40 has been moved is within the allowable range, the state observation unit 25 uses the position commands 53 in the X-axis direction and the Z-axis direction, the FB positions 54 in the X-axis direction and the Z-axis direction, and the FB currents 55 in the X-axis direction and the Z-axis direction. Alternatively, the state observation unit 25 may determine whether the position to which the workpiece 40 has been moved is within the allowable range based on the position command 53 in the X-axis direction, the FB position 54 in the X-axis direction, and the FB current 55 in the X-axis direction.Further alternatively, the state observation unit 25 may determine whether the position to which the workpiece 40 has been moved is within the allowable range based on the position command 53 in the Z-axis direction, the FB position 54 in the Z-axis direction, and the FB current 55 in the Z-axis direction. Further alternatively, the state observation unit 25 may determine whether the position to which the workpiece 40 has been moved is within the allowable range based on the FB position 54 in the Z-axis direction and the position command 53 in the Z-axis direction without using the FB current 55 in the Z-axis direction.
[0042] In addition, when the learning unit 21 learns the position commands 53 to the driving unit 37, the state observation unit 25 observes the position command 53 in the X-axis direction, the FB position 54 in the X-axis direction, and the FB current 55 in the X-axis direction as state variables 56 and sends the state variables 56, which are results of the observation, to the learning unit 21. Therefore, the state variables 56 transmitted from the state observation unit 25 to the learning unit 21 include the position command 53 in the X-axis direction, the FB position 54 in the X-axis direction, and the FB current 55 in the X-axis direction.
[0043] The transfer position of the workpiece 40 may be offset from the center position of the lathe chuck 31 in the X-axis direction depending on the shape of the workpiece 40. Furthermore, in a case where the transfer position at which the workpiece 40 is transferred by the loader chuck 32 is inappropriate, the transfer position of the workpiece 40 may be offset from the center position of the lathe chuck 31 in the X-axis direction. In these cases, the workpiece 40 is more likely to hit the lathe chuck 31 at a position different from the transfer position of the lathe chuck 31 and stop at a position where the workpiece 40 cannot be gripped by the lathe chuck 31.
[0044] In a case where the workpiece 40 hits the chuck 31 at a position different from the transfer position, the position to which the workpiece 40 was moved is outside the allowable range. In a case where the workpiece 40 rubs against the chuck 31 while moving to the transfer position, the position to which the workpiece 40 was moved is also outside the allowable range.
[0045] In an abnormal case where the position to which the workpiece 40 has been moved is outside the allowable range, the load on the loader 36 increases in the X-axis direction and / or the Z-axis direction. Therefore, in the case where the position to which the workpiece 40 has been moved is outside the allowable range, the current sent to the servo motor 38 increases, and the FB current 55 also increases. Therefore, the state observation unit 25 determines whether the position to which the workpiece 40 has been moved is within the allowable range or not based on a result of a comparison between the reference current value A and the FB current 55. In a normal case where the position to which the workpiece 40 has been moved is within the allowable range, the workpiece 40 stops at the transfer position of the transfer by the lathe chuck 31, and therefore the load on the loader 36 does not increase and the FB current 55 does not increase.
[0046] Furthermore, in an abnormal case where the position to which the workpiece 40 has been moved is outside the allowable range, the position of the loader 36 assigned to the position command 53 and the position of the loader 36 corresponding to the FB position 54 do not match within a certain period of time. Therefore, the state observation unit 25 determines whether the position to which the workpiece 40 has been moved is within the allowable range or not based on a comparison result between the position command 53 and the FB position 54. The state observation unit 25 sends the result of the determination of whether the position to which the workpiece 40 has been moved is within the allowable range or not to the control unit 13.
[0047] The learning unit 21 learns an action 58, which is the next position command 53, depending on the state variable 56. In other words, the learning unit 21 learns a position command 53 in which the offset value of the transfer position of the transfer of the workpiece 40 by the loader 36 is reduced.
[0048] In particular, the learning unit 21 learns an action 58 depending on a data set generated based on the state variables 56, which include the position command 53, the FB position 54, and the FB current 55. The learning unit 21 includes a function update unit 22 and a reward calculation unit 23.
[0049] The reward calculation unit 23 calculates a reward 57 based on the state variables 56. The reward calculation unit 23 calculates a difference between the position specified by the position command 53 and the FB position 54 based on the state variables 56 and extracts the FB current 55 from the state variables 56. The reward calculation unit 23 increases the reward 57 when the difference between the position specified by the position command 53 and the FB position 54 is equal to or less than a threshold value and the FB current 55 is equal to or less than the reference current value A. In this case, the reward calculation unit 23 increases the reward 57 as the difference between the position specified by the position command 53 and the FB position 54 decreases and as the FB current 55 decreases. The learning unit 21 sends the calculated reward 57 to the function update unit 22.In the following description, the difference between the position specified by the position command 53 and the FB position 54 may also be referred to as a position difference.
[0050] The function updating unit 22 stores a function for determining an action 58 and updates the function for determining an action 58 based on the reward 57. An example of the function for determining an action 58 is an action value function Q(s t ,a t), which will be described later. The function updating unit 22 of the embodiment updates the action value function Q(s,a) so that the offset value of the transfer position decreases each time the operation of transferring a workpiece 40 by the machine tool 2 is repeated. The function updating unit 22 calculates an action 58 using the updated action value function Q(s,a). The function updating unit 22 sends the calculated action 58 to the control unit 13 and sends previously learned data, data used for learning, and data necessary for controlling the loader 36 to the storage unit 14. An example of the learned data is a position command 53 for the next transfer calculated when a transfer was successful, and an example of the data used for learning is the action value function Q(s,a) used for learning by the learning unit 21.An example of the data used to control the loader 36 is a retraction value Lz. The storage unit 14 stores the previously learned data, the data used for learning, and the data required to control the loader 36 by the control unit 13.
[0051] Next, operation procedures of the machining system 1 are explained. Fig. 3 is a flowchart showing the operation procedures of the machining system according to the embodiment. After reading a loading command 61, when the transfer of the workpiece 40 is the first transfer, that is, in an untrained case, the numerical control device 10 starts moving the loader 36 to the transfer position specified by the command argument of the above-described (A1) (step ST1). In this case, the control unit 13 transmits a position command 53 to the drive unit 37 and the state observation unit 25. As a result, the loader 36 moves in the X-axis direction and then moves in the Z-axis direction, and the loader chuck 32 grips the workpiece 40.
[0052] When the workpiece 40 starts to move, the driver unit 37 receives the FB position 54 and the FB current 55 at certain intervals and sends the FB position 54 and the FB current 55 to the state observation unit 25. The state observation unit 25 therefore monitors the position command 53, the FB position 54, and the FB current 55.
[0053] The state observation unit 25 determines whether the FB current 55 of the X-axis and the Z-axis is equal to or less than the reference current value A (step ST2). Note that reference current values A of different values can be used for the X-axis and the Z-axis.
[0054] If the FB current 55 of the X-axis or Z-axis becomes greater than a reference current value A before the workpiece 40 reaches the transfer position (step ST2: No), the state monitoring unit 25 sends to the learning unit 21 state variables 56 including position commands 53 in the X-axis direction and Z-axis direction, FB positions 54 in the X-axis direction and Z-axis direction, and FB currents 55 in the X-axis direction and Z-axis direction. Furthermore, the state monitoring unit 25 informs the control unit 13 that the position to which the workpiece 40 has been moved is outside the allowable range.
[0055] When the FB current 55 of an axis is greater than the reference current value A, that is, when the movement to the transfer position fails, the learning unit 21 sets a reward 57 with a small value for the position command 53 used for the transfer. Thus, the learning unit 21 learns an appropriate position command 53 depending on the state variable 56 and determines an action 58, which is the next position command 53, such that the reward 57 becomes maximum (step ST4a). The control unit 13 moves the workpiece 40 back in the Z-axis direction by the retraction value Lz (step ST5). The control unit 13 then starts moving the loader 36 to the transfer position (step ST1). The control unit 13 therefore moves the workpiece 40 in the X-axis direction using the position command 53 generated in response to the action 58, and then moves the workpiece 40 in the Z-axis direction.
[0056] When, in the process of step ST2, the FB current 55 of the X-axis and the Z-axis each becomes equal to or smaller than the reference current value A (step ST2: Yes), the state observation unit 25 for the X-axis and the Z-axis each determines whether or not the position difference between the position specified by the position command 53 and the FB position 54 is equal to or smaller than the threshold value (step ST3).
[0057] If the FB position 54 and the position command 53 for the X-axis and Z-axis are different from each other (step ST3: No), the state observation unit 25 sends state variables 56 including the position command 53 in the X-axis direction, the FB position 54 in the X-axis direction, and the FB current 55 in the X-axis direction to the learning unit 21. Furthermore, the state observation unit 25 informs the control unit 13 that the position to which the workpiece 40 has been moved is outside the allowable range. Thus, the above-described processes in steps ST4a and ST5 are performed, and the process in step ST1 is further performed.
[0058] Now, a process of learning an action 58 through the learning unit 21 will be explained. A learning algorithm used for the learning unit 21 can be any learning algorithm. Here, a case in which reinforcement learning is applied to the learning algorithm will be explained. In reinforcement learning, an agent, which is a subject of an action in an environment, observes the current state indicated by state variables 56 and determines an action 58 to be performed based on the result of the observation. The agent receives a reward 57 from the environment as a result of selecting the action 58 and learns an action that can obtain the greatest reward 57 through a series of actions 58. Q-learning and TD learning are known as typical reinforcement learning techniques.For example, in the case of Q-learning, a typical formula (action value table) for updating the action value function Q(s,a) is expressed by the following formula (1) below. Therefore, an example of the action value table is the action value function Q(s,a) of formula (1). [Math. 1] Q(st,at)←Q(st,at)+α(rt+1+γ maxa Q(st+1,a)−Q(st,at))
[0059] In formula (1) s represents t an environment at time t and a t represents an action at time t. By the action a t the environment is t+1 changed. r t+1 represents a reward 57 given as a result of the change in the environment, γ represents a discount rate, and α represents a learning coefficient. When Q-learning is applied, the action a t the next position command 53 of the transfer operation.
[0060] The update formula expressed by formula (1) increases the action value Q if the action value of the best action a at time t+1 is greater than the action value Q of the action a performed at time t, or decreases the action value Q in the opposite case. In other words, the action value function Q(s,a) is updated so that the action value Q of action a at time t is closer to the best action value at time t+1. Thus, the best action value in one environment evolves sequentially to action values in previous environments.
[0061] The reward calculation unit 23 calculates a reward 57 based on the position difference between the position specified by the position command 53 and the FB position 54 and based on the FB current 55.
[0062] As described above, the reward calculation unit 23 increases the reward 57 when the position difference between the position specified by the position command 53 and the FB position 54 is equal to or less than a threshold value and the FB current 55 is equal to or less than the reference current value A. In this case, the reward calculation unit 23 gives a reward 57 of "1", for example.
[0063] In contrast, the reward calculation unit 23 decreases the reward 57 when the position difference between the position specified by the position command 53 and the FB position 54 is greater than the threshold value or when the FB current 55 is greater than the reference current value A. In this case, the reward calculation unit 23 gives a reward 57 of "-1", for example.
[0064] For example, when the position difference is zero and a change value of the FB current 55 is zero, the reward calculation unit 23 sets the reward 57 to a maximum reward. When the position difference is equal to or less than the threshold value and the FB current 55 is half of the reference current value A, the reward calculation unit 23 sets the reward 57 to half of the maximum reward. An example of the case where the FB current 55 is half of the reference current value A is a case where the transfer of the workpiece 40 is successful, but the transfer position is slightly offset from a desired position. In a case where the workpiece 40 rubs against the lathe chuck 31 before reaching the transfer position, the FB current 55 is large during the rub.Furthermore, in a case where the transfer position is slightly offset from a desired position, when the lathe chuck 31 attempts to grip the workpiece 40 in the center of the transfer position of the workpiece 40 after the workpiece 40 reaches the transfer position, the workpiece 40 is pressed against the lathe chuck 31, and the FB current 55 increases. In such cases, the transfer position is within the allowable range, however, the machine learning device 20 gives a reward 57 with a value that is between a case where the workpiece 40 does not hit the end 30 and a case where the workpiece 40 does hit the end 30.
[0065] If the position difference is greater than the threshold value or if the FB current 55 is greater than the reference current value A, the reward calculation unit 23 sets the reward 57 to a minimum reward. The reward calculation unit 23 sends the calculated reward 57 to the function update unit 22.
[0066] The function updating unit 22 updates the action determining function 58 depending on the reward 57 calculated by the reward calculating unit 23. In the case of Q-learning, the action value function Q(s t ,a t ) expressed by the formula (1), for example, the function for calculating an action 58, which is updated by the function updating unit 22.
[0067] Fig. 4 is a diagram for explaining a first learning example performed by the machine learning apparatus according to the embodiment. Fig. 4 shows positions P0 to P6 of the workpiece 40 in the X-axis direction. The position of the workpiece 40 in the X-axis direction when the workpiece 40 hits the end 30 of the lathe chuck 31 is assumed to be the position P0. In this case, the machine learning device 20 repeats a process of moving the workpiece 40 back along the positive Z-axis direction, a process of moving the workpiece 40 to the next position in the X-axis direction, and a process of inserting the workpiece 40 along the negative Z-axis direction until the workpiece 40 no longer hits the end 30 of the lathe chuck 31.
[0068] Specifically, the machine learning device 20 moves the workpiece 40 in the X-axis direction in an order according to position P1, position P2, position P3, position P4, position P5, and position P6, which are positions in the X-axis direction. The interval between position P0 and position P1 is the movement value Lx. Similarly, the intervals between position P1 and position P2, between position P2 and position P3, between position P0 and position P4, between position P4 and position P5, and between position P5 and position P6 are each the movement value Lx. Furthermore, the intervals between position P0 and position P3 and between position P0 and position P6 are each the movement distance Lmax.
[0069] For example, to move the workpiece 40 to position P1, the machine learning device 20 sends an action 58 to the control unit 13 to load the workpiece 40 to position P1. As a result, a position command 53 associated with the action 58 causes the loader 36 to move in the X-axis direction.
[0070] If the workpiece 40 is successfully moved to the lathe chuck 31 without hitting the end 30, the machine learning device 20 stops moving the workpiece 40 in the X-axis direction. The machine learning device 20 gives a small reward 57 to a position command 53 when the workpiece 40 hits the end 30, and gives a large reward 57 to a position command 53 when the workpiece 40 has not hit the end 30.
[0071] Note that the machine learning device 20 can move the workpiece 40 to the positions P1 to P6 in any order. For example, the machine learning device 20 can move the workpiece 40 in the X-axis direction in ascending order of distance from the position P0, such as according to position P1, position P4, position P2, position P5, position P3, and position P6. Furthermore, the machine learning device 20 is not limited to the case where six positions P1 to P6 are set as the positions in the X-axis direction, and can set five or fewer or seven or more positions in the X-axis direction.
[0072] When the FB current 55 of the X-axis and the Z-axis each becomes equal to or smaller than the reference current value A (step ST2: Yes), and the position difference between the position specified by the position command 53 and the FB position 54 for the X-axis and the Z-axis becomes equal to or smaller than the threshold value (step ST3: Yes), the learning unit 21 learns an action 58, which is the next position command 53, depending on the state variable 56 (step ST4b). Therefore, the learning unit 21 sets a reward 57 with a large value for the position command 53 used for the handover, and then determines the action 58 corresponding to the next position command 53.
[0073] After the workpiece 40 has been moved to the transfer position, the state monitoring unit 25 also informs the control unit 13 that the position to which the workpiece 40 has been moved is within the allowable range. The control unit 13 therefore performs the process of transferring the workpiece 40 between the chucks.
[0074] Specifically, the control unit 13 causes the machine tool 2 to perform the operations from steps s2 to s6 described above. Specifically, the control unit 13 starts the operation of closing the lathe chuck 31 (step ST6) and waits until the lathe chuck 31 is closed (step ST7). By closing the lathe chuck 31, the lathe chuck 31 grips the workpiece 40.
[0075] Thereafter, the control unit 13 starts the operation of opening the loader chuck 32 (step ST8) and waits until the loader chuck 32 is opened (step ST9). The control unit 13 then retracts the loader 36 along the positive Z-axis direction (step ST10).
[0076] When the workpiece 40 is unloaded from the lathe chuck 31 to the loader 36, the machine learning device 20 also learns position commands 53 through processes similar to those in the case where the workpiece 40 is loaded from the loader 36 onto the lathe chuck 31.
[0077] As described above, when a transfer of the workpiece 40 fails, the machine learning device 20 causes the transfer of the workpiece 40 to be performed again depending on the movement value Lx, thereby enabling correction of the transfer position. In addition, because the machine learning device 20 learns the transfer position of the workpiece 40, a transfer failure can be prevented. In addition, because the machine learning device 20 learns the transfer position of the workpiece 40, the machine learning device 20 is also applicable to an environment in which the transfer fails due to a slight offset, for example, in a case of a collet chuck. Furthermore, because the machine learning device 20 can cause the transfer of the workpiece 40 to be performed again and prevent a transfer failure, the productivity of the machine tool 2 is increased.
[0078] Furthermore, because the machine learning device 20 determines the transfer position of the workpiece 40 based on the position difference between the position specified by the position command 53 and the FB position 54, it is not necessary to provide a special jig or device such as a camera for checking the transfer of the workpiece 40. Therefore, the transfer can be checked at low cost.
[0079] Furthermore, if the transfer of the workpiece 40 fails, the machine learning device 20 performs the transfer of the workpiece 40 again depending on the movement value Lx, eliminating the need for manual recovery even if the workpiece 40 hits the end 30 of the lathe chuck 31. This can shorten the downtime when loading the workpiece 40 and prevent a decrease in productivity.
[0080] While the present embodiment explains the case where the machine tool 2 moves the workpiece 40 in the X-axis direction and the Z-axis direction, the machine tool 2 can move the workpiece 40 in the X-axis direction, the Y-axis direction, and the Z-axis direction. In this case, position commands 53 that the numerical control device 10 transmits to the drive unit 37 include a position command in the X-axis direction, a position command in the Y-axis direction, and a position command in the Z-axis direction. Furthermore, the servo motor 38 includes servo motors for moving the loader 36 in the X-axis direction, the Y-axis direction, and the Z-axis direction. The machine learning device 20 then learns position commands in the X-axis direction, position commands in the Y-axis direction, and position commands in the Z-axis direction.
[0081] Fig. 5 is a diagram for explaining a second learning example performed by the machine learning apparatus according to the embodiment. Fig. Fig. 6 is a diagram for explaining relative positions of the lathe chuck of the machine tool according to the embodiment and a workpiece. Here, a case will be explained in which the machine tool 2 moves a workpiece 40 in the X-axis direction, the Y-axis direction, and the Z-axis direction. Fig. 5 and Fig. 6 show positions of the workpiece 40 in an XY plane when the workpiece 40 is viewed in the Z-axis direction.
[0082] One in Fig. 6 is an example of the lathe chuck 31 described with reference to Fig. 1. The lathe chuck 31A is a three-jaw chuck. The three jaws of the lathe chuck 31A grip the workpiece 40 by each moving toward the center Q1 of a circular chuck area 45.
[0083] The numerical control device 10 generates a position command 53 for inserting the workpiece 40 at the center Q1, however, the actual workpiece 40 may be inserted at a position offset from the center Q1, such as a position P0, and hit the end 30. The machine learning device 20 therefore learns a position command 53 with which the workpiece 40 does not hit the end 30.
[0084] It is assumed that the position of the workpiece 40 when the workpiece 40 hits the end 30 of the lathe chuck 31A is the position P0. In this case, the machine learning device 20 repeats a process of moving the workpiece 40 back along the positive Z-axis direction, a process of moving the workpiece 40 to the next positions in the X-axis direction and the Y-axis direction, and a process of inserting the workpiece 40 along the negative Z-axis direction until the workpiece 40 no longer hits the end 30 of the lathe chuck 31A. In this case, the machine learning device 20 calculates the next position to which the workpiece 40 is moved based on the movement value Lx and the movement distance Lmax. During the learning of the movement positions, the workpiece 40 is therefore moved to a position which is limited by certain learned directions and certain learned movement values in the XY plane.
[0085] Specifically, the machine learning device 20 moves the workpiece 40 in an order according to position P11, position P12, position P13, position P14, position P15, position P16, position P17, position P18, and position P19, which are positions in the XY plane. The intervals between position P0 and position P11, between position P11 and position P12, and between position P12 and position P13 are each the movement value Lx. Similarly, the intervals between position P0 and position P14, between position P14 and position P15, and between position P15 and position P16 are each the movement value Lx. Similarly, the intervals between position P0 and position P17, between position P17 and position P18, and between position P18 and position P19 are each the movement value Lx.In addition, the intervals between position P0 and position P13, between position P0 and position P16 and between position P0 and position P19 are each the movement distance Lmax.
[0086] In addition, an angle between a direction from the position P0 to the position P13 and a direction from the position P0 to the position P16 is a learned angle θ, and an angle between a direction from the position P0 to the position P16 and a direction from the position P0 to the position P19 is the learned angle θ. The machine learning device 20 repeats a process of moving the workpiece 40 from the position P0 in a first direction by the movement amount Lx and, after the workpiece 40 is moved by the movement distance Lmax, a process of moving the workpiece 40 from the position P0 in a second direction, which is a direction with the learned angle θ to the first direction, by the movement amount Lx.Therefore, when performing a first search for a suitable movement position of the workpiece 40, the machine learning device 20 searches for a suitable movement position by moving the workpiece 40 in the first direction by the movement value Lx each time up to the maximum movement distance Lmax. If no suitable movement position is found in the first direction, the machine learning device 20 performs a search, which is the same as that in the first direction, in the second direction, which has the learned angle θ with respect to the first direction. The machine learning device 20 repeats a process of searching for a suitable movement position while moving the workpiece 40 by the movement value Lx and a process of rotating the learned angle θ each time until a sum of the learned angles θ within the chuck range 45 exceeds 360°.
[0087] For example, to move the workpiece 40 to position P11, the machine learning device 20 sends an action 58 to the control unit 13 to load the workpiece 40 at position P11. As a result, a position command 53 associated with the action 85 causes the loader 36 to move in the X-axis direction and the Y-axis direction.
[0088] When the workpiece 40 has been successfully moved to the lathe chuck 31A without hitting the end 30, the machine learning device 20 terminates the movements of the workpiece 40 in the X-axis direction and in the Y-axis direction.
[0089] Note that the machine learning device 20 can move the workpiece 40 to the positions P11 to P19 in any order. For example, the machine learning device 20 can move the workpiece 40 in ascending order of distance from the position P0, such as according to position P11, position P14, position P17, position P12, position P15, position P18, position P13, position P16, and position P19. Furthermore, the machine learning device 20 is not limited to the case where nine positions P11 to P19 are set as the positions on the XY plane, and can set eight or fewer or ten or more positions on the XY plane.
[0090] A hardware configuration of the numerical control device 10 will now be described. Fig. 7 is a diagram showing an example of a hardware configuration of the numerical control apparatus according to the embodiment.
[0091] The numerical control device 10 can be controlled by a Fig. 7, namely, a processor 301 and a memory 302. Examples of the processor 301 include a central processing unit (CPU; also referred to as a central processing device, a processing device, a computing device, a microprocessor, a microcomputer, a processor, or a digital signal processor (DSP)) and a highly integrated system. Examples of the memory 302 include random access memory (RAM) and read-only memory (ROM).
[0092] The numerical control device 10 is implemented by the processor 301, which reads and executes programs for performing operations of the numerical control device 10 stored in the memory 302. In other words, these programs cause a computer to execute the procedures or methods performed by the numerical control device 10. The memory 302 is also used as temporary storage when the processor 301 performs various processes.
[0093] It should be noted that some of the functions of the numerical control device 10 may be implemented by dedicated hardware, and others may be implemented by software or firmware. Furthermore, the machine learning device 20 may be implemented by the Fig. 7 shown control circuit 300 may be implemented.
[0094] While in the present embodiment, whether a handover position is within the allowable range or not is determined based on the FB current 55 and the position difference, the determination of whether a handover position is within the allowable range or not may be made based on the position difference without using the FB current 55.
[0095] Furthermore, while in the present embodiment, the position commands 53 for the handover position are learned based on the FB current 55 in the X-axis direction and the position difference in the X-axis direction, the position commands 53 for the handover position may be learned based on the position difference in the X-axis direction without using the FB current 55 in the X-axis direction.
[0096] Note that in a case where the lathe chuck 31 is an electric chuck, the displacement of the workpiece 40 may also be detected at the lathe chuck 31. In this case, the machine learning device 20 may learn a position command 53 based on an displacement detected at the lathe chuck 31.
[0097] While in the present embodiment, the case where the machine learning device 20 performs machine learning using reinforcement learning is explained, the machine learning device 20 may perform machine learning using other known methods, for example, a neural network, genetic programming, functional logic programming, or a support vector machine.
[0098] Furthermore, while the case where the control unit 13 controls the loader 36 based on the action 58 learned by the learning unit 21 is explained in the present embodiment, the control unit 13 may control the loader 36 without using the action 58. In this case, the control unit 13 determines whether the position to which the workpiece 40 has been moved is within the allowable range or not based on the position command 53, the FB position 54, and the FB current 55, and, if the offset value is not within the allowable range, outputs a new position command 53 in which a position is offset from that specified by the position command 53. The control unit 13 therefore searches for a suitable moving position of the workpiece 40 by offsetting the moving position of the workpiece 40 little by little.Specifically, the control unit 13 ensures that the offset value falls within the allowable range by executing once or more a process of determining the offset and, if the offset value is not within the allowable range, a process of issuing the new position command 53.
[0099] As described above, according to the embodiment, because a position command 53 which reduces or prevents a displacement of the transfer position of a workpiece 40 between the chucks is learned depending on a data set generated based on the state variables 56, the probability of failure of transferring the workpiece 40 by repeating the operation of transferring the workpiece 40 can be reduced.
[0100] The configurations shown in the embodiment described above are examples of the present invention and may be combined with other known technologies and / or may be partially omitted or modified without departing from the scope of the present invention. List of reference symbols
[0101] 1 Machining system; 2 Machine tool; 3 Control system; 10 Numerical control device; 11 Machining control program storage unit; 12 Analysis unit; 13 Control unit; 14 Storage unit; 20 Machine learning device; 21 Learning unit; 22 Function update unit; 23 Reward calculation unit; 25 State observation unit; 30 End; 31, 31A Lathe chuck; 32 Loader chuck; 35 Rotary unit; 36 Loader; 37 Driver unit; 38 Servo motor; 39 Encoder; 40 Workpiece; 52 Transfer position information; 53 Position command; 54 FB position; 55 FB current; 56 State variable; 57 Reward; 58 Action; 61 Loading command.
Claims
[1] A machine learning device (20) that learns a position command (53) to a driver device (37) that moves a first chuck (32) when a workpiece (40) is transferred between the first chuck (32) that grips and transfers the workpiece (40) and a second chuck (31) that grips to receive the workpiece (40), the machine learning device (20) comprising: a state observation unit (25) for observing, as a state variable (56), the position command (53) to the driver device (37) and feedback data from the driver device (37); and a learning unit (21) for learning the position command (53) which reduces an offset of a transfer position of the workpiece (40) between the first chuck (32) and the second chuck (31), depending on a data set which is generated based on the state variable (56), wherein the feedback data comprises a feedback position (54) indicating a position of a conveying device (36) for conveying the workpiece (40), and the learning unit (21) includes: a reward calculation unit (23) for calculating a reward (57) based on a difference between a position specified by the position command (53) to the driver device (37) and the feedback position (54); and a function updating unit (22) for updating, based on the reward (57), a function for determining the position command (53). [2] Numerical control device (10) comprising: the machine learning device (20) according to claim 1; and a control unit (13) for outputting the position command (53) learned by the learning unit (21) to the driver device (37). [3] Numerical control device (10) according to claim 2, wherein the feedback data comprises a feedback current (55), wherein the feedback current (55) is data indicating a current output by the driver device (37) for the driver device (37) to move the first chuck (32), and the reward calculation unit (23) calculates the reward (57) based on the difference and the feedback stream (55). [4] The numerical control device (10) according to claim 3, wherein the reward calculation unit (23) increases the reward (57) when the difference is equal to or less than a threshold value and the feedback current (55) is equal to or less than a reference current value (A), or decreases the reward (57) when the difference is greater than the threshold value or when the feedback current (55) is greater than the reference current value (A). [5] The numerical control apparatus (10) according to claim 3 or 4, wherein the function updating unit (22) updates an action value table expressing the function depending on the reward (57). [6] Numerical control device (10) according to claim 3, wherein the state observation unit (25) determines that the transfer has failed if the difference is greater than a limit value or if the feedback current (55) is greater than a reference current value (A), upon determining that the handover has failed, the learning unit (21) learns the position command (53), and upon determining that the handover has failed, the control unit (13) attempts the handover again according to the position command (53) learned by the learning unit (21). [7] The numerical control device (10) according to any one of claims 2 to 6, wherein the learning unit (21) of position commands (53) to the driving device (37) learns a position command (53) in a first direction which is perpendicular to a direction of transfer between the first chuck (32) and the second chuck (31). [8] A machine tool (2) controlled by the numerical control device (10) according to any one of claims 2 to 7 and driven by the driving device (37). [9] A machine learning method for learning a position command (53) to a driver device (37) that moves a first chuck (32) when a workpiece (40) is transferred between the first chuck (32) that grips and transfers the workpiece (40) and a second chuck (31) that grips to receive the workpiece (40), the machine learning method comprising: a state observation step of observing, as a state variable (56), the position command (53) to the driver device (37) and feedback data from the driver device (37); and a learning step of learning the position command (53) which reduces an offset of a transfer position of the workpiece (40) between the first chuck (32) and the second chuck (31) in dependence on a data set which is generated based on the state variable (56), wherein the feedback data comprises a feedback position (54) indicating a position of a conveying device (36) for conveying the workpiece (40), and the learning step includes: a reward calculation step of calculating a reward (57) based on a difference between a position specified by the position command (53) to the driver device (37) and the feedback position (54); and a function updating step of updating, based on the reward (57), a function for determining the position command (53).
Citation Information
Patent Citations
machine learning apparatus, robotic system and machine learning system for learning a workpiece picking operation
DE102016009030A1
Loader control device
JP2002187040A
Loader control device
US20020091461A1
JP002002187040A