Machine learning device, numerical control system, setting device, numerical control device, and machine learning method
By using machine learning devices to perform reinforcement learning on numerical control devices, the feed rate and cutting speed in the machining program are optimized, solving the problem of insufficient operator optimization and improving production efficiency and tool life.
Patent Information
- Application Number
- CN202180021050.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-03-17
- Filing Date
- 2021-03-10
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2041-03-10
AI Technical Summary
In multi-product, variable-quantity production environments, operators lack the time and experience to optimize machining processes, leading to excessively reduced cutting speeds, longer cycle times, and decreased production efficiency.
A machine learning device is used to perform reinforcement learning on the numerical control device. By acquiring state information, outputting behavioral information, calculating rewards, and updating the value function, the depth of cut and cutting speed in the machining program are optimized.
Optimize machining processes, improve production efficiency, and ensure tool life and machining quality without increasing operator time and effort.
Smart Images

Figure CN115280252B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to machine learning devices, numerical control systems, setting devices, numerical control devices, and machine learning methods. Background Technology
[0002] The depth of cut and cutting speed in a fixed cycle such as drilling, tapping, boring, and turning are determined by the operator based on experience through repeated trial processing, taking into account the material and shape of the workpiece and tool.
[0003] Regarding this, the following technique is known: based on state variables including machining condition data, cutting condition data, machining result data, and tool data, cluster analysis is used to create clusters. These clusters are then used to learn and complete a model. Based on newly input machining conditions, cutting conditions, and machining results, an appropriate tool is determined. Within a range that maintains a specified good result within the determined tool cluster, the maximum cutting speed is determined. For example, see Patent Document 1.
[0004] Existing technical documents
[0005] Patent documents
[0006] Patent Document 1: Japanese Patent Application Publication No. 2019-188558 Summary of the Invention
[0007] The problem that the invention aims to solve
[0008] For example, in multi-product, variable-quantity production sites, the following situations frequently occur: (1) a certain processing procedure is reused in other machines; (2) other processing procedures are made by slightly modifying the form of a certain processing procedure; (3) the material of the workpiece is changed to process a certain processing procedure.
[0009] In such situations, operators lack sufficient time to optimize each machining program based on experience. Therefore, machining sometimes has to proceed without adequate optimization of the program and cutting conditions. In these cases, for example, regardless of changes made, the cutting speed is sometimes excessively reduced for safety reasons. This leads to reduced cycle times and decreased production efficiency.
[0010] Therefore, it is desirable to optimize the processing procedure without increasing the operator's time and effort.
[0011] Methods for solving problems
[0012] (1) One aspect of the machine learning apparatus disclosed herein is to perform machine learning on a numerical control device that moves a machine tool according to a machining program. The machine learning apparatus comprises: a state information acquisition unit that executes the machining program, in which the numerical control device sets at least one depth of cut and cutting speed, to cause the machine tool to perform cutting machining, thereby acquiring state information including the depth of cut and the cutting speed; a behavior information output unit that outputs behavior information including adjustment information of the depth of cut and the cutting speed included in the state information; a reward calculation unit that acquires determination information and outputs a reward value in reinforcement learning corresponding to predetermined conditions based on the acquired determination information, wherein the determination information is information related to at least the following: the pressure intensity applied to the tool during cutting machining, the waveform shape of the pressure applied to the tool, and the time required for machining; and a value function update unit that updates a value function based on the reward value, the state information, and the behavior information.
[0013] (2) In one aspect of the setting device of this disclosure, a behavior obtained from the machine learning device of (1) is selected according to a preset threshold, and the selected behavior is set to the processing program.
[0014] (3) One aspect of the numerical control system of the present disclosure includes: (1) a machine learning device; (2) a setting device; and a numerical control device that executes the processing program set by the setting device.
[0015] (3) One aspect of the numerical control device of the present disclosure includes (1) a machine learning device and (2) a setting device, the numerical control device executing the processing program set by the setting device.
[0016] (4) One aspect of the numerical control method disclosed herein is a machine learning method using a machine learning device. This machine learning device performs machine learning on a numerical control device that causes a machine tool to move according to a machining program. The numerical control device executes the machining program, which sets at least one entry depth and cutting speed, causing the machine tool to perform cutting machining. As a result, state information including the one entry depth and the cutting speed is obtained, and behavioral information is output. This behavioral information includes adjustment information of the one entry depth and the cutting speed contained in the state information. Judgment information is obtained, and a reward value in reinforcement learning corresponding to a specified condition is output based on the obtained judgment information. The judgment information is information that is at least related to the following: the pressure intensity applied to the tool during the cutting machining, the waveform shape of the pressure applied to the tool, and the time required for machining. The value function is updated based on the reward value, the state information, and the behavioral information.
[0017] Invention Effects
[0018] According to one method, the processing procedure can be optimized without increasing the operator's time and effort. Attached Figure Description
[0019] Figure 1 This is a functional block diagram illustrating an example of the functional structure of the numerical control system in the first embodiment.
[0020] Figure 2 This is a functional block diagram representing a functional structure example of a machine learning device.
[0021] Figure 3 This is a flowchart illustrating the operation of the machine learning device during Q-learning in the first embodiment.
[0022] Figure 4 Yes Figure 3 The flowchart below describes the detailed processing steps of the return calculation process shown in step S16.
[0023] Figure 5 It is a flowchart representing the actions taken by the optimal behavior information output unit when generating optimal behavior information.
[0024] Figure 6 This is a functional block diagram illustrating an example of the functional structure of the numerical control system in the second embodiment.
[0025] Figure 7 This is a functional block diagram representing a functional structure example of a machine learning device.
[0026] Figure 8 This is a flowchart illustrating the operation of the machine learning device during Q-learning in the second embodiment.
[0027] Figure 9 This is a diagram illustrating an example of the structure of a numerical control system.
[0028] Figure 10 This is a diagram illustrating an example of the structure of a numerical control system. Detailed Implementation
[0029] Hereinafter, the first embodiment of the present disclosure will be described with reference to the accompanying drawings. Here, an example is given where the machining process includes a fixed cycle such as drilling and tapping, and learning is performed according to the machining process, i.e., each time a workpiece is machined.
[0030] Therefore, the depth of cut and cutting speed set in this fixed cycle can be determined as the behavior for this machining program.
[0031] <First Implementation Method>
[0032] Figure 1 This is a functional block diagram illustrating an example of the functional structure of the numerical control system in the first embodiment.
[0033] like Figure 1 As shown, the numerical control system 1 includes a machine tool 10 and a machine learning device 20.
[0034] Machine tool 10 and machine learning device 20 can be directly connected to each other via a connection interface not shown. Alternatively, machine tool 10 and machine learning device 20 can also be connected to each other via a network not shown, such as a LAN (Local Area Network) or the Internet. In this case, machine tool 10 and machine learning device 20 have a communication unit (not shown) for communicating with each other via this connection. Furthermore, as described later, numerical control device 101 can be included in machine tool 10 or can be a different device from machine tool 10. Additionally, numerical control device 101 may also include machine learning device 20.
[0035] Machine tool 10 is a machine tool well known to those skilled in the art, and includes a numerical control device 101. Machine tool 10 operates according to motion commands from the numerical control device 101.
[0036] The numerical control device 101 is a known numerical control device to those skilled in the art, and includes a setting device 111. The numerical control device 101 generates motion commands based on a machining program obtained from an external device (not shown) such as a CAD / CAM device, and sends the generated motion commands to the machine tool 10. Thus, the numerical control device 101 controls the operation of the machine tool 10. Furthermore, during the operation of the machine tool 10, the numerical control device 101 can obtain, at predetermined time intervals such as a pre-set sampling time, the rotational speed, motor current value, and torque of motors such as the spindle motor (not shown) and the servo motors of the feed axes (not shown) included in the machine tool 10.
[0037] Furthermore, the numerical control device 101 can also obtain from the machine tool 10 the motor temperature, machine temperature, and ambient temperature, etc., measured by sensors such as a temperature sensor (not shown) included in the machine tool 10. Additionally, the numerical control device 101 can obtain from the machine tool 10 the axial and rotational pressures applied to the tool mounted on the spindle (not shown), measured by sensors such as a pressure sensor (not shown) included in the machine tool 10. Furthermore, the numerical control device 101 can also obtain the time required for a predetermined cutting operation when the machine tool 10 performs the cutting operation, measured by a cycle counter (not shown) included in the machine tool 10.
[0038] Furthermore, in this embodiment, as described above, the processing procedure consists of only one fixed cycle; therefore, the processing time is the same as the cycle time.
[0039] Furthermore, the numerical control device 101 can output, for example, information such as the tool material, tool shape, tool diameter, tool length, remaining tool life, workpiece material, and cutting conditions in the tool catalog of the spindle (not shown) mounted on the machine tool 10 to the machine learning device 20 described later. Additionally, the numerical control device 101 can output information obtained from the machine tool 10, such as spindle speed, motor current, machine temperature, ambient temperature, pressure applied to the tool (axial and rotational directions), waveform of the pressure applied to the tool (axial and rotational directions), torque applied to the feed axis, waveform of the torque applied to the feed axis, torque applied to the spindle, waveform of the torque applied to the spindle, and the machining time to the machine learning device 20 described later.
[0040] Furthermore, the numerical control device 101 may, for example, store a tool management table (not shown) in a storage unit such as an HDD (Hard Disk Drive) included in the numerical control device 101 to manage all tools mounted on the spindle (not shown) of the machine tool 100. The numerical control device 101 can also retrieve tool material, tool shape, tool diameter, tool length, and remaining tool life from the tool management table (not shown) based on the tool number set in the machining program. Here, the remaining tool life is obtained, for example, from the durability time calculated from a corresponding table listed in the catalog and considered as the tool life, and is calculated based on the usage time for machining each workpiece. The remaining tool life in the tool management table (not shown) can also be updated using the calculated value.
[0041] In addition, the numerical control device 101 can obtain the material of the workpiece to be processed, the cutting conditions of the tool catalog, etc., through input devices (not shown) such as keyboards and touch panels included in the numerical control device 101, by means of operator input.
[0042] In addition, the waveform of the pressure applied to the tool is time-series data of the pressure applied to the tool. Additionally, the waveform of the torque applied to the feed axis is time-series data of the torque applied to the feed axis. Furthermore, the waveform of the torque applied to the spindle is time-series data of the spindle torque.
[0043] The setting device 111 selects a behavior from the behaviors obtained from the machine learning device 20 (described later) according to a preset threshold, and sets the selected behavior to the processing program.
[0044] Specifically, the setting device 111 compares, for example, the remaining tool life of the tool currently in use in the machine tool 10 with a preset threshold (e.g., 10%). Thus, if the remaining tool life is greater than the threshold, it selects an action that prioritizes machining time; if the remaining tool life is less than the threshold, it selects an action that prioritizes tool life. The setting device 111 then sets the selected action to the machining program.
[0045] Furthermore, the setting device 111 may be composed of a computer such as a numerical control device 101 having a CPU or other arithmetic processing unit.
[0046] In addition, the setting device 111 may be a different device from the numerical control device 101.
[0047] <Machine Learning Device 20>
[0048] The machine learning device 20 is a device for reinforcement learning of the depth of cut and cutting speed of each workpiece when the machine tool 10 is operated by the numerical control device 101 executing the machining program.
[0049] Before describing the functional blocks included in the machine learning device 20, the basic structure of Q-learning, which serves as an example of reinforcement learning, will first be explained. Reinforcement learning is not limited to Q-learning. An agent (equivalent to the machine learning device 20 in this embodiment) observes the state of the environment (equivalent to the machine tool 10 and numerical control device 101 in this embodiment), selects a certain action, and the environment changes according to the selected action. As the environment changes, a reward is given, and based on the reward, the agent learns better action selections.
[0050] Supervised learning represents a completely correct answer, while the rewards in reinforcement learning are mostly based on fragmented values of partial changes in the environment. Therefore, agents learn to maximize the total rewards they receive in the future.
[0051] In this way, reinforcement learning learns appropriate behaviors based on the interaction between the behaviors and the environment, that is, it learns methods to maximize future rewards. This means that in this embodiment, it is possible to obtain future-influencing behaviors, such as optimizing fixed cycles of processing procedures in multi-product, variable-quantity production environments, without increasing operator time and effort.
[0052] Here, any learning method can be used as reinforcement learning. In the following description, we will take the case of using Q-learning in a certain environmental state s as an example. Q-learning is a method of learning the value function Q(s, a) of choosing behavior a.
[0053] Q-learning aims to select the behavior with the highest value of the value function Q(s, a) from the available behaviors a in a certain state s as the optimal behavior.
[0054] However, at the initial point in Q-learning, the correct value of the value function Q(s, a) is completely unknown for any combination of state s and action a. Therefore, the agent chooses various actions a in a given state s, and based on the reward given for that action a, chooses the better action, thus continuing to learn the correct value function Q(s, a).
[0055] Furthermore, to maximize the total future returns, the goal is to eventually achieve Q(s, a) = E[Σ(γ] t )r t Here, E[] represents the expected value, t represents time, γ represents the parameter referred to later as the discount rate, and r t Let Σ represent the reward at time t, and Σ be the total at time t. The expected value in this formula is the expected value when the optimal behavior changes. However, during Q-learning, since the optimal behavior is unknown, reinforcement learning is performed while exploring various behaviors. The update formula for such a value function Q(s, a) can be represented, for example, by the following mathematical formula 1.
[0056]
Mathematical Formula 1
[0057]
[0058] In the mathematical formula 1 above, s t a represents the environmental state at time t. t This represents the behavior at time t. Through behavior a... t The state change is s t+1 r t+1 This represents the reward obtained through the change of this state. Additionally, the term with `max` is: in state `s` t+1 The following is the Q-value obtained by multiplying γ by the Q-value when selecting the behavior with the highest known Q-value at that time, 'a'. Here, γ is a parameter of 0 < γ ≤ 1, called the discount rate. Additionally, α is the learning coefficient, assuming the range of α is 0 < α ≤ 1.
[0059] The above mathematical formula 1 represents the following method: based on trial a t The result and the feedback r t+1 Update status s t The following behavior a t Value function Q(s) t a t ).
[0060] This update expression indicates that if behavior a tThe resulting next state s t+1 The value of the best behavior under the following circumstances is max a Q(s t+1 a) Compared to state s t The following behavior a t Value function Q(s) t a t If ) is large, then Q(s) will increase. t a t If the value is small, then decrease Q(s). t a t That is, to make the value of a certain action in a certain state close to the value of the optimal action in the next state resulting from that action. Here, although this difference is due to the discount rate γ and the return r... t+1 The form of existence changes, but it is essentially a structure in which the optimal behavioral value in a certain state is propagated to the behavioral value in the previous state.
[0061] Here, Q-learning can be approached by creating a table of Q(s, a) for all state-action pairs (s, a) and then learning from it. However, sometimes the number of states becomes too large to calculate the Q(s, a) values for all state-action pairs, making Q-learning convergence time quite long.
[0062] Therefore, a well-known technique called DQN (Deep Q-Network) can be used. Specifically, a suitable neural network can be used to construct the value function Q, and the parameters of the neural network can be adjusted. Thus, the value of the value function Q(s, a) can be calculated by approximating the value function Q using the appropriate neural network. By utilizing DQN, the time required for Q to learn and converge can be shortened. Furthermore, DQN is described in detail in, for example, the following non-patent literature.
[0063] <Non-Patent Literature>
[0064] "Human-level control through deep reinforcement learning", by Volodymyr Mnih1 [online], [accessed January 17, 2009], Internet <URL: http: / / files.davidqiu.com / research / nature14236.pdf>
[0065] The machine learning device 20 performs the Q-learning described above. Specifically, the machine learning device 20 learns the following value Q: taking information related to the tool and workpiece set in the machine tool 10, the depth of cut and cutting speed set in a fixed cycle, and the measured values obtained from the machine tool 10 by executing the machining program as state s, and selecting the setting and modification of the depth of cut and cutting speed set in the fixed cycle related to state s as behavior a for state s. Here, examples of information related to the tool and workpiece include tool material, tool shape, tool diameter, tool length, remaining tool life, material of the workpiece to be machined, and cutting conditions in the tool catalog. In addition, examples of measured values obtained from the machine tool 10 include spindle speed, motor current value, machine temperature, and ambient temperature.
[0066] The machine learning device 20 observes state information (state data) s and determines behavior a. This state information includes information related to the tool and workpiece set in the machine tool 10, the depth of cut and cutting speed set in a fixed cycle, and the measured values obtained from the machine tool 10 by executing the machining program. The machine learning device 20 returns a report each time behavior a is performed. The machine learning device 20 explores the optimal behavior a through trial and error to maximize the total future returns. Thus, the machine learning device 20 can select the optimal behavior a (i.e., "depth of cut" and "cutting speed") for a state s, which includes information related to the tool and workpiece set in the machine tool 10, the depth of cut and cutting speed set in a fixed cycle, and the measured values obtained from the machine tool 10 by executing the machining program.
[0067] Figure 2 This is a functional block diagram representing a functional structure example of the machine learning device 20.
[0068] In order to perform the above reinforcement learning, such as Figure 2 As shown, the machine learning device 20 includes: a state information acquisition unit 201, a learning unit 202, a behavior information output unit 203, a value function storage unit 204, an optimal behavior information output unit 205, and a control unit 206. The learning unit 202 includes: a reward calculation unit 221, a value function update unit 222, and a behavior information generation unit 223. The control unit 206 controls the operations of the state information acquisition unit 201, the learning unit 202, the behavior information output unit 203, and the optimal behavior information output unit 205.
[0069] The status information acquisition unit 201 acquires status data s from the numerical control device 101 as the status of the machine tool 10. The status data s includes information related to the tool and workpiece set in the machine tool 10, the depth of cut and cutting speed set for one cut in a fixed cycle, and the measured values obtained from the machine tool 10 by executing the machining program. This status data s corresponds to the environmental status s in Q-learning.
[0070] The status information acquisition unit 201 outputs the acquired status data s to the learning unit 202.
[0071] Furthermore, the state information acquisition unit 201 can store the acquired state data s in a storage unit (not shown) included in the machine learning device 20. At this time, the learning unit 202, described later, can read the state data s from the storage unit (not shown) of the machine learning device 20.
[0072] In addition, the status information acquisition unit 201 also acquires determination information for calculating the reward of Q-learning. Specifically, the determination information for calculating the reward of Q-learning is the pressure intensity (axial and rotational directions) applied to the tool, the waveform shape of the pressure applied to the tool (axial and rotational directions), the intensity of the torque applied to the feed axis, the waveform shape of the torque applied to the feed axis, the intensity of the torque applied to the spindle, the waveform shape of the torque applied to the spindle, and the machining time required to execute the machining program, which are obtained from the machine tool 10 by executing the machining program.
[0073] The learning unit 202 is the part that learns the value function Q(s, a) when choosing a certain behavior a under certain state data (environment state) s. Specifically, the learning unit 202 has: a reward calculation unit 221, a value function update unit 222, and a behavior information generation unit 223.
[0074] Furthermore, the Learning Department 202 determines whether to continue learning. For example, it can determine whether to continue learning based on whether the number of trials since the start of machine learning has reached the maximum number of trials, or whether the elapsed time since the start of machine learning has exceeded the specified time (or is more than the specified time).
[0075] The reward calculation unit 221 calculates the reward when behavior a is selected in a certain state s based on the determination information. The reward can be calculated based on multiple evaluation items included in the determination information. In this embodiment, for example, the reward is calculated based on items such as (1) the intensity of the pressure (torque) applied to the tool, feed axis, and spindle, (2) the waveform shape of the pressure (torque) applied to the tool, feed axis, and spindle, and (3) the processing time.
[0076] Therefore, the calculation results for (1) the intensity of the pressure (torque) applied to the tool, feed axis, and spindle, (2) the waveform shape of the pressure (torque) applied to the tool, feed axis, and spindle, and (3) the time required for machining are explained.
[0077] Regarding the return on the item concerning (1) the intensity of the pressure (torque) applied to the tool, feed axis, and spindle.
[0078] Let the intensity values of the pressure (torque) applied to the tool, feed axis, and spindle in state s and state s' when transitioning from state s to state s' through behavior a be set as value P respectively. t (s), P f (s), P m (s), and the value P t (s'), P f (s'), P m (s').
[0079] The return calculation unit 221 calculates the return based on the intensity of the pressure (torque) applied to the tool, feed axis, and spindle in the following manner.
[0080] In value P t (s') < value P t (s), and the value P f (s') < value P f (s), and the value P m (s') < value P m In the case of (s), the return r p Set to a positive value.
[0081] In the value P of state s' t (s'), P f (s'), P m At least one of (s') is greater than the value P of state s. t (s), P f (s), P m (s) In the case of a large value, the return will be r p Set to a negative value.
[0082] Furthermore, regarding negative and positive values, they can be pre-defined values (e.g., a first negative value and a first positive value).
[0083] Regarding the report on (2) the waveform shape of the pressure (torque) applied to the tool, feed axis, and spindle.
[0084] Let the waveform shape of the pressure (torque) applied to the tool, feed axis, and spindle in state s' when transitioning from state s to state s' through behavior a be set as WF. t (s'), WFf (s'), WF m (s').
[0085] The return calculation unit 221 calculates the return based on the waveform shape of the pressure (torque) applied to the tool, feed axis, and spindle in the following manner.
[0086] The waveform shape WF of the pressure (torque) applied to the tool, feed axis, and spindle. t (s'), WF f (s'), WF m If at least one of (s') is similar to a waveform indicating a precursor to tool damage or a waveform indicating a more drastic reduction in tool life, a return r will be made. w Set to a negative value.
[0087] The waveform shape WF of the pressure (torque) applied to the tool, feed axis, and spindle. t (s'), WF f (s'), WF m (s') If all waveforms are dissimilar to those indicating a precursor to tool failure or a more drastic reduction in tool life, then r will be returned. w Set to a positive value.
[0088] Furthermore, data on waveforms indicating signs of tool damage and waveforms indicating a more drastic reduction in tool life can be obtained in advance for each tool and stored in the storage unit (not shown) of the machine learning device 20.
[0089] Additionally, regarding negative and positive values, these can be pre-defined values (e.g., a second negative value and a second positive value).
[0090] Regarding (3) the return on the time required for processing
[0091] Let the values of state s and the processing time required in state s' when transitioning from state s to state s' through action a be denoted as T(s) and T(s'), respectively.
[0092] The return calculation unit 221 calculates the return based on the time required for processing in the following manner.
[0093] If the value T(s') > the value T(s), the reward will be r. c Set to a negative value.
[0094] When the value T(s') = the value T(s), the return will be r. c Set it to zero.
[0095] If the value T(s') < the value T(s), the return will be r. c Set to a positive value.
[0096] Furthermore, regarding negative and positive values, they can be pre-set values (e.g., a third negative value and a third positive value).
[0097] The return calculation unit 221 calculates the returns for machine learning that prioritizes processing time and machine learning that prioritizes tool lifespan using mathematical formula 2. These returns r are calculated separately for the aforementioned items based on the machine learning prioritizing processing time and the machine learning prioritizing tool lifespan, respectively. p r w r c It is obtained by weighted summation.
[0098]
Mathematical Formula 2
[0099] r = a w ·r p +b w ·r w +c w ·r c
[0100] In addition, coefficient a w b w c w This represents the weighting coefficient.
[0101] Additionally, the return calculation unit 221 calculates the return r (hereinafter also referred to as "return r") during machine learning that prioritizes the processing time. cycle In the calculation of "), for example, compared to machine learning that prioritizes tool lifespan, the coefficient c of mathematical formula 2 can be used. w The value can be set to a larger value, or the absolute value of the third negative value and the third positive value can be set to a larger value.
[0102] Additionally, the return calculation unit 221 calculates the return r (hereinafter also referred to as "return r") during machine learning that prioritizes the tool's lifespan. tool In the calculation of "), for example, compared with machine learning that prioritizes the time required for processing, the coefficient b of mathematical formula 2 can be used. w The value can be set to a larger value, or the absolute value of the second negative value and the second positive value can be set to a larger value.
[0103] Unless otherwise specified, machine learning that prioritizes processing time will be referred to as "machine learning in processing time priority mode". Additionally, unless otherwise specified, machine learning that prioritizes tool lifespan will be referred to as "machine learning in tool lifespan priority mode".
[0104] When performing machine learning in the processing time priority mode, the value function update unit 222 updates the value function based on the state s, the action a, the state s' when the action a is applied to the state s, and the reward r calculated as described above. cycle The value is processed using Q-learning in time-priority mode, thereby updating the value function Q stored in value function storage unit 204. cycle Furthermore, during machine learning in the tool lifetime priority mode, the value function update unit 222 updates the value based on the state s, the action a, the state s' when the action a is applied to the state s, and the reward r calculated as described above. tool The value is used for Q-learning in the tool lifetime priority mode, thereby updating the value function Q stored in the value function storage unit 204. tool .
[0105] The value function Q of the processing time priority mode cycle And the value function Q of the tool life priority mode tool Updates can be done through online learning, batch learning, or small-batch learning.
[0106] Online learning is a learning method that applies an action 'a' to the current state 's', and immediately updates the value function 'Q' whenever state 's' transitions to a new state 's'. Batch learning, on the other hand, involves repeatedly applying an action 'a' to the current state 's', thereby collecting learning data, and using all the collected learning data to update the value function 'Q'. Mini-batch learning, then, is an intermediate learning method between online and batch learning, updating the value function 'Q' only after a certain amount of learning data has been accumulated.
[0107] The behavior information generation unit 223 selects behavior a during the Q-learning process for the current state s. During the Q-learning process corresponding to the machining time priority mode or the tool life priority mode, the behavior information generation unit 223 generates behavior information a in order to correct the cutting depth and cutting speed set in a fixed cycle (equivalent to behavior a in Q-learning) and outputs the generated behavior information a to the behavior information output unit 203.
[0108] More specifically, the behavior information generation unit 223 can increase or decrease the depth of cut and cutting speed of behavior a in an incremental manner relative to the depth of cut and cutting speed of a fixed cycle in state s, according to the processing time priority mode and the tool life priority mode.
[0109] In this embodiment, for example, we illustrate the case where machine learning in a processing time-priority mode and machine learning in a tool lifespan-priority mode are performed alternately. Furthermore, in this case, known methods such as the greedy method or the ε-greedy method (described later) can be used randomly to perform machine learning without bias towards either mode. Alternatively, as described later, machine learning in a processing time-priority mode and machine learning in a tool lifespan-priority mode can also be performed separately.
[0110] The behavior information generation unit 223 can use machine learning in the machining time priority mode or tool life priority mode to adjust the single cut depth and cutting speed of the machining program according to behavior a. When transitioning to state s', it selects the single cut depth and cutting speed of the machining program for behavior a' for state s' based on the state of the force (torque) of the tool, feed axis, and spindle (whether it decreases), the state of the waveform shape of the force (torque) of the tool, feed axis, and spindle (whether it is similar), and the state of the machining time (increase, decrease, or maintain).
[0111] For example, the following strategy can be adopted: In machine learning under the processing time priority mode, when the reward r increases due to the increase in the depth of cut and / or cutting speed in one operation... cycle In cases where the force (torque) of all tools, feed axes, and spindle is increased, the waveforms of the force (torque) of all tools, feed axes, and spindle are dissimilar, and the machining time is reduced, as a behavior a' for state s', such as choosing to increase the depth of cut and / or cutting speed incrementally to shorten the machining time.
[0112] Alternatively, the following strategy can be adopted: In machine learning under the processing time priority mode, when the reward r increases due to the increase in the depth of cut and / or cutting speed in one operation... cycle In the case of reduction, as a behavior a' for state s', such as choosing to return the depth of cut and / or cutting speed of one cut to the previous one, thus shortening the processing time required, behavior a'.
[0113] Alternatively, the following strategy can be adopted: In machine learning under the tool life priority mode, when the reward r is reduced due to the decrease in the depth of cut and / or cutting speed in one pass... tool In cases where the force (torque) of all tools, feed axes, and spindle is increased, the waveforms of the force (torque) of all tools, feed axes, and spindle are dissimilar, and the machining time is increased, decreased, or maintained, the action a' for state s' may be, for example, choosing to delay the reduction of tool life by progressively decreasing the depth of cut and / or the cutting speed.
[0114] Alternatively, the following strategy can be adopted: In machine learning under tool life priority mode, when the return r is reduced due to the decrease in depth of cut and / or cutting speed... tool In the case of reduction, as a behavior a' for state s', such as choosing to return the depth of cut and / or cutting speed to the previous one, such as behavior a' that delays the reduction of tool life.
[0115] Alternatively, the behavior information generation unit 223 may also adopt the following strategies: a well-known method of selecting behavior a with the highest value function Q(s, a) by a greedy method, or selecting behavior a' randomly with a small probability ε and otherwise selecting behavior a with the highest value function Q(s, a) by an ε-greedy method.
[0116] The behavior information output unit 203 outputs the behavior information 'a' output from the learning unit 202 to the numerical control device 101. For example, the behavior information output unit 203 can output updated values of the feed rate and cutting speed for the first cut, which are considered as behavior information, to the numerical control device 101. Therefore, the numerical control device 101 updates the feed rate and cutting speed set in the fixed cycle based on the received updated values of the feed rate and cutting speed. Furthermore, the numerical control device 101 generates a motion command based on the updated feed rate and cutting speed set in the fixed cycle, and causes the machine tool 10 to perform cutting operations according to the generated motion command.
[0117] In addition, the behavior information output unit 203 can output a machining program, which is updated based on the updated values of the cutting depth and cutting speed, to the numerical control device 101 as behavior information.
[0118] Value function storage unit 204 stores the value function Q of the processing time priority mode. cycle And the value function Q of the tool life priority mode tool The storage device. Value function Q cycle Q tool For example, they can be stored as tables (hereinafter also referred to as "behavior value tables") based on state s and behavior a, respectively. The value function Q is stored in the value function storage unit 204. cycle Q tool Updated by value function update section 222.
[0119] The optimal behavior information output unit 205 updates the value function Q based on Q-learning performed by the value function update unit 222. cycle or value function Q tool Generate behavioral information a (hereinafter also referred to as "optimal behavioral information") for the numerical control device 101 to perform actions that maximize the value function value.
[0120] More specifically, the optimal behavior information output unit 205 obtains the value function Q of the processing time priority pattern stored in the value function storage unit 204. cycle And the value function Q of the tool life priority mode tool The value function Q cycle Q tool The function is updated by Q-learning through the value function update unit 222 as described above. Furthermore, the optimal behavior information output unit 205 generates a value function Q based on the acquired processing time priority pattern. cycle Behavioral information and the value function Q based on the acquired processing time priority pattern tool The generated behavior information is output to the numerical control device 101. This optimal behavior information, similar to the behavior information output by the behavior information output unit 203 during Q-learning, includes information representing the updated depth of cut and cutting speed values.
[0121] The functional blocks included in the machine learning device 20 have been described above.
[0122] To implement these functional blocks, the machine learning device 20 includes a computing processing unit such as a CPU. In addition, the machine learning device 20 also includes auxiliary storage devices such as HDDs that store application software, operating systems, and other control programs, as well as main storage devices such as RAM that store data temporarily needed by the computing processing unit when executing programs.
[0123] Furthermore, in the machine learning device 20, the processing unit reads application software and operating system from the auxiliary storage device, expands the read application software and operating system in the main storage device, and performs processing based on these application software and operating system. Additionally, based on the processing results, it controls various hardware components of the machine learning device 20. Thus, the functional blocks of this embodiment are implemented. In other words, this embodiment can be implemented through hardware and software cooperation.
[0124] Regarding the machine learning device 20, since the computational demands associated with machine learning increase, a technique using GPUs (Graphics Processing Units) in personal computers, known as GPGPUs (General-Purpose Computing on Graphics Processing Units), can be employed to perform high-speed processing when using GPUs for computational tasks accompanying machine learning. Furthermore, for even faster processing, multiple computers equipped with such GPUs can be used to construct a computer cluster, allowing parallel processing across the various computers within the cluster.
[0125] Next, refer to Figure 3 The flowchart illustrates the operation of the machine learning device 20 during Q-learning in this embodiment.
[0126] Figure 3 This is a flowchart illustrating the operation of the machine learning device 20 during Q-learning in the first embodiment.
[0127] In step S11, the control unit 206 sets the number of trials to the initial setting, i.e., "1", and instructs the status information acquisition unit 201 to acquire status information.
[0128] In step S12, the status information acquisition unit 201 acquires the initial status data from the numerical control device 101. The acquired status data is then output to the behavior information generation unit 223. As described above, this status data (status information) corresponds to the status s in Q-learning, including the depth of cut, cutting speed, tool material, tool shape, tool diameter, tool length, remaining tool life, workpiece material, cutting conditions in the tool catalog, spindle speed, motor current, machine temperature, and ambient temperature at the time of step S12. Furthermore, the status data at the initial Q-learning start time is pre-generated by the operator.
[0129] In step S13, the behavior information generation unit 223 generates new behavior information 'a' for either the machining time priority mode or the tool life priority mode through machine learning, and outputs the generated new behavior information 'a' to the numerical control device 101 via the behavior information output unit 203. The numerical control device 101 executes a machining program that updates the depth of cut and cutting speed set once in a fixed cycle, based on the behavior information 'a' selected by the setting device 111 from the received behavior information 'a' of the machining time priority mode and tool life priority mode. The numerical control device 101 generates motion commands based on the updated machining program, and causes the machine tool 10 to perform cutting operations according to the generated motion commands.
[0130] In step S14, the status information acquisition unit 201 acquires status data corresponding to the new status s' from the numerical control device 101. Here, the new status data includes the depth of cut, cutting speed, tool material, tool shape, tool diameter, tool length, remaining tool life, material of the workpiece being machined, cutting conditions in the tool catalog, spindle speed, motor current, machine temperature, and ambient temperature. The status information acquisition unit 201 outputs the acquired status data to the learning unit 202.
[0131] In step S15, the status information acquisition unit 201 acquires determination information for the new status s'. This determination information includes the following data acquired from the machine tool 10 during step S13 when executing the updated machining program: the pressure intensity applied to the tool (axial and rotational directions), the waveform shape of the pressure applied to the tool (axial and rotational directions), the intensity of the torque applied to the feed axis, the waveform shape of the torque applied to the feed axis, the intensity of the torque applied to the spindle, the waveform shape of the torque applied to the spindle, and the machining time required to execute the updated machining program. The acquired determination information is then output to the learning unit 202.
[0132] In step S16, the reward calculation unit 221 performs reward calculation processing based on the obtained determination information, and calculates the reward r for the processing time priority mode respectively. cycle And the reward r of the tool life priority mode tool Furthermore, the detailed process of reward calculation and processing is described later.
[0133] In step S17, the value function update unit 222 updates the value function based on the calculated return r. cycle and returns tool Update the value function Q stored in value function storage unit 204 respectively. cycle and value function Q tool .
[0134] In step S18, the control unit 206 determines whether the maximum number of trials has been reached since the start of machine learning. A maximum number of trials is preset. If the maximum number of trials has not been reached, the number of trials is counted in step S19, and the process returns to step S13. Steps S13 to S19 are repeated until the maximum number of trials is reached.
[0135] also, Figure 3 The process ends when the maximum number of trials is reached, but it can also end when the accumulated time from the start of machine learning for steps S13 to S19 exceeds a preset maximum elapsed time (or is greater than the preset maximum elapsed time).
[0136] In addition, step S17 illustrates online updates, but it can also be replaced by batch updates or small-batch updates.
[0137] Figure 4 Yes Figure 3 The flowchart below describes the detailed processing steps of the return calculation process shown in step S16.
[0138] In step S61, the report calculation unit 221 determines the value P of the intensity of the pressure (torque) applied to the tool, feed axis, and spindle included in the determination information of state s'. t (s'), P f (s'), P m (s') Whether all of the pressure (torque) applied to the tool, feed axis, and spindle contained in the determination information of state s is the value P. t (s), P f (s), P m (s) Small, i.e. weak. The value P of the intensity of the pressure (torque) applied to the tool, feed axis, and spindle in state s'. t (s'), P f (s'), P m If all conditions (s') are weaker than state s, the process proceeds to step S62. Additionally, the value P of the intensity of the pressure (torque) applied to the tool, feed axis, and spindle in state s' is determined. t (s'), P f (s'), P m If at least one of (s') is stronger than state s, the process proceeds to step S63.
[0139] In step S62, the reward calculation unit 221 calculates the reward r. p Set to a negative value.
[0140] In step S63, the reward calculation unit 221 calculates the reward r.p Set to a positive value.
[0141] In step S64, the report calculation unit 221 determines the waveform shape WF of the pressure (torque) applied to the tool, feed axis, and spindle included in the determination information of state s'. t (s'), WF f (s'), WF m (s') Are all waveforms similar to those indicating a precursor to tool failure or a further reduction in tool life? The waveform shape WF of the pressure (torque) applied to the tool, feed axis, and spindle in state s'. t (s'), WF f (s'), WF m If all (s') are dissimilar, the process proceeds to step S66. Additionally, the waveform shape WF of the pressure (torque) applied to the tool, feed axis, and spindle in state s' is... t (s'), WF f (s'), WF m If at least one of the similar cases (s') occurs, the process proceeds to step S65.
[0142] In step S65, the reward calculation unit 221 calculates the reward r. w Set to a negative value.
[0143] In step S66, the reward calculation unit 221 calculates the reward r. w Set to a positive value.
[0144] In step S67, the report calculation unit 221 determines whether the value T(s') of the processing time required in the determination information of state s' increases, decreases, or remains the same compared to the value T(s) of the processing time required in the determination information of state s. If the value T(s') of the processing time required in state s' increases compared to state s, the process proceeds to step S68. If the value T(s') of the processing time required in state s' decreases compared to state s, the process proceeds to step S70. If the value T(s') of the processing time required in state s' remains the same, the process proceeds to step S69.
[0145] In step S68, the reward calculation unit 221 calculates the reward r. c Set to a negative value.
[0146] In step S69, the reward calculation unit 221 calculates the reward r. c Set it to zero.
[0147] In step S70, the reward calculation unit 221 calculates the reward r. c Set to a positive value.
[0148] In step S71, the return calculation unit 221 uses the calculated return r p r w r c Using mathematical formula 2, calculate the reward r for the processing time priority mode. cycle And the reward r of the tool life priority mode tool Therefore, the reward calculation process ends, and the process proceeds to step S17.
[0149] The above is based on reference. Figure 3 as well as Figure 4 The described actions, in this embodiment, can generate a value function Q that optimizes fixed cycles of the processing procedure in multi-product, variable-quantity production environments without increasing operator time and effort. cycle Q tool .
[0150] Next, refer to Figure 5 The flowchart describes the actions of the optimal behavior information output unit 205 when generating optimal behavior information.
[0151] In step S21, the optimal behavior information output unit 205 obtains the value function Q of the processing time priority pattern stored in the value function storage unit 204. cycle And the value function Q of the tool life priority mode. tool .
[0152] In step S22, the optimal behavior information output unit 205 outputs the value function Q based on the obtained value function. cycle and the value function Q tool The optimal behavior information for the processing time priority mode and the tool life priority mode is generated respectively, and the generated optimal behavior information for the processing time priority mode and the tool life priority mode is output to the numerical control device 101.
[0153] As described above, the numerical control device 101 updates the machining program, which sets the depth of cut and cutting speed once in a fixed cycle, based on the behavior in either the machining time priority mode or the tool life priority mode selected by the setting device 111. This allows for optimization of the machining program in multi-product, variable-quantity production environments without increasing operator time and effort. Consequently, the numerical control device 101 can prioritize machining based on the required machining time (i.e., cycle time) or tool life.
[0154] In addition, the numerical control device 101 eliminates the need for the operator to set the input depth and cutting speed variables once, thus reducing the time and effort required to create machining programs.
[0155] The first embodiment has been described above.
[0156] <Second Implementation Method>
[0157] Next, the second embodiment will be described. In the second embodiment, the machine learning device 20A, in addition to the functions of the first embodiment, also has the following functions: for a machining program containing two or more (e.g., n) fixed cycles, whenever each fixed cycle (e.g., the i-th fixed cycle) is executed, the machining program is stopped, the state s(i), behavior a(i), decision information (i), reward r(i), and behavior a'(i) for the state s'(i) of the i-th fixed cycle are calculated, and the cut depth and cutting speed of one operation in the i-th fixed cycle are updated. Furthermore, n is an integer of 2 or more, and i is an integer from 1 to n.
[0158] Therefore, the cut depth and cutting speed set in the i-th fixed loop can be defined as the behavior for the i-th fixed loop. Hereinafter, the i-th fixed loop will also be referred to as "fixed loop (i)" (1≤i≤n).
[0159] The second embodiment will be described below.
[0160] <Second Implementation Method>
[0161] Figure 6 This is a functional block diagram illustrating an example of the functional structure of the numerical control system according to the second embodiment. Furthermore, for systems having... Figure 1 Elements with the same function in the numerical control system 1 are labeled with the same symbols, and detailed descriptions are omitted.
[0162] like Figure 6 As shown, the numerical control system 1 of the second embodiment includes a machine tool 10 and a machine learning device 20A.
[0163] The machine tool 10 is similar to that in the first embodiment, and is a machine tool well known to those skilled in the art, including a numerical control device 101a. The machine tool 10 operates according to action commands from the numerical control device 101a.
[0164] Similar to the first embodiment, the numerical control device 101a is a well-known numerical control device to those skilled in the art. It generates motion commands based on a machining program obtained from an external device (not shown) such as a CAD / CAM device, and sends the generated motion commands to the machine tool 10. Thus, the numerical control device 101a controls the operation of the machine tool 10.
[0165] Furthermore, in the case of executing a machining program, the numerical control device 101a of the second embodiment stops the machining program whenever n fixed cycles (i) such as drilling and tapping included in the machining program are completed, and outputs information related to the tool and workpiece set for the machine tool 10 in the fixed cycle, the depth of cut and cutting speed set in the fixed cycle (i), and the measured values obtained from the machine tool 10 by executing the machining program to the machine learning device 20A.
[0166] Furthermore, the setting device 111 has the same function as the setting device 111 in the first embodiment.
[0167] <Machine Learning Device 20A>
[0168] The machine learning device 20A is a device for reinforcement learning of the feed rate and cutting speed of each of the n fixed cycles in the machining program when the machine tool 10 is moved by the numerical control device 101a executing the machining program.
[0169] Figure 7 This is a functional block diagram representing a functional structure example of the machine learning device 20A.
[0170] like Figure 7 As shown, the machine learning device 20A includes: a state information acquisition unit 201a, a learning unit 202a, a behavior information output unit 203a, a value function storage unit 204a, an optimal behavior information output unit 205a, and a control unit 206. The learning unit 202a includes: a reward calculation unit 221a, a value function update unit 222a, and a behavior information generation unit 223a.
[0171] Furthermore, the control unit 206 has the same functions as the control unit 206 in the first embodiment.
[0172] The status information acquisition unit 201a, as the status of the machine tool 10, acquires status data s from the numerical control device 101 each time it executes each of the n fixed cycles contained in the machining program. The status data s includes information related to the tool and workpiece set in the machine tool 10, the depth of cut and cutting speed set once in each fixed cycle (i) (1≤i≤n), and the measured values obtained from the machine tool 10 by executing the machining program.
[0173] The status information acquisition unit 201a outputs the status data s(i) acquired according to a fixed cycle (i) to the learning unit 202a.
[0174] Furthermore, the state information acquisition unit 201a can store the state data s(i) acquired according to a fixed cycle (i) in a storage unit (not shown) included in the machine learning device 20A. In this case, the learning unit 202a, described later, can read the state data s(i) of each fixed cycle (i) from the storage unit (not shown) of the machine learning device 20A.
[0175] In addition, the status information acquisition unit 201a also acquires determination information for calculating the reward of Q-learning according to a fixed cycle (i). Specifically, the determination information for calculating the reward of Q-learning is obtained from the machine tool 10 by executing the fixed cycle (i) contained in the machining program related to the status information s (i), including the pressure intensity applied to the tool (axial and rotational directions), the waveform shape of the pressure applied to the tool (axial and rotational directions), the intensity of the torque applied to the feed axis, the waveform shape of the torque applied to the feed axis, the intensity of the torque applied to the spindle, the waveform shape of the torque applied to the spindle, and the machining time required to execute the fixed cycle (i).
[0176] The learning unit 202a is the part that learns the value function Q(s(i), a(i)) when choosing a certain behavior a(i) under a certain state data (environment state) s(i) in each fixed loop (i). Specifically, the learning unit 202a has: a reward calculation unit 221a, a value function update unit 222a, and a behavior information generation unit 223a.
[0177] Furthermore, the learning unit 202a determines whether to continue learning in the same way as the learning unit 202 in the first embodiment. For example, it can determine whether to continue learning based on whether the number of trials of the processing procedure since the start of machine learning has reached the maximum number of trials, or whether the elapsed time since the start of machine learning has exceeded a predetermined time (or is more than the predetermined time).
[0178] In each fixed cycle (i), the reward calculation unit 221a calculates the reward when behavior a(i) is selected in a certain state s(i) based on the determination information of each fixed cycle (i). Furthermore, the reward calculated in each fixed cycle (i) is similar to that in the first embodiment, calculated based on (1) the intensity of the pressure (torque) applied to the tool, feed axis, and spindle, (2) the waveform shape of the pressure (torque) applied to the tool, feed axis, and spindle, and (3) the processing time. That is, for example, the reward r in the first embodiment... p r w r c Similarly, calculate the return r for each item in the fixed loop (i). p (i), r w (i), r c (i).
[0179] Furthermore, the reward calculation unit 221a can, in the same manner as the reward calculation unit 221 in the first embodiment, use the reward r for each item. p (i), r w (i), r c (i) and mathematical formula 2 are used to calculate the reward r of the processing time priority mode in the fixed loop (i). cycle (i) and the return r of the processing life priority mode tool (i).
[0180] Similar to the value function update unit 222 in the first embodiment, the value function update unit 222a, during machine learning in the processing time priority mode, updates the value function based on the state s(i) in the fixed loop (i), the behavior a(i), the state s'(i) when the behavior a(i) is applied to the state s(i), and the reward r calculated as described above. cycle The value of (i) is used for Q-learning, thereby updating the value function Q of the fixed loop (i) stored in the value function storage unit 204a. cycle_i Furthermore, during machine learning in the tool lifetime priority mode, the value function update unit 222a updates the value based on the state s(i) in the fixed loop (i), the action a(i), the state s'(i) when the action a(i) is applied to the state s(i), and the reward r calculated as described above. tool The value of (i) is used for Q-learning, thereby updating the value function Q stored in the value function storage unit 204a. tool_i .
[0181] Similar to the behavior information generation unit 223 in the first embodiment, the behavior information generation unit 223a selects behavior a(i) during the Q-learning process for the current state s(i) in the fixed loop (i). During the Q-learning process corresponding to the machining time priority mode or the tool life priority mode, in order to perform the action of correcting the feed amount and cutting speed of the i-th fixed loop once (equivalent to behavior a in Q-learning), the behavior information generation unit 223a generates behavior information a for the i-th fixed loop and outputs the generated behavior information a for the i-th fixed loop to the behavior information output unit 203a.
[0182] Similar to the behavior information output unit 203 in the first embodiment, the behavior information output unit 203a outputs the behavior information a(i) of each fixed cycle (i) output from the learning unit 202a to the numerical control device 101a. For example, the behavior information output unit 203a can output updated values of the depth of cut and cutting speed, which are behavior information for each fixed cycle (i), to the numerical control device 101a. Therefore, the numerical control device 101a updates each of the n fixed cycles (i) included in the machining program based on the received updated values of the depth of cut and cutting speed. Furthermore, the numerical control device 101a generates motion commands based on the machining program including the updated fixed cycles (i), and causes the machine tool 10 to perform cutting operations based on the generated motion commands.
[0183] In addition, the behavior information output unit 203a can output the machining program of each fixed cycle (i), which is behavior information for each fixed cycle (i) and updated according to the updated values of the infeed and cutting speed, to the numerical control device 101a.
[0184] Value function storage unit 204a stores the value function Q of the processing time priority pattern for each fixed cycle (i). cycle_i And the value function Q of the tool life priority mode tool_i The storage device. Furthermore, the value function Q cycle_i The set of (1≤i≤n) and the value function Q cycle The relationship and the value function Q tool_i The set of (1≤i≤n) and the value function Q tool The relationship is shown in mathematical formula 3.
[0185]
Mathematical Expression 3
[0186]
[0187]
[0188] The value function Q of each fixed cycle (i) stored in the value function storage unit 204a cycle_i Q tool_i Updated by value function update section 222.
[0189] The optimal behavior information output unit 205a, like the optimal behavior information output unit 205 in the first embodiment, updates the value function Q of the processing time priority mode based on Q-learning performed by the value function update unit 222a. cycle Or the value function Q of the tool life priority mode toolGenerate behavioral information (optimal behavioral information) a in a fixed loop (i) that maximizes the value function of the numerical control device 101a.
[0190] More specifically, the optimal behavior information output unit 205a obtains the value function Q of the processing time priority pattern stored in the value function storage unit 204. cycle And the value function Q of the tool life priority mode tool Furthermore, the optimal behavior information output unit 205a generates a value function Q based on the acquired processing time priority pattern. cycle Behavioral information in the fixed loop (i) and the value function Q based on the acquired processing time priority pattern. tool The behavior information in the fixed loop (i) is generated and output to the numerical control device 101a. In this optimal behavior information, similar to the behavior information output by the behavior information output unit 203a during Q learning, information representing the updated depth of cut and cutting speed values is included.
[0191] The functional blocks included in the machine learning device 20A have been described above.
[0192] Next, refer to Figure 8 The flowchart illustrates the operation of the machine learning device 20A during Q-learning in this embodiment.
[0193] Figure 8 This is a flowchart illustrating the operation of the machine learning device 20A during Q-learning in the second embodiment. Furthermore, regarding... Figure 8 The flowchart with Figure 3 Processes that follow the same steps are labeled with the same step numbers, and detailed explanations are omitted.
[0194] In step S11a, the control unit 206 sets the number of trial runs j of the processing program to the initial setting, i.e., "1", and instructs the status information acquisition unit 201a to acquire status information.
[0195] In step S11b, the control unit 206 initializes i to "1".
[0196] In step S12a, the status information acquisition unit 201a acquires the status data s(i) of the fixed loop (i) from the numerical control device 101a. The acquired status data s(i) is output to the behavior information generation unit 223a. As described above, this status data (status information) s(i) is information equivalent to the status s(i) in the fixed loop (i) of Q-learning, including the depth of cut, cutting speed, tool material, tool shape, tool diameter, tool length, remaining tool life, material of the workpiece being processed, cutting conditions in the tool catalog, spindle speed, motor current value, machine temperature, and ambient temperature at the time of step S12a. Furthermore, the status data at the initial time of Q-learning is generated in advance by the operator.
[0197] In step S13a, the behavior information generation unit 223a generates new behavior information a(i) for the fixed loop (i) of the machining time priority mode or tool life priority mode through machine learning, and outputs the generated new behavior information a(i) of the machining time priority mode and tool life priority mode to the numerical control device 101a via the behavior information output unit 203a. The numerical control device 101a executes a machining program that updates the depth of cut and cutting speed set once in the fixed loop (i) based on the received behavior information a(i) of the machining time priority mode and tool life priority mode selected by the setting device 111. The numerical control device 101a generates motion commands based on the updated fixed loop (i) and causes the machine tool 10 to perform cutting operations based on the generated motion commands. Furthermore, the numerical control device 101a stops the machining program upon completion of the fixed loop (i).
[0198] In step S14, the status information acquisition unit 201a performs the same processing as step S14 in the first embodiment to acquire new status data s'(i) in the fixed cycle (i) obtained from the numerical control device 101a.
[0199] In step S15, the state information acquisition unit 201a performs the same processing as step S15 in the first embodiment to acquire determination information for the new state s'(i) in the fixed loop (i). The acquired determination information is then output to the learning unit 202a.
[0200] In step S16, the reward calculation unit 221a performs the same processing as step S16 in the first embodiment, and performs the calculation based on the obtained determination information. Figure 4 The reward calculation process calculates the reward r for the fixed loop (i) in the processing time priority mode. cycle (i) and the reward r of the fixed cycle (i) of the tool life priority mode. tool (i).
[0201] In step S17, the value function update unit 222a performs the same processing as step S17 in the first embodiment, and updates the value function based on the calculated return r of the fixed cycle (i). cycle (i) and the return r tool (i) respectively update the value function Q of the processing time priority mode of the fixed cycle (i) stored in the value function storage unit 204a. cycle_i And the value function Q of the tool life priority mode tool_i .
[0202] In step S17a, the control unit 206 determines whether i is less than n. If i is less than n, the process proceeds to step S17b. On the other hand, if i is greater than or equal to n, the process proceeds to step S18.
[0203] In step S17b, the control unit 206 increments i by "1". The process returns to step S12a.
[0204] In step S18, the control unit 206 performs the same processing as in step S18 of the first embodiment, determining whether the number of trials j of the processing procedure since the start of machine learning has reached the maximum number of trials. If the maximum number of trials has not been reached, the number of trials j is incremented by "1" in step S19, and the process returns to step S11b. The processing from step S11b to step S19 is repeated until the maximum number of trials is reached.
[0205] also, Figure 8 The process ends when the number of trials j of the processing procedure reaches the maximum number of trials. However, the process can also end when the time accumulated from the start of machine learning to the processing time of steps S11b to S19 exceeds the preset maximum elapsed time (or is greater than the maximum elapsed time).
[0206] In addition, step S17 illustrates online updates, but it can also be replaced by batch updates or small-batch updates.
[0207] The above is based on reference. Figure 8 The described action, in this embodiment, can generate a value function Q that optimizes fixed cycles of the processing procedure in multi-product, variable-quantity production environments without increasing operator time and effort. cycle Q tool .
[0208] Furthermore, regarding the actions of the optimal behavior information output unit 205a in generating optimal behavior information, apart from the fact that it generates optimal behavior information according to a fixed loop (i), it is similar to... Figure 5 process Figure 1 (The explanation is omitted.)
[0209] As described above, the numerical control device 101a updates the machining program, which sets the depth of cut and cutting speed once in each fixed cycle (i), based on the behavior of the machining time priority mode or tool life priority mode selected by the setting device 111. This allows for optimization of the machining program in multi-product, variable-quantity production environments without increasing operator time and effort. Thus, the numerical control device 101 can prioritize machining based on the required machining time (i.e., cycle time) or tool life.
[0210] In addition, the numerical control device 101a eliminates the need for the operator to set the input depth and cutting speed variables once, thus reducing the time and effort required to create machining programs.
[0211] The second embodiment has been described above.
[0212] The first and second embodiments have been described above, but the numerical control devices 101, 101a and the machine learning devices 20, 20A are not limited to the embodiments described above, and include variations and improvements within the scope of achieving the purpose.
[0213] <Variation Example 1>
[0214] In the first and second embodiments described above, the machine learning devices 20 and 20A alternately perform machine learning in a processing time priority mode and a tool life priority mode, but are not limited thereto. For example, the machine learning devices 20 and 20A may also perform machine learning in a processing time priority mode and a tool life priority mode, respectively.
[0215] <Variation Example 2>
[0216] In addition, for example, in the first and second embodiments described above, the setting device 111 selects either the behavior in the processing time priority mode or the behavior in the tool life priority mode based on a comparison between the remaining tool life of the tool being used in the machine tool 10 and a preset threshold, but is not limited to this.
[0217] For example, if the remaining tool life is 5%, the number of remaining machining parts is 3, and the tool life decreases by 0.1% per machining cycle, the remaining tool life after machining a workpiece with 3 remaining machining parts is 4.7%, which will not become 0%. Therefore, even if the remaining tool life is below the threshold, machining a workpiece with 3 remaining machining parts will not result in 0% remaining tool life. In this case, the setting device 111 can select the behavior in the machining time priority mode.
[0218] Therefore, even when the remaining tool life is short, as long as the remaining tool life is sufficient relative to the number of remaining machined parts, machining can be performed without reducing the machining time (cycle time).
[0219] <Variation Example 3>
[0220] In addition, for example, in the first and second embodiments described above, the machine learning devices 20 and 20A are shown to be different from the numerical control devices 101 and 101a, but the numerical control devices 101 and 101a may also have some or all of the functions of the machine learning devices 20 and 20A.
[0221] Alternatively, the server may include, for example, a state information acquisition unit 201, a learning unit 202, a behavior information output unit 203, a value function storage unit 204, an optimal behavior information output unit 205, and a control unit 206 of the machine learning device 20, or a portion or all of the state information acquisition unit 201a, learning unit 202a, behavior information output unit 203a, value function storage unit 204a, optimal behavior information output unit 205a, and control unit 206 of the machine learning device 20A. Furthermore, the functions of the machine learning devices 20 and 20A can be implemented in the cloud using virtual server functionality, etc.
[0222] Furthermore, the machine learning devices 20 and 20A can also be distributed processing systems that appropriately distribute the functions of the machine learning devices 20 and 20A to multiple servers.
[0223] <Variation Example 4>
[0224] Furthermore, for example, in the first and second embodiments described above, in the control system 1, one machine tool 10 and one machine learning device 20, 20A can be communicatively connected, but this is not a limitation. For example Figure 9As shown, the control system 1 may have m machine tools 10A(1)-10A(m) and m machine learning devices 20B(1)-20B(m) (m is an integer greater than or equal to 2). In this case, the machine learning device 20B(j) can be connected to the machine tool 10A(j) in a one-to-one communication manner via the network 50 to perform machine learning on the machine tool 10A(j) (j is an integer from 1 to m).
[0225] Furthermore, the value function Q stored in the value function storage unit 204 (204a) of the machine learning device 20B(j) cycle Q tool (Q cycle_i Q tool_i The value function Q can be shared among other machine learning devices 20B(k) (where k is an integer from 1 to m, k ≠ j). cycle Q tool (Q cycle_i Q tool_i This allows reinforcement learning to be distributed across various machine learning devices 20B, thereby improving the efficiency of reinforcement learning.
[0226] Furthermore, each of machine tools 10A(1)-10A(m) is related to Figure 1 or Figure 6 The machine tool 10 corresponds to it. Additionally, each of the machine learning devices 20B(1)-20B(m) is associated with... Figure 1 machine learning device 20 or Figure 6 The machine learning device 20A corresponds to this.
[0227] In addition, such as Figure 10 As shown, server 60 can act as machine learning device 20 (20A) and can communicate with m machine tools 10A (1)-10A (m) via network 50 to perform machine learning on each of the machine tools 10A (1)-10A (m).
[0228] Furthermore, the functions included in the numerical control devices 101, 101a and the machine learning devices 20, 20A in the first and second embodiments can be implemented respectively by hardware, software, or a combination thereof. Here, implementation by software means implementation by reading and executing a program into a computer.
[0229] The various structural units included in the numerical control devices 101 and 101a and the machine learning devices 20 and 20A can be implemented using hardware, software, or a combination thereof, including electronic circuits. In the case of software implementation, the program constituting the software is installed on a computer. Alternatively, these programs can be recorded on removable media and distributed to users, or downloaded to users' computers via a network. In the case of hardware implementation, for example, integrated circuits (ICs) such as ASICs (Application Specific Integrated Circuits), gate arrays, FPGAs (Field Programmable Gate Arrays), and CPLDs (Complex Programmable Logic Devices) can constitute part or all of the functionality of the various structural units included in the aforementioned devices.
[0230] Programs can be stored and provided to a computer using various types of non-transitory computer-readable media. Non-transitory computer-readable media include various types of tangible storage media. Examples of non-transitory computer-readable media include magnetic recording media (e.g., floppy disks, magnetic tapes, hard disks), optical-magnetic recording media (e.g., optical discs), CD-ROMs (Read Only Memory), CD-Rs, CD-R / Ws, and semiconductor memories (e.g., mask ROMs, PROMs (Programmable ROMs), EPROMs (Erasable PROMs), flash memory ROMs, and RAM). Alternatively, programs can also be provided to a computer using various types of transient computer-readable media. Examples of transient computer-readable media include electrical signals, optical signals, and electromagnetic waves. Transient computer-readable media can provide programs to a computer via wired communication paths such as wires and optical fibers, or via wireless communication paths.
[0231] Furthermore, the steps describing a program recorded in a recording medium naturally include processes performed in that order and in chronological order, as well as processes that are not necessarily performed in chronological order and processes that are performed in parallel or separately.
[0232] In other words, the machine learning apparatus, setting apparatus, numerical control system, numerical control apparatus, and machine learning method disclosed herein can be implemented in various ways having the following structures.
[0233] (1) The machine learning apparatus 20 of this disclosure performs machine learning on a numerical control device 101 that moves a machine tool 10 according to a machining program. It includes: a state information acquisition unit 201 that executes a machining program with a cutting depth and cutting speed set at least once by the numerical control device 101 to make the machine tool 10 perform cutting machining, thereby acquiring state information including the cutting depth and cutting speed of one cut; a behavior information output unit 203 that outputs behavior information including the adjustment information of the cutting depth and cutting speed of one cut contained in the state information; a reward calculation unit 221 that acquires determination information and outputs a reward value in reinforcement learning corresponding to a specified condition based on the acquired determination information, wherein the determination information is information related to at least the following: the pressure intensity applied to the tool during cutting machining, the waveform shape of the pressure applied to the tool, and the time required for machining; and a value function update unit 222 that updates the value function Q based on the reward value, the state information, and the behavior information.
[0234] According to the machine learning device 20, the processing procedure can be optimized without increasing the operator's time and effort.
[0235] (2) In the machine learning device 20 described in (1), the specified condition may be either a condition that prioritizes processing time or a condition that prioritizes tool life, and the reward calculation unit 221 outputs a reward r under the condition that prioritizes processing time. cycle Output return r while prioritizing tool lifespan tool Under the condition of prioritizing processing time, the value function update unit 222 updates the value function based on the return r. cycle The value function Q is updated using state information and behavioral information. cycle Under the condition of prioritizing the tool's lifespan, the value function update unit 222 updates according to the return r. tool The value function Q is updated using state information and behavioral information. tool .
[0236] Therefore, a value function Q for optimizing fixed cycles of the machining program can be generated without increasing the operator's time and effort. cycle Q tool .
[0237] (3) In the machine learning apparatus 20, 20A described in (2), machine learning may be performed each time the processing program is executed, or each fixed loop of the multiple fixed loops contained in the processing program is executed.
[0238] Therefore, the machining program can be optimized by processing the workpiece or by a fixed cycle.
[0239] (4) In the machine learning apparatus 20, 20A described in (2) or (3), the machine learning apparatus may also have: an optimal behavior information output unit 205, 205a, whose output is based on the reward r cycle Updated value function Q cycle The value is the largest behavioral information, and based on the reward r tool Updated value function Q tool The value is the largest behavioral information.
[0240] Therefore, machine learning devices 20 and 20A can optimize the machining program based on the tool status.
[0241] (5) In the machine learning device 20 described in (1), it is also possible that when the processing time required for determining the information contained therein is less than the processing time required for the previous processing, the reward calculation unit 221 will report r. cycle r tool If the processing time is set to a positive value, and the processing time is longer than the previous processing time, the reward calculation unit 221 will report r. cycle r tool Set to a negative value.
[0242] Therefore, the machine learning device 20 can optimize the processing procedure based on the processing time required.
[0243] (6) In the machine learning device 20 described in (1), it is also possible that, if the waveform shape of the pressure applied to the tool contained in the determination information is not similar to at least the waveform shape indicating a sign of tool damage and the waveform shape indicating a sharp decrease in the tool's lifespan, the report calculation unit 221 will report r. cycle r tool If the waveform of the pressure applied to the tool is set to a positive value, and the waveform shape is at least similar to a waveform shape indicating a sign of tool damage or a waveform shape indicating a sharp reduction in tool life, the report calculation unit 221 will report r. cycle r tool Set to a negative value.
[0244] Therefore, the machine learning device 20 can optimize the machining process while taking machining safety into account.
[0245] (7) In any of the machine learning devices 20 and 20A described in (1) to (6), the maximum number of machine learning trials can be set to perform machine learning.
[0246] Therefore, machine learning devices 20 and 20A can avoid long-term machine learning operations.
[0247] (8) The setting device 111 of this disclosure selects a behavior from the behaviors obtained from the machine learning device described in any one of (1) to (7) according to a preset threshold, and sets the selected behavior to the processing program.
[0248] According to the setting device 111, the same effect as (1) to (7) can be obtained.
[0249] (9) The numerical control system 1 of this disclosure includes: a machine learning device 20, 20A as described in any one of (1) to (7); a setting device 111 as described in (8); and numerical control devices 101, 101a, which execute a processing program set by the setting device 111.
[0250] According to the numerical control system 1, the same effect as (1) to (7) can be obtained.
[0251] (10) The numerical control devices 101 and 101a of this disclosure include: the machine learning devices 20 and 20A described in any one of (1) to (7); and the setting device 111 described in (8), wherein the numerical control device executes a processing program set by the setting device 111.
[0252] According to the numerical control devices 101 and 101a, the same effect as (1) to (7) can be obtained.
[0253] (11) The numerical control method disclosed herein is a machine learning method of machine learning devices 20 and 20A. Machine learning devices 20 and 20A perform machine learning on numerical control devices 101 and 101a that cause machine tool 10 to move according to machining program. The numerical control devices 101 and 101a execute machining program with at least one set depth of cut and cutting speed, so that machine tool 10 performs cutting machining. Thus, state information including depth of cut and cutting speed is obtained; behavioral information is output, which includes adjustment information of depth of cut and cutting speed included in state information; decision information is obtained, and a reward value in reinforcement learning corresponding to specified conditions is output according to the obtained decision information. The decision information is information related to at least the following: pressure intensity applied to the tool during cutting machining, waveform shape of pressure applied to the tool, and machining time; the value function Q is updated according to the reward value, state information, and behavioral information.
[0254] According to this numerical control method, the same effect as (1) can be obtained.
[0255] Symbol Explanation
[0256] 1 Numerical Control System
[0257] 10 machine tools
[0258] 101, 101a Numerical Control Device
[0259] 111 Setting device
[0260] 20, 20A Machine Learning Device
[0261] 201, 201a Status Information Acquisition Department
[0262] 202, 202a Study Department
[0263] 221, 221a Return Calculation Department
[0264] 222, 222a Value Function Update Section
[0265] 223, 223a Behavioral Information Generation Department
[0266] 203, 203a Behavioral Information Output Department
[0267] Value function storage section 204, 204a
[0268] 205, 205a Optimal Behavior Information Output Department
[0269] 206 Control Department.
Claims
1. A machine learning device that performs machine learning on a numerical control device that causes a machine tool to move according to a machining program, characterized in that, The machine learning device has: The status information acquisition unit executes the machining program with a cutting depth and cutting speed set at least once by the numerical control device, causing the machine tool to perform cutting machining, thereby acquiring status information including the cutting depth and cutting speed of the one cutting operation. The behavior information output unit outputs behavior information, which includes the cutting depth and cutting speed adjustment information contained in the status information. The reward calculation unit acquires decision information and outputs a reward value in reinforcement learning corresponding to specified conditions based on the acquired decision information. The decision information is information related to at least the following: the pressure intensity applied to the tool during the cutting process, the waveform shape of the pressure applied to the tool, and the processing time; and The value function update unit updates the value function based on the reward value, the state information, and the behavior information. The specified condition is either a condition that prioritizes processing time or a condition that prioritizes the lifespan of the tool. The reward calculation unit calculates reward values based on the pressure intensity applied to the tool, the waveform shape of the pressure applied to the tool, and the time required for the cutting process. When weighted and summing these reward values, it outputs a first reward value prioritizing the processing time and a second reward value prioritizing the tool's lifespan. The first reward value is obtained by summing the reward values based on the processing time with a larger weighting factor compared to the condition prioritizing tool lifespan. The second reward value is obtained by summing the reward values based on the waveform shape of the pressure applied to the tool with a larger weighting factor compared to the condition prioritizing processing time. Under the condition of prioritizing the processing time, the value function updating unit updates the first value function based on the first reward value, the state information, and the behavior information. Under the condition of prioritizing the tool's lifespan, the value function updating unit updates the second value function based on the second reward value, the state information, and the behavior information.
2. The machine learning apparatus according to claim 1, characterized in that, The machine learning is performed each time the processing procedure is executed, or each time a fixed loop of a plurality of fixed loops contained in the processing procedure is executed.
3. The machine learning apparatus according to claim 1 or 2, characterized in that, The machine learning device further comprises: an optimal behavior information output unit, which outputs first behavior information whose value of the first value function updated according to the first reward value is the largest, and second behavior information whose value of the second value function updated according to the second reward value is the largest.
4. The machine learning apparatus according to claim 1, characterized in that, If the processing time required, as indicated in the determination information, is less than the processing time required previously, the reward calculation unit sets the reward value to a positive value; if the processing time required is more than the processing time required previously, the reward calculation unit sets the reward value to a negative value.
5. The machine learning apparatus according to claim 1, characterized in that, If the waveform shape of the pressure applied to the tool contained in the determination information is at least dissimilar to the waveform shape indicating a sign of tool damage or the waveform shape indicating a sharp decrease in the tool's lifespan, the reward calculation unit sets the reward value to a positive value; if the waveform shape of the pressure applied to the tool is at least similar to the waveform shape indicating a sign of tool damage or the waveform shape indicating a sharp decrease in the tool's lifespan, the reward calculation unit sets the reward value to a negative value.
6. The machine learning apparatus according to claim 1 or 2, characterized in that, The maximum number of trials for the machine learning process is set.
7. A setting device, characterized in that, Select a behavior from the behaviors obtained from the machine learning device of claim 1 or 2 according to a preset threshold, and set the selected behavior to the processing program.
8. A numerical control system, characterized in that, have: The machine learning apparatus according to claim 1 or 2; The setting device as claimed in claim 7; and A numerical control device that executes the processing program set by the setting device.
9. A numerical control device, characterized in that, The numerical control device includes: The machine learning apparatus according to claim 1 or 2; and The setting device according to claim 7, The numerical control device executes the processing program set by the setting device.
10. A machine learning method for a machine learning device, the machine learning device performing machine learning on a numerical control device that causes a machine tool to move according to a machining program, characterized in that... The machining program, with a pre-set depth of cut and cutting speed, is executed by the numerical control device at least once, causing the machine tool to perform cutting operations. This process yields status information including the depth of cut and the cutting speed for each cut. Output behavioral information, which includes the cut depth and cutting speed adjustment information for the first cut, as contained in the status information. The system obtains decision information and outputs a reward value in reinforcement learning corresponding to specified conditions based on the obtained decision information. The decision information is information related to at least the following: the pressure intensity applied to the tool during the cutting process, the waveform shape of the pressure applied to the tool, and the processing time. The value function is updated based on the reward value, the state information, and the behavior information. The specified condition is either a condition that prioritizes processing time or a condition that prioritizes the lifespan of the tool. In the process of outputting the feedback, a feedback value based on the pressure intensity applied to the tool, a feedback value based on the waveform shape of the pressure applied to the tool, and a feedback value based on the time required for the machining are calculated. When these feedback values are weighted and summed, a first feedback value is output under the condition that machining time is prioritized, and a second feedback value is output under the condition that tool life is prioritized. The first feedback value is obtained by summing the feedback values based on the machining time with a larger weighting coefficient compared to the condition that tool life is prioritized; the second feedback value is obtained by summing the feedback values based on the waveform shape of the pressure applied to the tool with a larger weighting coefficient compared to the condition that machining time is prioritized. During the process of updating the value function, the first value function is updated based on the first reward value, the state information, and the behavior information, while prioritizing the processing time. The second value function is updated based on the second reward value, the state information, and the behavior information, while prioritizing the tool's lifespan.
Citation Information
Patent Citations
Tool selection device and machine learning device
JP2019188558A
Quality controlling method of press machine and its device
JP1995164199A
Machine learning device, numerical control device, numerical control system, and machine learning method
US20190025794A1