Machine learning device, computer device, control system, and machine learning method
By introducing machine learning devices into computer devices, monitoring and optimizing data communication commands of access control devices, overload and delay problems caused by frequent or large-scale application access are solved, and more efficient command issuance and system performance improvements are achieved.
Patent Information
- Application Number
- CN202180012298.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-02-05
- Filing Date
- 2021-02-01
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2041-02-01
AI Technical Summary
When the computer device is connected to the control device, frequent or large number of applications access causes excessive issuance of data communication commands, causing overload and access delays, and affecting overall performance.
Using machine learning devices, by monitoring the commands of the access control device, status data is obtained, including release schedule, acceptance time and release time, adjust the release sequence and interval, and use reinforcement learning algorithms to optimize behavior information, and update the value function to optimize the release timeline.
Effectively prevent over-issuance of data communication commands, reduce delay time for command issuance, and improve system performance and response speed.
Smart Images

Figure CN115066659B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a machine learning device, a computer device, a control system, and a machine learning method. Background Art
[0002] For example, in a computer device (e.g., a personal computer, a tablet terminal, a smartphone, etc.) connected to a control device that controls industrial machines such as machine tools and robots, in order for an application operating on the computer device to access data in the control device, there is a communication processing unit that serves as an interface with the control device.
[0003] Among the applications that access data in the control device, there are applications that access frequently with little interval, applications that access periodically, and applications that access sporadically.
[0004] In a state where a large number of such applications are operating simultaneously, access by other applications is often delayed due to being hindered by applications that frequently access data in the control device, and the overall operation of the applications becomes slow.
[0005] Regarding this, the following technique is known: the priority of data set in an application of a personal computer as a computer device is passed to a numerical control device as a control device, and when a plurality of data are requested from the application of the personal computer, the numerical control device first sends high-priority data and stores it in a buffer, and adjusts the transmission interval according to the load of the numerical control device and the response allowable time. For example, refer to Patent Document 1.
[0006] Prior Art Documents
[0007] Patent Documents
[0008] Patent Document 1: Japanese Patent No. 6517706 Summary of the Invention
[0009] Problems to be Solved by the Invention
[0010] In a control device connected to a computer device, if applications that do not consider overall performance frequently access the control device, or if a very large number of applications simultaneously access the control device, performance degradation due to access delay and processing delay occur.
[0011] Figure 10 is an example of a timing chart showing commands output by a plurality of applications operating on a personal computer as a computer device. In addition, Figure 10 shows a case where a personal computer as a computer device executes four applications A1 - A4. In addition, in Figure 10 the order of commands of circles, quadrilaterals, diamonds, and triangles indicates the decreasing order of urgency.
[0012] As Figure 10 shown, application A1 periodically outputs commands with a relatively high urgency level to access data within the access control device. Application A2 sporadically outputs commands with the highest urgency level to access data within the access control device. Application A3 frequently outputs commands and frequently accesses data within the access control device. Application A4 periodically outputs multiple commands to access data within the access control device.
[0013] In Figure 10 such a case, for example, at times T1 and T2, in the command sets of applications A1 - A4, excessive access to the control device occurs. As problems arising in such a state, there are issues such as data access that should be processed regularly becoming irregular, event processing that should be handled emergently being delayed even when an event occurs, and the overall operation of the applications becoming slow.
[0014] Patent Document 1 is limited to the optimization of data request commands returned by a numerical control device, and cannot achieve the optimization of command issuance from a personal computer, which is a computer device, to a numerical control device, which is a control device, cannot achieve load reduction, and has no effect on the transmitted data of a write request.
[0015] In addition, in the prior art, in order to adjust command issuance to the control device, it is necessary to correct each application.
[0016] Therefore, it is desired to prevent data communication commands from being overly issued to the control device and causing an overload, and it is desired to shorten the command issuance delay time.
[0017] Means for Solving the Problem
[0018] (1) One aspect of the machine learning device of the present disclosure is a machine learning device that performs machine learning on a computer device that issues commands for accessing a control device that can be communicatively connected. The machine learning device includes: a state data acquisition unit that monitors commands for accessing data within the control device and acquires state data, where the data within the control device is data respectively instructed by one or more applications operating on the computer device, and the state data at least includes: a command issuance schedule, a reception time and an issuance time of the commands issued according to the issuance schedule; a behavior information output unit that outputs behavior information to the computer device, and the behavior information includes correction information of the issuance schedule included in the state data; a reward calculation unit that calculates a reward for the behavior information based on the delay time of each command until the command is issued to the control device and the average issuance interval of all the issued commands; and a value function update unit that updates a value function related to the state data and the behavior information based on the reward calculated by the reward calculation unit.
[0019] (2) One mode of the computer device of the present disclosure has the machine learning device of (1), and machine learning is performed on the release schedule by the machine learning device.
[0020] (3) One mode of the control system of the present disclosure has: the machine learning device of (1); and a computer device that performs machine learning on the release schedule by the machine learning device.
[0021] (4) One mode of the machine learning method of the present disclosure is a machine learning method for performing machine learning on a computer device for an issuance command, the command being used to access a controllable device communicably connected, monitoring a command for accessing data in the controllable device, and obtaining status data, wherein the data in the controllable device is data respectively instructed by one or more applications operating on the computer device, the status data at least includes: the release schedule of the command, the reception time and release time of the command released according to the release schedule, outputting behavior information to the computer device, the behavior information including correction information of the release schedule included in the status data, calculating a reward for the behavior information based on the delay time of each command until the command is released to the controllable device and the average release interval of all the released commands, and updating a value function related to the status data and the behavior information based on the calculated reward.
[0022] Advantageous Effects of the Invention
[0023] According to one mode, it is possible to prevent data communication commands from being excessively issued to the controllable device and causing an overload, and it is possible to shorten the release delay time of the commands. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 It is a functional block diagram showing a functional structural example of a control system according to an embodiment.
[0025] Figure 2 It is a diagram showing an example of a command table including a release schedule.
[0026] Figure 3 It is a functional block diagram showing a functional structural example of a machine learning device.
[0027] Figure 4 It is a diagram showing an example of the reward of each command calculated by a reward calculation unit.
[0028] Figure 5 It is a diagram showing an example of a timing chart of commands before and after update.
[0029] Figure 6It is a flowchart showing the operation of the machine learning device 40 during Q learning in one embodiment.
[0030] Figure 7 It is a flowchart showing the operation during the generation of the optimized behavior information by the optimized behavior information output unit.
[0031] Figure 8 It is a diagram showing an example of the structure of the control system.
[0032] Figure 9 It is a diagram showing an example of the structure of the control system.
[0033] Figure 10 It is a diagram showing an example of a timing chart of commands output by multiple applications operating on a personal computer. Detailed Embodiment
[0034] Hereinafter, one embodiment of the present disclosure will be described with reference to the accompanying drawings.
[0035] <One Embodiment>
[0036] Figure 1 It is a functional block diagram showing a functional structure example of the control system of one embodiment. Here, a machine tool is exemplified as an industrial machine, and a numerical control device is exemplified as a control device. The present invention is not limited to a machine tool, and can also be applied to, for example, industrial robots, service robots, etc. In addition, when the industrial machine is a robot, the control device includes a robot control device, etc. In addition, a personal computer is exemplified as a computer device, but the present invention is not limited to a personal computer, and can be applied to so-called client terminals such as tablet terminals and smart phones.
[0037] As Figure 1 shown, the control system 1 has: a machine tool 10, a numerical control device 20, a personal computer 30, and a machine learning device 40.
[0038] The machine tool 10, the numerical control device 20, the personal computer 30, and the machine learning device 40 can be directly connected to each other via a connection interface (not shown). In addition, the machine tool 10, the numerical control device 20, the personal computer 30, and the machine learning device 40 can also be connected to each other via a network (not shown) such as a LAN (Local Area Network) or an interconnect. In this case, the machine tool 10, the numerical control device 20, the personal computer 30, and the machine learning device 40 have a communication unit (not shown) for communicating with each other through such a connection. In addition, as described later, the personal computer 30 may include the machine learning device 40. In addition, the numerical control device 20 may be included in the machine tool 10.
[0039] The machine tool 10 is a machine tool well-known to those skilled in the art, and operates according to control information from the numerical control device 20 described later.
[0040] The numerical control device 20 is a numerical control device well-known to those skilled in the art. It generates an operation command based on the control information and sends the generated operation command to the machine tool 10. Thus, the numerical control device 20 controls the operation of the machine tool 10. In addition, the numerical control device 20 receives data communication commands for each of the n applications AP1 - APn operating on the personal computer 30 described later to access the data in the numerical control device 20, and sends the data to the personal computer 30 in the order of the received commands. Here, n is an integer of 2 or more.
[0041] <Personal computer 30>
[0042] The personal computer 30 includes: a central processing unit 301, a data communication interface unit 302, a command processing unit 303, a communication processing unit 304, and a storage unit 305.
[0043] The central processing unit 301 includes: a CPU (Central Processing Unit), a ROM (ReadOnly Memory), a RAM (Random Access Memory), a CMOS (Complementary Metal - Oxide - Semiconductor) memory, etc. They are configured to be able to communicate with each other via a bus and are well-known to those skilled in the art.
[0044] The CPU is a processor that controls the personal computer 30 as a whole. The CPU reads out the system program stored in the ROM and the programs of the n applications AP1 - APn via the bus, and controls the entire personal computer 30 according to the system program and the programs of the applications AP1 - APn. Various data such as temporary calculation data and display data are stored in the RAM. In addition, the CMOS memory is configured as a non-volatile memory as follows: It is backed up by a battery (not shown) and maintains the storage state even when the power of the numerical control device 20 is turned off.
[0045] The data communication interface unit 302 is a general communication interface and has, for example, a buffer (not shown). The data communication interface unit 302 receives data communication commands for accessing the data in the numerical control device 20 and temporarily stores the received command data in the buffer (not shown).
[0046] For example, the command processing unit 303 obtains the commands stored in a buffer (not shown) of the data communication interface unit 302 according to the release schedule, and releases the obtained commands via the communication processing unit 304.
[0047] Here, the release schedule will be described. In the present embodiment, as the release schedule, for example, a release schedule that determines the "release order" and "release interval" of each command for the commands stored in the buffer of the data communication interface unit 302 is exemplified.
[0048] Therefore, in the present embodiment, as a structure for determining the release schedule, a command table CT is introduced. The command table CT refers to an array table indexed by the command number, which corresponds data such as "command number", "command acceptance number", "command priority Pa", "instruction process ID", "process priority Pb", "comprehensive priority Ps", "required processing time Tc", "delay coefficient Td", and "release schedule" for data communication commands used to access data in the numerical control device 20.
[0049] The "command number" in the command table CT is an identification number for respectively identifying the commands for the application AP1-APn instructions, and is an index in the command table CT. The command number is identified by the application APi (1≤i≤n) that issues the command.
[0050] The "command acceptance number" in the command table CT indicates the acceptance number of the commands received from each of the applications AP1-APn by the data communication interface unit 302 and stored in a buffer (not shown).
[0051] The "command priority Pa" in the command table CT is a value indicating the priority of the command, and is preset such that the higher the value, the higher the priority of execution. In addition, the "command priority Pa" can also be preset such that the smaller the value, the higher the priority of execution.
[0052] The "instruction process ID" in the command table CT is a process ID, which is assigned by the OS (Operating System) of the personal computer 30 at the start of the process, and is an identifier when the OS or other processes specify and process the process.
[0053] The "process priority Pb" in the command table CT is a coefficient indicating the process priority of the command. The "process priority Pb" is, for example, set to "1" as the initial value, and is at least one coefficient parameter included in the actions selected by the machine learning device 40 described later.
[0054] The "comprehensive priority Ps" in the command table CT is the value obtained by adding the "command priority Pa" and the "process priority Pb", and commands are issued in descending order of value. In addition, for example, when the "comprehensive priority Ps" is the same for two or more different commands, the command processing unit 303 can be preset to preferentially issue the command with the higher "command priority Pa".
[0055] In addition, the command processing unit 303 can also be preset to preferentially issue the command with the higher "process priority Pb".
[0056] The "required processing time Tc" in the command table CT represents the time required for processing the command, and it is preferable to issue the next command at an interval of more than this time.
[0057] The "delay coefficient Td" in the command table CT is a time coefficient for adjusting the command issuance interval according to the "required processing time Tc". Specifically, the value obtained by adding the "delay coefficient Td" of each command to the "required processing time Tc" of each command is set as the "issuance interval Ts" of the command. By adjusting the "delay coefficient Td", the optimal issuance interval of each command can be adjusted.
[0058] As described above, the "issuance order" in the "issuance schedule" in the command table CT represents the order in which the command processing unit 303 issues the commands stored in the buffer (not shown) of the data communication interface unit 302 according to the "comprehensive priority Ps".
[0059] The "issuance interval Ts" in the "issuance schedule" in the command table CT is the value obtained by adding the "required processing time Tc" and the "delay coefficient Td", and the command processing unit 303 issues commands at intervals of the "issuance interval Ts".
[0060] In addition, the command table CT at the start of learning by the machine learning device 40 can be arbitrarily set by the user.
[0061] Figure 2 is a diagram showing an example of the command table CT. Refer to Figure 2 For the sake of simplicity in explanation, the command table CT stores the data arrangements of 5 commands.
[0062] As Figure 2 shown, for these 5 commands, the issuance order is set in descending order of the comprehensive priority value. In addition, for command number 18 and command number 8 with the same comprehensive priority value, as described above, it can be seen that command number 8 with the higher "command priority Pa" value is preferred.
[0063] In addition, as described above, the issuance interval of each command is set to the value obtained by adding the required processing time of each command to the delay coefficient of each command.
[0064] Further, as described later, the machine learning device 40 uses the "process priority Pb" and the "delay coefficient Td" as actions, and for example, selects various actions according to a certain policy, and thus performs reinforcement learning while exploring, so as to be able to select the best release schedule.
[0065] The communication processing unit 304 is a well-known communication unit to those skilled in the art, and transmits and receives data, machining programs, etc. to and from the numerical control device 20.
[0066] Specifically, the communication processing unit 304 sequentially sends the commands received from the command processing unit 303 to the numerical control device 20 and receives the data for the sent commands.
[0067] The storage unit 305 is a RAM, HDD (Hard Disk Drive), etc. The storage unit 305 stores system programs, programs of n applications AP1 - APn, and a command table CT, etc.
[0068] <Machine learning device 40>
[0069] The machine learning device 40 is a device that performs reinforcement learning on the release schedule by executing the programs of the applications AP1 - APn on the personal computer 30, and the release schedule is for the unpublished commands of the applications AP1 - APn stored in the buffer (not shown) of the data communication interface unit 302.
[0070] Before explaining each functional block included in the machine learning device 40, first, the basic structure of reinforcement learning will be explained. The agent (corresponding to the machine learning device 40 in the present embodiment) observes the state of the environment (corresponding to the numerical control device 20 and the personal computer 30 in the present embodiment), selects a certain action, and the environment changes according to the selected action. As the environment changes, a certain reward is provided, and based on the provided reward, the agent learns to be able to select better actions.
[0071] Supervised learning represents a complete correct answer, while the rewards in reinforcement learning are mostly fragmentary values based on partial changes in the environment. Therefore, the agent learns to maximize the sum of future rewards.
[0072] In this way, in reinforcement learning, by learning actions, appropriate actions are learned based on the interaction between the actions and the environment, that is, a method for learning to maximize the future rewards obtained is learned. This means that in the present embodiment, it is possible to obtain behavior information that, for example, selects actions to prevent data communication commands from being overly released to the numerical control device 20 and causing overload, and to shorten the release delay time of the commands, which has an impact on the future.
[0073] Here, any learning method can be used for reinforcement learning. In the following description, the case of using Q-learning in a certain environmental state s is taken as an example for explanation. The Q-learning is a method for learning the value function Q(s, a) for selecting an action a.
[0074] The purpose of Q-learning is to select, as the optimal action, the action a with the highest value function Q(s, a) from the actions a that can be taken in a certain state s.
[0075] However, at the time of initially starting Q-learning, the correct value of the value function Q(s, a) for the combination of the state s and the action a is completely unknown. Therefore, the agent selects various actions a in a certain state S, and for the action a at that time, based on the given reward, selects a better action, and thus continues to learn the correct value function Q(s, a).
[0076] In addition, in order to maximize the sum of the rewards obtained in the future, the goal is ultimately to make Q(s, a) = E[Σ(γ t )r t . Here, E[] represents the expected value, t represents the time, γ represents a parameter called the discount rate described later, and r t represents the reward at time t, and Σ is the sum at time t. The expected value in this formula is the expected value when the optimal action state changes. However, during the process of Q-learning, since the optimal action is unknown, various actions are performed, and exploration and reinforcement learning are carried out simultaneously. The update formula for such a value function Q(s, a) can be expressed, for example, by the following mathematical formula 1.
[0077] [Mathematical formula 1]
[0078]
[0079] In the above mathematical formula 1, s t represents the environmental state at time t, and a t represents the action at time t. By the action a t , the state changes to s t+1 . r t+1 represents the reward obtained through this state change. In addition, the term with max is: in the state s t+1 , the value obtained by multiplying γ by the Q value when selecting the action a with the highest known Q value at that time. Here, γ is a parameter where 0 < γ ≤ 1, called the discount rate. In addition, α is the learning coefficient, and the range of α is set to 0 < α ≤ 1.
[0080] The above mathematical formula 1 represents the following method: Try a t , and as a result, according to the feedback reward r t+1 , update the state st the behavior a below t of the value function Q(s t , a t ).
[0081] This update formula indicates that: if the next state s t resulting from the behavior a t+1 the value of the best behavior max a Q(s t+1 , a) is greater than the value function Q(s t under the behavior a t , a t , a t ), then increase Q(s t , a t ), and conversely if it is small, then decrease Q(s t , a t ). That is to say, make the value of a certain behavior in a certain state approach the value of the best behavior in the next state resulting from that behavior. Among them, although this difference varies due to the existence form of the discount rate γ and the reward r t+1 , basically the value of the best behavior in a certain state is a structure that spreads to the value of the behavior in its previous state.
[0082] Here, there is the following method for Q-learning: Make a table of Q(s, a) for all state-action pairs (s, a) to perform learning. However, sometimes in order to obtain the values of Q(s, a) for all state-action pairs, the number of states is too large, making it take a long time for Q-learning to converge.
[0083] Therefore, a well-known technique called DQN (Deep Q-Network) can be utilized. Specifically, an appropriate neural network can be used to construct the value function Q, and the parameters of the neural network can be adjusted. Thus, the value function Q(s, a) is approximated by an appropriate neural network to calculate its value. By using DQN, the time required for Q-learning to converge can be shortened. In addition, regarding DQN, for example, it is described in detail in the following non-patent literature.
[0084] <Non-patent literature>
[0085] “Human-level control through deep reinforcement learning”, by Volodymyr Mnih1 [online], [searched on January 17, 2017], Internet <URL: http: / / files.davidqiu.com / research / nature14236.pdf>
[0086] The machine learning device 40 performs the Q - learning described above. Specifically, the machine learning device 40 sets the command table CT, the reception time of each command received by the data communication interface unit 302, and the issuance time of each command issued by the command processing unit 303 via the communication processing unit 304 as the state s, and sets the setting and change of the parameters for adjusting the issuance schedule included in the command table CT in the state s as the action a, to learn the value function Q to be selected. Here, the command table CT is a table of unissued commands stored in a buffer (not shown) of the data communication interface unit 302. Here, as parameters, “process priority Pb” and “delay coefficient Td” are exemplified.
[0087] The machine learning device 40 monitors the commands instructed by each of the applications AP1 - APn, observes the state information (state data) s including the command table CT and the reception time and issuance time of each command issued according to the “issuance schedule” of the command table CT, and determines the action a. The machine learning device 40 returns a reward every time it determines the action a. The machine learning device 40, for example, explores the optimal action a by trial and error so that the total future reward is maximized. Thus, the machine learning device 40 can select the optimal action a (i.e., “process priority Pb” and “delay coefficient Td”) for the state s, where the state s includes: the command table CT obtained by executing the applications AP1 - APn by the personal computer 30, and the reception time and issuance time of each command issued according to the “issuance schedule” of the command table CT.
[0088] Figure 3 It is a functional block diagram showing a functional structural example of the machine learning device 40.
[0089] To perform the above - mentioned reinforcement learning, as Figure 3 shown, the machine learning device 40 has: a state data acquisition unit 401, a determination data acquisition unit 402, a learning unit 403, a behavior information output unit 404, a value function storage unit 405, an optimal behavior information output unit 406, and a control unit 407. The learning unit 403 has: a reward calculation unit 431, a value function update unit 432, and a behavior information generation unit 433. The control unit 407 controls the operations of the state data acquisition unit 401, the determination data acquisition unit 402, the learning unit 403, the behavior information output unit 404, and the optimal behavior information output unit 406.
[0090] The state data acquisition unit 401 acquires the state data s from the personal computer 30 as the state of data communication from the personal computer 30 to the numerical control device 20. The state data s includes: the command table CT, and the reception time and issuance time of all commands received within a preset specific time as described later according to the “issuance schedule” of the command table CT. This state data s corresponds to the environmental state s in Q - learning.
[0091] The status data acquisition unit 401 outputs the acquired status data s to the determination data acquisition unit 402 and the learning unit 403.
[0092] In addition, the command table CT at the time of initially starting Q learning can be set by the user as described above.
[0093] The status data acquisition unit 401 can store the acquired status data s in a storage unit (not shown) included in the machine learning device 40. In this case, the determination data acquisition unit 402 and the learning unit 403 described later can read the status data s from the storage unit (not shown) of the machine learning device 40.
[0094] The determination data acquisition unit 402 periodically analyzes the command table CT received from the status data acquisition unit 401, the reception times and issuance times of all commands accepted within a preset specific time, to acquire determination data.
[0095] Specifically, the determination data acquisition unit 402 acquires, as determination data, the average issuance interval of all commands accepted by the data communication interface unit 302 at a preset specified time (e.g., 1 minute), the issuance delay time of each command, the command priority, etc. for all commands accepted within a specific time. The determination data acquisition unit 402 outputs the acquired determination data to the learning unit 403.
[0096] In addition, the average issuance interval of commands is the average of the issuance intervals of commands accepted at a preset specified time (e.g., 1 minute). Also, the issuance delay time of each command is the difference between the reception time and the issuance time of each command accepted at a preset specified time (e.g., 1 minute).
[0097] The learning unit 403 is a part that learns the value function Q(s, a) when selecting a certain action a in a certain status data (environmental status) s. Specifically, the learning unit 403 includes: a reward calculation unit 431, a value function update unit 432, and an action information generation unit 433.
[0098] In addition, the learning unit 403 determines whether to continue learning. Whether to continue learning can be determined, for example, based on whether the number of trials since the start of machine learning has reached the maximum number of trials, or whether the elapsed time since the start of machine learning has exceeded a specified time (or more).
[0099] The reward calculation unit 431 is a part that calculates the reward when selecting the adjustment of the "process priority Pb" and the "delay coefficient Td" of the command table CT, that is, the action a, in a certain status s.
[0100] Here, an example of calculating the reward for the action a is described.
[0101] Specifically, first, the reward calculation unit 431 calculates the evaluation value V for each command according to the average release interval Ta, release delay time Tb, and command priority Pa obtained by the determination data acquisition unit 402 for all commands accepted within a preset specific time, as described above. In addition, as the preset specific time, it is preferable to set the time when the applications AP1 - APn executed on the personal computer 30 are executed in parallel. Further, the specific time may be the same as the aforementioned specified time (e.g., 1 minute) or may include the specified time (e.g., 1 minute).
[0102] As an example of the calculation of the evaluation value, the following mathematical formula (Mathematical Formula 2) is exemplified.
[0103] [Equation 2]
[0104] V = average release interval Ta × a 1 - release delay time Tb × command priority Pa × a 2
[0105] Here, a 1 and a 2 are coefficients, which are set to "20" and "1" respectively, for example. In addition, the values of a 1 and a 2 are not limited to this and can be determined according to the required accuracy of machine learning, etc.
[0106] Moreover, the reward calculation unit 431 calculates the evaluation value V for all commands accepted within the specific time, and takes the average value of all the calculated evaluation values as the reward r for the action a. Thus, regarding the action a, the smaller the release delay time of the determination target command, the greater the reward that can be obtained. In addition, the larger the average release interval of the determination target command, the greater the reward that can be obtained.
[0107] Figure 4 is a diagram showing an example of the evaluation value V for each command (command number) calculated by the reward calculation unit 431. In addition, the average release interval Ta in Mathematical Formula 2 is the average value of the release intervals of each command (average release interval), which is "21" in the case of Figure 4 . And, as shown in Figure 4 , the evaluation values of each command are calculated, and the average value of all the calculated evaluation values (= 176) is set as the reward r.
[0108] The value function update unit 432 performs Q - learning based on the state s, action a, the state s' when the action a is applied to the state s, and the value r of the reward calculated as described above, thereby updating the value function Q stored in the value function storage unit 405.
[0109] The update of the value function Q can be performed through online learning, batch learning, or mini-batch learning.
[0110] Online learning is a learning method as follows: by applying a certain action a to the current state s, whenever the state s transfers to a new state s’, the update of the value function Q is immediately performed. Additionally, batch learning is a learning method as follows: by applying a certain action a to the current state s, the state s transfers to a new state s’, learning data is collected by repeating the above actions, and all the collected learning data is used to perform the update of the value function Q. Furthermore, mini-batch learning is a learning method between online learning and batch learning, which is a learning method that performs the update of the value function Q whenever a certain amount of learning data is accumulated.
[0111] The action information generation unit 433 selects an action a during the Q-learning process for the current state s. During the Q-learning process, the action information generation unit 433 generates action information a for the purpose of performing actions equivalent to the action a in Q-learning, i.e., adjusting the “process priority Pb” and “delay coefficient Td” of the correction command table CT, and outputs the generated action information a to the action information output unit 404.
[0112] More specifically, for example, the action information generation unit 433 can increase or decrease the “process priority Pb” and “delay coefficient Td” included in the action a with respect to the “process priority Pb” and “delay coefficient Td” of the command table CT included in the state s.
[0113] The action information generation unit 433 can adjust the “process priority Pb” and “delay coefficient Td” of the command table CT through the action a. When transferring to the state s’, the “process priority Pb” and “delay coefficient Td” of the command table CT for selecting the next action a’ are selected according to the state of the “release schedule” of the command table CT (whether the “release order” and “release interval Ts” are appropriate).
[0114] For example, when the reward r increases due to an increase in the “process priority Pb” and / or “delay coefficient Td” and the “release order” and “release interval Ts” of the “release schedule” are appropriate, as the next action a’, for example, the following strategy can be adopted: select an action a’ such as increasing the “process priority Pb” and / or “delay coefficient Td” to shorten the release delay time of the priority command and optimize the release interval.
[0115] Alternatively, when the return r decreases due to an increase in "process priority Pb" and / or "delay coefficient Td", as the next action a', for example, the following strategy can be adopted: Select an action a' such as returning "process priority Pb" and / or "delay coefficient Td" to the previous level, shortening the release delay time of the priority command, and optimizing the release interval.
[0116] In addition, for each of "process priority Pb" and "delay coefficient Td", for example, it can be incremented by +1 when the return r increases as they increase, and returned to the previous level when the return r decreases.
[0117] In addition, the action information generation unit 433 can adopt the following strategy: By using a well-known method such as a greedy algorithm that selects the action a' with the highest value Q(s, a) among the currently estimated actions a, or randomly selecting the action a' with a certain small probability ε, and otherwise selecting the action a' with the highest value function Q(s, a), to select the action a'.
[0118] The action information output unit 404 is the part that outputs the action information a output from the learning unit 403 to the personal computer 30. The action information output unit 404 can, for example, output the updated values of "process priority Pb" and "delay coefficient Td" as the action information to the personal computer 30. Thereby, the personal computer 30 updates the command table CT according to the received updated values of "process priority Pb" and "delay coefficient Td". And the command processing unit 303 issues a data communication command to the communication processing unit 304 according to the "release schedule" of the updated command table CT.
[0119] In addition, the action information output unit 404 can output the command table CT updated according to the updated values of "process priority Pb" and "delay coefficient Td" as the action information to the personal computer 30.
[0120] The value function storage unit 405 is a storage device that stores the value function Q. The value function Q can be stored, for example, as a table by state s and action a (hereinafter, also referred to as "action value table"). The value function Q stored in the value function storage unit 405 is updated by the value function update unit 432.
[0121] The optimized action information output unit 406 generates action information a (hereinafter, referred to as "optimized action information") for causing the personal computer 30 to perform an action that maximizes the value function Q(s, a) according to the value function Q updated by Q-learning by the value function update unit 432.
[0122] More specifically, the optimization behavior information output unit 406 acquires the value function Q stored in the value function storage unit 405. As described above, this value function Q is a function updated through Q-learning by the value function update unit 432. Further, the optimization behavior information output unit 406 generates behavior information based on the value function Q and outputs the generated behavior information to the personal computer 30. In this optimization behavior information, similar to the behavior information output by the behavior information output unit 404 during the process of Q-learning, it includes information indicating the values of the updated "process priority Pb" and "delay coefficient Td".
[0123] Figure 5 FIG. is an example of a timing chart showing commands before and after update. Figure 5 The upper part of Figure 10 is, similarly to the case of Figure 5 an example of a timing chart showing the commands before update output by the four applications AP1 - AP4 operating on the personal computer 30. Figure 10 The lower part of Figure 5 is an example of a timing chart showing the commands after update output by the four applications AP1 - AP4. Further, similarly to the case of
[0124] As Figure 5 shown in the lower part of Figure 5 the command processing unit 303 issues the unsent commands according to the "issue schedule" of the updated command table CT in which the command issue order is adjusted according to the comprehensive priority Ps. Thus, the command processing unit 303 can average out the command issue intervals at times T1', T2' and time T3' corresponding to times T1 and T2 in the upper part of
[0125] so as not to cause excessive access.
[0126] As described above, the functional blocks included in the machine learning device 40 have been described.
[0127] To implement these functional blocks, the machine learning device 40 has an arithmetic processing device such as a CPU. In addition, the machine learning device 40 also has an auxiliary storage device such as an HDD that stores various control programs such as application software and an OS (Operating System), and a main storage device such as a RAM that stores data temporarily required when the arithmetic processing device executes programs.
[0128] Further, in the machine learning device 40, the arithmetic processing unit reads the application software and the OS from the auxiliary storage device, expands the read application software and OS in the main storage device, and performs arithmetic processing based on these application software and OS. Further, based on the arithmetic result, various hardware included in the machine learning device 40 is controlled. Thereby, the functional blocks of the present embodiment are realized. That is, the present embodiment can be realized by the cooperation of hardware and software.
[0129] Regarding the machine learning device 40, since the amount of arithmetic operations accompanying machine learning increases, for example, a technique called GPGPU (General-Purpose computing on Graphics Processing Units), which uses a GPU (Graphics Processing Units) mounted on a personal computer, can achieve high-speed processing when the GPU is used for arithmetic processing accompanying machine learning. Further, for even faster processing, a computer cluster can be constructed using multiple computers each mounted with such a GPU, and parallel processing can be performed by the multiple computers included in the computer cluster.
[0130] Next, with reference to Figure 6 the flowchart of, the operation of the machine learning device 40 during Q-learning in the present embodiment will be described.
[0131] Figure 6 FIG. is a flowchart showing the operation of the machine learning device 40 during Q-learning in one embodiment.
[0132] In step S11, the control unit 407 sets the trial number to "1" and instructs the state data acquisition unit 401 to acquire state data.
[0133] In step S12, the state data acquisition unit 401 acquires the initial state data from the personal computer 30. The acquired state data is output to the behavior information generation unit 433. As described above, this state data (state information) is information corresponding to the state s in Q-learning, and includes the command table CT at the time point of step S12, the reception time and the transmission time of each command transmitted according to the "transmission schedule" of the command table CT. In addition, the command table CT at the time point when Q-learning starts for the first time is generated in advance by the user.
[0134] In step S13, the behavior information generation unit 433 generates new behavior information a, and outputs the generated new behavior information a to the personal computer 30 via the behavior information output unit 404. The personal computer 30 that receives the behavior information updates the "process priority Pb" and "delay coefficient Td" of the current state s to set them as state s'. The personal computer 30 updates state s to state s' according to the updated behavior a. Specifically, the personal computer 30 updates the command table CT. The command processing unit 303 issues the unsent commands stored in the buffer (not shown) of the data communication interface unit 302 according to the "issue schedule" of the updated command table CT.
[0135] In step S14, the state data acquisition unit 401 acquires state data corresponding to the new state s' obtained from the personal computer 30. Here, the new state data includes the command table CT of state s', the reception time and issue time of each command issued according to the "issue schedule" of the command table CT. The state data acquisition unit 401 outputs the acquired state data to the determination data acquisition unit 402 and the learning unit 403.
[0136] In step S15, the determination data acquisition unit 402 acquires determination data at a specified time (for example, 1 minute) according to the command table CT included in the new state data received by the state data acquisition unit 401, the reception time and issue time of each command for all commands received within a preset specific time. The determination data acquisition unit 402 outputs the acquired determination data to the learning unit 403. This determination data includes, for example, the average issue interval Ta of the commands received by the data communication interface unit 302 at a specified time such as 1 minute, the issue delay time Tb of each command, the command priority Pa, etc.
[0137] In step S16, the reward calculation unit 431 calculates the evaluation value V of each command for all commands received within a preset specific time according to the acquired determination data, that is, the average issue interval Ta of the commands, the issue delay time Tb of each command, and the command priority Pa and mathematical formula 2. The reward calculation unit 431 takes the average value of the evaluation values V of each command as the reward r.
[0138] In step S17, the value function update unit 432 updates the value function Q stored in the value function storage unit 405 according to the calculated reward r.
[0139] In step S18, the control unit 306 determines whether the number of trials since the start of machine learning has reached the maximum number of trials. The maximum number of trials is preset. If the maximum number of trials has not been reached, the number of trials is counted in step S19, and the process returns to step S13. The processing from step S13 to step S19 is repeated until the maximum number of trials is reached.
[0140] In addition, Figure 6 the process of Figure 6 ends the processing when the number of trial executions reaches the maximum number of trial executions, but it is also possible to end the processing on the condition that the time obtained by accumulating the time of the processing from step S13 to step S19 since the start of machine learning exceeds a preset maximum elapsed time (or more).
[0141] In addition, step S17 exemplifies online update, but it is also possible to replace online update with batch update or mini-batch update.
[0142] As described above, by referring to Figure 6 the operations described above, in the present embodiment, it is possible to generate a value function Q for generating behavior information that prevents data communication commands from being excessively issued to the numerical control device 20 and causing an overload, and shortens the command issuance delay time.
[0143] Next, with reference to Figure 7 the flowchart of Figure 7 , the operations when generating the optimized behavior information based on the optimized behavior information output unit 406 will be described.
[0144] In step S21, the optimized behavior information output unit 406 acquires the value function Q stored in the value function storage unit 405. The value function Q is a function updated by Q-learning through the value function update unit 432 as described above.
[0145] In step S22, the optimized behavior information output unit 406 generates optimized behavior information based on the value function Q and outputs the generated optimized behavior information to the personal computer 30.
[0146] As described above, the personal computer 30 can prevent data communication commands from being excessively issued to the control device and causing an overload, and can shorten the command issuance delay time by updating the command table CT.
[0147] As described above, one embodiment has been described, but the personal computer 30 and the machine learning device 40 are not limited to the above-described embodiment, and include variations, improvements, etc. within the scope that can achieve the purpose.
[0148] <Modification Example 1>
[0149] In the above-described embodiment, it is exemplified that the machine learning device 40 is a device different from the personal computer 30, but the personal computer 30 may also have a part or all of the functions of the machine learning device 40.
[0150] Alternatively, for example, the server may include some or all of the state data acquisition unit 401, determination data acquisition unit 402, learning unit 403, behavior information output unit 404, value function storage unit 405, optimized behavior information output unit 406, and control unit 407 of the machine learning device 40. Additionally, each function of the machine learning device 40 may be implemented using virtual server functions or the like on the cloud.
[0151] Moreover, the machine learning device 40 may be a distributed processing system that appropriately distributes each function of the machine learning device 40 across multiple servers.
[0152] <Variant Example 2>
[0153] Furthermore, for example, in the above-described embodiment, in the control system 1, one personal computer 30 is communicably connected to one machine learning device 40, but this is not limiting. For example, as Figure 8 shown, the control system 1 may include m personal computers 30A(1)-30A(m) and m machine learning devices 40A(1)-40A(m) (m is an integer greater than or equal to 2). In this case, the machine learning device 40A(j) may be communicably connected to the personal computer 30A(j) in a one-to-one manner via the network 50, and perform machine learning on the personal computer 30A(j) (j is an integer from 1 to m).
[0154] In addition, the value function Q stored in the value function storage unit 405 of the machine learning device 40A(j) may be shared with other machine learning devices 40A(k) (k is an integer from 1 to m, k≠j). If the value function Q is shared among the machine learning devices 40A(1)-40A(m), reinforcement learning can be performed dispersedly in each machine learning device 40A, thereby improving the efficiency of reinforcement learning.
[0155] In addition, each of the personal computers 30A(1)-30A(m) is connected to each of the numerical control devices 20A(1)-20A(m), and each of the numerical control devices 20A(1)-20A(m) is connected to each of the machine tools 10A(1)-10A(m).
[0156] Furthermore, each of the machine tools 10A(1)-10A(m) corresponds to Figure 1 the machine tool 10. Each of the numerical control devices 20A(1)-20A(m) corresponds to Figure 1 the numerical control device 20. Each of the personal computers 30A(1)-30A(m) corresponds to Figure 1 the personal computer 30. Each of the machine learning devices 40A(1)-40A(m) corresponds to Figure 1 the machine learning device 40.
[0157] In addition, as Figure 9 shown, the server 60 can operate as the machine learning device 40, is communicably connected to m personal computers 30A(1)-30A(m) via the network 50, and performs machine learning on each of the personal computers 30A(1)-30A(m).
[0158] <Modification Example 3>
[0159] In addition, for example, in the above-described embodiment, as parameters for adjusting the release schedule, the process priority Pb and the delay coefficient Td are applied, but parameters other than the process priority Pb and the delay coefficient Td may also be used.
[0160] Furthermore, each function included in the personal computer 30 and the machine learning device 40 in one embodiment can be implemented separately by hardware, software, or a combination thereof. Here, implementing by software means implementing by a computer reading and executing a program.
[0161] Each structural part included in the personal computer 30 and the machine learning device 40 can be implemented by hardware including an electronic circuit, etc., software, or a combination thereof. When implemented by software, the program constituting the software is installed in the computer. In addition, these programs can be distributed by recording them on a removable medium and distributing them to users, or by downloading them to the users' computers via the network. In addition, when constituted by hardware, for example, each structural part included in the above-described device can be constituted by an integrated circuit (IC) such as an ASIC (Application Specific Integrated Circuit), a gate array, an FPGA (Field Programmable Gate Array), or a CPLD (Complex Programmable Logic Device) to implement part or all of the functions.
[0162] A variety of non-transitory computer-readable media can be used to store programs and provide them to a computer. Non-transitory computer-readable media include various types of tangible storage media. Examples of non-transitory computer-readable media include: magnetic storage media (e.g., floppy disks, magnetic tapes, hard disk drives), magneto-optical storage media (e.g., optical disks), CD-ROM (Read Only Memory), CD-R, CD-R / W, semiconductor memories (e.g., mask ROM, PROM (Programmable ROM), EPROM (Erasable PROM), flash ROM, RAM). Additionally, programs can be supplied to a computer via various types of transitory computer-readable media. Examples of transitory computer-readable media include electrical signals, optical signals, and electromagnetic waves. Transitory computer-readable media can supply programs to a computer via wired communication paths such as wires and optical fibers or wireless communication paths.
[0163] In addition, the steps of describing a program recorded in a recording medium, of course, include processing that proceeds chronologically in that order, and also include processing that does not necessarily proceed chronologically, as well as processing that is executed in parallel or individually.
[0164] In other words, the machine learning device, computer device, control system, and machine learning method of the present disclosure can adopt various embodiments having the following structures.
[0165] (1) The machine learning device 40 of the present disclosure performs machine learning on a personal computer 30 that issues commands for accessing a numerically controlled device 20 that can be communicatively connected. The machine learning device 40 includes: a status data acquisition unit 401 that monitors commands for accessing data in the numerically controlled device 20 and acquires status data, where the data in the numerically controlled device 20 is data respectively instructed by one or more applications AP1-APn operating on the personal computer 30, and the status data at least includes: a command release schedule, the reception time and release time of the commands released according to the release schedule; a behavior information output unit 404 that outputs behavior information a to the personal computer 30, and the behavior information a includes correction information of the release schedule included in the status data; a reward calculation unit 431 that calculates a reward r for the behavior information a based on the release delay time Tb of each command until the command is issued to the numerically controlled device 20 and the average release interval Ta of all the issued commands; and a value function update unit 432 that updates the value function Q related to the status data and the behavior information a based on the reward r calculated by the reward calculation unit 431.
[0166] According to the machine learning device 40, it is possible to prevent data communication commands from being excessively issued to the control device and causing an overload, and it is possible to shorten the command release delay time.
[0167] (2) In the machine learning device 40 described in (1), it may also be that the correction information a of the release schedule includes: a process priority Pb indicating the priority of the process that has instructed the command, and a delay coefficient Td that delays the release of the command.
[0168] Thereby, the machine learning device 40 can adjust the release schedule to be optimal.
[0169] (3) In the machine learning device 40 described in (1) or (2), it may also be that the reward calculation unit 431 calculates an evaluation value V for each command based on the release delay time Tb and the average release interval Ta of each command, and uses the average value of the calculated evaluation values of each command as the reward r.
[0170] Thereby, the machine learning device 40 can accurately calculate the reward.
[0171] (4) In the machine learning device 40 described in any one of (1) to (3), it may also be that the machine learning device 40 further includes: an optimized behavior information output unit 406 that outputs the behavior information a with the maximum value of the value function Q based on the value function Q updated by the value function update unit 432.
[0172] Thereby, the machine learning device 40 can obtain a more appropriate release schedule.
[0173] In the machine learning device 40 described in any one of (1) to (4), the numerical control device 20 may be a control device for an industrial machine.
[0174] Accordingly, the machine learning device 40 can be applied to control devices such as machine tools and robots.
[0175] In the machine learning device 40 described in any one of (1) to (5), machine learning may be performed by setting a maximum number of trial runs of machine learning.
[0176] Accordingly, the machine learning device 40 can avoid performing machine learning for a long time.
[0177] The personal computer 30 of the present disclosure has the machine learning device 40 described in any one of (1) to (6), and the release schedule is machine-learned by the machine learning device 40.
[0178] According to this personal computer 30, the same effects as those in (1) to (6) can be obtained.
[0179] The control system 1 of the present disclosure includes: the machine learning device 40 described in any one of (1) to (6); and a computer device that machine-learns the release schedule by the machine learning device 40.
[0180] According to this control system 1, the same effects as those in (1) to (6) can be obtained.
[0181] The machine learning method of the present disclosure is used to perform machine learning on a personal computer 30 that issues commands for accessing a numerically controlled device 20 that can be communicatively connected, monitors commands for accessing data in the numerically controlled device 20, and obtains status data. The data in the numerically controlled device 20 is data respectively instructed by one or more applications AP1 - APn operating on the personal computer 30. The status data at least includes: the release schedule of the commands, the reception time and release time of the commands released according to the release schedule. Behavior information is output to the personal computer 30, and the behavior information includes correction information of the release schedule included in the status data. According to the release delay time Tb of each command until the command is released to the numerically controlled device 20 and the average release interval Ta of all the released commands, the reward r for the behavior information is calculated, and according to the calculated reward r, the value function Q related to the status data and the behavior information is updated.
[0182] According to this machine learning method, the same effect as that in (1) can be obtained.
[0183] Symbol Explanation
[0184] 1 Control System
[0185] 10 Machine Tool
[0186] 20 Numerical Control Device
[0187] 30 Personal Computer
[0188] 301 Central Processing Unit
[0189] 302 Data Communication Interface Unit
[0190] 303 Command Processing Unit
[0191] 304 Communication Processing Unit
[0192] 305 Storage Unit
[0193] 40 Machine Learning Device
[0194] 401 Status Data Acquisition Unit
[0195] 402 Judgment Data Acquisition Unit
[0196] 403 Learning Unit
[0197] 404 Behavior Information Output Unit
[0198] 405 Value Function Storage Unit
[0199] 406 Optimized Behavior Information Output Unit.
Claims
1. A machine learning device that performs machine learning on a computer device that issues commands for accessing a controllable device communicably connected thereto, characterized in that, the machine learning device has: a status data acquisition unit that monitors commands for accessing data in the control device and acquires status data, wherein the data in the control device is data respectively instructed by one or more applications operating on the computer device, and the status data at least includes: a release schedule of the commands, a reception time and a release time of the commands released according to the release schedule; a behavior information output unit that outputs behavior information to the computer device, the behavior information including correction information of the release schedule included in the status data; a reward calculation unit that calculates a reward for the behavior information based on a delay time of each command until the command is released to the control device and an average release interval of all the released commands; and a value function update unit that updates a value function related to the status data and the behavior information based on the reward calculated by the reward calculation unit.
2. The machine learning device according to claim 1, characterized in that, the correction information of the release schedule includes: a process priority indicating a priority of a process that has instructed the command, and a delay coefficient that delays the release of the command.
3. The machine learning device according to claim 1 or 2, characterized in that, the reward calculation unit calculates an evaluation value of each command based on the delay time and the average release interval of each command, and takes an average value of the calculated evaluation values of each command as the reward.
4. The machine learning device according to claim 1, characterized in that, the machine learning device further has: an optimized behavior information output unit that outputs behavior information with the maximum value of the value function based on the value function updated by the value function update unit.
5. The machine learning device according to claim 1, characterized in that, the control device is a control device of an industrial machine.
6. The machine learning device according to claim 1, characterized in that, the maximum number of trials for setting the machine learning is performed for the machine learning.
7. A computer device, characterized in that, the computer device has the machine learning device according to claim 1, and performs machine learning on the release schedule through the machine learning device.
8. A control system, characterized in that, it has: the machine learning device according to claim 1; and a computer device that performs machine learning on the release schedule through the machine learning device.
9. A machine learning method for performing machine learning on a computer device that issues commands for accessing a controllable device communicably connected thereto, characterized in that, Monitor commands for accessing data within the control device and obtain status data, where the data within the control device is data respectively instructed by one or more applications operating on the computer device, and the status data at least includes: the release schedule of the commands, the reception time and release time of the commands released according to the release schedule, Output behavior information to the computer device, where the behavior information includes correction information of the release schedule included in the status data, Calculate a reward for the behavior information based on the delay time of each command until the command is released to the control device and the average release interval of all the released commands, Update the value function related to the status data and the behavior information according to the calculated reward.
Citation Information
Patent Citations
Servo control device, servo control system, machine learning device, and machine learning method
CN108628355A
Whole equipment control device, rolling mill control device, control method, and storage medium
CN108687137A