Reinforcement learning method, program product, reinforcement learning device, molding machine, and storage medium

By employing a dual-agent model reinforcement learning approach, which combines deep neural networks and function models, the problem of inappropriate molding condition settings in reinforcement learning is solved. This enables safe and efficient search for optimal molding conditions, improving the adjustment efficiency and safety of injection molding machines.

CN116997913BActive Publication Date: 2026-04-24THE JAPAN STEEL WORKS LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
THE JAPAN STEEL WORKS LTD
Filing Date
2022-03-17
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

In the prior art, reinforcement learning is prone to setting inappropriate molding conditions when adjusting the molding conditions of injection molding machines, which can lead to unpredictable adverse situations for the equipment and operators.

Method used

A reinforcement learning method employing a dual-agent model is proposed. In this model, the first agent searches for a wider range of shaping conditions through a deep neural network, while the second agent searches within a safe range using a function model or function approximator. The shaping conditions are adjusted using reward data to ensure the safety and effectiveness of the learning process.

Benefits of technology

It enables the safe search for and finding of optimal molding conditions without limiting the search range, avoiding adverse situations for equipment and operators, and improving the adjustment efficiency and safety of the learning machine.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116997913B_ABST
    Figure CN116997913B_ABST
Patent Text Reader

Abstract

The present application provides a reinforcement learning method of learning machine, the learning machine is equipped with: the first agent, the manufacturing condition of manufacturing device is adjusted based on the observation data obtained by observing the state of manufacturing device; and the second agent, the function model or function approximator that has the relationship of observation data and manufacturing condition expressed in a different way from the first agent, the manufacturing condition searched by the first agent in reinforcement learning is adjusted using observation data and the function model or function approximator of the second agent, the reward data corresponding to the state of the product manufactured by the manufacturing device under the adjusted manufacturing condition is calculated, and the first agent and the second agent are reinforced learning based on observation data and calculated reward data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to reinforcement learning methods, program products, reinforcement learning devices, molding machines, and storage media. Background Technology

[0002] There are injection molding machine systems that can appropriately adjust the molding conditions of the injection molding machine using reinforcement learning (e.g., Patent Document 1).

[0003] Existing technical documents

[0004] Patent documents

[0005] Patent Document 1: JP 2019-166702 Summary of the Invention

[0006] However, using reinforcement learning to search for molding conditions raises concerns that setting inappropriate molding conditions as behaviors could lead to abnormal actions of the injection molding machine, causing unpredictable adverse effects on the equipment and operators. This is a common problem with manufacturing equipment.

[0007] The purpose of this disclosure is to provide a reinforcement learning method, computer program, reinforcement learning device, and molding machine that, in reinforcement learning of a learning machine that adjusts the manufacturing conditions of a manufacturing apparatus, can safely search for the optimal manufacturing conditions without limiting the search range to a certain range, thereby enabling the learning machine to perform reinforcement learning.

[0008] The reinforcement learning method of this aspect is a reinforcement learning method for a learning machine, the learning machine comprising: a first agent that adjusts the manufacturing conditions of the manufacturing device based on observation data obtained from observing the state of the manufacturing device; and a second agent that has a function model or function approximator that represents the relationship between the observation data and the manufacturing conditions in a manner different from that of the first agent. In this reinforcement learning method, the manufacturing conditions output by the first agent in reinforcement learning are adjusted using the observation data and the function model or function approximator of the second agent, and reward data corresponding to the state of the product manufactured by the manufacturing device under the adjusted manufacturing conditions is calculated. Based on the observation data and the calculated reward data, the first agent and the second agent are subjected to reinforcement learning.

[0009] The computer program of this aspect is used to enable a learning machine to perform reinforcement learning via a computer. The learning machine comprises: a first agent that adjusts the manufacturing conditions of a manufacturing apparatus based on observation data obtained from observing the state of that apparatus; and a second agent that has a function model or function approximator representing the relationship between the observation data and the manufacturing conditions in a manner different from that of the first agent. The computer program causes the computer to perform the following processing: adjusting the manufacturing conditions output by the first agent in reinforcement learning using the observation data and the function model or function approximator of the second agent; calculating reward data corresponding to the state of a product manufactured by the manufacturing apparatus under the adjusted manufacturing conditions; and enabling the first and second agents to perform reinforcement learning based on the observation data and the calculated reward data.

[0010] The reinforcement learning apparatus of this aspect enables a learning machine to perform reinforcement learning by adjusting the manufacturing conditions of a manufacturing apparatus based on observation data obtained from observing the state of the manufacturing apparatus. The learning machine comprises: a first agent that adjusts the manufacturing conditions of the manufacturing apparatus based on the observation data; a second agent having a function model or function approximator that represents the relationship between the observation data and the manufacturing conditions in a manner different from that of the first agent; and an adjustment unit that adjusts the manufacturing conditions searched by the first agent in the reinforcement learning using the observation data and the function model or function approximator of the second agent. The reinforcement learning apparatus further comprises a reward calculation unit that calculates reward data corresponding to the state of a product manufactured by the manufacturing apparatus under the adjusted manufacturing conditions. The learning machine performs reinforcement learning on the first agent and the second agent based on the observation data and the reward data calculated by the reward calculation unit.

[0011] The molding machine of this aspect includes: the aforementioned reinforcement learning device; and a manufacturing device using the manufacturing conditions action adjusted by the first agent.

[0012] Invention Effects

[0013] According to this disclosure, in the reinforcement learning of a learning machine that adjusts the manufacturing conditions of a manufacturing apparatus, the optimal manufacturing conditions are searched safely without limiting the search range to a certain range, thereby enabling the learning machine to perform reinforcement learning. Attached Figure Description

[0014] Figure 1 This is a schematic diagram illustrating an example of the configuration of the molding machine system according to Embodiment 1.

[0015] Figure 2 This is a block diagram showing an example of the configuration of the molding machine system according to Embodiment 1.

[0016] Figure 3 This is a functional block diagram of the molding machine system according to Implementation Method 1.

[0017] Figure 4 This is a conceptual diagram illustrating the function model and the search range.

[0018] Figure 5 This is a flowchart illustrating the processing sequence of the processor.

[0019] Figure 6 This is a flowchart illustrating the order of the search range adjustment process in Implementation 2. Detailed Implementation

[0020] The following description, with reference to the accompanying drawings, illustrates specific examples of reinforcement learning methods, computer programs, reinforcement learning devices, and manufacturing apparatuses according to embodiments of the present invention. At least some of the embodiments described below can be combined in any way. Furthermore, the present invention is not limited to these examples and is intended to include all modifications shown in the claims, within the meaning and scope equivalent to the claims.

[0021] Figure 1 This is a schematic diagram illustrating an example of the configuration of the molding machine system in Embodiment 1. Figure 2 This is a block diagram illustrating an example of the configuration of the molding machine system according to Embodiment 1. Figure 3 This is a functional block diagram of the molding machine system of Embodiment 1. The molding machine system of Embodiment 1 includes a molding machine (manufacturing device) 2 with a manufacturing condition adjustment device 1, and a measuring unit 3.

[0022] The molding machine 2 can be, for example, an injection molding machine, a blow molding machine, a film molding machine, an extruder, a twin-screw extruder, a spinning extruder, a granulator, a magnesium injection molding machine, etc. Hereinafter, in this embodiment 1, the molding machine 2 will be described as an injection molding machine. The molding machine 2 includes an injection unit 21, a mold clamping device 22 disposed in front of the injection unit 21, and a control device 23 for controlling the operation of the molding machine 2.

[0023] The injection device 21 includes: a heating cylinder; a screw configured to be driven in the heating cylinder in both the rotational and axial directions; a rotary motor that drives the screw in the rotational direction; and a motor that drives the screw in the axial direction.

[0024] The mold clamping device 22 includes a toggle mechanism that secures the mold without opening it when the mold is opened and closed to fill the mold with molten resin injected from the injection device 21, and a motor that drives the toggle mechanism.

[0025] The control device 23 controls the operation of the injection unit 21 and the mold closing unit 22. In this embodiment 1, the control device 23 includes a manufacturing condition adjustment device 1. The manufacturing condition adjustment device 1 is a device for adjusting multiple parameters of the molding conditions of the molding machine 2. In particular, the manufacturing condition adjustment device 1 of this embodiment 1 has the function of adjusting the parameters in a way that reduces the degree of defects in the molded product.

[0026] The molding machine 2 is configured with parameters specifying molding conditions, including resin temperature inside the mold, nozzle temperature, cylinder temperature, hopper temperature, clamping force, injection speed, injection acceleration, peak injection pressure, injection stroke, resin pressure at the cylinder front end, check ring seat status, holding pressure switching pressure, holding pressure switching speed, holding pressure switching position, holding pressure completion position, buffer position, metering back pressure, metering torque, metering completion position, screw retraction speed, cycle time, mold closing time, injection time, holding pressure time, metering time, and mold opening time. The optimal parameters vary depending on the environment of the molding machine 2 and the molded product.

[0027] The measuring unit 3 is a device that measures the physical quantities actually formed during molding based on the molding machine 2. The measuring unit 3 outputs the physical quantity data obtained by the measurement processing to the manufacturing condition adjustment device 1. The physical quantities include temperature, position, speed, acceleration, current, voltage, pressure, time, image data, torque, force, strain, power consumption, etc.

[0028] The information measured by the measuring unit 3 includes, for example, information about the molded product, molding conditions (measured values), peripheral equipment settings (measured values), and ambient gas information. This peripheral equipment is equipment that forms part of a system linked to the molding machine 2, including a mold clamping device 22 and a mold. Examples of peripheral equipment include a molded product removal device (robot or robotic arm), an insert product insertion device, a nesting insertion device, an in-mold foil feeding device, a tape feeding device for tape forming, a gas injection device for gas-assisted molding, a gas injection device and a long fiber injection device for foaming molding using supercritical fluids, a material mixing device for LIM molding, a deburring device for the molded product, a runner cutting device, a molded product weighing instrument, a molded product strength testing machine, an optical inspection device for the molded product, a molded product imaging device and image processing device, and a robotic arm for handling the molded product.

[0029] Molded product information includes, for example, camera images of the molded product, deformation data obtained using a laser displacement sensor, optical measurements of the product's color and brightness obtained using optical measuring instruments, weight measured using a weighing instrument, and strength measured using a strength gauge. This information indicates whether the molded product is normal, its defect type, and the severity of the defect, and is also used in reward calculations.

[0030] Molding conditions include information measured using thermometers, pressure gauges, velocity measuring devices, acceleration measuring devices, position sensors, timers, weighing instruments, etc., such as the resin temperature inside the mold, nozzle temperature, cylinder temperature, hopper temperature, clamping force, injection speed, injection acceleration, peak injection pressure, injection stroke, resin pressure at the front end of the cylinder, check ring seat status, holding pressure switching pressure, holding pressure switching speed, holding pressure switching position, holding pressure completion position, buffer position, metering back pressure, metering torque, metering completion position, screw retraction speed, cycle time, mold closing time, injection time, holding pressure time, metering time, and mold opening time.

[0031] Peripheral equipment settings include information such as mold temperature (set to a fixed value) and mold temperature (set to a variable value) and particle supply rate, which are obtained by using thermometers, measuring instruments, etc.

[0032] Ambient gas information includes information such as ambient gas temperature, ambient gas humidity, and convection-related information (Reynolds number, etc.) obtained using thermometers, hygrometers, flow meters, etc.

[0033] In addition to these, the measuring unit 3 can also measure the mold opening amount, reflux flow rate, tie rod deformation amount, and heater heating rate.

[0034] Manufacturing condition adjustment device 1 is a computer, such as Figure 2As shown, the hardware configuration includes a processor 11 (reinforcement learning device), a storage unit (memory) 12, and an operation unit 13. The processor 11 has arithmetic circuits for CPUs (Central Processing Units), multi-core CPUs, GPUs (Graphics Processing Units), GPGPUs (General-purpose computing on graphics processing units), TPUs (Tensor Processing Units), ASICs (Application Specific Integrated Circuits), FPGAs (Field-Programmable Gate Arrays), and NPUs (Neural Processing Units), internal storage devices such as ROMs (Read Only Memory) and RAMs (Random Access Memory), and I / O terminals. The processor 11 functions as a physical quantity acquisition unit 14, a control unit 15, and a learning machine 16 by executing a computer program (program product) 12a stored in the storage unit 12 (described later). Furthermore, each functional part of the manufacturing condition adjustment device 1 can be implemented in software or in hardware, either partially or entirely.

[0035] Storage unit 12 is a non-volatile memory such as a hard disk, EEPROM (Electrically Erasable Programmable ROM), or flash memory. Storage unit 12 stores a computer program 12a for enabling the computer to perform reinforcement learning processing and parameter adjustment processing of the learning machine 16.

[0036] The computer program 12a of this embodiment 1 can be in a form that can be readable by a computer and recorded in the storage medium 4. The storage unit 12 stores the computer program 12a read from the storage medium 4 by a reading device (not shown). The storage medium 4 is a semiconductor memory such as flash memory. Alternatively, the storage medium 4 can also be an optical disc such as a CD (Compact Disc)-ROM, DVD (Digital Versatile Disc)-ROM, or BD (Blu-ray Disc). Moreover, the storage medium 4 can also be a disk such as a floppy disk or a hard disk, or a magneto-optical disk. In addition, the computer program 12a of this embodiment 1 can also be downloaded from an external server (not shown) connected to a communication network (not shown) and stored in the storage unit 12.

[0037] The operation unit 13 is an input device such as a touch panel, soft keys, hard keys, keyboard, mouse, etc.

[0038] The physical quantity acquisition unit 14 acquires physical quantity data measured and output by the measurement unit 3 during molding based on the molding machine 2. The physical quantity acquisition unit 14 outputs the acquired physical quantity data to the control unit 15.

[0039] Control unit 15 Figure 3 The device shown includes an observation unit 15a and a calculation unit 15b. Physical quantity data output from the measurement unit 3 is input into the observation unit 15a.

[0040] The observation unit 15a observes the state of the molding machine 2 and the molded product by analyzing the physical quantity data, and outputs the observed data to the first agent (also known as intelligent agent) 16a and the second agent 16b of the learning machine 16. Since the physical quantity data contains a large amount of information, the observation unit 15a only needs to generate observation data that is a compressed version of the physical quantity data. The observation data represents information such as the state of the molding machine 2 and the state of the molded product.

[0041] For example, the observation unit 15a calculates observation data based on camera images and measurements from a laser displacement sensor. This observation data displays feature quantities representing the appearance of the molded article, the dimensions, area, volume, and optical axis offset of the optical components (molded article). Furthermore, the observation unit 15a can perform preprocessing on time-series waveform data relative to injection speed, injection pressure, holding pressure, etc., and extract feature quantities from this time-series waveform data as observation data. Alternatively, time-series waveform data and image data representing the time-series waveforms can also be used as observation data.

[0042] Furthermore, the observation unit 15a calculates the degree of defect in the molded product by analyzing physical quantity data, and outputs the calculated degree of defect to the reward calculation unit 15b. The degree of defect may include, for example, the area of ​​burrs, the area of ​​short sides, the amount of deformation such as shrinkage marks / warping / twisting, the length of welds, the size of silver streaks, the degree of spraying, the size of flow marks, and the amount of color change caused by uneven coloring. Alternatively, the degree of defect may be set as the amount of change between the observation data obtained from the molding machine and the baseline observation data when the product becomes a qualified product.

[0043] The reward calculation unit 15b calculates the reward data, which serves as a benchmark for the quality of the defect output from the observation unit 15a, and outputs the calculated reward data to the first agent 16a and the second agent 16b of the learning machine 16.

[0044] Alternatively, as described later, the reward can be configured such that if the action a1 output from the first agent 16a falls outside the search range output from the second agent 16b, a negative reward is added corresponding to the degree of deviation. In other words, the reward data can be calculated by adding a larger negative reward (a larger absolute value negative reward) as the deviation of the action a1 output from the first agent 16a from the search range output from the second agent 16b is greater.

[0045] Learning Machine 16 Figure 3 As shown, it includes a first agent 16a, a second agent 16b, and a regulation unit 16c. The first agent 16a and the second agent 16b are agents of different types. The first agent 16a is a more complex model compared to the second agent 16b. The first agent 16a is a model with richer expressiveness compared to the second agent 16b. In other words, the first agent 16a is a model that can achieve better parameter regulation using reinforcement learning compared to the second agent 16b.

[0046] The search range for molding conditions obtained using the first agent 16a is larger than that of the second agent 16b, but there is a concern that abnormal operation of the molding machine 2 may cause unpredictable adverse situations for the molding machine 2 and the operator. On the other hand, the search range of the second agent 16b is narrower than that of the first agent 16a, but the probability of abnormal operation of the molding machine 2 is very low.

[0047] The first agent 16a includes, for example, reinforcement learning models with deep neural networks such as DQN, A3C, and D4PG, as well as reinforcement learning models from model libraries such as PlaNet and SLAC.

[0048] In the case of a reinforcement learning model with a deep neural network, the first agent 16a possesses a DQN (Deep Q-Network), which determines the action a1 corresponding to the state s of the molding machine 2 based on the observed data. The DQN is a neural network model that outputs the value of multiple actions a1 given the state s represented by the observed data. These multiple actions a1 correspond to molding conditions. A high-value action a1 indicates an appropriate molding condition to be set for the molding machine 2. The molding machine 2 transitions to other states based on the actions a1. After a state transition, the first agent 16a receives the reward calculated by the reward calculation unit 15b, and learns to maximize the reward, that is, the accumulation of rewards.

[0049] More specifically, DQN has an input layer, an intermediate layer, and an output layer. The input layer has multiple nodes that receive the input state s (i.e., the observed data). The output layer has multiple nodes that output the value Q(s, a1) of each behavior a1 in the input state s. Behavior a1 can correspond to the value of a parameter of the shaping condition and can be a variable. Here, behavior a1 is set as the parameter value.

[0050] Based on the state s, behavior a1, and the reward r obtained using that behavior, the value Q, expressed by the following equation (1), is used as training data to adjust various weight coefficients representing DQN, thereby enabling the DQN reinforcement learning of the first agent 16a.

[0051] Q(s, a1)←Q(s, a1)+α(r+γmaxQ(snext, a1next)-Q(s, a1))(1)

[0052] Here,

[0053] s: state

[0054] a1: Behavior

[0055] α: Learning coefficient

[0056] r: reward

[0057] γ: Discount rate

[0058] maxQ(snext, a1next): The maximum value of Q for the next action.

[0059] In the case of a reinforcement learning model in the model library, the first agent 16a has a state representation graph and uses the state representation graph as a guideline for behavior determination to determine parameters (behavior a1). The first agent 16a uses the state representation graph to determine the parameters (behavior a1) corresponding to the state of the molding machine 2, based on the state s represented by the observation data. The state representation graph is, for example, a model that, given the input observation data (state s) and parameters (behavior a1), outputs a reward r for adopting parameters (behavior a1) in state s, and a state transition probability (confidence) Pt to the next state s′. The reward r can be considered as information about whether the molded product obtained in state s is normal when a certain parameter (behavior a) is set. Behavior a1 is the parameter set for the molding machine 2 in that state. Based on behavior a1, the molding machine 2 transitions to another state. After the state transition, the first agent 16a receives the reward calculated by the reward calculation unit 15b and updates the state representation graph.

[0060] The second agent 16b has a functional model or function approximator representing the relationship between observed data and parameters of shaping conditions. The functional model can be, for example, a functional model that can be defined using interpretable domain knowledge. The functional model can be implemented based on approximations of multinomial functions, exponential functions, logarithmic functions, trigonometric functions, etc., and approximations of probability distributions based on uniform distributions, multinomial distributions, Gaussian distributions, Gaussian mixture models (GGM). The functional model can be a linear function or a nonlinear function. Alternatively, the distribution can be defined using histograms, kernel density estimations, or function approximators using neighborhood methods, decision trees, shallow neural network functions, etc., to construct the second agent 16b.

[0061] Figure 4 This is a conceptual diagram illustrating the function model and search range. The function model of Agent 2 16b takes as input, for example, observed data (state s), and parameters of the shaping conditions (behavior a2), and returns the optimal probability. The optimal probability is the probability that behavior a2 is optimal under state s, calculated based on the degree of defect or reward. Figure 4 The horizontal axis of the graph shows a parameter of the forming condition (fixed observation data and other parameters), and the vertical axis shows the state represented by the observation data and the optimal probability of the parameter. By substituting the observation data and reward into the function model of Agent 16b, the range of candidate parameters that would be the optimal forming condition can be calculated and used as the search range. There is no particular limitation on the method for setting the search range, but it can be, for example, a specified confidence interval, such as a 95% confidence interval. Alternatively, if the curve representing the optimal probability relative to a parameter (fixed observation data and other parameters) can be empirically defined as a Gaussian distribution, the search range for that parameter can also be set using a confidence interval expressed in 2σ.

[0062] The search range can also be set when using a function approximator to construct the second agent 16b.

[0063] By replacing the first agent 16a and causing the second agent 16b to act randomly within a specified search range, the learning of the second agent 16b can also be performed before the learning of the first agent 16a. By allowing only the second agent 16b to learn in advance, the first agent 16a can learn within a safer and wider range.

[0064] The adjustment unit 16c adjusts the parameters (behavior a1) searched by the first agent 16a in reinforcement learning based on the search range calculated using the second agent 16b, and outputs the adjusted parameters (behavior a).

[0065] The following describes the reinforcement learning method of Embodiment 1 in detail.

[0066] [Reinforcement learning processing]

[0067] Figure 5 This is a flowchart illustrating the processing sequence of processor 11. The initial values ​​of the parameters set for molding machine 2 are assumed to indicate that actual molding has been performed.

[0068] First, when the molding machine 2 is molding, the measuring unit 3 measures the physical quantities of the molding machine 2 and the molded product, and outputs the measured physical quantity data to the control unit 15 (step S11).

[0069] The control unit 15 acquires physical quantity data output from the measurement unit 3, generates observation data based on the acquired physical quantity data, and outputs the generated observation data to the first agent 16a and the second agent 16b of the learning machine 16 (step S12).

[0070] The first agent 16a of the learning machine 16 acquires the observation data output from the observation unit 15a, calculates the parameter (behavior a1) for adjusting the parameters of the molding machine 2 based on the observation data (step S13), and outputs the calculated parameter (behavior a1) to the adjustment unit 16c (step S14). The first agent 16a selects the optimal behavior a1 during application (inference), and determines a searchable behavior a1 during learning to enhance the learning of the first agent 16a. Furthermore, the first agent 16a can use a objective function that prioritizes behaviors with higher values, smaller values ​​for behaviors that are less frequently searched, and larger values ​​for behaviors with greater changes in the current molding conditions, and selects the behavior a1 with a smaller value based on this objective function.

[0071] The second agent 16b of the learning machine 16 acquires the observation data output from the observation unit 15a, calculates the search range data representing the search range of the parameters based on the observation data (step S15), and outputs the calculated search range data to the adjustment unit 16c (step S16).

[0072] The adjustment unit 16c of the learning machine 16 adjusts the parameter output from the first agent 16a to fall within the search range output from the second agent 16b (step S17). That is, the adjustment unit 16c determines whether the parameter output from the first agent 16a is within the search range output from the second agent 16b. Then, if it determines that the parameter is outside the search range, it changes the parameter to fall within the search range. If the parameter is within the search range, the parameter output from the first agent 16a is directly used.

[0073] The adjustment unit 16c outputs the adjusted parameters (behavior a) to the molding machine 2 (step S18).

[0074] The molding machine 2 adjusts the molding conditions using parameters and performs molding processing according to the adjusted molding conditions. The operation of the molding machine 2 and the physical quantities of the molded product are input to the measuring unit 3. The molding process can also be repeated multiple times. When the molding machine 2 performs molding, the measuring unit 3 measures the physical quantities of the molding machine 2 and the molded product, and outputs the measured physical quantity data to the observation unit 15a of the control unit 15 (step S19).

[0075] The observation unit 15a of the control unit 15 acquires physical quantity data output from the measurement unit 3, generates observation data based on the acquired physical quantity data, and outputs the generated observation data to the first agent 16a and the second agent 16b of the learning machine 16 (step S20). Additionally, the reward calculation unit 15b calculates predetermined reward data based on the physical quantity data measured by the measurement unit 3 and according to the degree of defect in the molded product, and outputs the calculated reward data to the learning machine 16 (step S21). However, if the behavior a1 output from the first agent 16a falls outside the search range, a negative reward is added corresponding to the degree of deviation. That is, the greater the deviation of the behavior a1 output from the first agent 16a from the search range output from the second agent 16b, the greater the negative reward (the larger the absolute value of the negative reward), thereby calculating the reward data.

[0076] The first agent 16a updates the model based on the observation data output from the observation unit 15a and the reward data output from the reward calculation unit 15b (step S22). In the case that the first agent 16a is a DQN, the value expressed by the above formula (1) is used as training data to enable the DQN to learn.

[0077] The second agent 16b updates the model based on the observation data output from the observation unit 15a and the reward data output from the reward calculation unit 15b (step S23). The second agent 16b can update the function model or function approximator by using, for example, the least squares method, the maximum likelihood estimation method, or the Bayesian estimation method.

[0078] According to the reinforcement learning method of Embodiment 1 configured as described above, in the reinforcement learning of the learning machine 16 that adjusts the molding conditions of the molding machine 2, the optimal molding conditions can be safely searched without limiting the search range to a certain range, so that the learning machine 16 can perform reinforcement learning.

[0079] Specifically, compared to the second agent 16b, the learning machine 16 of this embodiment 1 is able to perform reinforcement learning on the optimal molding conditions using the first agent 16a, which has a higher ability to learn the optimal molding conditions.

[0080] Furthermore, the search range of molding conditions obtained using the first agent 16a is larger than that of the second agent 16b. There is concern that abnormal operation of the molding machine 2 may cause unpredictable adverse situations for the molding machine 2 and the operator. However, the adjustment unit 16c can limit the search range to a safe range represented by the second agent 16b, which reflects the function or distribution specified by the user's prior knowledge. Therefore, the first agent 16a can safely search for the optimal molding conditions for reinforcement learning.

[0081] Furthermore, in this embodiment 1, an example of adjusting the molding conditions of an injection molding machine using reinforcement learning was described, but the application scope of the present invention is not limited thereto. For example, it is also possible to use the manufacturing condition adjustment, reinforcement learning method, and computer program 12a of the present invention to adjust the manufacturing conditions of the molding machine 2 of an extruder, a film molding machine, or other manufacturing apparatus using reinforcement learning.

[0082] Furthermore, in this embodiment 1, an example was described in which the molding machine 2 is equipped with the manufacturing condition adjustment device 1 and the reinforcement learning device. However, the manufacturing condition adjustment device 1 or the reinforcement learning device may also be configured separately from the molding machine 2. Additionally, it may be configured to utilize the cloud to execute the reinforcement learning method and parameter adjustment processing.

[0083] The example of the learning machine 16 having two agents is also explained, but it can also have three or more agents. It can also be configured as having a first agent 16a and multiple second agents 16b, 16b... with different function models or function approximators. The adjustment unit 16c adjusts the parameters output by the first agent 16a in reinforcement learning based on the search range calculated using the multiple second agents 16b, 16b... Furthermore, the adjustment can be performed by calculating the search range using the logical sum or logical product of the search ranges calculated by the multiple second agents 16b, 16b... and then converging the parameters output by the first agent 16a within that search range.

[0084] (Implementation Method 2)

[0085] The method for adjusting the search range of the molding machine system parameters in Embodiment 2 differs from that in Embodiment 2. The other components of the molding machine system are the same as those in Embodiment 1; therefore, the same reference numerals are used to label the same parts, and detailed descriptions are omitted.

[0086] Figure 6 This is a flowchart illustrating the adjustment process sequence of the search range in Implementation Method 2. Figure 5 In step S17 shown, processor 11 performs the following processing. Processor 11 obtains a threshold for adjusting the search range (step S31). The threshold is, for example, a specified value. Figure 4 The values ​​(%) of this confidence interval, σ interval, etc., are shown. The control unit 15 or the adjustment unit 16c acquires the threshold, for example, via the operation unit 13. By operating the operation unit 13, the operator can input the threshold and adjust the tolerance of the search range.

[0087] Next, agent 16a calculates the parameters of the forming conditions using the observation data (step S32). Furthermore, agent 16b calculates the search range defined by the threshold obtained in step S31 (step S33).

[0088] Next, the adjustment unit 16c determines whether the parameter calculated using the first agent 16a is within the search range calculated in step S33 (step S34). If it is determined that the parameter falls outside the search range calculated in step S33 (step S34: No), the adjustment unit 16c adjusts the parameter to fall within the search range (step S35). For example, the adjustment unit 16c changes the value to fall within the search range and be closest to the parameter calculated in step S32.

[0089] If, in step S34, it is determined that the parameter is within the search range (step S34: Yes), or if the processing in step S35 has ended, the adjustment unit 16c determines whether the parameter calculated in step S32 is within the specified search range (step S36). The specified search range is a pre-defined numerical range, stored by the storage unit 12. The specified search range is the set of values ​​that the parameter can take; values ​​outside the specified search range are unsettable numerical ranges.

[0090] If the parameter is determined to be within the specified search range (step S36: Yes), the adjustment unit 16c performs the processing of step S18. If the parameter is determined to be outside the specified search range (step S36: No), the adjustment unit 16c adjusts the parameter to fall within the specified search range (step S37). For example, the adjustment unit 16c changes the parameter value to be closest to the value calculated in step S32 within the range of the search range calculated in step S33 and the specified search range.

[0091] According to the reinforcement learning method of Implementation Method 2, the restriction strength based on the search range of the second agent 16b can be freely adjusted. That is, it is possible to select and adjust whether to allow a certain degree of abnormal operation of the molding machine 2 to actively search for better molding conditions so that the first agent 16a can strengthen learning, or to give priority to the normal operation of the molding machine 2 so that the first agent 16a can strengthen learning.

[0092] Furthermore, although there are concerns that the search range calculated by the second agent 16b may be inappropriate due to the learning results of the second agent 16b or the threshold used for adjusting the search range, by setting a specified search range, it is possible to safely search for shaping conditions and enable the learning machine 16 to reinforce its learning.

[0093] (Modified Example)

[0094] In Embodiment 2, an example was mainly described whereby the intensity of the restriction on the search range based on the second agent 16b was adjusted by having the operator set a threshold. However, the threshold could also be automatically adjusted by the adjustment unit 16c. For example, if the learning progress of the first agent 16a is such that rewards above a predetermined value account for a predetermined proportion or more, the adjustment unit 16c could also be configured to change the threshold in a way that expands the search range calculated using the second agent 16b. Conversely, if rewards below a predetermined value account for a predetermined proportion or more, the adjustment unit 16c could also be configured to change the threshold in a way that reduces the search range calculated using the second agent 16b.

[0095] It can also be configured to change the threshold periodically by altering the search range calculated using the second agent 16b. For example, the adjustment unit 16c changes the threshold once out of 10 times to expand the search range, and nine times out of 10 times to reduce the search range to prioritize safety.

[0096] Furthermore, in Embodiment 2, an example was described of adjusting the restriction strength of the search range based on the second agent 16b using a threshold. However, the adjustment unit 16c can be configured to remove the restriction on the search range based on the second agent 16b when the operator's operation meets predetermined conditions. For example, if the learning progress of the first agent 16a is advanced and the reward above a predetermined value accounts for a predetermined proportion or more, the adjustment unit 16c can also remove the restriction on the search range based on the second agent 16b. Additionally, the adjustment unit 16c can remove the restriction on the search range based on the second agent 16b at a predetermined frequency.

[0097] Explanation of reference numerals in the attached figures

[0098] 1 Manufacturing Condition Adjustment Device

[0099] 2 Molding Machine

[0100] 3 Measurement Section

[0101] 4 storage media

[0102] 11 processor

[0103] 12 Storage Units

[0104] 12a Computer Program

[0105] 13 Operations Department

[0106] 14 Physical Quantity Acquisition Department

[0107] 15 Control Department

[0108] 15a Observation Department

[0109] 15b Reward Calculation Department

[0110] 16 learning machines

[0111] 16a First Agent

[0112] 16b Second Agent

[0113] 16c Adjustment Section.

Claims

1. A reinforcement learning method, characterized in that, The reinforcement learning method described is a reinforcement learning method for learning machines. The learning machine has the following features: The first agent adjusts the manufacturing conditions of a molding machine based on observation data obtained from observing the state of the manufacturing device; and The second agent has a functional model or functional approximator that represents the relationship between the observed data and the manufacturing conditions in a manner different from that of the first agent. In the reinforcement learning method, The manufacturing conditions for the first agent search in reinforcement learning are adjusted using the observed data and the function model or function approximator of the second agent. Calculate the reward data corresponding to the state of the product manufactured by the manufacturing apparatus under the adjusted manufacturing conditions. Based on the observed data and the calculated reward data, the first agent and the second agent are subjected to reinforcement learning.

2. The reinforcement learning method according to claim 1, characterized in that, Using the observed data and the function model or function approximator of the second agent, the search range of the manufacturing conditions is calculated. If the manufacturing conditions of the first agent search in reinforcement learning fall outside the calculated search range, the manufacturing conditions of the search are changed to the manufacturing conditions within the search range.

3. The reinforcement learning method according to claim 2, characterized in that, Using the observed data and the function model or function approximator of the second agent, a threshold for calculating the search range of the manufacturing conditions is obtained. Using the acquired threshold, the observed data, and the function model or function approximator of the second agent, the search range of the manufacturing conditions is calculated.

4. The reinforcement learning method according to claim 2 or 3, characterized in that, If the manufacturing conditions of the first agent search in reinforcement learning fall outside the range of the specified search range, the manufacturing conditions of the search are changed to the specified search range and the manufacturing conditions within the calculated search range.

5. The reinforcement learning method according to any one of claims 1 to 3, characterized in that, When the manufacturing conditions searched by the first agent are adjusted by the second agent, the reward data is calculated by adding a negative reward corresponding to the degree of deviation of the first agent from the search range.

6. The reinforcement learning method according to claim 1, characterized in that, The manufacturing apparatus is an injection molding machine. The manufacturing conditions include the following parameters: resin temperature inside the mold, nozzle temperature, cylinder temperature, hopper temperature, clamping force, injection speed, injection acceleration, peak injection pressure, injection stroke, resin pressure at the cylinder front end, check ring seat status, holding pressure switching pressure, holding pressure switching speed, holding pressure switching position, holding pressure completion position, buffer position, metering back pressure, metering torque, metering completion position, screw retraction speed, cycle time, mold closing time, injection time, holding pressure time, metering time, or mold opening time. The reward data is based on observation data from the injection molding machine or data calculated from the degree of defects in the molded articles manufactured by the injection molding machine.

7. A program product, characterized in that, The program product has a computer program for enabling learning machines to reinforce learning via a computer. The learning machine has the following features: The first agent adjusts the manufacturing conditions of a molding machine based on observation data obtained from observing the state of the manufacturing device; and The second agent has a functional model or functional approximator that represents the relationship between the observed data and the manufacturing conditions in a manner different from that of the first agent. The computer program causes the computer to perform the following processes: The manufacturing conditions for the first agent search in reinforcement learning are adjusted using the observed data and the function model or function approximator of the second agent. Calculate the reward data corresponding to the state of the product manufactured by the manufacturing apparatus under the adjusted manufacturing conditions. Based on the observed data and the calculated reward data, the first agent and the second agent are subjected to reinforcement learning.

8. A reinforcement learning device, characterized in that, To enable learning machines to reinforce their learning, The learning machine adjusts the manufacturing conditions of a molding machine based on observation data obtained from observing the state of the manufacturing device. The learning machine has the following features: The first agent adjusts the manufacturing conditions of the manufacturing apparatus based on the observed data; A second agent, having a function model or function approximator that uses the relationship between the observed data and the manufacturing conditions in a manner different from that of the first agent; and The adjustment unit uses the observed data and the function model or function approximator of the second agent to adjust the manufacturing conditions of the first agent search in reinforcement learning. The reinforcement learning device further includes a reward calculation unit that calculates reward data corresponding to the state of the product manufactured by the manufacturing device under the adjusted manufacturing conditions. The learning machine uses the observation data and the reward data calculated by the reward calculation unit to enable the first agent and the second agent to learn reinforcement learning.

9. A molding machine, characterized in that, have: The reinforcement learning device as claimed in claim 8; and A manufacturing apparatus that operates using the manufacturing conditions regulated by the first agent.

10. A storage medium, characterized in that, The storage medium is readable by a computer and contains computer programs used to enable the learning machine to reinforce its learning. The learning machine has the following features: The first agent adjusts the manufacturing conditions of a molding machine based on observation data obtained from observing the state of the manufacturing device; and The second agent has a functional model or functional approximator that represents the relationship between the observed data and the manufacturing conditions in a manner different from that of the first agent. The computer program causes the computer to perform the following processes: The manufacturing conditions for the first agent search in reinforcement learning are adjusted using the observed data and the function model or function approximator of the second agent. Calculate the reward data corresponding to the state of the product manufactured by the manufacturing apparatus under the adjusted manufacturing conditions. Based on the observed data and the calculated reward data, the first agent and the second agent are subjected to reinforcement learning.

Citation Information

Patent Citations

  • Injection molding machine system that adjusts molding conditions by machine learning device

    JP2019166702A

  • Machining equipment system and manufacturing system

    CN109308048A

  • Reward function modeling method under multi-agent reinforcement learning cooperation task

    CN112084721A