Reinforcement learning device, injection molding machine, reinforcement learning method, and computer program

The reinforcement learning device optimizes cooling time in injection molding by using temperature and pressure data to prevent defects, enhancing product quality.

JP2026017235APending Publication Date: 2026-02-04THE JAPAN STEEL WORKS LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024117983
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-23
Publication Date
2026-02-04

AI Technical Summary

Technical Problem

Existing injection molding systems lack the ability to automatically optimize cooling time, leading to potential defects in molded products due to insufficient or excessive cooling.

Method used

A reinforcement learning device that acquires cavity surface temperature and in-mold pressure data, selects the optimal cooling stop timing based on reinforcement learning, and adjusts cooling time to prevent defects.

Benefits of technology

Automatically optimizes cooling time in injection molding, preventing defects by learning the relationship between temperature, pressure, and cooling time, ensuring high-quality molded products.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026017235000001_ABST
    Figure 2026017235000001_ABST
Patent Text Reader

Abstract

To provide a reinforcement learning device capable of automatically optimizing a cooling time in injection molding.SOLUTION: The reinforcement learning device includes an acquisition unit that acquires temperature data indicating at least a cavity surface temperature of a mold in a cooling process of a molded article, an action selector that selects a cooling stop timing in injection molding based on the acquired temperature data, an evaluator that calculates reward data based on quality data indicating quality of a molded article obtained by terminating the cooling process at the selected cooling stop timing and a cooling time of the molded article, and a learning device that causes the action selector to perform reinforcement learning based on the temperature data of the cavity surface temperature and the calculated reward data.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a reinforcement learning device, an injection molding machine, a reinforcement learning method, and a computer program. [Background technology]

[0002] There is an injection molding machine system that can automatically adjust molding conditions of an injection molding machine by using reinforcement learning (for example, Patent Document 1). [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Publication No. 2019-166702 Summary of the Invention [Problem to be solved by the invention]

[0004] An object of the present disclosure is to provide a reinforcement learning device, an injection molding machine, a reinforcement learning method, and a computer program that can automatically optimize the cooling time in injection molding. [Means for solving the problem]

[0005] The reinforcement learning device according to this embodiment comprises an acquisition unit that acquires temperature data indicating at least the cavity surface temperature of the mold during the cooling process of the molded product; a behavior selector that selects the timing to stop cooling in injection molding based on the temperature data acquired by the acquisition unit; an evaluator that calculates reward data based on pass / fail data indicating the pass / fail of the molded product obtained by completing the cooling process at the cooling stop timing selected by the behavior selector and the cooling time of the molded product; and a learning unit that causes the behavior selector to perform reinforcement learning based on the temperature data of the cavity surface temperature and the calculated reward data.

[0006] The injection molding machine of this embodiment comprises an acquisition unit that acquires temperature data indicating at least the cavity surface temperature of the mold during the cooling process of the molded product; a behavior selector that selects the timing to stop cooling in injection molding based on the temperature data acquired by the acquisition unit; an evaluator that calculates reward data based on pass / fail data indicating the pass / fail of the molded product obtained by completing the cooling process at the cooling stop timing selected by the behavior selector and the cooling time of the molded product; and a learning unit that causes the behavior selector to perform reinforcement learning based on the temperature data of the cavity surface temperature and the calculated reward data.

[0007] The reinforcement learning method of this embodiment acquires temperature data indicating at least the cavity surface temperature of a mold during the cooling process of a molded product, selects the timing to stop cooling in injection molding based on the acquired temperature data, calculates reward data based on pass / fail data indicating the pass / fail of the molded product obtained by completing the cooling process at the selected cooling stop timing and the cooling time of the molded product, and reinforces learning the timing to stop cooling in injection molding based on the temperature data of the cavity surface temperature and the calculated reward data.

[0008] The computer program of this embodiment acquires temperature data indicating at least the cavity surface temperature of the mold during the cooling process of the molded product, selects the timing to stop cooling in injection molding based on the acquired temperature data, calculates reward data based on pass / fail data indicating the pass / fail of the molded product obtained by completing the cooling process at the selected cooling stop timing and the cooling time of the molded product, and causes the computer to execute a process of reinforcement learning to determine the timing to stop cooling in injection molding based on the temperature data of the cavity surface temperature and the calculated reward data. [Effects of the Invention]

[0009] According to the present disclosure, the cooling time in injection molding can be automatically optimized. [Brief explanation of the drawings]

[0010] [Figure 1]1 is a schematic diagram illustrating an example of the configuration of an injection molding machine according to a first embodiment. [Figure 2] FIG. 2 is a cross-sectional view showing a mold portion in which a resin temperature sensor and an in-mold pressure sensor are provided. [Figure 3] 1 is a block diagram showing an example of the configuration of a reinforcement learning device according to a first embodiment. [Figure 4] FIG. 1 is a functional block diagram of an injection molding machine and a reinforcement learning device according to a first embodiment. [Figure 5] 1 is a flowchart showing a processing procedure of reinforcement learning according to the first embodiment. [Figure 6] 4 is a flowchart showing the procedure of a cooling time adjustment process according to the first embodiment. [Figure 7] FIG. 10 is a functional block diagram of an injection molding machine and a reinforcement learning device according to a second embodiment. [Figure 8] 10 is a flowchart showing the procedure of a cooling time adjustment process according to the second embodiment. [Figure 9] FIG. 10 is a functional block diagram of an injection molding machine and a reinforcement learning device according to a third embodiment. [Figure 10] 11 is a flowchart showing the procedure of a cooling time adjustment process according to the third embodiment. [Figure 11] FIG. 10 is a functional block diagram of an injection molding machine and a reinforcement learning device according to a fourth embodiment. [Figure 12] 10 is a flowchart showing the procedure of a cooling time adjustment process according to the fourth embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0011] Specific examples of a reinforcement learning device, an injection molding machine, a reinforcement learning method, and a computer program according to embodiments of the present invention will be described below with reference to the drawings. At least some of the embodiments described below may be combined in any manner. Note that the present invention is not limited to these examples, but is defined by the claims, and is intended to include all modifications within the meaning and scope equivalent to the claims.

[0012] (Embodiment 1) 1 is a schematic diagram illustrating an example configuration of an injection molding machine 1 according to embodiment 1. The injection molding machine 1 according to embodiment 1 includes a mold clamping device 2 that clamps a mold 21, an injection device 3 that melts and injects resin, which is a molding material, a control device 4, a reinforcement learning device 5, a resin temperature sensor (temperature sensor) 6, an in-mold pressure sensor (pressure sensor) 7, and a molded product inspection device 8. Note that, although embodiment 1 describes an example in which the molding material is resin, the technology according to the present disclosure can also be applied to an injection molding machine 1 that uses a molding material other than resin, such as magnesium.

[0013] The mold clamping device 2 includes a fixed platen 22 fixed on a bed 20, a mold clamping housing 23 slidably provided on the bed 20, and a movable platen 24 that similarly slides on the bed 20. The fixed platen 22 and the mold clamping housing 23 are connected by a plurality of, for example, four tie bars 25, 25, .... The movable platen 24 is configured to be slidable between the fixed platen 22 and the mold clamping housing 23. A mold clamping mechanism 26 is provided between the mold clamping housing 23 and the movable platen 24. The mold clamping mechanism 26 is configured, for example, by a toggle mechanism. Note that the mold clamping mechanism 26 may also be configured as a direct pressure type mold clamping mechanism, i.e., a mold clamping cylinder. A fixed mold 27 and a movable mold 28 are provided on the fixed platen 22 and the movable platen 24, respectively, and the mold 21 is opened and closed when the mold clamping mechanism 26 is driven.

[0014] 2 is a cross-sectional view showing a mold 21 portion provided with a resin temperature sensor 6 and an in-mold pressure sensor 7. The mold 21 includes a fixed mold 27 and a movable mold 28 opposed to the fixed mold 27.

[0015] The fixed mold 27 includes a fixed-side mold plate 27a and a fixed-side mounting plate 27b, and the fixed-side mounting plate 27b is fixed to the fixed platen 22. A sprue bushing 27c that defines a sprue is formed on the fixed-side mold plate 27a and the fixed-side mounting plate 27b. On the other hand, the movable mold 28 is clamped to the fixed mold 27 to define a cavity 27d, and includes a movable side mold plate 28a, a movable side mounting plate 28b, a receiving plate 28c, a spacer block 28d, an ejector plate 28e, and an ejector pin 28f.

[0016] The fixed platen 22 has a flow path 22a that runs from one surface, to which the fixed-side mounting plate 27b is fixed, to the other surface. An opening (valve gate) at one end of the flow path 22a is connected to a sprue bushing 27c, and an opening at the other end is formed on the other surface. An annular contact portion 22b, to which a nozzle 31a of the injection device 3 is in close contact, is provided around the opening formed on the other surface. The molten molding material injected from the nozzle 31a of the injection device 3 is supplied to the cavity 27d of the mold 21 through the flow path 22a and the sprue. A heating device 22c is provided around the flow path 22a to keep the molten molding material flowing through the flow path 22a warm. A resin temperature sensor 6 and an in-mold pressure sensor 7 are also provided at appropriate locations on the mold 21.

[0017] The above-described configurations of the fixed mold 27 and the movable mold 28 are merely examples, and the shape of the cavity 27d is not particularly limited.

[0018] The resin temperature sensor 6 is, for example, an infrared temperature sensor that detects the cavity surface temperature by detecting infrared rays emitted from the resin in the mold 21. The cavity surface temperature is the temperature near the cavity surface of the resin filled in the cavity 27d or the mold 21. The resin temperature sensor 6 is connected to a temperature measurement amplifier 61. The temperature measurement amplifier 61 is a device that receives temperature-related signals output from the infrared temperature sensor at a set sampling period and sequentially transmits the received temperature data to the reinforcement learning device 5. The temperature measurement amplifier 61 transmits temperature data measured, for example, at intervals of several milliseconds to several tens of milliseconds, to the reinforcement learning device 5. The reinforcement learning device 5 receives and stores the temperature data sequentially transmitted from the temperature measurement amplifier 61. The temperature sampling period may be set appropriately depending on the processing capacity of the reinforcement learning device 5. The infrared temperature sensor is an example of the resin temperature sensor 6, and the temperature detection method is not particularly limited, and may be a thermocouple-type non-contact temperature sensor, for example.

[0019] The in-mold pressure sensor 7 is a pressure sensor that detects the pressure exerted by the resin in the mold 21. The in-mold pressure sensor 7 includes, for example, a pin that advances and retreats in response to the pressure of the resin in the mold 21, a diaphragm that bends when the pin is pressed, and a strain gauge whose electrical resistance value changes with the bending of the diaphragm, and outputs a signal corresponding to the resistance value of the strain gauge. The in-mold pressure sensor 7 is connected to a pressure measurement amplifier 71. The pressure measurement amplifier 71 receives pressure-related signals output from the pressure sensor at a set sampling period and sequentially transmits the received pressure data to the reinforcement learning device 5. The pressure measurement amplifier 71 transmits pressure data measured, for example, at intervals of several milliseconds to several tens of milliseconds, to the reinforcement learning device 5. The reinforcement learning device 5 receives and stores the pressure data sequentially transmitted from the pressure measurement amplifier 71. The pressure sampling period may be set appropriately depending on the processing capability of the reinforcement learning device 5. The in-mold pressure sensor 7 using a strain gauge is just one example, and the pressure detection method is not particularly limited, and may be a capacitance type pressure sensor, a piezoelectric voltage type pressure sensor, a semiconductor pressure sensor, or the like.

[0020] The locations where the resin temperature sensor 6 and in-mold pressure sensor 7 are provided are not particularly limited, but it is preferable to provide the resin temperature sensor 6 and in-mold pressure sensor 7 in at least one of the locations of the mold 21 corresponding to the thick wall portions, corner portions, edge portions, ribs, and bosses of the molded product, so as to detect the cavity surface temperature and in-mold pressure at that location. It is also preferable to provide the resin temperature sensor 6 and in-mold pressure sensor 7 in locations of the mold 21 corresponding to multiple locations of the molded product with different thicknesses, so as to detect the cavity surface temperature and in-mold pressure at that location. The resin temperature sensor 6 and in-mold pressure sensor 7 may be provided in the same location or in different locations. The number of resin temperature sensors 6 and in-mold pressure sensors 7 provided is not particularly limited, and they may be provided in multiple locations on the mold 21 depending on the shape of the cavity 27d.

[0021] When resin temperature sensors 6 are provided at multiple locations, it is advisable to provide them at multiple locations where the temperature difference of the molded product during the cooling process is large.It is also advisable to provide resin temperature sensors 6 at multiple locations where the temperature change process of the molded product during the cooling process is different. Similarly, when installing the in-mold pressure sensors 7 at multiple locations, it is advisable to install them at multiple locations where the pressure difference inside the mold is large during the cooling process.It is also advisable to install the in-mold pressure sensors 7 at multiple locations where the process of the in-mold pressure change during the cooling process is different.

[0022] The molded product inspection device 8 is, for example, an imaging device that captures an image of the molded product. The imaging device captures an image of the molded product and transmits image data showing the appearance of the molded product as inspection data. The imaging device is a still image camera or a video camera. The imaging device may be a smartphone, tablet terminal, laptop-type personal computer (PC) or the like that has a light-receiving lens and an imaging element. The molded product inspection device 8 may also be a laser displacement sensor that measures the amount of deformation of the molded product, an optical measuring instrument that measures the color of the molded product, a weighing scale that measures the weight of the molded product, a strength measuring instrument that measures the strength of the molded product, or the like.

[0023] Returning to FIG. 1, the injection device 3 will be described. The injection device 3 is provided on a base 30. The injection device 3 includes a heating cylinder 31 having a nozzle 31a at its tip, and a screw 32 disposed within the heating cylinder 31 so as to be rotatable in both the circumferential and axial directions. The screw 32 is driven in both the rotational and axial directions by a drive mechanism 33. The drive mechanism 33 is composed of a rotary motor that drives the screw 32 in the rotational direction, a motor that drives the screw 32 in the axial direction, and the like. Note that the drive mechanism 33 shown in FIG. 1 is covered with a cover, and therefore the internal configuration is not shown.

[0024] A hopper 34 into which molding material is poured is provided near the rear end of the heating cylinder 31. The injection molding machine 1 also includes a nozzle touch device 35 that moves the injection unit 3 in the front-to-rear direction (the left-to-right direction in FIG. 1). When the nozzle touch device 35 is driven, the injection unit 3 moves forward and the nozzle 31a of the heating cylinder 31 touches the contact portion 22b.

[0025] The control device 4 sets parameters that determine molding conditions such as the injection start time, injection speed, injection acceleration, injection peak pressure, injection stroke, VP switching position, dwell time, dwell pressure, cylinder temperature, nozzle temperature, hopper temperature, resin temperature inside the mold, clamping force, and cooling time of the molded product, and controls the operation of the injection device 3 and the clamping device 2. The cooling time is the time required to cool and solidify the molding material inside the mold 21 after it has been filled into the mold 21. The various parameters that determine the operating conditions of the injection molding machine 1 are called operating data.

[0026] 3 is a block diagram showing an example of the configuration of a reinforcement learning device 5 according to the first embodiment. The reinforcement learning device 5 is a computer, and includes, as its hardware configuration, a processing unit 51, a storage unit 52, a display unit 53, an operation unit 54, and a communication unit (transmission unit) 55. The reinforcement learning device 5 may be a server device connected to a network. The reinforcement learning device 5 may also be configured to perform distributed processing using multiple computers, may be realized by multiple virtual machines provided in a single server, or may be realized using a cloud server.

[0027] The processing unit 51 is a processor having an arithmetic circuit such as a CPU (Central Processing Unit), a multi-core CPU, a GPU (Graphics Processing Unit), a GPGPU (General-purpose computing on graphics processing units), a TPU (Tensor Processing Unit), an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programmable Gate Array), or an NPU (Neural Processing Unit), an internal storage device such as a ROM (Read Only Memory) or a RAM (Random Access Memory), an I / O terminal, etc. The processing unit 51 functions as the reinforcement learning device 5 by executing a computer program (program product) P1 stored in a storage unit 52 described below. Note that each functional unit of the reinforcement learning device 5 may be realized by software, or some or all of them may be realized by hardware.

[0028] The storage unit 52 is a non-volatile memory such as a hard disk, an EEPROM (Electrically Erasable Programmable ROM), a flash memory, etc. The storage unit 52 stores a computer program P1 for causing a computer to execute a process for adjusting the cooling stop timing or the cooling time in injection molding.

[0029] The computer program P1 according to the first embodiment may be recorded on a recording medium M1 in a computer-readable manner. The storage unit 52 stores the computer program P1 read from the recording medium M1 by a reading device. The recording medium M1 is a semiconductor memory such as a flash memory. The recording medium M1 may also be an optical disc such as a CD (Compact Disc)-ROM, a DVD (Digital Versatile Disc)-ROM, or a BD (Blu-ray (registered trademark) Disc). The recording medium M1 may also be a magnetic disc such as a flexible disk or a hard disk, a magneto-optical disc, or the like. Furthermore, the computer program P1 according to the first embodiment may be downloaded from an external server connected to a communication network and stored in the storage unit 52.

[0030] The display unit 53 is, for example, a liquid crystal panel, an organic EL display, electronic paper, a plasma display, etc. The display unit 53 displays various information according to the image data provided from the processing unit 51.

[0031] The operation unit 54 has input devices such as a touch panel, soft keys, hard keys, a keyboard, and a mouse.

[0032] The communication unit 55 is a communication circuit that transmits and receives information in accordance with a predetermined communication protocol. The communication unit 55 is connected to the control device 4, the temperature measurement amplifier 61, the pressure measurement amplifier 71, and the molded product inspection device 8 via a communication network, and the processing unit 51 transmits and receives various data between the control device 4, the temperature measurement amplifier 61, the pressure measurement amplifier 71, and the molded product inspection device 8 via the communication unit 55. Specifically, the communication unit 55 receives cooling time data transmitted from the control device 4, temperature data transmitted from the temperature measurement amplifier 61, pressure data transmitted from the pressure measurement amplifier 71, and inspection data transmitted from the molded product inspection device 8. The temperature data is data indicating the surface temperature of the resin cavity, the pressure data is data indicating the pressure inside the mold, and the inspection data is data indicating the appearance or shape of the molded product.

[0033] Fig. 4 is a functional block diagram of the injection molding machine 1 and the reinforcement learning device 5 according to the first embodiment. As shown in Fig. 4, the reinforcement learning device 5 has, as functional units, an observation unit 5a, an action selector 5b, a classifier 5c, an evaluator 5d, and a learner 5e. The action selector 5b, the evaluator 5d, and the learner 5e constitute an agent of the reinforcement learning device 5. Note that each functional unit of the reinforcement learning device 5 may be realized by software, or some or all of them may be realized by hardware.

[0034] The observation unit 5a observes the state of the injection molding machine 1 and the molded product by receiving cooling time data, temperature data, pressure data, and inspection data transmitted from the control device 4, the temperature measurement amplifier 61, the pressure measurement amplifier 71, and the molded product inspection device 8. In particular, the temperature data and pressure data are transmitted sequentially from the temperature measurement amplifier 61 and the pressure measurement amplifier 71, and the observation unit 5a outputs each piece of temperature data and pressure data to the action selector 5b each time it receives them. Note that, since the cooling time is adjusted by the reinforcement learning device 5 and set in the control device 4, reinforcement learning may be performed using the cooling time data held by the reinforcement learning device 5.

[0035] The observation unit 5a outputs the temperature data and pressure data to the action selector 5b, the cooling time data to the evaluator 5d, and the inspection data to the classifier 5c. The observation unit 5a also outputs data necessary for reinforcement learning of the action selector 5b, such as the temperature data, pressure data, and cooling time data, to the learning unit 5e.

[0036] The action selector 5b is, for example, a reinforcement learning model having a deep neural network such as DQN, A3C, or D4PG, or a model-based reinforcement learning model such as PlaNet or SLAC.

[0037] In the case of a reinforcement learning model having a deep neural network, the action selector 5b is equipped with a DQN (Deep Q-Network) and determines an action a corresponding to the state s of the injection molding machine 1 indicated by the observation data, based on the state s. Here, reinforcement learning using an action-value function will be mainly described, but reinforcement learning may also be configured to be performed using other known methods, such as reinforcement learning using a policy gradient method.

[0038] The state s includes, for example, the acquired time-series temperature data and pressure data. The state s may be, for example, image data of a graph that represents the acquired temperature data and pressure data as time-series changes in temperature and pressure, or may be numerical data in which the temperature and pressure values ​​are arranged in chronological order.

[0039] DQN is a neural network model that outputs the value of each of multiple actions a when a state s indicated by observed data is input. The multiple actions a include, for example, stopping cooling in an injection molding process and continuing cooling. An action a with a high value indicates the appropriate timing to stop cooling.

[0040] When the action selector 5b selects to stop cooling, it transmits a cooling stop signal to the control device 4 of the injection molding machine 1. The control device 4 receives the cooling stop signal transmitted from the reinforcement learning device 5 and stops cooling the molded product. Specifically, the control device 4 opens the mold 21 and drives the ejector pins 28f to remove the molded product from the mold 21.

[0041] The molded product removed from the mold 21 is inspected by a molded product inspection device 8. For example, the molded product inspection device 8 captures an image of the molded product and transmits the image data obtained by capturing the image to the reinforcement learning device 5 as inspection data. If the cooling time is insufficient, deformation will occur in the molded product.

[0042] The classifier 5c classifies the molded product as either a good product or a defective product based on image data obtained by capturing an image of the molded product's appearance, and outputs pass / fail data indicating the classification result to the evaluator 5d. For example, if the molded product has appearance defects including deformation and cracks, the molded product is judged to be defective.

[0043] The evaluator 5d calculates reward data based on the pass / fail data and the cooling time, and outputs the calculated reward data to the learning device 5e. The shorter the cooling time, the larger the reward value indicated by the reward data. However, if the molded product is defective, the reward value indicated by the reward data is set to zero. More specifically, the reward value is expressed by the following formula (1).

[0044] r = (100 - t) × n … (1) however, r: reward value t: cooling time n: Inspection result (n=1 if the molded product is good, n=0 if it is defective)

[0045] The learning unit 5e is a functional unit that performs reinforcement learning on the action selector 5b. The learning unit 5e receives the reward calculated by the evaluator 5d and trains the action selector 5b so as to maximize the profit, i.e., the accumulation of the reward.

[0046] More specifically, DQN has an input layer, a hidden layer, and an output layer. The input layer has a plurality of nodes to which a state s, i.e., observed data, is input. The output layer has a plurality of nodes that correspond to a plurality of actions a, respectively, and output the value Q(s, a) of the action a in the input state s.

[0047] Based on the state s, action a, and the reward r obtained from the action, the value Q expressed by the following formula (2) is used as training data, and various weighting coefficients that characterize the DQN can be adjusted to enable reinforcement learning of the DQN of the action selector 5b. Q(s,a)←Q(s,a)+α(r+γmaxQ(snext,anext)-Q(s,a))…(2) however, s:Status a:Action α: learning coefficient r:Reward γ: discount rate maxQ(snext,anext): The maximum Q value for the next possible action

[0048] <Reinforcement learning processing> 5 is a flowchart showing the processing procedure of reinforcement learning according to embodiment 1. The processing unit 51 of the reinforcement learning device 5 initially sets a cooling time in the control device 4 (step S111). The control device 4 operates based on the initially set cooling time.

[0049] When the cooling step of injection molding begins, the processing unit 51 receives temperature data sequentially transmitted from the temperature measurement amplifier 61 (step S112). The processing unit 51 also receives pressure data sequentially transmitted from the pressure measurement amplifier 71 (step S113). The processing unit 51 that executes the processes of steps S112 and S113 functions as an acquisition unit that acquires temperature data indicating the cavity surface temperature of the mold 21 and pressure data indicating the internal mold pressure during the cooling step of the molded product.

[0050] The processing unit 51 executes an action selection process to determine whether or not to stop cooling of the molded product based on the received time-series temperature data and pressure data (step S114), and determines whether or not to stop cooling (step S115). If it determines that cooling should be continued (step S115: NO), the processing unit 51 returns the process to step S112 and continues monitoring the temperature and pressure inside the mold 21. The processing unit 51 may be configured to execute the processes of steps S112 to S115 after waiting for the minimum required cooling time to elapse, or may be configured to execute the processes of steps S112 to S115 from a time point a predetermined time before the initially set cooling time.

[0051] When it is determined that the cooling should be stopped (step S115: YES), the processing unit 51 transmits a cooling stop signal to the control device 4 of the injection molding machine 1 (step S116). Upon receiving the cooling stop signal, the control device 4 stops the cooling of the molded product.

[0052] Next, the processing unit 51 receives the inspection data of the molded product obtained after stopping the cooling in the process of step S116 (step S117) and judges whether the molded product is good or bad (step S118). For example, the processing unit 51 judges whether the molded product is good or bad by detecting the presence or absence of deformation, cracks, etc. in the molded product based on image data obtained by capturing an image of the molded product. The processing unit 51 may also be configured to judge whether the molded product is good or bad using an image recognition model that has been trained on the appearances of good and bad molded products. Then, the processing unit 51 causes the behavior selector 5b to perform reinforcement learning based on the temperature data, pressure data, and reward data obtained by the evaluator 5d (step S119).

[0053] Next, the processing unit 51 determines whether learning of the action selector 5b has been completed (step S120). For example, it determines whether reinforcement learning has been performed a predetermined number of times. If it is determined that learning has not been completed (step S120: NO), the processing unit 51 returns the process to step S112 and continues reinforcement learning of the action selector 5b. If it is determined that learning has been completed (step S120: YES), the processing unit 51 ends the reinforcement learning process.

[0054] <Cooling stop timing adjustment process> 6 is a flowchart showing the procedure of the cooling time adjustment process according to embodiment 1. The processing unit 51 of the reinforcement learning device 5 initially sets the cooling time in the control device 4 (step S131). The control device 4 operates based on the initially set cooling time.

[0055] When the cooling step of injection molding begins, the processing unit 51 receives temperature data sequentially transmitted from the temperature measurement amplifier 61 (step S132). The processing unit 51 also receives pressure data sequentially transmitted from the pressure measurement amplifier 71 (step S133). Based on the received time-series temperature data and pressure data, the processing unit 51 executes an action selection process to determine whether or not cooling of the molded product should be stopped (step S134), and determines whether or not cooling should be stopped (step S135). If it is determined that cooling should be continued (step S135: NO), the processing unit 51 returns the process to step S132 and continues monitoring the temperature and pressure inside the mold 21.

[0056] When it is determined that the cooling should be stopped (step S135: YES), the processing unit 51 transmits a cooling stop signal to the control device 4 of the injection molding machine 1 (step S136). Upon receiving the cooling stop signal, the control device 4 stops the cooling of the molded product.

[0057] With the injection molding machine 1 according to the first embodiment configured as described above, the reinforcement learning device 5 can perform reinforcement learning on the relationship between the time series changes in the resin cavity surface temperature and in-mold pressure and the cooling time that will prevent molding defects, and can automatically optimize the cooling time in injection molding. In other words, the cooling time can be adjusted so that it is as short as possible within a range that will prevent molded product defects.

[0058] Furthermore, a resin temperature sensor 6 and an in-mold pressure sensor 7 are provided in at least one of the areas of the mold 21 that correspond to the thick wall portions, corner portions, edge portions, ribs, and bosses of the molded product, and by using reinforcement learning to learn the relationship between the time series changes in the cavity surface temperature and in-mold pressure in those areas and the cooling time that does not cause molding defects, it is possible to adjust the cooling time to a more appropriate value. In addition, resin temperature sensors 6 and mold pressure sensors 7 are installed in areas of the mold 21 that correspond to multiple areas where the molded product has different thicknesses, and by using reinforcement learning to learn the relationship between the time-series changes in the cavity surface temperature and mold pressure at those areas and the cooling time that does not cause molding defects, it is possible to adjust the cooling time to a more appropriate value. In other words, it is possible to prevent defects from occurring in parts of molded products due to insufficient cooling. For example, even for molded products with complex shapes that have thick and thin parts, corners, edges, ribs, bosses, etc., the cooling time can be adjusted to prevent molding defects.

[0059] In the first embodiment, an example has been described in which the reinforcement learning device 5 is a device separate from the control device 4, but the reinforcement learning device 5 may be incorporated into the control device 4. The reinforcement learning device 5 may also be constructed as a cloud server. Furthermore, the cooling time adjustment process using the action selector 5b and the reinforcement learning process of the action selector 5b may be configured to be executed by different computers.

[0060] In the first embodiment, an example has been described in which the cooling time is adjusted by monitoring both the resin temperature and the pressure inside the mold, but the cooling time may be adjusted by monitoring the resin temperature without using the pressure inside the mold.

[0061] (Embodiment 2) The injection molding machine 1 according to the second embodiment differs from the first embodiment in that it includes an adjuster 5f that adjusts the cooling stop timing selected by the action selector 5b. Since the other configurations of the injection molding machine 1 are the same as those of the injection molding machine 1 according to the first embodiment, the same reference numerals are used for the same parts and detailed description will be omitted.

[0062] FIG. 7 is a functional block diagram of an injection molding machine 1 and a reinforcement learning device 5 according to a second embodiment. The reinforcement learning device 5 according to the second embodiment further includes an adjuster 5f in addition to the same functional units as those in the first embodiment. The adjuster 5f adjusts the cooling timing selected by the action selector 5b. Specifically, the adjuster 5f performs a process of delaying the cooling timing by a predetermined time. In other words, the adjuster 5f extends the cooling time by the predetermined time. The cooling timing determined by the action selector 5b is not necessarily optimal; if the cooling timing is too short, defects such as deformation of the molded product may occur. On the other hand, a slightly longer cooling time does not cause defects in the molded product. For this reason, the reinforcement learning device 5 according to the second embodiment adjusts the cooling time selected by the action selector 5b so that it is extended by the predetermined time.

[0063] 8 is a flowchart showing the procedure of the cooling time adjustment process according to the second embodiment. The processing unit 51 of the reinforcement learning device 5 executes the same processes as steps S131 to S135 in the first embodiment to determine the timing of cooling (steps S231 to S235). When the processing unit 51 determines that it is time to stop cooling (step S235: YES), it delays the cooling stop timing determined by the action selector 5b by adding a predetermined time (step S236).

[0064] Then, the processing unit 51 transmits a cooling stop signal to the control device 4 of the injection molding machine 1 at the cooling timing delayed by adding the predetermined time (step S237), and ends the processing.

[0065] According to the injection molding machine 1 of the second embodiment configured in this manner, as in the first embodiment, the cooling time in injection molding can be automatically optimized, and further, by providing a predetermined time margin in the cooling time, the cooling time can be shortened while reliably maintaining the quality of the molded product.

[0066] The adjuster 5f according to the second embodiment may be provided in the reinforcement learning device 5 according to the third and fourth embodiments.

[0067] (Embodiment 3) The injection molding machine 1 according to the third embodiment differs from the first embodiment in that it is equipped with a plurality of behavior selectors 5b that have undergone reinforcement learning for each type of mold 21 and resin characteristics of the molding material. Since the other configurations of the injection molding machine 1 are the same as those of the injection molding machine 1 according to the first embodiment, the same reference numerals are used for the same parts and detailed description will be omitted.

[0068] FIG. 9 is a functional block diagram of an injection molding machine 1 and a reinforcement learning device 5 according to a third embodiment. The reinforcement learning device 5 according to the third embodiment includes multiple behavior selectors 5b. Molding conditions vary depending on the type of mold 21 and the molding material, and temperature and pressure changes in the molded product during the cooling process also vary. Therefore, in the third embodiment, reinforcement learning is performed on each behavior selector 5b according to the type of mold 21 and the resin characteristics of the molding material. For example, if there are three types of molds 21 and three different resin characteristics of the molding materials, the reinforcement learning device 5 includes 3 × 3 = 9 behavior selectors 5b. Specifically, the reinforcement learning device 5 associates various parameters characterizing the multiple behavior selectors 5b with data indicating the type of mold 21 and resin characteristic data and stores them in the storage unit 52. The reinforcement learning method for the multiple behavior selectors 5b is the same as that in the first embodiment.

[0069] In the third embodiment, the control device 4 of the injection molding machine 1 transmits the mold type data and the resin characteristic data to the reinforcement learning device 5. The processing unit 51 of the reinforcement learning device 5 receives the mold type data and the resin characteristic data transmitted from the control device 4.

[0070] 10 is a flowchart showing the procedure for adjusting the cooling time according to the third embodiment. The processing unit 51 of the reinforcement learning device 5 initially sets the cooling time in the control device 4 (step S331). Next, the processing unit 51 receives mold type data indicating the type of the mold 21 and resin characteristic data indicating the resin characteristics of the molding material from the control device 4 (step S332). Then, the processing unit 51 selects the behavior selector 5b corresponding to the received mold type data and resin identification data (step S333).

[0071] Thereafter, the processing unit 51 uses the selected action selector 5b to perform processing similar to steps S132 to S136 in embodiment 1 to determine the timing to stop cooling, and transmits a cooling stop signal to the control device 4 of the injection molding machine 1 (steps S334 to S338).

[0072] According to the injection molding machine 1 of embodiment 3 configured in this manner, by using a behavior selector 5b according to the type of mold 21 and the resin properties of the molding material, it is possible to determine the optimal timing to stop cooling according to the type of mold 21 and the molding material.

[0073] (Embodiment 4) The injection molding machine 1 according to the fourth embodiment differs from the first and third embodiments in that it uses reinforcement learning to determine the optimal timing to stop cooling, taking into account the type of mold 21, the resin properties of the molding material, and the operating data of the injection molding machine 1 as the environment. The other configurations of the injection molding machine 1 are the same as those of the injection molding machines 1 according to the first and third embodiments, and therefore the same reference numerals are used for the same parts, and detailed descriptions thereof will be omitted.

[0074] FIG. 11 is a functional block diagram of an injection molding machine 1 and a reinforcement learning device 5 according to a fourth embodiment. In the fourth embodiment, the control device 4 of the injection molding machine 1 transmits cooling time data, mold type data, resin property data, and operation data to the reinforcement learning device 5. The processing unit 51 of the reinforcement learning device 5 receives the cooling time data, mold type data, resin property data, and operation data transmitted from the control device 4. The observation unit 5a, which serves as a functional unit, outputs the temperature data, pressure data, mold type data, resin property data, and operation data acquired from the control device 4, the temperature measurement amplifier 61, and the pressure measurement amplifier 71 to the action selector 5b as data representing the environment. The action selector 5b determines whether to stop cooling in the injection molding process based on the temperature data, pressure data, mold type data, resin property data, and operation data.

[0075] In addition, in the reinforcement learning process, the observation unit 5a outputs the temperature data, pressure data, mold type data, resin property data, and operation data acquired from the control device 4, the temperature measuring amplifier 61, and the pressure measuring amplifier 71 to the learning device 5e as data representing the environment. The learning device 5e uses the temperature data, pressure data, mold type data, resin property data, and operation data as data representing the state s to cause the action selector 5b to perform reinforcement learning. The reinforcement learning method is the same as in the first embodiment.

[0076] 12 is a flowchart showing the procedure for adjusting the cooling time according to embodiment 4. The processing unit 51 of the reinforcement learning device 5 initially sets the cooling time in the control device 4 (step S431). Next, the processing unit 51 receives mold type data indicating the type of the mold 21, resin property data indicating the resin property of the molding material, and operation data from the control device 4 (step S432).

[0077] The processing unit 51 then receives the temperature data sequentially transmitted from the temperature measuring amplifier 61 (step S433), and receives the pressure data sequentially transmitted from the pressure measuring amplifier 71 (step S434). Based on the received time-series temperature and pressure data, mold type data, resin property data, and operation data, the processing unit 51 executes an action selection process to determine whether or not to stop cooling of the molded product (step S435). Thereafter, the processing unit 51 executes the same processes as steps S135 and S136 in the first embodiment, and transmits a cooling stop signal to the control device 4 (steps S436 to S437).

[0078] With the injection molding machine 1 of embodiment 4 configured in this manner, the optimal timing to stop cooling can be determined based on the changes in temperature and pressure within the mold 21 over time, taking into account the type of mold 21, the resin properties of the molding material, and operating data.

[0079] In addition, the optimum timing for stopping cooling can be determined based on the time-dependent changes in temperature and pressure inside the mold 21, taking into consideration operating conditions such as injection speed, VP switching position, dwell time, and dwell pressure.

[0080] The means for solving the problems of the present disclosure are described below. (Appendix 1) an acquisition unit that acquires temperature data indicating at least the cavity surface temperature of the mold during the cooling process of the molded product; an action selector that selects a timing to stop cooling in injection molding based on the temperature data acquired by the acquisition unit; an evaluator that calculates reward data based on quality data indicating the quality of the molded product obtained by completing the cooling process at the cooling stop timing selected by the action selector and the cooling time of the molded product; a learning device that causes the behavior selector to perform reinforcement learning based on temperature data of the cavity surface temperature and calculated reward data; A reinforcement learning device comprising: (Appendix 2) the acquisition unit sequentially acquires temperature data during the cooling process, The action selector selects the timing to stop cooling in the injection molding based on the time-series temperature data acquired sequentially. 2. The reinforcement learning apparatus of claim 1. (Appendix 3) The acquisition unit further acquires pressure data indicating an internal pressure in the cooling step, the learning unit causes the action selector to undergo reinforcement learning based on the temperature data and pressure data acquired by the acquisition unit and the calculated reward data; The action selector selects a timing to stop cooling in the injection molding based on the temperature data and pressure data acquired by the acquisition unit. 3. The reinforcement learning device according to claim 1 or 2. (Appendix 4) the acquisition unit sequentially acquires temperature data and pressure data during the cooling process, The action selector selects the timing to stop cooling in the injection molding based on the time-series temperature data and pressure data acquired sequentially. 4. The reinforcement learning apparatus of claim 3. (Appendix 5) The evaluator If there are appearance defects including deformation and cracks in the molded product, the reward data is calculated so that the reward is low, and the shorter the cooling time, the higher the reward. 5. The reinforcement learning device according to claim 1. (Appendix 6) an adjuster that adds a predetermined time to the cooling stop timing selected by the action selector; 6. The reinforcement learning device according to claim 1. (Appendix 7) The acquisition unit Obtain temperature data indicating the cavity surface temperature of at least one portion corresponding to the thickness portion, corner portion, edge portion, rib, and boss of the molded product. 7. The reinforcement learning device according to claim 1. (Appendix 8) The acquisition unit Obtain temperature data showing the cavity surface temperatures at multiple locations corresponding to the molded part's wall thickness, corners, edges, ribs, and bosses. 8. The reinforcement learning device according to claim 1. (Appendix 9) Obtain temperature data showing the cavity surface temperatures at multiple locations with different part thicknesses 9. The reinforcement learning device according to any one of Supplementary Note 1 to Supplementary Note 8. (Appendix 10) The system is provided with a plurality of the behavior selectors that have undergone reinforcement learning for each type of mold and resin characteristics of the molding material. 10. The reinforcement learning device according to any one of Supplementary Note 1 to Supplementary Note 9. (Appendix 11) the learning device performs reinforcement learning on the behavior selector based on temperature data of the cavity surface temperature, the type of mold, resin properties of the molding material, and calculated reward data; The action selector selects the timing to stop cooling in the injection molding based on the temperature data acquired by the acquisition unit, the type of mold, and the resin properties of the molding material. 11. The reinforcement learning device according to claim 1. (Appendix 12) the learning device performs reinforcement learning on the behavior selector based on temperature data of the cavity surface temperature, operation data including a holding pressure and an injection speed set in the injection molding machine, and calculated reward data; The action selector selects a timing to stop cooling in the injection molding based on the temperature data acquired by the acquisition unit and the operation data. 12. The reinforcement learning device according to claim 1. (Appendix 13) a transmitting unit that outputs a cooling stop signal to the injection molding machine at the cooling stop timing selected by the action selector; 13. The reinforcement learning device according to any one of claims 1 to 12. (Appendix 14) Equipped with a classifier that classifies the quality of molded products based on data showing the appearance of the molded products, including deformation and cracks. 14. The reinforcement learning device according to any one of claims 1 to 13. (Appendix 15) The acquisition unit Temperature data is obtained from a temperature sensor that detects the cavity surface temperature by detecting infrared rays emitted from the resin inside the mold. 15. The reinforcement learning device according to any one of claims 1 to 14. (Appendix 16) The acquisition unit Obtain pressure data from a pressure sensor that detects the pressure from the resin inside the mold 16. The reinforcement learning device according to any one of claims 1 to 15. (Appendix 17) The behavior selector: During the selection process, it is selected whether or not to terminate the cooling process. 17. The reinforcement learning device according to any one of claims 1 to 16. [Explanation of symbols]

[0081] 1: Injection molding machine 2: Mold clamping device 3: Injection device 4: Control device 5: Reinforcement learning device 5a: Observation section 5b: Action selector 5c:Classifier 5d: Evaluator 5e: Learner 6: Resin temperature sensor 7: Mold pressure sensor 8: Molded product inspection equipment 21: Mold 27d: Cavity 51: Processing section 52: Storage section 53:Display section 54:Operation unit 55: Communications Department 61: Temperature measurement amplifier 71: Pressure measurement amplifier P1: Computer Program M1: Recording medium

Claims

1. an acquisition unit that acquires temperature data indicating at least the cavity surface temperature of the mold during the cooling process of the molded product; an action selector that selects a timing to stop cooling in injection molding based on the temperature data acquired by the acquisition unit; an evaluator that calculates reward data based on quality data indicating the quality of the molded product obtained by completing the cooling process at the cooling stop timing selected by the action selector and the cooling time of the molded product; a learning device that causes the behavior selector to perform reinforcement learning based on temperature data of the cavity surface temperature and calculated reward data; A reinforcement learning device comprising:

2. the acquisition unit sequentially acquires temperature data during the cooling process, The action selector selects the timing to stop cooling in the injection molding based on the time-series temperature data acquired sequentially. The reinforcement learning device according to claim 1 .

3. The acquisition unit further acquires pressure data indicating an internal pressure in the cooling step, the learning unit causes the action selector to undergo reinforcement learning based on the temperature data and pressure data acquired by the acquisition unit and the calculated reward data; The action selector selects a timing to stop cooling in the injection molding based on the temperature data and pressure data acquired by the acquisition unit. The reinforcement learning device according to claim 1 .

4. the acquisition unit sequentially acquires temperature data and pressure data during the cooling process, The action selector selects the timing to stop cooling in the injection molding based on the time-series temperature data and pressure data acquired sequentially. The reinforcement learning device according to claim 3 .

5. The evaluator If there are appearance defects including deformation and cracks in the molded product, the reward data is calculated so that the reward is low, and the shorter the cooling time, the higher the reward. The reinforcement learning device according to claim 1 .

6. an adjuster that adds a predetermined time to the cooling stop timing selected by the action selector; The reinforcement learning device according to claim 1 .

7. The acquisition unit Obtain temperature data indicating the cavity surface temperature of at least one portion corresponding to the thickness portion, corner portion, edge portion, rib, and boss of the molded product. The reinforcement learning device according to claim 1 .

8. The acquisition unit Obtain temperature data showing the cavity surface temperatures at multiple locations corresponding to the molded part's wall thickness, corners, edges, ribs, and bosses. The reinforcement learning device according to claim 1 .

9. Obtain temperature data showing the cavity surface temperatures at multiple locations with different part thicknesses The reinforcement learning device according to claim 1 .

10. The system is provided with a plurality of the behavior selectors that have undergone reinforcement learning for each type of mold and resin characteristics of the molding material. The reinforcement learning device according to claim 1 .

11. the learning device performs reinforcement learning on the behavior selector based on temperature data of the cavity surface temperature, the type of mold, resin properties of the molding material, and calculated reward data; The action selector selects the timing to stop cooling in the injection molding based on the temperature data acquired by the acquisition unit, the type of mold, and the resin properties of the molding material. The reinforcement learning device according to claim 1 .

12. the learning device performs reinforcement learning on the behavior selector based on temperature data of the cavity surface temperature, operation data including a holding pressure and an injection speed set in the injection molding machine, and calculated reward data; The action selector selects a timing to stop cooling in the injection molding based on the temperature data acquired by the acquisition unit and the operation data. The reinforcement learning device according to claim 1 .

13. a transmitting unit that outputs a cooling stop signal to the injection molding machine at the cooling stop timing selected by the action selector; The reinforcement learning device according to claim 1 .

14. Equipped with a classifier that classifies the quality of molded products based on data showing the appearance of the molded products, including deformation and cracks. The reinforcement learning device according to claim 1 .

15. The acquisition unit Temperature data is obtained from a temperature sensor that detects the cavity surface temperature by detecting infrared rays emitted from the resin inside the mold. The reinforcement learning device according to claim 1 .

16. The acquisition unit Obtain pressure data from a pressure sensor that detects the pressure from the resin inside the mold The reinforcement learning device according to claim 1 .

17. The behavior selector: During the selection process, it is selected whether or not to terminate the cooling process. The reinforcement learning device according to claim 1 .

18. an acquisition unit that acquires temperature data indicating at least the cavity surface temperature of the mold during the cooling process of the molded product; an action selector that selects a timing to stop cooling in injection molding based on the temperature data acquired by the acquisition unit; an evaluator that calculates reward data based on quality data indicating the quality of the molded product obtained by completing the cooling process at the cooling stop timing selected by the action selector and the cooling time of the molded product; a learning device that causes the behavior selector to perform reinforcement learning based on temperature data of the cavity surface temperature and calculated reward data; An injection molding machine comprising:

19. Obtaining temperature data indicating at least the cavity surface temperature of the mold during the cooling process of the molded product; Based on the acquired temperature data, the timing to stop cooling during injection molding is selected. Calculate the reward data based on the quality data indicating the quality of the molded product obtained by completing the cooling process at the selected cooling stop timing and the cooling time of the molded product, Based on the temperature data of the cavity surface temperature and the calculated reward data, the timing to stop cooling in the injection molding is learned by reinforcement learning. Reinforcement learning methods.

20. Obtaining temperature data indicating at least the cavity surface temperature of the mold during the cooling process of the molded product; Based on the acquired temperature data, the timing to stop cooling during injection molding is selected. Calculate the reward data based on the quality data indicating the quality of the molded product obtained by completing the cooling process at the selected cooling stop timing and the cooling time of the molded product, Based on the temperature data of the cavity surface temperature and the calculated reward data, the timing to stop cooling in the injection molding is learned by reinforcement learning. A computer program that causes a computer to execute a process.

Citation Information

Patent Citations

  • Injection molding machine system that adjusts molding conditions by machine learning device

    JP2019166702A