Continuous mixing apparatus and control method thereof

By employing reinforcement learning algorithms in continuous mixing equipment and updating control conditions based on rewards, the problems of large control errors and long parameter adjustment times in existing technologies are solved, achieving more efficient temperature control.

CN114379038BActive Publication Date: 2026-03-24THE JAPAN STEEL WORKS LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-19
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Existing continuous mixing equipment based on measurement-based temperature feedback control suffers from problems such as large control errors, long parameter adjustment times, and material waste in multiple annular heaters.

Method used

By employing a reinforcement learning algorithm, the control error is calculated through a state observation unit. Based on the reward, the control conditions are updated, and the optimal action is selected to control the annular heater, thereby reducing parameter adjustment time and material waste.

Benefits of technology

This reduces parameter adjustment time and material consumption when process conditions change, thereby improving control accuracy and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114379038B_ABST
    Figure CN114379038B_ABST
Patent Text Reader

Abstract

The present invention relates to a continuous mixing apparatus and a control method thereof, and in the continuous mixing apparatus according to the embodiment, for each of a plurality of annular heaters, a control unit determines a current state and a reward for a past selected action based on a control error calculated from an acquired temperature; updates a control condition based on the reward, and determines an optimal action corresponding to the current state under the updated control condition, the control condition being a combination of a state and an action; and controls a target annular heater based on the optimal action.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to a continuous kneading apparatus and a control method thereof. BACKGROUND

[0002] Injection molding apparatuses and extrusion molding apparatuses for resins are equipped with continuous kneading apparatuses that heat pellets by using a heater and knead resin pellets charged into a cylinder by using a screw. For example, Japanese Unexamined Patent Application Publication No. 2009-172822 discloses an injection molding apparatus equipped with a continuous kneading apparatus that performs feedback control of a heater based on a measured temperature. SUMMARY

[0003] The present inventors have found various problems in developing a continuous kneading apparatus that performs feedback control of a heater based on a measured temperature.

[0004] Other problems and novel features will be clarified from the description of this specification and the attached drawings.

[0005] In the continuous kneading apparatus according to the embodiment, for each of the plurality of annular heaters, the control unit determines a current state and a reward for a past selected action based on a control error calculated from a measured temperature; updates a control condition based on the reward, and determines an optimal action corresponding to the current state under the updated control condition, the control condition being a combination of a state and an action; and controls the target annular heater based on the optimal action.

[0006] According to the above-described embodiment, an excellent continuous kneading apparatus can be provided.

[0007] The above and other objects, features and advantages of the present disclosure will be more fully understood from the following detailed description taken in conjunction with the accompanying drawings, which are to be considered illustrative only, and thus are not to be considered as restricting the present disclosure. BRIEF DESCRIPTION OF DRAWINGS

[0008] Figure 1 is a schematic cross-sectional view showing a configuration of a continuous kneading apparatus and an injection molding apparatus including the same according to a first embodiment;

[0009] Figure 2 is a schematic cross-sectional view showing a configuration of a continuous kneading apparatus and an injection molding apparatus including the same according to a first embodiment;

[0010] Figure 3 is a schematic cross-sectional view showing a configuration of a continuous kneading apparatus and an injection molding apparatus including the same according to a first embodiment;

[0011] is a schematic cross-sectional view showing a configuration of a continuous kneading apparatus and an injection molding apparatus including the same according to a first embodiment;Figure 4 is a block diagram showing a configuration of the control unit 70 according to the first embodiment;

[0012] Figure 5 is a flowchart showing a method for controlling the continuous kneading apparatus according to the first embodiment; and

[0013] Figure 6 is a block diagram showing a configuration of the control unit 70 according to the second embodiment. DETAILED DESCRIPTION

[0014] Hereinafter, specific embodiments will be described in detail with reference to the accompanying drawings. However, the present disclosure is not limited to the embodiments shown below. In addition, the following description and drawings are appropriately simplified in order to make the explanation clear.

[0015] (First Embodiment)

[0016] <Configuration of Continuous Kneading Apparatus>

[0017] First, the configuration of the continuous kneading apparatus and the injection molding apparatus including the same according to the first embodiment will be described with reference to Figures 1 to 3 Figures 1 to 3 is a schematic cross-sectional view showing the configuration of the continuous kneading apparatus and the injection molding apparatus including the same according to the first embodiment.

[0018] Note that, needless to say, Figures 1 to 3 the right-hand xyz orthogonal coordinates shown in, are shown for the convenience of explanation of positional relationships between members. In general, in all the drawings, the z-axis positive direction is the vertically upward direction, and the xy plane is the horizontal plane.

[0019] As shown in Figures 1 to 3 , the continuous kneading apparatus 10 according to the first embodiment includes a barrel 11, a screw 12, a hopper 13, a ring heater 14, a temperature sensor 60, and a control unit 70. In addition to the continuous kneading apparatus 10, the injection molding apparatus includes a stationary mold 21 and a movable mold 22.

[0020] Figure 1 The injection molding apparatus in a state before the molten resin 82 is about to be injected into the cavity C formed by the molds (the stationary mold 21 and the movable mold 22) is shown.

[0021] Figure 2 The injection molding apparatus in a state after the molten resin 82 has been injected into the cavity C of the mold is shown.

[0022] Figure 3 The injection molding apparatus in a state when the resin molded product 83 is removed from the mold is shown.

[0023] The cylinder 11 is a cylindrical member extending in the x-axis direction.

[0024] The screw 12 is configured to extend in the x-axis direction and is rotatably housed inside the cylinder 11. Although not shown in the drawing, for example, a motor is connected to the screw 12 with a speed reducer as a rotational drive source. In addition, the screw 12 can be moved in the x-axis direction by an actuator (not shown). As shown in the drawing, when the screw 12 is moved forward in the negative direction of the x-axis, the molten resin 82 is injected into the inside of the mold (the stationary mold 21 and the movable mold 22). Figure 2

[0025] The hopper 13 is a cylindrical member for filling the inside of the cylinder 11 with the resin pellets 81, which are raw materials for the resin molded article 83 shown in the drawing. The hopper 13 is disposed on the upper side of the end portion on the positive side of the x-axis direction of the cylinder 11. Figure 3

[0026] The annular heaters 14 are arranged so as to cover the outer peripheral surface of the cylinder 11 along the longitudinal direction (x-axis direction) of the cylinder 11. In the example shown in the drawing, four annular heaters 14 are provided on the distal side (negative side of the x-axis direction) of the hopper 13. Each of the plurality of annular heaters 14 is individually controlled by the control unit 70. Figures 1 to 3

[0027] Each of the temperature sensors 60 measures the temperature of a portion of the cylinder 11 heated by a corresponding one of the plurality of annular heaters 14. Each of the temperature sensors 60 is, for example, a thermocouple. In the example shown in the drawing, each of the temperature sensors 60 is inserted into a through-hole formed on a corresponding one of the annular heaters 14 and positioned in contact with the cylinder 11. Figures 1 to 3

[0028] The control unit 70 learns the control condition(s) for each of the annular heaters 14 while performing feedback control on a corresponding one of the annular heaters 14 based on the temperature measured by a corresponding one of the temperature sensors 60. More specifically, the control unit 70 controls the output to each of the annular heaters 14 so that the temperature measured by a corresponding one of the temperature sensors 60 approaches a set temperature (target temperature).

[0029] Note that the configuration and operation of the control unit 70 will be described in more detail later.

[0030] ​​​​In the continuous mixing apparatus 10 according to the first embodiment, resin particles 81 supplied from the hopper 13 are mixed by a rotating screw 12 inside the barrel 11 while being heated by an annular heater 14. As the resin particles 81 are heated and pushed (i.e. extruded) from the base of the screw 12 toward its tip (along the negative x-axis direction), they are compressed and transformed into molten resin 82.

[0031] The fixed mold 21 is a mold fixed to the top of the continuous mixing apparatus 10. Meanwhile, the moving mold 22 is a mold driven by a drive source (not shown) and capable of sliding along the x-axis. When the moving mold 22 moves along the positive x-axis and comes into close contact with the fixed mold 21, as... Figure 1 As shown, a cavity C is formed between the fixed mold 21 and the moving mold 22, the shape of which corresponds to the shape of the resin molded article 83 to be manufactured (see Figure 83). Figure 3 The shapes are consistent.

[0032] Next, as Figure 2 As shown, screw 12 moves forward along the negative x-axis and fills cavity C with molten resin 82, so that resin molded article 83 (see Figure 3 It can be molded into shape.

[0033] Then, as Figure 3 As shown, the screw 12 retracts in the positive x-axis direction, and the moving mold 22 moves in the negative x-axis direction and is thus released from the fixed mold 21 (i.e., separated), so that the resin molded article 83 can be removed.

[0034] <Construction of control unit 70 according to the comparative example>

[0035] The continuous mixing equipment according to the comparative example has the same characteristics as that according to... Figures 1 to 3 The overall construction of the continuous mixing apparatus shown in the first embodiment is similar to that of the previous one. In this comparative example, the control unit 70 performs feedback control on a corresponding annular heater in the annular heater 14 using PID control based on the temperature obtained from the various temperature sensors in the temperature sensor 60. In the case of PID control, parameters(s)(s)(s) need to be adjusted each time the process conditions(s) change. Generally, the operator adjusts parameters(s)(s)(s) through trial and error, thus resulting in a significant amount of time and resin material being required for parameter(s)(s).

[0036] <Construction of the control unit 70 according to the first embodiment>

[0037] Next, we will refer to Figure 4 The construction of the control unit 70 according to the first embodiment will be described in more detail. Figure 4 This is a block diagram illustrating the construction of the control unit 70 according to the first embodiment. Figure 4As shown, the control unit 70 according to the first embodiment includes a state observation unit 71, a control condition learning unit 72, a storage unit 73, and a control signal output unit 74.

[0038] Note that each of the functional blocks constituting the control unit 70 can be realized by hardware, such as a CPU (Central Processing Unit), a memory, and other circuits, or by software, such as a program(s) loaded in a memory or the like. Thus, each functional block can be realized in various forms by computer hardware, software, or a combination thereof.

[0039] The state observation unit 71 calculates a control error for a corresponding one of the annular heaters 14 from a measured temperature value pv acquired from each of the temperature sensors 60. The control error is a difference between a target value and the measured value pv. Note that the target value is a target temperature set for each of the annular heaters 14. Meanwhile, the measured value pv is a measured temperature value acquired from the temperature sensor 60 corresponding to the target annular heater 14.

[0040] Then, the state observation unit 71 determines a current state st and a reward rw for the action ac selected in the past (e.g., the last time) for each of the annular heaters 14 based on the calculated control error.

[0041] The states st are defined in advance so as to divide the values of the control error, which can take any of an infinite number of values, into a finite number of groups. As a simple example for explanatory purposes, when the control error is denoted by err, for example, the range "-4.0°C ≤ err < -3.0°C" is defined as a state stl; the range "-3.0°C ≤ err < -2.0°C" is defined as a state st2; the range "-2.0°C ≤ err < -1.0°C" is defined as a state st3; the range "-1.0°C ≤ err < 1.0°C" is defined as a state st4; the range "1.0°C ≤ err < 2.0°C" is defined as a state st5; the range "2.0°C ≤ err < 3.0°C" is defined as a state st6; the range "3.0°C ≤ err < 4.0°C" is defined as a state st7; and the range "4.0°C ≤ err < 5.0°C" is defined as a state st8. In practice, in many cases, more states st with narrower ranges can be defined.

[0042] The reward rw is an index for evaluating the action ac selected in the past state st.

[0043] Specifically, when the absolute value of the calculated current control error is smaller than the absolute value of the past control error, the state observation unit 71 determines that the past selected action ac is appropriate, and sets the reward rw to a positive value, for example. In other words, the reward rw is determined so that the previously selected action ac is more likely to be selected again in the same state st as the past state.

[0044] On the other hand, when the absolute value of the calculated current control error is larger than the absolute value of the past control error, the state observation unit 71 determines that the past selected action ac is inappropriate, and sets the reward rw to a negative value, for example. In other words, the reward rw is determined so that the previously selected action ac is less likely to be selected again in the same state st as the past state.

[0045] Note that specific examples of the reward rw will be described later. Further, the value of the reward rw can be determined as appropriate. For example, the reward rw can always have a positive value, or the reward rw can always have a negative value.

[0046] The control condition learning unit 72 performs reinforcement learning on each of the annular heaters 14. Specifically, the control condition learning unit 72 updates the control condition (learning result) based on the reward rw, and selects the best action ac corresponding to the current state st under the updated control condition. The control condition is a combination of the state st and the action ac. Table 1 shows simple control conditions (learning results) corresponding to the above-described states st1 to st8. In Figure 4 In the example shown, the control condition learning unit 72 stores the updated control condition cc in the storage unit 73, which is a memory, for example, and updates it by reading the control condition cc from the storage unit 73.

[0047] [Table 1]

[0048]

[0049]

[0050] Table 1 shows control conditions (learning results) based on Q-learning, which is an example of reinforcement learning. The above-described eight states st1 to st8 are shown in the top row of Table 1. That is, the eight states st1 to st8 are shown in the second to ninth columns, respectively. Meanwhile, the five actions ac1 to ac5 are shown in the leftmost column of Table 1. That is, the five actions ac1 to ac5 are shown in the second to sixth rows, respectively.

[0051] Note that, in the example shown in Table 1, an action for reducing the output (e.g., voltage) to the target ring heater 14 by 1.0% is defined as action ac1 (output change: -1%). An action for reducing the output (e.g., voltage) to the target ring heater 14 by 0.5% is defined as action ac2 (output change: -0.5%). An action for maintaining the output to the target ring heater 14 is defined as action ac3 (output change: 0%). An action for increasing the output to the target ring heater 14 by 0.5% is defined as action ac4 (output change: +0.5%). An action for increasing the output to the target ring heater 14 by 1.0% is defined as action ac5 (output change: +1.0%). The example shown in Table 1 is a simple example for explanatory purposes. That is, in actuality, in many cases, more detailed actions ac can be defined.

[0052] A value determined by the combination of the state st and the action ac in Table 1 is referred to as a quality Q(st, ac). The quality Q is continuously updated by using a known update formula based on a reward rw after a given initial value. The initial value of the quality Q is included in, for example, a learning condition shown in Table 2. Figure 4 The learning condition is, for example, input by an operator. The initial value of the quality Q can be stored in the storage unit 73, and, for example, a past learning result can be used as the initial value. Further, for example, the states st1 to st8 and the actions ac1 to ac5 shown in Table 1 are included in the learning condition shown in Table 2. Figure 4

[0053] The quality Q will be described by taking the state st7 in Table 1 as an example. In the state st7, since the control error is not lower than 3.0°C and lower than 4.0°C, the heating temperature of the target ring heater 14 is too high. Therefore, it is necessary to reduce the output to the target ring heater 14. Therefore, as a result of learning by the control condition learning unit 72, the qualities Q of the actions ac1 and ac2 for reducing the output to the target ring heater 14 are larger. Meanwhile, the qualities Q of the actions ac4 and ac5 for increasing the output to the target ring heater 14 are smaller.

[0054] For example, in the example shown in Table 1, when the control error is 3.5°C, the state st falls in the state st7. Therefore, the control condition learning unit 72 selects the optimal action ac2 having the largest quality Q in the state st7, and outputs the selected action ac2 to the control signal output unit 74. The control signal output unit 74 reduces the control signal ctr output to the ring heater 14 by 0.5% based on the action ac2 received from the control condition learning unit 72.

[0055] The control signal ctr is, for example, a voltage signal.

[0056] ​Then, if the absolute value of the next control error is smaller than the absolute value of the current control error by 3.5°C, the state observation unit 71 determines that it is appropriate to select the action ac2 in the current state st7, and outputs a reward rw having a positive value. Accordingly, the control condition learning unit 72 updates the control condition in such a manner that the quality of the action ac2 in the state st7 is increased by +3.6 according to the reward rw. As a result, in the case of the state st7, the control condition learning unit 72 continues to select the action ac2.

[0057] On the other hand, if the absolute value of the next control error is larger than the absolute value of the current control error by 3.5°C, the state observation unit 71 determines that it is inappropriate to select the action ac2 in the current state st7, and outputs a reward rw having a negative value. Accordingly, the control condition learning unit 72 updates the control condition in such a manner that the quality of the action ac2 in the state st7 is decreased by +3.6 according to the reward rw. As a result, in the case of the state st7, when the quality of the action ac2 in the state st7 becomes smaller than the quality of the action ac1 +2.6, the control condition learning unit 72 selects the action ac1 instead of the action ac2.

[0058] Note that the timing of the control condition update is not limited to the next time (for example, not limited to the time when the control error is next calculated). That is, the timing of the update can be determined as appropriate taking into account time lag and the like. Furthermore, in the initial stage of learning, in order to speed up learning, the action ac can be randomly selected. Also, although the reinforcement learning by the simple Q-learning is described above with reference to Table 1, there are various types of learning algorithms (such as Q-learning, an AC (Actor-Critic) method, TD learning, and a Monte Carlo method), and the learning algorithm is not limited to any type of algorithm. For example, when the number of states st and actions ac increases and the number of combinations thereof increases explosively, the algorithm can be selected as appropriate, such as using the AC method.

[0059] Furthermore, in the AC method, a probability distribution function is used as the policy function in many cases. The probability distribution function is not limited to a normal distribution function. For example, for the purpose of simplification, a sigmoid function, a soft max function, or the like can be used. The sigmoid function is the most commonly used function in neural networks. Because reinforcement learning is a type of machine learning like neural networks, it can use the sigmoid function. Furthermore, the sigmoid function has another advantage that the function itself is simple and easy to handle.

[0060] As described above, there are various learning algorithms and functions available for use, and the best algorithm and the best function appropriate for the process can be selected.

[0061] As described above, the PID control is not used in the continuous kneading apparatus according to the first embodiment. Therefore, first, the parameter adjustment(s) required when the process conditions change is not required. Further, the control unit 70 updates the control conditions (learning result) based on the reward rw by the reinforcement learning, and selects the optimal action ac corresponding to the current state st under the updated control conditions. Therefore, even when the process condition(s) change, the time required for the adjustment and the amount of the resin material required therefor can be reduced compared to the conditions in the comparative example.

[0062] Note that the application of the continuous kneading apparatus 10 according to the first embodiment is not limited to the injection molding apparatus. That is, the continuous kneading apparatus 10 can also be used in an extrusion molding apparatus. In the case of the extrusion molding apparatus, since the injection operation in the continuous kneading apparatus 10 is unnecessary, the screw 12 does not have to be able to move in the x-axis direction. The continuous kneading apparatus 10 in the injection molding apparatus and the continuous kneading apparatus 10 in the extrusion molding apparatus are generally similar to each other in the remaining configurations.

[0063] <Control method of continuous kneading apparatus>

[0064] Next, the method for controlling the continuous kneading apparatus according to the first embodiment will be described with reference to Figure 5 The method for controlling the continuous kneading apparatus according to the first embodiment will be described in detail. Figure 5 is a flowchart showing the method for controlling the continuous kneading apparatus according to the first embodiment. Appropriate reference will be made to Figure 4 and Figure 5 The following description will be made.

[0065] First, as Figure 5 indicated, Figure 4 The state observation unit 71 of the control unit 70 shown in

[0066] Next, as Figure 5As shown, the control condition learning unit 72 of the control unit 70 updates the control condition, which is a combination of the state st and the action ac, based on the reward rw. Then, the control condition learning unit 72 selects the optimal action ac corresponding to the current state st under the updated control condition (step S2). Note that at the start of the control, the control condition is not updated and remains the initial value, but the optimal action ac corresponding to the state st at the start of the control is selected.

[0067] Then, as shown in FIG. 6, the control signal output unit 74 of the control unit 70 outputs the control signal ctr to the annular heater 14 based on the optimal action ac selected by the control condition learning unit 72 (step S3). Figure 5

[0068] When the manufacturing of the resin molded article 83 has not been completed (step S4: No), the process returns to step S1 and continues the control. On the other hand, when the manufacturing of the resin molded article 83 has been completed (step S4: Yes), the control ends. That is, steps S1 to S3 are repeated until the manufacturing of the resin molded article 83 is completed.

[0069] As described above, in the continuous kneading apparatus 10 according to the first embodiment, the PID control is not used. Therefore, first, the adjustment of the parameter(s) required when the process condition is changed is not required. Further, by the reinforcement learning by the computer, the control condition (learning result) is updated based on the reward rw, and the optimal action ac corresponding to the current state st is selected under the updated control condition. Therefore, even when the process condition(s) is changed, the time required for the adjustment and the amount of the resin material required therefor can be reduced compared to the condition in the comparative example.

[0070] (Second Embodiment)

[0071] Next, a continuous kneading apparatus according to a second embodiment will be described with reference to Figure 6 The overall configuration of the continuous kneading apparatus according to the second embodiment is similar to that of the continuous kneading apparatus according to the first embodiment shown in FIG. 1, and thus the description thereof will be omitted. The configuration of the control unit 70 in the continuous kneading apparatus according to the second embodiment is different from that of the control unit 70 in the continuous kneading apparatus according to the first embodiment. Figures 1 to 3

[0072] Figure 6 is a block diagram showing the configuration of the control unit 70 according to the second embodiment. As shown in Figure 6 The control unit 70 according to the second embodiment includes the state observation unit 71, the control condition learning unit 72, the storage unit 73, and the PID controller 74a. That is, the control unit 70 according to the second embodiment includes the PID controller 74a as the control signal output unit 74 according to the first embodiment. Figure 4 ​​The control signal output unit 74 in the control unit 70 of the first embodiment shown. The PID controller 74a is also an example of a control signal output unit.

[0073] Similar to the first embodiment, the state observation unit 71 determines the current state st and the reward rw for the previously selected action ac for each annular heater 14 based on the calculated control error err. Then, the state observation unit 71 outputs the current state st and the reward rw to the control condition learning unit 72. Furthermore, according to the second embodiment, the state observation unit 71 outputs the calculated control error err to the PID controller 74a.

[0074] Similar to the first embodiment, the control condition learning unit 72 also performs reinforcement learning with respect to each annular heater 14. Specifically, the control condition learning unit 72 updates the control conditions (learning results) based on the reward rw, and selects the optimal action ac corresponding to the current state st under the updated control conditions. Note that in the first embodiment, the output to the annular heater 14 is directly changed according to the content (i.e., details) of the action ac selected by the control condition learning unit 72. In contrast, in the second embodiment, one or more parameters of the PID controller 74a are changed according to the content (e.g., details) of the action ac selected by the control condition learning unit 72.

[0075] like Figure 6 As shown, the parameters of the PID controller 74a are continuously changed based on the action ac output from the control condition learning unit 72. Simultaneously, the PID controller 74a outputs a control signal ctr to the annular heater 14 based on the control error err received from the state observation unit 71. This control signal ctr is, for example, a voltage signal.

[0076] The remaining construction is similar to that of the first embodiment, so its description will be omitted.

[0077] As described above, PID control is used in the continuous mixing apparatus according to the second embodiment, so that when one or more process conditions change, one or more parameters need to be adjusted. In the continuous mixing apparatus according to the second embodiment, the control unit 70 updates the control conditions (learning results) based on the reward rw through reinforcement learning, and selects the optimal action ac corresponding to the current state st under the updated control conditions. Note that the action ac in reinforcement learning is to change the parameters of the PID controller 74a. Therefore, when one or more process conditions change, the time required for parameter adjustment and the amount of resin material required can be reduced compared to the conditions in the comparative example.

[0078] In the above examples, the program includes instructions (or software code) that, when loaded into a computer, cause the computer to perform one or more of the functions described in the present embodiments. The program can be stored in a non-transitory computer-readable medium or a tangible storage medium. By way of example, and not limitation, a non-transitory computer-readable medium or a tangible storage medium can include random access memory (RAM), read-only memory (ROM), flash memory, solid state drive (SSD), or other type of storage technology, a CD-ROM, a digital versatile disc (DVD), a Blu-ray disc, or other type of optical disc storage, and a magnetic cassette, magnetic tape, magnetic disk storage or other type of magnetic storage device. The program can be transmitted on a transitory computer-readable medium or a communication medium. By way of example, and not limitation, a transitory computer-readable medium or a communication medium can include an electrical, optical, acoustical or other form of propagated signals.

[0079] From the disclosure thus far described, it will be evident that embodiments of the present disclosure can vary in many ways. Such variations should not be considered as a departure from the spirit and scope of the present invention, and all such modifications as would be obvious to one skilled in the art are intended to be included within the scope of the following claims.

Claims

1. A continuous mixing apparatus, comprising: cylindrical body; A screw housed within the cylinder; Multiple annular heaters are arranged along the longitudinal direction of the cylinder to cover the outer peripheral surface of the cylinder; Multiple temperature sensors, each of which is configured to measure the temperature of a portion of the cylinder heated by a respective annular heater of the multiple annular heaters; and A control unit configured to perform feedback control on a specific annular heater among the plurality of temperature sensors based on temperatures measured by each of the plurality of temperature sensors. The resin particles filled into the cylinder are heated by the plurality of annular heaters and mixed by the screw. Wherein, for each of the plurality of annular heaters, the control unit The current state and reward for past actions are determined based on the control error calculated from the measured temperature. The control conditions are updated based on the reward using reinforcement learning, and the optimal action corresponding to the current state is selected under the updated control conditions. The control conditions are a combination of state and action. The target annular heater is controlled based on the aforementioned optimal action. In the reinforcement learning, the state represents a predefined range of control error values, and the current state is the state that includes the calculated control error. When the absolute value of the current control error is less than the absolute value of the past control error, the reward is increased, making the previously selected action more likely to be selected again under the same state as the past. When the absolute value of the current control error is greater than the absolute value of the past control error, the reward is reduced so that the previously selected action is less likely to be selected again under the same state as in the past.

2. The continuous mixing equipment according to claim 1, wherein, The action is to change the output to the target annular heater.

3. The continuous mixing equipment according to claim 1, wherein, The action involves changing the parameters of a PID controller configured to control the output to the target annular heater.

4. The continuous mixing equipment according to claim 1, wherein, Each of the multiple temperature sensors is a thermocouple.

5. The continuous mixing equipment according to claim 4, wherein, Each of the thermocouples is inserted into a through hole formed on a corresponding annular heater among the plurality of annular heaters and positioned to contact the cylinder.

6. A control method for controlling a continuous mixing equipment, the continuous mixing equipment comprising: cylindrical body; A screw housed within the cylinder; Multiple annular heaters are arranged along the longitudinal direction of the cylinder to cover the outer peripheral surface of the cylinder; Multiple temperature sensors, each of which is configured to measure the temperature of a portion of the cylinder heated by a respective annular heater of the multiple annular heaters; and A control unit configured to perform feedback control on a specific annular heater among the plurality of temperature sensors based on the temperature measured by each of the plurality of temperature sensors. The resin particles filled into the cylinder are heated by the plurality of annular heaters and mixed by the screw. For each of the plurality of annular heaters, the control method comprises the following steps: (a) The current state and reward for previously selected actions are determined by the control unit based on the control error calculated according to the measured temperature; (b) The control unit updates the control conditions based on the reward through reinforcement learning, and selects the optimal action corresponding to the current state under the updated control conditions, wherein the control conditions are a combination of state and action; and (c) The target annular heater is controlled by the control unit based on the optimal action. In the reinforcement learning involved in steps (a) and (b), The state represents a predefined range of control error values, and the current state is the state that includes the calculated control error. When the absolute value of the current control error is less than the absolute value of the past control error, the reward is increased, making the previously selected action more likely to be selected again under the same state as the past. When the absolute value of the current control error is greater than the absolute value of the past control error, the reward is reduced so that the previously selected action is less likely to be selected again under the same state as in the past.

7. The control method for controlling a continuous mixing equipment according to claim 6, wherein, The action selected in step (b) is to change the output of the target annular heater.

8. The control method for controlling a continuous mixing equipment according to claim 6, wherein, The action selected in step (b) is to change the parameters of a PID controller configured to control the output to the target annular heater.

9. The control method for controlling a continuous mixing equipment according to claim 6, wherein, Each of the multiple temperature sensors is a thermocouple.

10. The control method for controlling a continuous mixing equipment according to claim 9, wherein, Each of the thermocouples is inserted into a through hole formed on a corresponding annular heater among the plurality of annular heaters and positioned to contact the cylinder.

Citation Information

Patent Citations

  • Temperature control method of injection moulding machine

    JP2007276189A