A centralized material supply and air supply coordinated plastic molding system
By using a closed-loop design of the vacuum pump and the return material feeding pipeline, and the coordinated adjustment of the air supply valve of the air supply pipeline and the material supply valve of the feeding pipeline, combined with the SAC reinforcement learning model of the central controller, the problem of unbalanced material and air supply was solved, and efficient and reliable production of plastic molded products was achieved.
Patent Information
- Application Number
- CN202610132366.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-30
- Publication Date
- 2026-04-03
- Estimated Expiration
- 2046-01-30
AI Technical Summary
Existing centralized feeding plastic molding systems lack a coordinated control structure for material and gas supply, resulting in an imbalance between material and gas supply in the extruder, which affects the quality stability of plastic molded products.
By using a closed-loop design of the vacuum pump and the return feed pipeline, combined with the air supply valve of the air supply pipeline and the feed valve of the feed pipeline, the material and air supply in the hopper, dryer and meter weighing equipment are coordinated and coordinated. The central controller is equipped with a SAC reinforcement learning model to carry out coordinated control of material and air supply and optimize valve action.
It achieves efficient recycling of raw materials, reduces raw material waste, ensures the reliability of plastic molded products, improves the intelligent and precise control of the material supply process, and reduces energy consumption and equipment wear.
Smart Images

Figure CN121608369B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of plastic molding, and more particularly to a centralized material supply and air supply coordinated plastic molding system. Background Technology
[0002] Plastic extrusion molding is the mainstream process in plastic molding and processing, widely used in the large-scale production of pipes, profiles, sheets, and other products. To ensure the quality stability and production efficiency of plastic molded products, the industry has begun to widely adopt centralized material supply to replace the traditional single-machine decentralized feeding mode. Centralized material supply integrates the mixing, stirring, conveying, and drying processes of plastic raw materials to achieve continuous and stable material supply to the extruder.
[0003] Currently, conventional centralized plastic feeding molding processes typically include core equipment such as mixers, hoppers, dryers, and extruders. The mixer is responsible for mixing the virgin material with additives evenly. The mixed plastic raw material is then transported to the hopper through a feeding pipe. In the hopper, the plastic raw material is mixed with high-pressure hot air. After being dehumidified and dried by the dryer, it is sent to the extruder to complete the plasticizing and molding process.
[0004] However, for example, the invention patent with publication number CN110405987A discloses a plastic central feeding system. However, its extruder for plastic molding does not have an exhaust return pipe. Therefore, it does not need to consider whether the high-pressure hot airflow entering the hopper, dryer, and metering device of the extruder matches the plastic raw material. That is, it does not need to consider that excessive high-pressure hot airflow would lead to excessive airflow pressure in the metering device, causing excessive backflow of plastic raw material within the metering device, or that insufficient high-pressure hot airflow would result in inadequate drying of the plastic raw material in the hopper and dryer, leading to defects such as bubbles and silver streaks in the plastic products. Therefore, it only needs to consider whether the amount of high-pressure hot airflow is sufficient, and thus does not require valves on the extruder's feeding pipe to regulate the extruder's feeding.
[0005] Therefore, the existing centralized feeding plastic molding process lacks a coordinated control structure for material and air supply, as well as a vacuum and material return structure on the extruder. It cannot combine the overall state of the centralized feeding system to ensure the balance of material and air supply, and thus ensure the dynamic balance of return material and the quality of plastic molded products.
[0006] Therefore, how to achieve coordinated regulation of extruder feeding and air supply in a centralized feeding plastic molding system with a vacuum reflux structure to ensure the reliable quality of plastic molded products is a technical problem that needs to be solved. Summary of the Invention
[0007] Therefore, the present invention provides a centralized material supply and air supply coordinated plastic molding system. The closed-loop design of the vacuum pump and the return material feeding pipeline enables the recycling and reuse of plastic residue from the metering machine. The air supply valve of the air supply pipeline and the material supply valve of the feeding pipeline enable the coordinated allocation of material and air supply in the hopper, dryer and metering machine, thereby ensuring the quality of material supply to the extruder and ensuring the reliability of the quality of plastic molded products.
[0008] To achieve the above objectives, the present invention proposes a centralized material supply and air supply coordinated plastic molding system, comprising:
[0009] The plastic molding subsystem includes a hopper, a dryer, a metering device, and an extruder connected in sequence in a closed manner, used to mix and dry plastic raw materials and produce plastic molded products;
[0010] The air supply subsystem is connected to the hopper and the mixer respectively through air supply pipelines to introduce high-pressure hot air flow;
[0011] The mixer has its inlet connected to one end of the return feed pipe and its outlet connected to one end of the feed pipe. The other end of the return feed pipe is connected to a meter weight via a vacuum pump, and the other end of the feed pipe is connected to the inlet of the hopper, for centralized feeding and circulation of plastic raw materials.
[0012] The air supply pipe is equipped with an air supply valve at the connection point with the hopper, and the material feeding pipe is equipped with a material feeding valve at the connection point with the hopper. The central controller communicates with the plastic molding subsystem, the mixer, the air supply valve, and the material feeding valve to perform coordinated control of the material and air supply based on the status of the plastic molding subsystem and the mixer through the SAC reinforcement learning model, so as to control the drying quality of the plastic molded products and the amount of material returned by the return material feeding pipe.
[0013] Furthermore, a return material selection station is also provided on the return material feeding pipe, and a storage tank and a discharge material selection station are also provided on the feeding pipe;
[0014] The gas supply subsystem also includes a Roots blower and a cyclone pulse filter.
[0015] Furthermore, the central controller is equipped with:
[0016] The state function acquisition module is used to acquire the load current collected by the current sensor of the mixer, the current feed flow rate collected by the flow sensor of the feed pipe, the hopper pressure collected by the pressure sensor of the hopper, the real-time material level collected by the material level sensor of the hopper, the target hopper material level stored in the central controller, the output exhaust temperature of the dryer, the real-time discharge rate collected by the flow sensor of the extruder, and the real-time rotation speed collected by the screw speed sensor of the extruder when the air supply valve and the material supply valve execute the valve action function, so as to generate a centralized feeding state function, which is used to combine the state of the plastic molding subsystem and the mixer.
[0017] The Actor network calculation module is used to generate an output vector by passing the centralized feeding state function through the Actor network, and to update the valve action function by passing the output vector through the valve opening constraint function. The valve action function includes the opening of the air supply valve for air supply to the hopper and the opening of the material supply valve for material supply to the hopper, so as to perform coordinated control of material supply and air supply.
[0018] The reward function calculation module is used to calculate the multi-feed target reward function based on the centralized feeding state function and the valve action function, wherein the multi-feed target reward function includes a material level pressure tracking reward, a mixing uniformity reward, an energy consumption reduction reward, a valve safety reward, and a valve smoothing reward.
[0019] The Actor network training module is used to generate action values by passing the valve action function, centralized feeding state function, and multi-feeding target reward function through the Critic network. The Actor network is trained based on the action values, wherein the Critic network is trained based on a multi-objective loss function. The Actor network and the Critic network together form a SAC reinforcement learning model.
[0020] Furthermore, the reward function calculation module includes:
[0021] The material level pressure tracking reward calculation submodule is used to generate material level pressure tracking reward items by passing the centralized material supply state function through a dead zone function;
[0022] The mixing uniformity reward calculation submodule is used to calculate the mixing uniformity of the material in the mixer and the material residence time in the hopper based on the centralized feeding state function, and to calculate the mixing uniformity reward based on the mixing uniformity and the material residence time in the hopper;
[0023] The energy consumption reduction reward calculation submodule is used to calculate the energy consumption reduction reward by fitting the valve action function to the Roots blower power and scheduling energy consumption;
[0024] The valve safety reward calculation submodule is used to calculate the valve safety reward based on the opening adjustment amount of the valve action function when the number of times the valve opening exceeds the limit adjustment is greater than the number of safe valve lifespan adjustments.
[0025] The valve smoothing reward calculation submodule is used to calculate the valve smoothing reward based on the frequency domain value of the opening adjustment amount;
[0026] The adaptive weight calculation submodule is used to generate adaptive weights by passing the production stage encoded vectors through a multilayer perceptron model.
[0027] The reward function weighted calculation submodule is used to perform weighted calculations on the material level pressure tracking reward item, the mixing uniformity reward item, the energy consumption reduction reward item, the valve safety reward item, and the valve smoothing reward item based on the adaptive weights, so as to generate the multi-feed target reward function.
[0028] Furthermore, the adaptive weight calculation submodule includes:
[0029] The inner loop update unit is used to generate initial adaptive weights based on the initial model parameters of the multilayer perceptron model, calculate the multi-feed target reward function based on the initial adaptive weights and empirical data, calculate the model update loss function, and optimize and update the model parameters of the multilayer perceptron model through the model update loss function when the Actor network updates the set number of steps.
[0030] The outer loop update unit is used to calculate the model update loss function based on the verification data when the number of times the valve action function is executed equals the set number of times, and to optimize and update the model parameters of the multilayer perceptron model through the model update loss function.
[0031] The adaptive weights include the initial adaptive weights.
[0032] Furthermore, the adaptive weight calculation submodule also includes:
[0033] A multi-level fully connected layer feature extraction unit is used to pass the production stage encoding vector through a multi-level fully connected layer to generate weighted mapping features;
[0034] An output layer mapping unit is used to pass the weight mapping features through an output layer based on a Softplus function to generate the adaptive weights.
[0035] Furthermore, the energy consumption reduction reward calculation submodule includes:
[0036] The fan power fitting unit is used to fit the Roots blower power based on the air supply valve opening degree and the fan reference power coefficient of the valve action function.
[0037] The material conveying power fitting unit is used to fit the feeding power based on the opening degree of the feeding valve and the feeding reference power coefficient of the valve action function.
[0038] The scheduling energy consumption fitting unit is used to calculate the scheduling energy consumption based on the opening change of the air supply valve and the average opening during the control cycle.
[0039] The energy consumption reduction reward accumulation calculation unit is used to add the Roots blower power, feeding power and scheduling energy consumption to generate the energy consumption reduction reward.
[0040] Furthermore, the Actor network training module includes:
[0041] The state value calculation submodule is used to generate a state value function by passing the centralized material supply state function through a state value branch.
[0042] The action advantage calculation submodule is used to generate an action advantage function by passing the valve action function and the centralized feeding state function through the action advantage branch;
[0043] The action value calculation subunit is used to calculate the action value based on the state value function, the action advantage function, and the batch sample mean of the action advantage function;
[0044] The Critic network includes a state value branch and an action advantage branch.
[0045] Furthermore, the Actor network training module also includes:
[0046] The material balance constraint calculation submodule is used to calculate material balance constraint terms based on the first gradient value of the action value with respect to the opening of the feed valve, the second gradient value of the action value with respect to the real-time material level, the third gradient value of the action value with respect to the opening of the air supply valve, and the fourth gradient value of the action value with respect to the hopper pressure.
[0047] The multi-objective loss function calculation submodule is used to calculate the multi-objective loss function based on the material balance constraint and the Bellman loss term.
[0048] Furthermore, the Actor network computing module also includes:
[0049] The mean and standard deviation generation submodule is used to generate the mean opening vector and the standard deviation opening vector by passing the centralized feeding state function through the Actor network, wherein the output vector includes the mean opening vector and the standard deviation opening vector.
[0050] The valve action function update submodule is used to update the valve action functions executed by the air supply valve and the material supply valve by passing the mean opening vector and the standard deviation of opening vector through a valve opening constraint function based on the tanh function and the safe opening range.
[0051] Compared with the prior art, the beneficial effects of the present invention are as follows: the present invention forms a closed-loop material supply cycle structure through a mixer, a return material feeding pipe, a vacuum pump, a feeding pipe, and a hopper. Combined with the optimized configuration of the return material selection station, storage tank, and discharge material selection station, it realizes the efficient recycling and graded management of raw materials. The residual material separated by the metering device can be transported to the mixer through the return material feeding pipe via the vacuum pump. After secondary mixing, it is fed back into the hopper through the feeding pipe to participate in the molding process, which greatly reduces the raw material waste rate. It is suitable for the molding production of high-value plastic raw materials. By setting an air supply valve in the air supply pipe and a feeding valve in the feeding pipe, the material and air supply in the hopper, dryer, and metering device can be coordinated and coordinated, thereby ensuring the quality of the extruder supply and ensuring the reliability of the quality of the plastic molded products.
[0052] In particular, the central controller of this invention is equipped with deep collaboration of state function acquisition, Actor network calculation, reward function calculation, Critic network calculation and Actor network training, constructing a fully closed-loop intelligent control system for centralized feeding valves. It realizes multi-objective collaborative optimization of gas and material supply valves by SAC reinforcement learning model, effectively avoiding the low control accuracy, poor multi-objective balancing ability, high energy consumption and insufficient equipment operation stability of traditional centralized plastic feeding systems, and realizing intelligent and precise control of the feeding process.
[0053] In particular, this invention enhances the stability and tracking accuracy of material level and pressure by utilizing dead-zone functions. Combined with the mixing uniformity of materials in the mixer and the residence time of materials in the hopper, it ensures the mixing quality of plastic raw materials. By fitting the power of the Roots blower and the scheduling energy consumption, it accurately controls the energy consumption of material supply. Based on the frequency domain characteristics of the number of times the valve opening exceeds the standard and the amount of valve opening adjustment, it reduces valve wear and suppresses valve oscillation. The adaptive weight calculation submodule generates adaptive weights by processing the encoding vectors of the production stage through a multilayer perceptron model. With the help of the dual-loop update mechanism of inner and outer loops, it realizes the dynamic adaptation of the weights of various reward items under different production stages, which greatly improves the comprehensive operating performance of the material supply system throughout the entire production cycle, ensuring the quality of raw material supply while minimizing energy consumption and equipment wear. Attached Figure Description
[0054] Figure 1 This is a schematic diagram of the overall structure of the centralized material supply and air supply coordinated plastic molding system according to an embodiment of the present invention;
[0055] Figure 2 This invention relates to a centralized material supply and air supply coordinated plastic molding system. Figure 1 Detailed structural diagram of Part A;
[0056] Figure 3 This invention relates to a centralized material supply and air supply coordinated plastic molding system. Figure 2 Detailed structural diagram of Part B;
[0057] Figure 4 This is a schematic diagram of the central controller module of the centralized material supply and air supply coordinated plastic molding system according to an embodiment of the present invention.
[0058] Figure 5 This is a flowchart illustrating the central controller module of the centralized material supply and air supply coordinated plastic molding system according to an embodiment of the present invention.
[0059] Figure 6 This is a flowchart illustrating the multi-feeding objective reward function of the centralized feeding and gas supply coordinated plastic molding system according to an embodiment of the present invention.
[0060] In the diagram: 1. Roots blower; 2. Central controller; 3. Air supply pipeline; 4. Mixer; 5. Air supply valve; 6. Feed valve; 7. Hopper; 8. Dryer; 9. Extruder; 10. Feeding pipeline; 11. Meter weigher; 12. Return material feed pipeline; 13. Vacuum pump; 14. Storage tank; 15. Discharge sorting station; 16. Return material sorting station; 17. Cyclone pulse filter. Detailed Implementation
[0061] To make the objectives and advantages of the present invention clearer, the present invention will be further described below with reference to embodiments; it should be understood that the specific embodiments described herein are merely for explaining the present invention and are not intended to limit the present invention.
[0062] Preferred embodiments of the present invention will now be described with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are merely illustrative of the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.
[0063] It should be noted that in the description of this invention, the terms "upper", "lower", "left", "right", "inner", "outer", etc., which indicate directions or positional relationships, are based on the directions or positional relationships shown in the accompanying drawings. This is only for the convenience of description and is not intended to indicate or imply that the device or element must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation of this invention.
[0064] Furthermore, it should be noted that, in the description of this invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.
[0065] like Figures 1 to 6 As shown, the present invention provides a centralized material supply and air supply coordinated plastic molding system. The closed-loop design of vacuum pump 13 and return material feeding pipe 12 realizes the recycling and reuse of plastic residue in meter weigher 11. The air supply valve 5 of air supply pipe 3 and the material supply valve 6 of feeding pipe 10 realize the coordinated allocation of material and air supply in hopper 7, dryer 8 and meter weigher 11, thereby ensuring the quality of material supplied by extruder 9 and ensuring the reliability of plastic molded products.
[0066] like Figures 1 to 3 As shown, this embodiment proposes a centralized material supply and air supply coordinated plastic molding system, including:
[0067] The plastic molding subsystem includes a hopper 7, a dryer 8, a metering device 11 and an extruder 9 connected in sequence and sealed, which are used to mix and dry plastic raw materials and produce plastic molded products.
[0068] The air supply subsystem is connected to the hopper 7 and the mixer 4 respectively through the air supply pipe 3 to introduce high-pressure hot air flow;
[0069] The mixer 4 has its inlet connected to one end of the return feed pipe 12 and its outlet connected to one end of the feed pipe 10. The other end of the return feed pipe 12 is connected to the meter weight 11 via the vacuum pump 13, and the other end of the feed pipe 10 is connected to the inlet of the hopper 7 for centralized feeding and circulation of plastic raw materials.
[0070] The air supply pipe 3 is connected to the hopper 7 by an air supply valve 5, and the feeding pipe 10 is connected to the hopper 7 by a feeding valve 6. The central controller 2 communicates with the plastic molding subsystem, the mixer 4, the air supply valve 5, and the feeding valve 6 to perform coordinated control of the feeding and air supply in combination with the state of the plastic molding subsystem and the mixer 4 through the SAC reinforcement learning model, so as to control the drying quality of the plastic molded products and the amount of return material from the return feeding pipe 12.
[0071] Understandably, see Figure 2 and 3 The hopper 7, dryer 8, and weighing device 11 are vertically and sealed together to form a closed chamber. This ensures that the high-pressure hot airflow from the air supply pipe 3 entering the chamber is only discharged after the dryer 8 has finished drying the plastic raw material and the weighing device 11 has finished weighing. The inlet on the upper side of the weighing device 11, connected to the dryer 8, is closed when the material is being fed into the extruder 9. Part of the high-pressure hot airflow extracted by the vacuum pump 13 is discharged externally, and part is used for transporting residual material from the weighing device 11 into the return feed pipe 12.
[0072] Understandably, see Figure 3The hopper dryer 8 can only control the temperature of the high-pressure hot airflow in the chamber. It cannot control the flow rate of hot air used for drying plastic raw materials in the closed chamber. Therefore, the supply air valve 5 and the material supply valve 6 are used to ensure the ratio of material supply and air supply in the closed chamber. This ensures that the supply of high-pressure hot airflow is not too large, thereby avoiding excessive backflow of plastic raw materials in the hopper 11. It also ensures that the supply of high-pressure hot airflow is not too small, thereby avoiding insufficient drying of plastic raw materials in the hopper 7 and the dryer 8, and thus avoiding defects such as bubbles and silver streaks in plastic products.
[0073] Furthermore, a return material selection station 16 is also provided on the return material feeding pipe 12, and a storage tank 14 and a discharge material selection station 15 are also provided on the feeding pipe 10;
[0074] The air supply subsystem also includes a Roots blower 1 and a cyclone pulsating filter 17.
[0075] Specifically, the mixer 4 is preferably a vertical mixer, and each vertical mixer is connected to the corresponding hopper 7 through a feeding pipe 10. The return material sorting station 16 and the discharge material sorting station 15 are only used to screen unqualified plastic raw materials and are not used to control the amount of return material and the amount of supply material. The Roots blower 1 is a three-in-one standby Roots blower, and the cyclone pulse filter is used for dust filtration of high-pressure hot airflow.
[0076] like Figures 4 to 5 As shown, the central controller 2 is further equipped with:
[0077] The state function acquisition module is used to acquire the load current collected by the current sensor of the mixer 4, the current feed flow rate collected by the flow sensor of the feed pipe 10, the hopper pressure collected by the pressure sensor of the hopper 7, the real-time material level collected by the material level sensor of the hopper 7, the target hopper material level stored in the central controller 2, the output exhaust temperature of the dryer 8, the real-time discharge rate collected by the flow sensor of the extruder 9, and the real-time rotation speed collected by the screw speed sensor of the extruder 9 when the air supply valve 5 and the material supply valve 6 execute the valve action function, so as to generate a centralized feeding state function, which is used to combine the state of the plastic molding subsystem and the mixer 4.
[0078] The Actor network calculation module is used to generate an output vector by passing the centralized feeding state function through the Actor network, and to update the valve action functions executed by the air supply valve 5 and the material supply valve 6 by passing the output vector through the valve opening constraint function. The valve action functions include the air supply valve opening of the air supply valve 5 for supplying air to the hopper 7 and the material supply valve opening of the material supply valve 6 for supplying material to the hopper 7, so as to perform coordinated control of material and air supply.
[0079] The reward function calculation module is used to calculate the multi-feed target reward function based on the centralized feeding state function and the valve action function, wherein the multi-feed target reward function includes a material level pressure tracking reward, a mixing uniformity reward, an energy consumption reduction reward, a valve safety reward, and a valve smoothing reward.
[0080] The Actor network training module is used to generate action values by passing the valve action function, centralized feeding state function, and multi-feeding target reward function through the Critic network. The Actor network is trained based on the action values, wherein the Critic network is trained based on a multi-objective loss function. The Actor network and the Critic network together form a SAC reinforcement learning model.
[0081] Specifically, in the SAC (Soft Actor-Critic) reinforcement learning model, the Actor network and the Critic network have a bidirectional dependency, alternating training, and collaborative goals, which specifically includes:
[0082] Step S1: The Critic network relies on the Actor network to obtain training data: The Actor network obtains training data in the centralized feeding state function s t Downsampling generates valve action function a t The valve executes the valve action function a. t Then, obtain the new centralized feeding state function s. t+1 and reward r t This forms training samples (s) stored in the experience replay pool D. t ,a t ,r t ,s t+1 The Critic network optimizes its network parameters through training using multiple sets of training samples from the experience replay pool D.
[0083] Step S2: The Actor network relies on the Critic network's value-guided updates: The core of the Actor network's Actor loss function is to maximize the soft Q-value (Critic score, action value) and entropy regularization (exploration). The Actor loss function can be expressed as:
[0084]
[0085] In the formula, This represents the network parameters for the Actor network. The Actor loss function, where E represents the expected value, is the mean of the expression within the parentheses under the specified distribution. The centralized feeding state function s is sampled from the empirical playback pool D. This indicates that the valve action function 'a' is sampled from the policy distribution output by the Actor network. , The action value generated by the Critic network with parameter θ represents the expected benefit of executing the valve action function a under the centralized feeding state function s. This represents a temperature coefficient used to adjust the weighting of the natural logarithm of the strategy distribution with the action value. The constant representing twice the value of pi corresponds to two types of variables in the action space and is a standard component of the Gaussian probability density function. This represents the standard deviation vector of the opening of the i-th action dimension generated by the Actor network. Let represent the mean opening vector of the i-th action dimension generated by the Actor network. This represents the valve action function updated by the valve opening constraint function. Therefore, the Criti network is optimized by using the Actor loss function to guide the Actor network update.
[0086] This embodiment is used for, for example Figure 1 In the plastic molding subsystem and mixer 4 shown, the plastic molding subsystem and mixer 4 can be represented as follows:
[0087]
[0088] In the formula, This represents the feeding status function. This indicates the hopper pressure (Pa) collected by the hopper pressure sensor. This indicates the real-time material level (%) collected by the hopper level sensor. This indicates the deviation between the real-time hopper level and the target hopper level, which is a determined value calculated based on the amount of plastic supplied by the extruder. This indicates the current feed flow rate (kg / h) collected by the flow sensor in the upstream feed pipeline. This indicates the real-time discharge rate (kg / h) collected by the downstream extruder flow sensor. This represents the load current (A, reaction load) collected by the upstream mixer current sensor. This indicates the real-time rotational speed (RPM, the main disturbance source) collected by the extruder screw speed sensor. This indicates the output exhaust temperature of the dryer (°C, which affects the dryness of PE plastic).
[0089] The action function of the centralized valve in the PE plastic system can be expressed as: In the formula, This indicates the opening degree of the air supply valve, controlling the amount of drying air delivered. This indicates the opening degree of the feed valve, which controls the amount of PE material entering the valve. The values are all between 0 and 1.
[0090] like Figure 6 As shown, the reward function calculation module further includes:
[0091] The material level pressure tracking reward calculation submodule is used to generate material level pressure tracking reward items by passing the centralized material supply state function through a dead zone function;
[0092] The mixing uniformity reward calculation submodule is used to calculate the mixing uniformity of the material in the mixer and the material residence time in the hopper based on the centralized feeding state function, and to calculate the mixing uniformity reward based on the mixing uniformity and the material residence time in the hopper;
[0093] The energy consumption reduction reward calculation submodule is used to calculate the energy consumption reduction reward by fitting the valve action function to the Roots blower power and scheduling energy consumption of the Roots blower 1.
[0094] The valve safety reward calculation submodule is used to calculate the valve safety reward based on the opening adjustment amount of the valve action function when the number of times the valve opening exceeds the limit adjustment is greater than the number of safe valve lifespan adjustments.
[0095] The valve smoothing reward calculation submodule is used to calculate the valve smoothing reward based on the frequency domain value of the opening adjustment amount;
[0096] The adaptive weight calculation submodule is used to generate adaptive weights by passing the production stage encoded vectors through a multilayer perceptron model.
[0097] The reward function weighted calculation submodule is used to perform weighted calculations on the material level pressure tracking reward item, the mixing uniformity reward item, the energy consumption reduction reward item, the valve safety reward item, and the valve smoothing reward item based on the adaptive weights, so as to generate the multi-feed target reward function.
[0098] Specifically, the production stage coding vector includes the thermal coding of the material level standard deviation, the pressure standard deviation, and the plastic formulation category over the past 3 minutes.
[0099] Furthermore, the energy consumption reduction reward calculation submodule includes:
[0100] The fan power fitting unit is used to fit the Roots blower power based on the air supply valve opening degree and the fan reference power coefficient of the valve action function.
[0101] The material conveying power fitting unit is used to fit the feeding power based on the opening degree of the feeding valve and the feeding reference power coefficient of the valve action function.
[0102] The scheduling energy consumption fitting unit is used to calculate the scheduling energy consumption based on the opening change of the air supply valve and the average opening during the control cycle.
[0103] The energy consumption reduction reward accumulation calculation unit is used to add the Roots blower power, feeding power and scheduling energy consumption to generate the energy consumption reduction reward.
[0104] Specifically, the reward function for the multi-supply target can be expressed as:
[0105]
[0106] In the formula, This represents the reward function for the multi-supply target. Both represent adaptive weights. These represent the rewards for material level and pressure tracking, uniform mixing, energy consumption reduction, valve safety, and valve smoothing, respectively. This represents the dead zone function based on the real-time change in hopper level, i.e., the change is less than or equal to a set range. (Preferably 10%) the value is 0, otherwise it is 0. This represents the dead zone function based on the change in hopper pressure, i.e., when the change in hopper pressure is less than or equal to a set range. (Preferably 50 Pa) the value is 0, otherwise it is 0. This indicates the rate of change of the material level in the hopper in real time. Therefore, the material level pressure tracking bonus can prevent unnecessary frequent valve operations near the set point. These represent the material mixing uniformity of the mixer and the material mixing uniformity threshold (preferably 0.03), respectively. This represents the variance of the ratio of mixer current to feed flow rate, used to indirectly estimate the uniformity of material mixing. Indicates the residence time of material in the hopper. This indicates a penalty intensity coefficient (preferably 0.5) to prevent insufficient drying due to excessively short material residence time. Indicates the opening degree of the air supply valve. Indicates the opening degree of the feed valve. This represents the fan reference power coefficient, which records the fan power at different air supply valve openings. Nonlinear regression fitting was used to determine the relationship between actual power and the 3 / 2 power of valve opening. This represents the feed reference power coefficient, determined by measuring the idling power of the conveyor mechanism and its rated power at the rated feed rate, based on a nonlinear regression fit of the rated power and idling power. This represents the adjustment loss parameter, estimated by analyzing historical data: A period of frequent valve adjustment is selected, the total energy consumption during that period is calculated, and the baseline energy consumption predicted based on the first two parameters is subtracted. The difference is then divided by the accumulated... , These represent the change in the opening degree of the air supply valve and the average opening degree during the control cycle, respectively. This indicates the indicator function, specifically the number of adjustments needed when the opening exceeds the limit. More than the valve's safe lifespan cycles The value is 1 if the condition is met, and 0 otherwise. This represents the absolute value of the opening adjustment raised to the power of 1.5, to emphasize the damage caused by large movements. This represents the frequency domain value of the opening adjustment amount, in order to suppress oscillations at dangerous frequencies.
[0107] Furthermore, the adaptive weight calculation submodule includes:
[0108] The inner loop update unit is used to generate initial adaptive weights based on the initial model parameters of the multilayer perceptron model, calculate the multi-feed target reward function based on the initial adaptive weights and empirical data, calculate the model update loss function, and optimize and update the model parameters of the multilayer perceptron model through the model update loss function when the Actor network updates the set number of steps.
[0109] The outer loop update unit is used to calculate the model update loss function based on the verification data when the number of times the valve action function is executed equals the set number of times, and to optimize and update the model parameters of the multilayer perceptron model through the model update loss function.
[0110] The adaptive weights include the initial adaptive weights.
[0111] Specifically, the model update loss function and its update process can be expressed as:
[0112]
[0113] In the formula, This represents the model update loss function. This indicates that the centralized material supply state function and valve action function are collected from empirical data or verification data. Discount factor The preferred discount factor is 0.99. For time steps, This represents the reward function for the multi-supply target. These represent the model parameters of the multilayer perceptron model before and after the update, respectively. This indicates that the gradient descent value is calculated based on the model update loss function.
[0114] Specifically, the process of updating the multilayer perceptron model includes: for the current task, using the current policy distribution. The system interacts with the current multilayer perceptron model and the environment, storing data in a replay pool. Data is periodically sampled from the replay pool, and the inner loop described above is executed to obtain adaptive parameters for the current task. The weights generated by the adapted weight generator are used to calculate rewards and update the SAC's policy and Critic networks. The outer loop is executed periodically to give the weight generator better and faster adaptability.
[0115] Furthermore, the adaptive weight calculation submodule also includes:
[0116] A multi-level fully connected layer feature extraction unit is used to pass the production stage encoding vector through a multi-level fully connected layer to generate weighted mapping features;
[0117] An output layer mapping unit is used to pass the weight mapping features through an output layer based on a Softplus function to generate the adaptive weights.
[0118] Specifically, the structure of the multilayer perceptron model can be represented as:
[0119]
[0120] In the formula, These represent the first and second level outputs of a multi-level fully connected layer, respectively. express Activation function Representation layer normalization, This represents the encoding vector for the production stage. This represents the model parameters of a multi-level fully connected layer. Indicates the initial output. express function, Indicates the model parameters of the output layer. Indicates adaptive weights, This indicates that the initial output is normalized. The sum of the total weights is represented by the denominator, which represents the L1 norm calculation.
[0121] Furthermore, the Actor network training module includes:
[0122] The state value calculation submodule is used to generate a state value function by passing the centralized material supply state function through a state value branch.
[0123] The action advantage calculation submodule is used to generate an action advantage function by passing the valve action function and the centralized feeding state function through the action advantage branch;
[0124] The action value calculation subunit is used to calculate the action value based on the state value function, the action advantage function, and the batch sample mean of the action advantage function;
[0125] The Critic network includes a state value branch and an action advantage branch.
[0126] Specifically, the process of calculating the value of the action can be expressed as follows:
[0127]
[0128] In the formula, Indicates the value of an action. The state value function represents the state value branch generated by the state value branch. This represents the action advantage function generated by the action advantage branch. Let N represent the batch sample mean of the action advantage function, where N represents the total number of batch samples. This represents the valve action function sampled for sample i.
[0129] Furthermore, the Actor network training module also includes:
[0130] The material balance constraint calculation submodule is used to calculate material balance constraint terms based on the first gradient value of the action value with respect to the opening of the feed valve, the second gradient value of the action value with respect to the real-time material level, the third gradient value of the action value with respect to the opening of the air supply valve, and the fourth gradient value of the action value with respect to the hopper pressure.
[0131] The multi-objective loss function calculation submodule is used to calculate the multi-objective loss function based on the material balance constraint and the Bellman loss term.
[0132] Furthermore, the multi-objective loss function calculation submodule is used to calculate the Bellman loss term based on the action value, strategy distribution, and multi-feeding objective reward function, and to add the material balance constraint term and the Bellman loss term to calculate the multi-objective loss function.
[0133] Specifically, the process of calculating the multi-objective loss function can be expressed as:
[0134]
[0135] In the formula, This represents the material balance constraint. This represents the mathematical expectation, which is the solution for the mean of the expression within parentheses under a specified distribution. The centralized feeding state function s is sampled from the empirical playback pool D. These represent the first gradient value of the partial derivative of the action value with respect to the opening of the feed valve, the second gradient value of the partial derivative of the action value with respect to the real-time material level, the third gradient value of the partial derivative of the action value with respect to the opening of the air supply valve, and the fourth gradient value of the partial derivative of the action value with respect to the hopper pressure, respectively. This represents the fan's reference power coefficient, recording the fan power at different air supply valve openings. This represents the feed reference power coefficient, and its fitting process is the same as described above. The Critic network learns that the benefit of opening the feed valve more fully must be proportional to the benefit of increasing the material level, and that the judgments regarding the value of the damper opening and the pressure value must remain consistent. This represents the target value, which is based on the value of the action. Strategy distribution Multi-supply target reward function The calculations, in which action value and strategy distribution are both based on the updated centralized feeding state function and valve action function, This indicates that the minimum value of the two actions generated by the two objective Critic networks in the SAC reinforcement learning model architecture is taken. This represents the discount factor, preferably 0.99. Indicates the use of this to determine the current state. Whether it is the termination flag of the terminated round state, that is, when The termination flag is set to 1 when the round is to be terminated, and 0 otherwise. This represents the mathematical expectation, which is the solution for the mean of the expression within parentheses under a specified distribution. Represents training samples (s) t ,a t ,r t ,s t+1 ) Taken from the experience replay pool D, where s t ,a t ,r t ,s t+1 The functions are, in order: the current centralized material supply status function, the valve action function, the reward function, and the updated centralized material supply status function. This represents the calculated action value based on the current centralized feeding state function and valve action function. This represents the Bellman loss term. This represents a multi-objective loss function.
[0136] Furthermore, the Actor network computing module also includes:
[0137] The mean and standard deviation generation submodule is used to generate the mean opening vector and the standard deviation opening vector by passing the centralized feeding state function through the Actor network, wherein the output vector includes the mean opening vector and the standard deviation opening vector.
[0138] The valve action function update submodule is used to update the valve action functions executed by the air supply valve 5 and the material supply valve 6 by passing the mean opening vector and the standard deviation of opening vector through the valve opening constraint function based on the tanh function and the safe opening range.
[0139] Specifically, the process of calculating the valve action function can be expressed as:
[0140]
[0141] In the formula, This represents the initial valve action function. This represents the mean valve opening vector generated by the Actor network, which is the most typical optimal valve opening in the current state. This represents the standard deviation vector of the opening value generated by the Actor network. The larger the standard deviation vector of the opening value, the more the algorithm tends to try actions that deviate from the mean. This represents element-wise multiplication. Sampling noise representing a standard Gaussian distribution. This represents the valve action function. The safe opening range is represented by tanh, which represents the tanh function. Therefore, the standard output vector of the Actor network is converted to a range of 0 to 1 using the tanh function, and the actual valve opening is constrained by the safe opening range to ensure the safety of the valve opening for the feeding process.
[0142] In this embodiment, a closed-loop material supply and circulation structure is formed by the mixer 4, the return material feeding pipe 12, the vacuum pump 13, the feeding pipe 10, and the hopper 7. Combined with the optimized configuration of the return material selection station 16, the storage tank 14, and the discharge material selection station 15, the efficient recycling and graded management of raw materials are realized. The residual material separated by the meter weigher 11 can be transported to the mixer 4 through the return material feeding pipe 12 via the vacuum pump 13. After secondary mixing, it is fed back into the hopper 7 by the feeding pipe 10 to participate in the molding process, which greatly reduces the raw material waste rate and is suitable for the molding production of high-value plastic raw materials. By setting the air supply valve 5 in the air supply pipe 3 and the feeding valve 6 in the feeding pipe 10, the material and air supply in the hopper 7, the dryer 8, and the meter weigher 11 can be coordinated and coordinated, thereby ensuring the quality of the material supplied by the extruder 9 and ensuring the reliability of the quality of the plastic molded products. The central controller 2 is equipped with a deep collaboration of state function acquisition, Actor network calculation, reward function calculation, Critic network calculation and Actor network training to build a fully closed-loop intelligent control system for centralized feeding valves. It realizes multi-objective collaborative optimization of air and material supply valves by SAC reinforcement learning model, effectively avoiding the low control accuracy, poor multi-objective balancing ability, high energy consumption and insufficient equipment operation stability of traditional PE plastic centralized feeding systems, and realizing intelligent and precise control of the feeding process. By leveraging dead-zone functions to enhance the stability and accuracy of material level and pressure tracking, and combining the mixing uniformity of materials in the mixer with the material residence time in the hopper, the mixing quality of PE plastic raw materials is ensured. By fitting the power of the Roots blower and the scheduling energy consumption, the energy consumption of material supply is precisely controlled. Based on the frequency domain characteristics of the number of times the valve opening exceeds the standard and the amount of valve opening adjustment, valve wear is reduced and valve oscillation is suppressed. The adaptive weight calculation submodule generates adaptive weights by processing the encoding vectors of the production stage through a multilayer perceptron model. With the help of a dual-loop update mechanism of inner loop (triggered by the number of steps updated by the Actor network) and outer loop (triggered by the number of times the valve action is executed), the weights of each reward item under different production stages are dynamically adapted, which greatly improves the overall operating performance of the material supply system throughout the entire production cycle, ensuring the quality of raw material supply while minimizing energy consumption and equipment wear. By employing a dual-branch architecture of state value branch and action advantage branch through the Critic network, and combining the batch sample mean of the action advantage function to accurately calculate the action value, the accuracy of value assessment is ensured. Based on the gradient values of action value for valve opening, material level, and pressure, material balance constraints are calculated, and a multi-objective loss function is constructed. This allows the SAC model to balance multiple objectives such as material level and pressure tracking, uniform mixing, and energy consumption reduction during training, while also ensuring material balance in the feeding process. This significantly improves the overall operating performance of the feeding system throughout the entire production cycle, ensuring the quality of raw material supply while minimizing energy consumption and equipment wear.
[0143] Those skilled in the art will recognize that the modules and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0144] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A centralized material supply and air supply coordinated plastic molding system, characterized in that, include: The plastic molding subsystem includes a hopper (7), a dryer (8), a meter weigher (11), and an extruder (9) connected in sequence in a closed manner, for mixing and drying plastic raw materials and producing plastic molded products; The gas supply subsystem is connected to the hopper (7) and the mixer (4) respectively through the gas supply pipe (3) to introduce high-pressure hot gas flow; The mixer (4) has its inlet connected to one end of the return feed pipe (12) and its outlet connected to one end of the feed pipe (10). The other end of the return feed pipe (12) is connected to the meter weight (11) via a vacuum pump (13). The other end of the feed pipe (10) is connected to the inlet of the hopper (7) for centralized feeding of plastic raw materials. Among them, the air supply pipe (3) is connected to the hopper (7) with an air supply valve (5), and the feeding pipe (10) is connected to the hopper (7) with a feeding valve (6). The central controller (2) communicates with the plastic molding subsystem, the mixer (4), the air supply valve (5) and the feeding valve (6) to perform coordinated control of the feeding and air supply in combination with the state of the plastic molding subsystem and the mixer (4) through the SAC reinforcement learning model, so as to control the drying quality of the plastic molded products and the amount of return material in the return feeding pipe (12). The central controller (2) is equipped with: The state function acquisition module is used to acquire the load current collected by the current sensor of the mixer (4), the current feed flow collected by the flow sensor of the feed pipe (10), the hopper pressure collected by the pressure sensor of the hopper (7), the real-time material level collected by the material level sensor of the hopper (7), the target hopper material level stored in the central controller (2), the output exhaust temperature of the dryer (8), the real-time discharge rate collected by the flow sensor of the extruder (9), and the real-time rotation speed collected by the screw speed sensor of the extruder (9) when the air supply valve (5) and the material supply valve (6) execute the valve action function, so as to generate a centralized feeding state function, which is used to combine the state of the plastic molding subsystem and the mixer (4); The Actor network calculation module is used to generate an output vector by passing the centralized feeding state function through the Actor network, and to update the valve action function by passing the output vector through the valve opening constraint function. The valve action function includes the air supply valve opening of the air supply valve (5) for supplying air to the hopper (7) and the material supply valve opening of the material supply valve (6) for supplying material to the hopper (7), so as to perform coordinated control of material supply and air supply. The reward function calculation module is used to calculate the multi-feed target reward function based on the centralized feeding state function and the valve action function, wherein the multi-feed target reward function includes a material level pressure tracking reward, a mixing uniformity reward, an energy consumption reduction reward, a valve safety reward, and a valve smoothing reward. The Actor network training module is used to generate action values by passing the valve action function, centralized feeding state function, and multi-feeding target reward function through the Critic network. The Actor network is trained based on the action values, wherein the Critic network is trained based on a multi-objective loss function. The Actor network and the Critic network together form a SAC reinforcement learning model.
2. The centralized material supply and air supply coordinated plastic molding system according to claim 1, characterized in that, The return material feeding pipe (12) is also equipped with a return material selection station (16), and the feeding pipe (10) is also equipped with a storage tank (14) and a discharge material selection station (15); The gas supply subsystem also includes a Roots blower (1) and a cyclone pulsating filter (17).
3. The centralized material supply and air supply coordinated plastic molding system according to claim 1, characterized in that, The reward function calculation module includes: The material level pressure tracking reward calculation submodule is used to generate material level pressure tracking reward items by passing the centralized material supply state function through a dead zone function; The mixing uniformity reward calculation submodule is used to calculate the mixing uniformity of the material in the mixer and the material residence time in the hopper based on the centralized feeding state function, and to calculate the mixing uniformity reward based on the mixing uniformity and the material residence time in the hopper; The energy consumption reduction reward calculation submodule is used to calculate the energy consumption reduction reward by fitting the valve action function to the Roots blower power and scheduling energy consumption of the Roots blower (1); The valve safety reward calculation submodule is used to calculate the valve safety reward based on the opening adjustment amount of the valve action function when the number of times the valve opening exceeds the limit adjustment is greater than the number of safe valve lifespan adjustments. The valve smoothing reward calculation submodule is used to calculate the valve smoothing reward based on the frequency domain value of the opening adjustment amount; The adaptive weight calculation submodule is used to generate adaptive weights by passing the production stage encoded vectors through a multilayer perceptron model. The reward function weighted calculation submodule is used to perform weighted calculations on the material level pressure tracking reward item, the mixing uniformity reward item, the energy consumption reduction reward item, the valve safety reward item, and the valve smoothing reward item based on the adaptive weights, so as to generate the multi-feed target reward function.
4. The centralized material supply and air supply coordinated plastic molding system according to claim 3, characterized in that, The adaptive weight calculation submodule includes: The inner loop update unit is used to generate initial adaptive weights based on the initial model parameters of the multilayer perceptron model, calculate the model update loss function of the multi-feed target reward function based on the initial adaptive weights and empirical data, and optimize and update the model parameters of the multilayer perceptron model through the model update loss function when the Actor network updates the set number of steps. The outer loop update unit is used to calculate the model update loss function based on the verification data when the number of times the valve action function is executed equals the set number of times, and to optimize and update the model parameters of the multilayer perceptron model through the model update loss function. The adaptive weights include the initial adaptive weights.
5. The centralized material supply and air supply coordinated plastic molding system according to claim 3, characterized in that, The adaptive weight calculation submodule also includes: A multi-level fully connected layer feature extraction unit is used to pass the production stage encoding vector through a multi-level fully connected layer to generate weighted mapping features; An output layer mapping unit is used to pass the weight mapping features through an output layer based on a Softplus function to generate the adaptive weights.
6. The centralized material supply and air supply coordinated plastic molding system according to claim 3, characterized in that, The energy consumption reduction reward calculation submodule includes: The fan power fitting unit is used to fit the Roots blower power based on the air supply valve opening degree and the fan reference power coefficient of the valve action function. The material conveying power fitting unit is used to fit the feeding power based on the opening degree of the feeding valve and the feeding reference power coefficient of the valve action function. The scheduling energy consumption fitting unit is used to calculate the scheduling energy consumption based on the opening change of the air supply valve and the average opening during the control cycle. The energy consumption reduction reward accumulation calculation unit is used to add the Roots blower power, feeding power and scheduling energy consumption to generate the energy consumption reduction reward.
7. The centralized material supply and air supply coordinated plastic molding system according to claim 1, characterized in that, The Actor network training module includes: The state value calculation submodule is used to generate a state value function by passing the centralized material supply state function through a state value branch. The action advantage calculation submodule is used to generate an action advantage function by passing the valve action function and the centralized feeding state function through the action advantage branch; The action value calculation subunit is used to calculate the action value based on the state value function, the action advantage function, and the batch sample mean of the action advantage function; The Critic network includes a state value branch and an action advantage branch.
8. The centralized material supply and air supply coordinated plastic molding system according to claim 1, characterized in that, The Actor network training module also includes: The material balance constraint calculation submodule is used to calculate material balance constraint terms based on the first gradient value of the action value with respect to the opening of the feed valve, the second gradient value of the action value with respect to the real-time material level, the third gradient value of the action value with respect to the opening of the air supply valve, and the fourth gradient value of the action value with respect to the hopper pressure. The multi-objective loss function calculation submodule is used to calculate the multi-objective loss function based on the material balance constraint and the Bellman loss term.
9. The centralized material supply and air supply coordinated plastic molding system according to claim 1, characterized in that, The Actor network computing module also includes: The mean and standard deviation generation submodule is used to generate the mean opening vector and the standard deviation opening vector by passing the centralized feeding state function through the Actor network, wherein the output vector includes the mean opening vector and the standard deviation opening vector. The valve action function update submodule is used to update the valve action functions executed by the air supply valve (5) and the material supply valve (6) by passing the mean opening vector and the standard deviation of opening vector through the valve opening constraint function based on the tanh function and the safe opening range.
Citation Information
Patent Citations
Plastic particle feeding machine with drying function
CN105328817A
Plastic central supply system
CN110405987A