Liquid mass flow control method
Patent Information
- Application Number
- CN202610953362.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-30
- Publication Date
- 2026-09-29
AI Technical Summary
[0005]为了解决现有技术中液体质量流量控制器难以抑制并联液路上其他流量计启停切换所引发的瞬态流量串扰的问题,本发明提供一种液体质量流量控制方法,该液体质量流量控制方法无需依托精准系统机理模型,可通过智能体与受控环境自主交互、实时学习最优闭环控制策略,实现对扰动的动态补偿,完美适配非线性、强耦合、时变扰动的工业流量控制场景,解决了现有技术中液体质量流量控制器难以抑制并联液路上其他流量计启停切换所引发的瞬态流量串扰的问题
本发明提供的液体质量流量控制方法,抗扰鲁棒性极强,采用无模型强化学习架构,无需精准系统机理模型,可自主适配时变、非线性、未知多重耦合干扰,旁路流量计启停切换所引发的瞬态流量串扰(即“抢液”效应)下流量无大幅波动,抗扰稳定性优于传统PID、模糊PID;自主学习与自适应能力突出,通过智能体与环境交互自主优化控制策略,无需人工设计规则、无需定期整定参数,适配突变工况,免后期维护;此外,该控制方法的工程适配性强,采用轻量化算法设计,可直接嵌入现有液体质量流量控制器,不改动硬件架构,泛化性强、通用性高,可广泛应用于各类工业场景。
Smart Images

Figure CN122837510A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of liquid mass flow control technology, and more particularly to a liquid mass flow control method. Background Technology
[0002] Liquid mass flow controllers are core actuators in industrial applications such as high-end chemicals, precision pharmaceuticals, food processing, semiconductor cooling, and new energy fluid transmission. Their core function is to achieve constant closed-loop control of the mass flow rate of liquid media. Their control accuracy, dynamic response speed, and anti-interference robustness directly determine the stability of the production line, the product qualification rate, and the safety of system operation.
[0003] In actual industrial operations, pipeline systems with multiple branches supplying liquid in parallel and multiple devices simultaneously drawing liquid commonly experience multi-branch parallel coupling interference. This type of interference is mainly caused by the simultaneous start-up and shutdown of multiple devices in upstream and downstream pipelines, on-demand switching of branch flow rates, and the mutual coupling and compression of pressure from multiple pumps coordinating liquid supply. It is a typical nonlinear, time-varying, sudden, and unknown disturbance. Such disturbances can cause instantaneous changes in the fluid driving force of each branch and drastic fluctuations in pipeline pressure, completely disrupting the dynamic balance of the flow control closed loop. This leads to numerous problems such as severe oscillations in liquid flow, continuous exceedance of steady-state errors, system regulation lag, and large and difficult-to-converge flow overshoot. It can easily cause deviations in precision fluid process parameters and unstable liquid supply to equipment, seriously affecting the process stability and production consistency in precision fluid control scenarios.
[0004] Currently, commercial liquid mass flow controllers generally employ traditional fixed-parameter PID control algorithms. Some high-end models incorporate fuzzy PID and adaptive PID improved algorithms. However, these methods all have inherent technical limitations when dealing with multi-liquid-path parallel coupling interference: Traditional PID control parameters are fixed after factory tuning. The disturbance intensity, timing of abrupt changes, and range of influence of coupling interference change in real time with the operating status of the field equipment. Fixed control parameters cannot adapt to these time-varying and sudden disturbance characteristics, easily leading to parameter mismatch problems under multi-liquid-path parallel coupling conditions, resulting in a precipitous drop in control performance and an inability to effectively suppress drastic flow fluctuations. Fuzzy control and conventional adaptive control rely on manually tuned control rules, precise pipeline fluid coupling mechanism models, and industry expert experience. They can only adapt to slight flow disturbances under fixed operating conditions and cannot autonomously adapt to irregular and sudden multi-path coupling disturbances and operating condition changes in industrial settings. Their generalization ability is weak, making it difficult to achieve globally optimal anti-disturbance control effects against multi-liquid-path parallel coupling interference. Therefore, existing liquid mass flow controllers are unable to suppress transient flow crosstalk caused by the start-stop switching of other flowmeters on parallel liquid paths. Summary of the Invention
[0005] To address the problem that existing liquid mass flow controllers struggle to suppress transient flow crosstalk caused by the start-stop switching of other flow meters on parallel liquid lines, this invention provides a liquid mass flow control method. This method does not rely on a precise system mechanism model; instead, it enables dynamic compensation for disturbances through autonomous interaction between an intelligent agent and the controlled environment, and real-time learning of the optimal closed-loop control strategy. It perfectly adapts to industrial flow control scenarios characterized by nonlinearity, strong coupling, and time-varying disturbances, thus solving the problem of transient flow crosstalk caused by the start-stop switching of other flow meters on parallel liquid lines in existing liquid mass flow controllers.
[0006] The technical solution adopted by this invention to solve its technical problem is: A method for controlling the mass flow rate of a liquid includes the following steps: S1: Collect multi-source operating condition data of the target liquid circuit and its parallel liquid circuit systems; S2: Normalize and preprocess the multi-source operating condition data to obtain the current state vector; S3: Input the current state vector into the reinforcement learning agent, and have the reinforcement learning agent output the action vector used to correct the PID control parameters; S4: Based on the action vector, the original PID control parameters are corrected online to obtain adaptive PID control parameters; S5: Generate valve control signals based on the adaptive PID control parameters and drive the regulating valve of the liquid mass flow controller to compensate for transient flow crosstalk caused by the start-up or flow switching of other branches in the parallel liquid circuit.
[0007] Optionally, it also includes: S6: Collect the actual flow feedback data after compensation, and construct a reward signal based on the flow control accuracy, dynamic response speed and steady-state stability; S7: Based on the current state vector, the action vector, the reward signal, and the next state vector, iteratively update the reinforcement learning agent.
[0008] Optionally, the reinforcement learning agent adopts an Actor-Critic dual-network structure, wherein the Actor network is used to output the action vector based on the current state vector, and the Critic network is used to evaluate the control effect corresponding to the action vector.
[0009] Optionally, the multi-source operating condition data includes at least the flow rate setpoint, real-time detected flow rate, medium temperature, pipeline static pressure, controller output control quantity, and equivalent comprehensive disturbance value.
[0010] Optionally, the equivalent integrated disturbance value is calculated based on at least one of the following: the flow deviation of the target liquid path, the rate of change of pipeline pressure, the rate of change of controller output, and the operating status of parallel branches.
[0011] Optionally, the normalization preprocessing of the multi-source operating condition data includes: normalizing the multi-source operating condition data using the Min-Max linear normalization method.
[0012] Optionally, online correction of the original PID control parameters based on the action vector includes: limiting the action vector before correcting the original PID control parameters.
[0013] Optionally, the limiting process uses the tanh activation function to limit the proportional parameter correction, integral parameter correction, and derivative parameter correction within a preset range.
[0014] Optionally, the reward signal is calculated by a multi-objective coupled reward function, which includes a flow accuracy reward term, a dynamic response reward term, and a steady-state stability reward term.
[0015] Optionally, iteratively updating the reinforcement learning agent includes: S71: Store the current state vector, the action vector, the reward signal, and the next state vector as interaction samples into the experience replay pool; S72: Randomly sample interactive samples from the experience replay pool; S73: Update the Critic value network based on the sampled interaction samples; S74: Update the Actor policy network based on the evaluation results of the action value by the Critic value network; S75: Update the Actor target network and the Critic target network separately using a soft update method.
[0016] The beneficial effects of this invention are: The liquid mass flow control method provided by this invention exhibits extremely strong anti-disturbance robustness. Employing a model-free reinforcement learning architecture, it requires no precise system mechanism model and can autonomously adapt to time-varying, nonlinear, and unknown multi-coupled disturbances. Even under transient flow crosstalk (i.e., the "liquid grabbing" effect) caused by bypass flowmeter start-stop switching, the flow rate exhibits no significant fluctuations, demonstrating superior anti-disturbance stability compared to traditional PID and fuzzy PID. Its outstanding autonomous learning and adaptive capabilities allow for autonomous optimization of the control strategy through interaction between the intelligent agent and the environment, eliminating the need for manual rule design or periodic parameter tuning. It adapts to sudden changes in operating conditions and requires no subsequent maintenance. Furthermore, this control method boasts strong engineering adaptability, employing a lightweight algorithm design that can be directly embedded into existing liquid mass flow controllers without altering the hardware architecture. Its strong generalization and versatility make it widely applicable to various industrial scenarios. Attached Figure Description
[0017] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0018] Figure 1 This is a closed-loop logic block diagram of the liquid mass flow control method in this invention; Figure 2 This is a flowchart of the internal decision-making process of the reinforcement learning agent in this invention; Figure 3 This is a schematic diagram of the actual working condition of the multi-fluid path parallel coupling in this invention; Figure 4 This is the recipe generated by the liquid grabbing effect in the multi-liquid-path parallel coupling operation condition of this invention; Figure 5 This is a comparison diagram of flow control under the condition of liquid competition in the multi-liquid-path parallel coupling working condition of this invention; Figure 6 This is a comparison diagram of valve control voltages under the condition of liquid competition in the multi-liquid-path parallel coupling working condition of this invention; Figure 7 This is the reward function curve of the present invention under the condition of liquid competition in a multi-liquid-path parallel coupling operation; Figure 8 This is the PID parameter adjustment curve of the present invention under the condition of liquid competition in a multi-liquid-path parallel coupling working condition. Detailed Implementation
[0019] The present invention will now be described in further detail. The embodiments described below are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0020] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0021] To address the problem in existing liquid mass flow controllers of suppressing transient flow crosstalk caused by the start-stop switching of other flow meters in parallel liquid lines, this invention provides a liquid mass flow control method, see [link to relevant documentation]. Figure 1 As shown, the control method includes the following steps: S1: Collect multi-source operating condition data of the target liquid circuit and its parallel liquid circuit systems; This invention provides a reinforcement learning adaptive liquid flow control algorithm for multi-liquid-path parallel coupling conditions. Specifically, the target liquid path in this invention refers to the controlled liquid path that needs to maintain a predetermined mass flow rate; the parallel liquid path system refers to the set of other branches that share at least one section of supply pipeline, supply source, pump set or pressure regulating unit with the target liquid path; the start-up and shutdown of other branches, changes in set values, changes in valve opening degree or switching of pump sets may all change the pressure and flow distribution of the common pipeline section, thereby constituting an external coupling disturbance of the target liquid path.
[0022] S2: Normalize and preprocess the multi-source operating condition data to obtain the current state vector; In this step, the influence of dimensions is eliminated through normalization preprocessing, and the parameters are mapped to the [0,1] interval.
[0023] S3: Input the current state vector into the reinforcement learning agent, and the reinforcement learning agent outputs the action vector used to correct the PID control parameters; By introducing a reinforcement learning agent, the agent learns the nonlinear dynamic characteristics of transient pressure fluctuations in parallel liquid circuits through continuous interaction with the environment, extracts the crosstalk patterns of multi-channel flow, and generates action commands based on these patterns.
[0024] S4: Correct the original PID control parameters online based on the action vector to obtain adaptive PID control parameters; The limited motion vector is superimposed onto the original PID control parameters or the PID control parameters of the previous control cycle to obtain the adaptive PID control parameters for the current control cycle.
[0025] When the target fluid flow rate is detected to be rapidly decreasing due to the activation of other branches, the agent can appropriately increase the proportional action or adjust the integral action to shorten the flow recovery time. When the target fluid flow shows an overshoot or oscillation trend, the agent can reduce the proportional action, suppress integral accumulation, or enhance appropriate differential damping. These adjustment rules are not pre-written into fixed rules, but are learned by the agent from interaction samples through reward feedback.
[0026] S5: Generates valve control signals based on adaptive PID control parameters and drives the regulating valve of the liquid mass flow controller to compensate for transient flow crosstalk caused by the start-up or flow switching of other branches in the parallel liquid circuit.
[0027] The control output is calculated based on adaptive PID control parameters and converted into a valve control signal that the regulating valve can receive. The valve control signal can be analog voltage, analog current, PWM duty cycle, digital opening command, or other signals matched to the regulating valve driver. The liquid mass flow controller drives the regulating valve to change its opening to counteract the effects of common pipeline pressure changes and flow redistribution on the target liquid path.
[0028] In practical applications, when a parallel branch suddenly opens, the common liquid supply capacity is redistributed in a short time, reducing the inlet pressure or differential pressure of the target liquid path, resulting in a decrease in the actual flow rate of the target liquid path. When using the liquid mass flow control method provided by this invention, pressure changes, flow deviations, and control output changes are captured by the state vector. The intelligent agent outputs corresponding PID parameter corrections, causing the regulating valve to increase its effective opening more quickly and prompting the actual flow rate to return to the set value. When the parallel branch stops or reduces its flow, the liquid mass flow control method provided by this invention adjusts the PID parameters and valve control signal in the opposite direction to reduce the sudden increase and overshoot of the target liquid path's flow rate.
[0029] The liquid mass flow control method provided by this invention exhibits extremely strong anti-disturbance robustness. Employing a model-free reinforcement learning architecture, it requires no precise system mechanism model and can autonomously adapt to time-varying, nonlinear, and unknown multi-coupled disturbances. Even under transient flow crosstalk (i.e., the "liquid grabbing" effect) caused by bypass flowmeter start-stop switching, the flow rate exhibits no significant fluctuations, demonstrating superior anti-disturbance stability compared to traditional PID and fuzzy PID. Its outstanding autonomous learning and adaptive capabilities allow for autonomous optimization of the control strategy through interaction between the intelligent agent and the environment, eliminating the need for manual rule design or periodic parameter tuning. It adapts to sudden changes in operating conditions and requires no subsequent maintenance. Furthermore, this control method boasts strong engineering adaptability, employing a lightweight algorithm design that can be directly embedded into existing liquid mass flow controllers without altering the hardware architecture. Its strong generalization and versatility make it widely applicable to various industrial scenarios.
[0030] Furthermore, the liquid mass flow control method provided by the present invention also includes: S6: Collect the actual flow feedback data after compensation, and construct a reward signal based on the flow control accuracy, dynamic response speed and steady-state stability; After executing the control action, the actual flow rate and other state data at the next control moment are collected, and a reward signal is constructed based on the flow control accuracy, dynamic response speed and steady-state stability. S7: Iteratively update the reinforcement learning agent based on the current state vector, action vector, reward signal, and next state vector.
[0031] Specifically, the present invention preferably includes at least the multi-source operating condition data, such as the flow rate setpoint, real-time detected flow rate, medium temperature, pipeline static pressure, controller output control quantity, and equivalent comprehensive disturbance value.
[0032] In the specific implementation process, at each control moment, the overall operating status parameters are collected, that is, multi-source operating condition data of the target fluid circuit and parallel fluid circuit systems are collected, and the state space is constructed to build a six-dimensional state vector, as shown in the following formula: ; In the formula: q r Set a value for the flow rate, q f To detect the flow rate in real time, T is the medium temperature, P is the pipeline static pressure, u is the controller output control quantity, and d is the equivalent comprehensive disturbance value.
[0033] The equivalent comprehensive disturbance value can be constructed in different ways depending on the available signal; preferably, the equivalent comprehensive disturbance value is calculated based on the flow deviation of the target liquid path, the rate of change of medium temperature, the rate of change of pipeline pressure, and the rate of change of controller output. The formula is as follows: ; In the formula, q r Set a value for the flow rate, q f To monitor flow rate in real time, T is the medium temperature, P is the pipeline static pressure, u is the controller output control quantity, d is the equivalent comprehensive disturbance value, and k q k is the disturbance weighted correction factor for the flow deviation. T k is the perturbation-weighted correction factor for the rate of temperature change. P k is the disturbance-weighted correction factor for the rate of pressure change. u dT / dt is the disturbance weighted correction coefficient for the controller output rate of change, dP / dt is the rate of change of medium temperature, dP / dt is the rate of change of pipeline pressure, and du / dt is the rate of change of controller output.
[0034] When the state of parallel branches cannot be directly obtained, an equivalent comprehensive disturbance value can be constructed using only the operating data of the target fluid circuit.
[0035] Furthermore, the present invention preferably performs normalization preprocessing on multi-source operating condition data by: normalizing the multi-source operating condition data using the Min-Max linear normalization method.
[0036] The specific formula is as follows: ; In the formula: These are the normalized state parameters. These are the original collected parameters. The lower limit threshold of the parameter. This is the upper limit threshold for the parameter.
[0037] The lower and upper thresholds can be determined based on the device's range, historical operating data, or training dataset. To prevent real-time states from exceeding the training range, truncation can be performed after normalization. When the value is less than 0, the value is 0; when the value is greater than 1, the value is 1.
[0038] The present invention preferably corrects the original PID control parameters online based on the action vector by: limiting the action vector before correcting the original PID control parameters.
[0039] Specifically, the limited motion vector is superimposed onto the original PID control parameters or the PID control parameters of the previous control cycle to obtain the adaptive PID control parameters for the current control cycle.
[0040] The preferred action space construction process of this invention involves continuously fine-tuning the increments of the three main PID parameters. The action vector formula is as follows: ; In the formula, , , These are the original PID parameters.
[0041] The present invention further preferably uses the tanh activation function for the amplitude limiting process, so that the correction amounts of the proportional parameter, integral parameter, and derivative parameter are limited to a preset range.
[0042] To avoid long-term cumulative parameter deviation, a physically feasible range can be set for the corrected PID parameters. Meanwhile, to suppress extreme sudden changes in behavior, incremental correction relative to the nominal parameters or smooth updates using a moving average method can be adopted.
[0043] Specifically, the action vector is limited to the range [-0.5, 0.5] by the tanh activation function, and the real-time control parameter update formula is as follows: ; In the formula, , , These are the real-time control parameters after adaptive correction.
[0044] See Figure 2As shown, the present invention preferably employs a dual-network strategy for reinforcement learning agent iteration and updating, specifically an Actor-Critic dual-network structure, wherein the Actor network is used to output action vectors based on the current state vector, and the Critic network is used to evaluate the control effect corresponding to the action vectors.
[0045] Specifically, the Actor policy network takes the current state vector as input and outputs a continuous action vector; the Critic value network takes the state-action pair as input and outputs the action value of that action in the current state. To improve iterative stability, the agent also sets up an Actor target network and a Critic target network corresponding to the current Actor network and the current Critic network.
[0046] The Actor network output layer uses the tanh activation function to restrict the original network output to the range [-1, 1], and then scales it according to the maximum allowable correction magnitude of each PID parameter. For example, each action component can be further restricted to a preset range of [-0.5, 0.5]. For different objects, the proportional, integral, and derivative parameters can be set with different scaling factors to adapt to the magnitude of each parameter.
[0047] During the training phase, exploration noise can be superimposed on the actions output by the Actor network to increase the exploration range of the state-action space; during the actual control or verification phase, the exploration noise is removed, and the Actor network outputs action vectors deterministically.
[0048] The preferred reward signal of this invention is calculated by a multi-objective coupled reward function, wherein the multi-objective coupled reward function includes a flow accuracy reward term, a dynamic response reward term, and a steady-state stability reward term.
[0049] Specifically, considering three key indicators—flow control accuracy, dynamic response speed, and steady-state stability—a weighted continuous reward function is constructed, as shown in the following formula: ; Satisfy the weight normalization constraint: ; , , These are adjustable weighting coefficients; Flow accuracy reward sub-function: ; Dynamic response reward subfunction: ; Steady-state stable reward subfunction: ; In the formula, , , This is the proportional adjustment coefficient. For disturbance adjustment response time, The standard deviation of steady-state flow rate fluctuation. Set a value for the flow rate. To monitor traffic flow values in real time.
[0050] The present invention preferably includes iterative updates to the reinforcement learning agent, comprising: S71: Store the current state vector, action vector, reward signal and the next state vector as interaction samples into the experience replay pool; S72: Randomly sample interactive samples from the experience replay pool; The experience replay pool is pre-configured with a fixed preset storage capacity. It adopts a cyclic overlay first-in-first-out method to maintain the preset capacity while ensuring the stability of the temporal distribution of samples. A batch of samples are randomly selected from it for network training to reduce the correlation between samples at adjacent control times.
[0051] S73: Update the Critic value network based on the sampled interaction samples; S74: Update the Actor policy network based on the evaluation results of action value from the Critic value network; S75: Update the Actor target network and the Critic target network separately using a soft update method.
[0052] Specifically, the preferred value function update process of this invention is as follows: Critic network objective value function: ; In the formula, For the target Critic network parameters, For the target Actor network parameters, For actions obtained based on the target Actor network, For the next state, The value function is obtained based on the target Critic network. This is the discount factor.
[0053] The goal of the Critic network is to minimize the loss function: ; In the formula, These are the current Critic network parameters. For window length, Let i represent the i-th state and action within the window.
[0054] Strategy Update: The objective of the Actor network is to maximize the Q-function of the Critic network to optimize the policy. Its objective function is: ; These are the current Actor network parameters. This refers to the i-th action within the window, obtained based on the current Actor network.
[0055] Update the Actor network using gradient ascent: ; Target network update: A soft update method is used to gradually bring the target network parameters closer to the current network: ; ; In the formula It is a soft update coefficient.
[0056] Combined with an experience replay mechanism, it stores interaction samples between the agent and the environment, performs random sampling training, breaks sample correlation, reduces data waste, and improves training stability.
[0057] In summary, this invention proposes a reinforcement learning-based adaptive liquid flow control algorithm for multi-liquid-path parallel coupled operating conditions. Reinforcement learning, a model-free intelligent decision-making algorithm, does not rely on a precise system mechanism model. It enables the agent to autonomously interact with the controlled environment and learn the optimal closed-loop control strategy in real time, achieving dynamic compensation for disturbances. This algorithm is perfectly suited for industrial flow control scenarios with nonlinear, strongly coupled, and time-varying disturbances. This invention focuses on multi-liquid-path parallel coupled disturbance scenarios, achieving model-free adaptive decision-making, real-time coupled disturbance perception, and online control parameter self-tuning. It improves the controller's adaptive suppression capability against coupled disturbances from the algorithmic level, significantly enhancing the stability and accuracy of flow control under multi-liquid-path parallel variable operating conditions.
[0058] Compared with existing technologies, this invention features closed-loop control throughout the entire process, precise guidance from a multi-objective reward function, significantly improved response speed under disturbances, reduced flow control errors, high control precision, and fast dynamic response, thus meeting the high-precision flow control requirements of industry.
[0059] Another objective of this invention is to provide a liquid mass flow controller that employs the liquid mass flow control method described above. This liquid mass flow controller is a closed-loop control architecture designed for multi-liquid-path coupling conditions, aiming to suppress transient flow crosstalk (i.e., the "liquid-grabbing" effect) caused by the start-stop switching of other flowmeters on parallel liquid paths. The architecture includes: a multi-source operating condition data acquisition module, a state normalization preprocessing module, a reinforcement learning agent module, a PID parameter adaptive correction module, a liquid mass flow controller execution module, a real-time disturbance monitoring module, and a closed-loop flow feedback module. These modules work together in series and parallel to form a complete closed loop, achieving end-to-end adaptive disturbance rejection control encompassing perception coupling, intelligent decision-making, dynamic decoupling, and feedback optimization.
[0060] Among them, the multi-source operating condition data acquisition module focuses on collecting high-frequency data across the entire domain, such as transient flow rate of main / branch lines, pipeline pressure drop jump, fluid temperature and control input of neighboring areas, to accurately capture the imbalance characteristics caused by branch line actions; State normalization preprocessing module: Performs dimensionless standardization processing on the collected heterogeneous multi-source data to eliminate the difference in feature magnitude caused by sudden pressure changes and accelerate the subsequent network convergence. The reinforcement learning agent module includes an Actor policy network, a Critic value network, an Actor target network, a Critic target network, an experience replay pool, and a reward calculation unit. This module learns the nonlinear dynamic characteristics of transient pressure fluctuations in parallel liquid circuits through continuous interaction with the environment, extracts the crosstalk patterns of multi-channel flow, and generates action commands based on these patterns. PID parameter adaptive correction module: Receives decoupling action commands from the intelligent agent, and adjusts the underlying control gain in advance or in real time for conditions such as pressure drop or flow surge in the liquid circuit, so as to realize dynamic feedforward compensation of control parameters. Liquid mass flow controller execution module: As a control actuator, it quickly responds to adaptive PID commands and drives the fluid pipeline regulating valve to counteract the sudden flow changes caused by transient pressure field changes; Real-time disturbance monitoring module: integrates control input, output and observer information to monitor local hydraulic coupling abrupt changes and global flow field disturbances in the parallel network in real time; Closed-loop flow feedback module: Real-time sampling of actual flow and control signal after compensation, evaluation of anti-disturbance effect and generation of reward signal, forming a global decoupled control closed loop.
[0061] For ease of understanding, the present invention provides the following specific embodiments: 1. State Space Construction: Collect global operational state parameters and construct a six-dimensional state vector, as shown in the following formula: ; In the formula: qr Set a value for the flow rate, q f To detect the flow rate in real time, T is the medium temperature, P is the pipeline static pressure, u is the controller output control quantity, and d is the equivalent comprehensive disturbance value.
[0062] State normalization: Min-Max linear normalization is used to eliminate the influence of dimensions and map the parameters to the [0,1] interval, as shown in the following formula: ; In the formula: These are the normalized state parameters. These are the original collected parameters. The lower limit threshold of the parameter. This is the upper limit threshold for the parameter.
[0063] Action space construction: Select the three main parameters of the PID and continuously fine-tune the increments. The action vector formula is as follows: ; The action is limited to the range [-0.5, 0.5] by the tanh activation function. The real-time control parameter update formula is as follows: ; In the formula, , , These are the original PID parameters; , , These are the real-time control parameters after adaptive correction.
[0064] 4. Design of multi-objective coupled reward function Considering three key performance indicators—flow control accuracy, dynamic response speed, and steady-state stability—a weighted continuous reward function is constructed, as shown in the following formula: ; Satisfy the weight normalization constraint: ; , , These are adjustable weighting coefficients; Flow accuracy reward sub-function: ; Dynamic response reward subfunction: ; Steady-state stable reward subfunction: ; In the formula, , , This is the proportional adjustment coefficient. For disturbance adjustment response time, The standard deviation of steady-state flow rate fluctuation. Set a value for the flow rate. To monitor traffic flow values in real time.
[0065] 5. Dual-network strategy iteration and update The agent adopts an Actor-Critic dual-network structure: the Actor network is responsible for outputting the optimal control action, and the Critic network is responsible for evaluating the quality of the action, specifically divided into value function update, policy update and target network update.
[0066] Value function update: Critic network objective value function: ; In the formula, For the target Critic network parameters, For the target Actor network parameters, For actions obtained based on the target Actor network, For the next state, The value function is obtained based on the target Critic network. This is the discount factor.
[0067] The goal of the Critic network is to minimize the loss function: ; In the formula, These are the current Critic network parameters. For window length, Let i represent the i-th state and action within the window.
[0068] Strategy Update: The objective of the Actor network is to maximize the Q-function of the Critic network to optimize the policy. Its objective function is: ; These are the current Actor network parameters. This refers to the i-th action within the window, obtained based on the current Actor network.
[0069] Update the Actor network using gradient ascent: ; Target network update: A soft update method is used to gradually bring the target network parameters closer to the current network: ; ; In the formula It is a soft update coefficient.
[0070] A schematic diagram of the actual operating conditions under multi-fluid path coupling is shown below. Figure 3 As shown, the recipe generated by the liquid grabbing effect is as follows: Figure 4 As shown, the control effect after adopting the liquid mass flow rate control method provided by this invention is as follows: Figures 5-8 As shown in the figure; it can be seen from the figure that after adopting the liquid mass flow control method provided by the present invention, when the liquid grabbing effect occurs, the reward function will change with the change of operating conditions (e.g. Figure 7 The control parameters will be adjusted in real time according to changes (e.g.) Figure 8 ), thereby ensuring that the traffic remains stable (e.g. Figure 5 ), and the control output is smooth (e.g. Figure 6 ).
[0071] Based on the above-described preferred embodiments of the present invention, and through the foregoing description, those skilled in the art can make various changes and modifications without departing from the inventive concept. The technical scope of this invention is not limited to the contents of the specification, but must be determined according to the scope of the claims.
Claims
1. A method for controlling the mass flow rate of a liquid, characterized in that, Includes the following steps: S1: Collect multi-source operating condition data of the target liquid circuit and its parallel liquid circuit systems; S2: Normalize and preprocess the multi-source operating condition data to obtain the current state vector; S3: Input the current state vector into the reinforcement learning agent, and have the reinforcement learning agent output the action vector used to correct the PID control parameters; S4: Based on the action vector, the original PID control parameters are corrected online to obtain adaptive PID control parameters; S5: Generate valve control signals based on the adaptive PID control parameters and drive the regulating valve of the liquid mass flow controller to compensate for transient flow crosstalk caused by the start-up or flow switching of other branches in the parallel liquid circuit.
2. The liquid mass flow rate control method as described in claim 1, characterized in that, Also includes: S6: Collect the actual flow feedback data after compensation, and construct a reward signal based on the flow control accuracy, dynamic response speed and steady-state stability; S7: Based on the current state vector, the action vector, the reward signal, and the next state vector, iteratively update the reinforcement learning agent.
3. The liquid mass flow rate control method as described in claim 2, characterized in that, The reinforcement learning agent adopts an Actor-Critic dual-network structure, where the Actor network is used to output the action vector based on the current state vector, and the Critic network is used to evaluate the control effect corresponding to the action vector.
4. The liquid mass flow rate control method as described in claim 1, characterized in that, The multi-source operating condition data includes at least the flow setpoint, real-time detected flow value, medium temperature, pipeline static pressure, controller output control quantity, and equivalent comprehensive disturbance value.
5. The liquid mass flow rate control method as described in claim 4, characterized in that, The equivalent comprehensive disturbance value is calculated based on at least one of the following: the flow deviation of the target liquid path, the rate of change of pipeline pressure, the rate of change of controller output, and the operating status of parallel branches.
6. The liquid mass flow rate control method as described in claim 1, characterized in that, The normalization preprocessing of the multi-source operating condition data includes: normalizing the multi-source operating condition data using the Min-Max linear normalization method.
7. The liquid mass flow rate control method as described in claim 1, characterized in that, Online correction of the original PID control parameters based on the action vector includes: limiting the action vector before correcting the original PID control parameters.
8. The liquid mass flow rate control method as described in claim 7, characterized in that, The amplitude limiting process uses the tanh activation function to limit the proportional parameter correction, integral parameter correction, and derivative parameter correction within a preset range.
9. The liquid mass flow rate control method as described in claim 2, characterized in that, The reward signal is calculated by a multi-objective coupled reward function, which includes a flow accuracy reward item, a dynamic response reward item, and a steady-state stability reward item.
10. The liquid mass flow rate control method as described in claim 2, characterized in that, Iterative updates to the reinforcement learning agent include: S71: Store the current state vector, the action vector, the reward signal, and the next state vector as interaction samples into the experience replay pool; S72: Randomly sample interactive samples from the experience replay pool; S73: Update the Critic value network based on the sampled interaction samples; S74: Update the Actor policy network based on the evaluation results of the action value by the Critic value network; S75: Update the Actor target network and the Critic target network separately using a soft update method.