Pneumatic valve intelligent control method based on deep reinforcement learning

By constructing a multi-physical coupled pneumatic valve dynamic model and dual-depth Q network control strategy, the problems of response hysteresis and high energy consumption of traditional pneumatic valves under complex operating conditions are solved, and intelligent control with high precision, low energy consumption and strong robustness are achieved.

CN120428561APending Publication Date: 2025-08-05WENZHOU POLYTECHNIC
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510563542.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

Traditional pneumatic valve control strategies have lagged responses, difficult to adapt to parameters, high energy consumption under complex operating conditions, and difficult to accurately describe the true behavior of pneumatic valves in non-steady environments, resulting in poor stability and convergence of control strategies.

Method used

Establish a pneumatic valve dynamic model covering mechanical, electromagnetic, fluid and thermal coupling characteristics, define multi-dimensional state space, and adopt a dual-depth Q network design control strategy to output electromagnetic control currents regulated by nonlinear compression and pressure difference to achieve high-precision, low energy consumption, and strong robustness control.

Benefits of technology

It significantly improves the intelligence level and adaptability of the pneumatic valve control system, realizes high-precision control of complex working conditions, reduces energy consumption, and improves the robustness and engineering practicality of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120428561A_ABST
    Figure CN120428561A_ABST
Patent Text Reader

Abstract

The invention discloses a pneumatic valve intelligent control method based on deep reinforcement learning, and relates to the technical field of automatic control. The method comprises the following steps: step 1, establishing a dynamic model of the pneumatic valve according to real-time operation parameters in a multi-physical coupling behavior of the pneumatic valve; 2, defining a state space of the pneumatic valve control system according to the kinetic model; 3, inputting the state space into a pre-established deep reinforcement learning model, and outputting a real-time electromagnetic control current for the pneumatic valve; the deep reinforcement learning model is a double-deep Q network and comprises two same deep Q networks so as to relieve an over-estimation problem; each deep Q network is provided with a single hidden layer, and the weight and the bias of the hidden layer are correspondingly equal to the weight and the bias of the output layer respectively. According to the method, the problems of response lag, difficulty in parameter self-adaption and high energy consumption of a traditional control strategy under a complex working condition are solved, and the intelligent level of a control system is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of automatic control technology, and in particular to an intelligent control method for pneumatic valves based on deep reinforcement learning. Background Art

[0002] As key executive units in industrial automation control systems, pneumatic valves are widely used in a variety of fields, including fluid transportation, process regulation, and environmental control. In particular, they play core roles in the chemical, food, semiconductor, energy, and manufacturing industries, including flow regulation, pressure control, and start-stop switching. Traditional pneumatic valves primarily rely on proportional control, on-off control, or analog control based on electromagnetic drive. Their control strategies are often based on classical control theory, such as proportional-integral-derivative (PID) control, fuzzy control, and empirical mapping tables. While these methods have clear structures and low implementation costs, and exhibit good control effects under static or low-complexity operating conditions, they still present numerous issues in complex operating conditions, high dynamic response, and strongly coupled systems, including response lag, difficulty in parameter tuning, poor system robustness, and inability to adapt.

[0003] In pneumatic control systems, the dynamic behavior of pneumatic valves is influenced not only by structural parameters (such as spring stiffness, damping coefficient, and valve core mass), but also by the interactive coupling of multiple physical quantities, including fluid physical properties (such as gas density, specific heat ratio, and temperature), upstream and downstream pressure differentials, flow area changes, control chamber pressure evolution, and valve core electromagnetic response. Especially under conditions of high-frequency start-stop or severe load disturbances, the valve core experiences motion inertia, nonlinear friction, and gas thermal expansion, making dynamic modeling and real-time control of the system a significant challenge. Existing literature on pneumatic valve modeling often employs linear approximations or simplifies control chamber behavior. For example, these approaches treat the relationship between valve core displacement and flow rate as a single-input, single-output linear mapping, ignoring the effects of temperature on pressure and the delay of heat conduction on dynamic response. This makes it difficult to accurately describe the true behavior of pneumatic valves in non-steady-state environments, leading to the accumulation of model errors and compromising the stability and convergence of control strategies. Summary of the Invention

[0004] The present application provides an intelligent control method for pneumatic valves based on deep reinforcement learning. By constructing a pneumatic valve dynamics model encompassing mechanical, electromagnetic, fluid, and thermodynamic coupling characteristics, a multidimensional state space is defined that includes both real-time state and historical information. A control strategy is designed based on a dual-deep Q network, outputting an electromagnetic control current regulated by nonlinear compression and pressure differential, achieving high-precision, low-energy, and robust control of the pneumatic valve. This method addresses the issues of traditional control strategies such as delayed response, difficulty in parameter adaptation, and high energy consumption under complex operating conditions, significantly improving the control system's intelligence, adaptability, and engineering practicality.

[0005] The present invention provides a method for intelligently controlling a pneumatic valve based on deep reinforcement learning. The method includes:

[0006] Step 1: Establish a dynamic model of the pneumatic valve based on the real-time operating parameters in the multi-physics coupled behavior of the pneumatic valve;

[0007] Step 2: Define the state space of the pneumatic valve control system based on the dynamic model;

[0008] Step 3: Input the state space into a pre-established deep reinforcement learning model, and output the real-time electromagnetic control current for the pneumatic valve; the deep reinforcement learning model is a dual deep Q network, including two identical deep Q networks to alleviate the over-estimation problem; each deep Q network has a single hidden layer, and the weights and biases of the hidden layer are respectively equal to the weights and biases of the output layer; the output layer of one deep Q network is connected to the hidden layer of the other deep Q network; a nonlinear adjustment mechanism is added to the dual deep Q network to compress the action space of the pneumatic valve control system through a hyperbolic tangent function, and then adjust the amplitude according to the real-time pressure difference, so that the electromagnetic control current maintains the response sensitivity to the pneumatic valve control system without exceeding the working limit of the pneumatic valve control system.

[0009] Furthermore, the real-time operating parameters include: valve core displacement x(t); control chamber pressure p c (t); valve flow Q v (t); Gas temperature in valve T v (t); valve core mass m; spring stiffness k s (t); damping coefficient c d ;Effective piston area A e ;Electromagnetic driving force F mag (t); fluid force F flow (t); gas specific heat ratio γ; gas constant R; control chamber volume V c ; Flow coefficient C d ; Flow area A flow ; Upstream gas density ρ1; Upstream pressure p1(t); Downstream pressure p2(t); Heat transfer coefficient α; Convection heat transfer coefficient h c ;Heat exchange area A heat Specific heat at constant volume c v Specific heat capacity at constant pressure c p ; Gas mass in the cavity m gas ; Inlet temperature T1; Wall temperature T w (t); t is time.

[0010] Furthermore, the convective heat transfer coefficient h c The value range of is 50 to 200; the value range of heat transfer coefficient α is 30 to 200; the flow coefficient Cd The value range is 0.6 to 0.9.

[0011] Furthermore, the dynamic model of the pneumatic valve is expressed using the following formula:

[0012]

[0013] in, is the second derivative of x(t) with respect to time; is the first derivative of x(t) with respect to time; For p c (t) first derivative with respect to time; Q v (t) first derivative with respect to time; T v (t) The first derivative with respect to time.

[0014] Furthermore, the state space of the pneumatic valve control system is:

[0015]

[0016] Among them, H x (t) is the historical data vector of the valve core position, H x (t) = [x(t-Δt), x(t-

[0017] 2Δt), ..., x(t-nΔt)]; Δt is the time step; n is the number of historical time steps; H p (t) is the pressure history data vector, H p (t)=[p c (t-Δt), p c (t-2Δt), ..., p c (t-nΔt)];s t is the state space at time t.

[0018] Furthermore, the reward function R(s) of the deep reinforcement learning model of the dual-depth Q network t , a t , s t+Δt )for:

[0019]

[0020] Among them, i(t) is the electromagnetic control current; i max is the maximum allowable current; p max is the maximum pressure difference between upstream and downstream; s t+Δt is the state space of the next time step; p +λ x +λ e +λo =1; where λ p is a set positive number, indicating the pressure control weight; x is a set positive number, indicating the position control weight; e is a set positive number, indicating the energy consumption control weight; o is a set positive number, indicating the pressure difference control weight; a t is the action space at time t, a t =i(t)+Δi(t); where Δi(t) is the electromagnetic control current adjustment amount.

[0021] Furthermore, the loss function of the deep reinforcement learning model of the dual deep Q network is for:

[0022]

[0023] Where θ is the network parameter of the deep reinforcement learning model of the dual deep Q network, including the weights and biases of the hidden layer or output layer of each deep Q network; Ω(s t ,p1(t)-p2(t)) is the pressure difference weight; γ is the discount factor; Q θ (·) represents the state-action value function of the deep reinforcement learning model; Indicates that in the sample (s t , a t ,r,s t+Δt ) From the experience replay buffer Find the expectation of this expression under the condition of sampling in .

[0024] Furthermore, the pressure difference weight Ω(s, p1(t)-p2(t)) is expressed using the following formula:

[0025]

[0026] Furthermore, the real-time electromagnetic control current I(t) is calculated using the following formula:

[0027]

[0028] This application provides a pneumatic valve intelligent control method based on deep reinforcement learning, which has the following beneficial effects: By establishing a pneumatic valve dynamics model based on multi-physics coupling characteristics, the present invention achieves joint modeling of valve core motion, electromagnetic actuation, fluid dynamics, and thermodynamic behavior. Unlike traditional models that only consider changes in air pressure or flow, this model constructs a complete system of nonlinear differential equations from multiple perspectives, including structural mechanics, electromagnetic mechanics, compressible gas flow, and heat exchange, effectively characterizing the dynamic behavior of the pneumatic valve under different operating conditions. By modeling multiple variables, including valve core displacement, pressure, flow, and temperature, the system modeling accuracy is improved, making the reinforcement learning model training more closely aligned with real-world physical behavior, laying a solid foundation for the convergence and generalization of subsequent intelligent control strategies. Secondly, based on this physical model, the present invention constructs a well-structured and comprehensive state-space representation, covering key variables such as current valve core position, velocity, control chamber pressure, upstream and downstream pressure differential, flow, and gas temperature. It also introduces a time series vector containing historical states, enhancing the ability to perceive system inertia and temporal changes. By fusing historical displacement and pressure data, the state space constructed by the present invention not only provides a richer information dimension, but also achieves effective modeling of system hysteresis and non-Markov characteristics, enabling deep reinforcement learning strategies to obtain stronger predictive and regulatory capabilities in non-stationary environments. In terms of control strategy, the present invention innovatively designs a dual-depth Q network structure, using two neural networks with symmetrical structures, responsible for action selection and Q value estimation respectively, thereby effectively alleviating the problem of over-estimation of action value in traditional deep Q networks. Furthermore, a cross-linking mechanism between the output layer and the hidden layer is introduced into the network structure, which improves the information fusion capability of the network, allowing the network to simultaneously combine feature information at different levels when evaluating action value, thereby improving the stability and accuracy of the strategy output. In addition, in order to improve the engineering feasibility of controlling current output, the present invention introduces a hyperbolic tangent function and a pressure difference normalization adjustment mechanism to perform nonlinear compression and dynamic amplitude adjustment on the action output by reinforcement learning, ensuring that the current output can meet the response speed requirements without exceeding the physical limits of the electromagnetic coil, thereby achieving physical controllability and safety boundary consistency of the control action. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] The following detailed description of the specific embodiments of the present application in conjunction with the accompanying drawings will make the technical solutions and other beneficial effects of the present application apparent.

[0030] Figure 1 A schematic diagram of a method flow chart of a pneumatic valve intelligent control method based on deep reinforcement learning provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0031] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts are within the scope of protection of this application.

[0032] Example 1: Reference Figure 1 , a pneumatic valve intelligent control method based on deep reinforcement learning, the method comprising:

[0033] Step 1: Establish a dynamic model of the pneumatic valve based on the real-time operating parameters in the multi-physics coupled behavior of the pneumatic valve;

[0034] Specifically, the actuator structure of a pneumatic valve includes components such as the valve core, spring, electromagnet, and valve body flow channel. Its core control behavior stems from the axial movement of the valve core within the valve cavity. This movement is determined by the combined action of multiple forces, including the electromagnetic driving force, spring force, fluid pressure, and friction damping. Therefore, these forces must be incorporated into the model in the form of dynamic differential equations. Furthermore, the electromagnetic control current exhibits highly nonlinear responses to the electromagnetic driving force, particularly in the low-current range, exhibiting hysteresis. This makes traditional linear models unsuitable for high-precision control requirements. To this end, the present invention incorporates real-time electromagnetic control current as a key input variable in the modeling process, enabling the model to couple electronic control responses. Furthermore, the state of the gas within the valve cavity is not constant; it is affected by temperature fluctuations, mass flow, and cavity volume changes, exhibiting typical thermodynamic and fluid dynamic behavior. Therefore, a gas equation of state is introduced into the model, and the gas compression behavior is described using the evolution equation for the control cavity pressure. Parameters such as the heat transfer area, heat transfer coefficient, and gas specific heat capacity during heat transfer are also considered to reflect the impact of valve cavity temperature on control performance.

[0035] The fluid control behavior of the pneumatic valve is reflected in the change of mass flow driven by the upstream and downstream pressure difference. This process shows strong nonlinear characteristics and is affected by factors such as flow area, flow coefficient, gas density and thermal expansion. In the modeling, the present invention not only considers the classical gas compressible fluid equation, but also introduces a bidirectional flow expression, so that the model can maintain accuracy under both high pressure difference and low pressure difference conditions, thereby improving the adaptability and robustness of the subsequent control strategy. By uniformly expressing the above-mentioned mechanical, electromagnetic, fluid and thermal coupling behaviors in a set of dynamic nonlinear differential equations, the pneumatic valve dynamics model established by the present invention has good physical interpretation and dynamic response consistency, providing a high-dimensional, real-time, continuous and differentiable state input space for the subsequent reinforcement learning model, so that the intelligent control strategy can achieve the learning of the optimal control path with minimal trial and error cost.

[0036] Compared with the simplified models commonly used in the prior art (such as conventional linear controllers or fuzzy control models) that ignore the complex physical processes within the system and only focus on macro variables such as pressure or flow, the modeling method of the present invention starts from the microscopic mechanism and comprehensively covers the real multi-field coupling process inside the pneumatic valve. It not only greatly improves the accuracy and timeliness of the model prediction, but also greatly enhances the adaptability and stability of the control system under complex working conditions.

[0037] Step 2: Define the state space of the pneumatic valve control system based on the dynamic model;

[0038] Specifically, the core control objects of the pneumatic valve are the valve core displacement and the control chamber pressure, which together determine the gas flow path and flow distribution, and thus affect the system's regulation accuracy and response speed. Based on this, the present invention characterizes the valve core displacement, speed and pressure in the control chamber at the current moment as the main variables of the state space. These variables not only reflect the motion state of the actuator, but also reflect the direct driving result of the electromagnetic control behavior on the pneumatic system. At the same time, in order to enhance the control model's ability to predict future trends, the present invention introduces a historical state memory mechanism. By constructing a historical state vector with time steps, the valve core displacement and control pressure data of several previous moments are collected and spliced into a time-series state subspace, so that the reinforcement learning agent can perceive the dynamic change trend of the system when selecting a strategy, rather than the static response characteristics. This historical state completion mechanism effectively improves the model's ability to recognize the "inertia" and "delay" phenomena in complex dynamic systems, making the control strategy more robust and forward-looking.

[0039] In addition, as a typical fluid control unit, the function of the pneumatic valve is highly dependent on the change of the upstream and downstream pressure difference. The present invention specifically considers introducing the real-time pressure difference as an independent feature into the state space to characterize the basis of fluid energy transfer. In practical applications, pressure difference fluctuations are often a direct manifestation of gas source instability, pipeline interference or load mutation. Therefore, the introduction of this parameter not only enhances the system's ability to perceive external disturbances, but also provides data support for energy consumption optimization and pressure difference constraints in subsequent reward functions. At the same time, the valve output flow and cavity temperature are also important reflections of the system's energy conservation and thermal response. They have a complex coupling relationship with the valve core movement and control current, so they are also included in the state space to enhance the system's coordinated control capabilities among multiple physical variables.

[0040] On the basis of the above, the state space structure finally formed by the present invention is a multi-dimensional composite vector covering position, speed, pressure, temperature, flow and historical trajectory information. The construction of this state space fully integrates the ontological characteristics of the pneumatic valve structure and the physical laws of the operation process. It not only reflects the precise expression of the control target, but also takes into account the sensitive perception of disturbance factors, so that the subsequent deep reinforcement learning algorithm can have a sufficient information basis when performing strategy optimization, thereby effectively achieving high-precision, low-energy consumption, and strong robustness intelligent control goals. Furthermore, compared with the traditional pneumatic system that relies on simplified state models for PID adjustment or fuzzy control, the present invention dynamically constructs a high-dimensional state space and guides the deep Q network to perform strategy evaluation and selection based on the global state distribution. It not only breaks through the bottleneck that traditional algorithms cannot handle nonlinear strongly coupled systems, but also realizes the control model's generalized approximation ability for complex aerodynamic behaviors.

[0041] Step 3: Input the state space into a pre-established deep reinforcement learning model, and output the real-time electromagnetic control current for the pneumatic valve; the deep reinforcement learning model is a dual deep Q network, including two identical deep Q networks to alleviate the over-estimation problem; each deep Q network has a single hidden layer, and the weights and biases of the hidden layer are respectively equal to the weights and biases of the output layer; the output layer of one deep Q network is connected to the hidden layer of the other deep Q network; a nonlinear adjustment mechanism is added to the dual deep Q network to compress the action space of the pneumatic valve control system through a hyperbolic tangent function, and then adjust the amplitude according to the real-time pressure difference, so that the electromagnetic control current maintains the response sensitivity to the pneumatic valve control system without exceeding the working limit of the pneumatic valve control system.

[0042] Specifically, the present invention adopts a structurally optimized dual-depth Q network as the main control model of reinforcement learning. This structure has been systematically improved on the basis of the traditional deep Q network, aiming to solve the common problem of over-estimation of action value in reinforcement learning in high-dimensional nonlinear systems. In the present invention, the dual Q network is composed of two identical neural networks, each of which contains a single hidden layer and an output layer, and the parameter configuration of the two networks follows a strict symmetry principle to ensure that the weights and biases of the hidden layer and the output layer correspond one to one, thereby effectively reducing the deviation and noise in the policy update process during training. At the same time, in order to further enhance the feature fusion ability of the network and the stability of action value judgment, the present invention introduces a structural cascade mechanism, that is, the output layer of a Q network is directly connected to the hidden layer of another Q network to form a cross-layer feature interaction path, so that the state-action value function not only depends on the current state during training, but also can obtain auxiliary information from the complementary network, thereby improving the robustness of action evaluation and the accuracy of policy output.

[0043] To meet the requirements of pneumatic valve systems for rapid and precise control responses, the present invention also introduces a nonlinear compression and amplitude adjustment mechanism into the action output of the dual-depth Q network. Considering that in actual control systems, the solenoid coil's drive current must be limited by the hardware current limit and the valve body's structural load capacity, directly mapping the action space into raw current commands can easily lead to current saturation and even damage the actuator. To this end, the present invention compresses the action space using the hyperbolic tangent function (tanh), ensuring that it always remains within a bounded interval, thereby ensuring the stability and controllability of the output action. Furthermore, considering that the control behavior of pneumatic valves is highly dependent on the changes in the upstream and downstream pressure differential, which, in actual fluid control, determines the evolution of the valve core force state, flow rate, and gas state within the cavity, a dynamic scaling factor based on the real-time pressure differential is introduced into the compressed action output. The pressure differential value is used to perform secondary adjustment of the output action amplitude. This ensures that the control system has sufficient refinement and response smoothness when the pressure differential is small, while maintaining rapid response and control strength when the pressure differential is large, thereby achieving the goal of "consistent control effect over a wide dynamic range."

[0044] In terms of strategy learning, the present invention designs a customized reward function for the physical characteristics of pneumatic valves. This reward function comprehensively considers multiple dimensions such as the error between the system output state and the target value, current energy consumption, pressure difference balance, and the stability of the control system. Among them, the control chamber pressure and valve core displacement are the core indicators of the control target. Their errors are quantified through an exponential decay function to ensure that the network priority learning can quickly and accurately achieve the control strategy of pressure and displacement targets; the energy consumption item realizes the energy optimization of the electromagnetic system by penalizing the square term of the current, and the pressure difference item reflects the penalty mechanism for the steady-state operation of the system, avoiding fluid agitation and system instability caused by excessive upstream and downstream pressure differences. By integrating multi-objective optimization into a single reward function structure and designing adjustable weight coefficients on the parameters, this method achieves a comprehensive trade-off between the system's safety, energy consumption, and response accuracy, so that the control strategy not only pursues optimal performance, but also takes into account hardware life and energy efficiency.

[0045] During the training process, the present invention uses an experience replay mechanism to store historical state-action-reward data in a buffer, and breaks data correlation through random sampling to improve training stability; at the same time, a target network delayed update strategy is used to avoid policy oscillation, further enhancing the stability and convergence of the learning process. The design of the loss function fully considers the important influence of pressure difference on training convergence. On the basis of traditional mean square error loss, a pressure difference weight factor is introduced to increase the sensitivity to policy errors under high pressure difference conditions during training, thereby converging faster to the key state area most prone to instability in actual control, thereby improving the effectiveness of the control strategy under key conditions.

[0046] In general, step 3 not only provides an algorithmic basis for reinforcement learning in pneumatic valve control by introducing a structurally innovative dual-depth Q network model, but also realizes the intelligent agent's fine control strategy learning of high-dimensional, nonlinear, and dynamic pneumatic systems through the design of a combined compression-scaling action space and the construction of a physical-oriented reward function. At the same time, the integration of multi-level stability enhancement mechanisms in the network training and strategy optimization process enables this method to have good real-time performance, robustness, and industrial adaptability when actually deployed. Compared with the static PID control or fuzzy control logic commonly used in the prior art, the adaptive control architecture of the present invention with deep reinforcement learning as the core has completely gotten rid of the limitations of model dependence and manual parameter adjustment, and has the ability to continuously self-optimize and adapt to new working conditions based on operating data.

[0047] Set the sampling period to Δt = 10ms, the rated valve body diameter D = 12mm, the valve core mass m = 0.42kg, and the effective piston area A e =1.13×10 -4 m 2 , initial spring stiffness k s (0) = 1.8 × 10 4 N\,m -1 , damping coefficient c d =68N\,s\,m -1 , flow coefficient C d =0.78, control cavity volume V c =9.5×10 -6 m 3 , heat transfer coefficient α=110 and convection heat transfer coefficient h c =85, gas constant R = 287 J\,kg -1 K -1 , specific heat ratio γ=1.4, constant volume specific heat c v =718J\,kg -1 K -1 , constant pressure specific heat c p =1005J\,kg -1 K -1 , heat exchange area A heat =6.2×10 -4 m 2 ; Initial upstream absolute pressure p1(0) = 0.60 MPa, downstream pressure p2(0) = 0.50 MPa, upstream temperature T1 = 298 K, wall temperature T w (0) = 298K, valve core displacement x(0) = 0.25mm, which corresponds to the electromagnetic coil current i(0) = 0.30A; the desired target outlet mass flow rate is set to Q ref =3.0×10 -3 kg\,s-1 In each cycle, the controller first estimates the unmeasurable state x(t) and where y t =[Q v (t),p c (t)] T , The prediction model is based on the dynamic Discretization is given, and the measurement equation h(·) is given by the ISO6358 flow formula Closed; then construct the state vector And fed into the double-depth Q network Q θ (s t , a t ), whose reward function Take weight λ Q =0.55,λ E =0.25,λ O =0.20,i max =0.8A, p max =0.4MPa, discount factor γ=0.98, network output action a t =Δi(t)∈[-0.05,0.05]A. The observation at the first iteration t=0 gives Q v (0) = 1.1 × 10 -3 kg\,s -1 , the network calculates a0 = +0.048A, so the current is updated I(0) = 0.348A and through the tanh amplitude compression formula The actual applied I(0) = 0.346A. According to the discrete integration of the dynamic model, Δx(0) = 0.055mm is obtained, which increases the flow area to A. flow (0) = 41.5 mm 2 , the next sampling point Q v (1) = 2.55 × 10 -3 kg\,s -1 , the pressure difference dropped to 0.085MPa, and the network gave a1=+0.012A according to the new state, which increased the current to 0.358A, and the flow rate increased to 2.97×10 -3 kg\,s -1 ; After the third cycle action a2=+0.004A is fine-tuned to 0.362A, the flow rate reaches 3.02×10 -3 kg\,s -1 The set value is met, and the upstream and downstream pressure difference is stable at 0.078MPa. After that, the network output jitters within the range of ±0.002A to maintain a steady state. The average flow rate in 10s is 3.01×10 -3 kg\,s-1 The standard deviation is 0.11%, compared with the traditional PID (parameter K p =1.5, K i =12) of 0.38% fluctuation converges significantly; when t = 5s artificially introduces an inlet pressure drop of Δp1 = -0.12MPa, the algorithm reduces the flow error |Q v -Q ref |Convergence return <0.5%, while PID requires 34 cycles; average power consumption ∑i(t) for the entire 60s operation 2 R c Δt is reduced by 16.8% compared to PID (coil resistance R c =6Ω), this numerical example shows that even if the target variable is only the required flow rate Q ref Although the complete target state cannot be measured directly, the reinforcement learning controller still forms a real-time closed loop through online observation → decision-making → feedback, which can quickly approach and stably maintain the outlet flow or pressure to the target value within 0.1s, meeting industrial feasibility.

[0048] As long as the controller collects the output of the controlled object or its full representation in real time during operation, and uses this information for the next decision through feedback, a closed loop is formed. The proposed method obtains the estimated value of the valve core displacement x(t) and the displacement change rate in each sampling period. The pressure difference before and after the valve p1(t)-p2(t), the control chamber pressure p c (t), instantaneous mass flow rate Q v (t) and valve internal temperature T v (t), then combined with the historical window H x (t) and H p (t) Form the state vector s t , and immediately sent to the double-depth Q network to obtain the optimal current action a t , then applied in real time Since these observations are continuously updated, the decision in the next control cycle must depend on the system response caused by the action just applied. This information feedback link is essentially the same as the classic PID closed loop, except that the mapping from error to control quantity is done by the deep reinforcement learning strategy function Q. θ (·) replaces the linear gain formula.

[0049] In the initial state, Q v (0) = 1.1 × 10 -3 kg·s -1 Far below the set target Q ref =3.0×10 -3 kg·s -1 , the network is based on the reward function Given an increase of a0 = +0.048A, the valve core opening increases by Δx(0) = 0.055mm within Δt = 10ms after the current is updated, and the flow area is expanded to A flow (0) = 41.5 mm 2 , directly pushing the flow rate to 2.55×10 -3 kg·s -1 In the next cycle, a1 = +0.012A is generated based on the new feedback state, and the flow rate is locked to 2.97×10 -3 kg·s -1 After the third cycle fine-tuning Δi=+0.004A, the flow rate reaches and maintains 3.02×10 -3 kg·s -1 The error is less than 0.2%. The controller never requires external "target state" x in the whole process. - or (p1-p2) * Instead, it relies on the immediate deviation between the actual state quantity measured or estimated in real time and the target output, and internalizes the deviation into rewards through the reinforcement learning strategy function and converts it into executable current actions, fully reflecting the closed-loop adaptive capability. Again, the review opinion is concerned that "the valve core displacement x(t) cannot be observed, resulting in the inability to form a complete feedback." Here, the extended Kalman filter (EKF) has been used in the embodiment to convert the estimated value Incorporate into the state vector and pass the covariance matrix P t|t Online update error bounds, actual measurements show that the absolute value of the estimated error is within the typical working range It is much smaller than the full stroke of the valve core 2.5mm, and its influence on the control accuracy can be ignored. v In the control experiment of simplifying the state space by constructing the same easy-to-measure quantity, the dual DQN can still pull the flow back to the set value within 0.25 seconds, verifying that the method is robust to observation loss. Looking at interference suppression, when the embodiment artificially introduces Δp1 = -0.12MPa, the flow deviation quickly expands to 12.5%, but the reinforcement learning controller attenuates the error to within 0.5% within 12 sampling cycles, that is, 0.12 seconds, while the traditional PID requires 34 cycles. This digital quantization reflects the rapid adaptive advantage of the feedback loop. Finally, the beneficial effects are enhanced from the two dimensions of energy consumption and device life: the coil resistance R c =6Ω, integrated energy consumption ∑i(t) during the 60s test period 2 R c Δt saves 16.8% compared with PID, and the current jitter σ I=0.002A, significantly lower than the PID's 0.009A, reducing solenoid valve heating and mechanical shock. Strain gauge testing showed an approximately 9.4% increase in spring fatigue life. In summary, through explicit feedback sampling, reward-penalty coupling, real-time current action adjustment, state estimation compensation, and disturbance recovery, the dual-network collaborative learning intelligent control method for pneumatic valves fully complies with the scope of closed-loop control technology.

[0050] Example 2: Real-time operating parameters include: valve core displacement x(t); control chamber pressure p c (t); valve flow Q v (t); Gas temperature in valve T v (t); valve core mass m; spring stiffness k s (t); damping coefficient c d ;Effective piston area A e ;Electromagnetic driving force F mag (t); fluid force F flow (t); gas specific heat ratio γ; gas constant R; control chamber volume V c ; Flow coefficient C d ; Flow area A flow ; Upstream gas density ρ1; Upstream pressure p1(t); Downstream pressure p2(t); Heat transfer coefficient α; Convection heat transfer coefficient h c ;Heat exchange area A heat Specific heat at constant volume c v Specific heat capacity at constant pressure c p ; Gas mass in the cavity m gas ; Inlet temperature T1; Wall temperature T w (t); t is time.

[0051] Specifically, the valve core displacement x(t) is the most core parameter of the control execution result, which directly determines the opening of the pneumatic valve, and thus affects the flow cross-section, flow rate and output pressure response. c (t) is the main driving force of the valve core. It changes with the change of electromagnetic control current in the dynamic process. Its change rate determines the acceleration response and regulation efficiency of the valve core. v (t) is the most intuitive output performance indicator of the valve. It is affected by the valve core displacement control as well as factors such as gas pressure difference, flow area and temperature density. Introducing this parameter in the state space helps the deep Q network perceive the final control effect of the entire system and optimize the strategy function through error inverse.

[0052] The present invention also takes into account the dynamic evolution characteristics of the system structure parameters, such as the spring stiffness k s(t) is the source of the reaction force. It may change due to long-term fatigue or thermal expansion and contraction during actual operation. Introducing this parameter helps enhance the model's ability to perceive changes in the system structure. Damping coefficient c d It reflects the friction resistance between the valve core and the cavity, affecting its response stability and steady-state maintenance ability. e It is the carrier of the pneumatic force applied to the valve core, which determines the efficiency of converting pressure into mechanical displacement. mag (t) is the driving effect generated by the external input control quantity. Its size is closely related to the control current and is the direct execution force for achieving action regulation. The fluid force F flow (t) reflects the disturbance caused by the airflow impacting the valve core, which constitutes a disturbance input to the balance and dynamic response of the valve core. This disturbance is particularly significant under high pressure difference or high flow rate.

[0053] In terms of thermodynamic parameters, gas specific heat ratio γ, gas constant R, control cavity volume V c , heat transfer coefficient α, convection heat transfer coefficient h c , heat exchange area A heat , constant volume specific heat capacity c v Specific heat capacity at constant pressure c p Variables such as ∠ and ∠ jointly determine the energy change of the cavity gas during the flow, compression and heat exchange process. The mass of the gas in the cavity m gas , intake air temperature T1, wall temperature T w (t) and the current valve temperature T v (t) constitutes the internal thermal balance relationship and is a key factor affecting the rate of pressure change and the thermal inertia of the system. In high-frequency control scenarios, these parameters have a significant impact on the stability of system response and the safety of the control strategy. Especially in the case of ambient temperature fluctuations or unstable gas source temperature, ignoring the thermal effect will seriously weaken the control accuracy. In addition, the upstream gas density ρ1, upstream pressure p1(t), and downstream pressure p2(t) jointly determine the pressure difference at both ends of the system. It is the potential energy basis for driving fluid flow and is also the core parameter for pressure difference regulation and energy consumption management. Flow area A flow and flow coefficient C d It determines the maximum achievable flow capacity under the action of pressure differential and is the parameter support of the flow model. In complex control environments, dynamic adjustment is especially required to adapt to different valve structures and working conditions.

[0054] Example 3: Convective heat transfer coefficient h c The value range of is 50 to 200; the value range of heat transfer coefficient α is 30 to 200; the flow coefficient C d The value range is 0.6 to 0.9.

[0055] Specifically, the convective heat transfer coefficient h cIt is a key indicator to measure the heat exchange capacity between the gas in the valve cavity and the valve body wall. During the operation of the pneumatic valve, due to the frequent changes between high-speed compression and expansion of the gas, its temperature will fluctuate significantly with the flow rate and pressure difference, resulting in unstable heat exchange between the inner wall of the valve body and the cavity gas. Set h c The value range is 50 to 200W / (m 2 ·K), covering the typical operating range from low-speed natural convection (<100) to high-speed forced convection (≥200), ensuring the universal adaptability of the model in actual industrial environments. c The h value is suitable for the state of high gas viscosity or slow flow rate, reflecting that the heat conduction process is mainly molecular heat diffusion; while the higher h c The value corresponds to high-speed turbulent conditions, reflecting the dominant characteristics of convective heat transfer, and is a typical thermal response performance under high-frequency switching or large pressure difference regulation conditions.

[0056] Secondly, the heat transfer coefficient α is a more macroscopic composite heat conduction performance indicator. It not only depends on the convective heat transfer capacity, but also takes into account the combined effects of material thermal conductivity, wall thickness, and internal and external thermal resistance. In the thermal model of the pneumatic valve, α directly determines whether the system can quickly respond to changes in gas temperature and complete thermal balance regulation. It is set in the range of 30 to 200 W / (m 2 ·K), and h c This method forms an effective pairing and maintains consistency with the heat transfer capacity of common industrial metal valve bodies (such as aluminum alloys and stainless steel) in the range of room temperature to medium and high temperatures, effectively ensuring the equivalence of the heat exchange model and the real structure. In particular, during deep reinforcement learning training, if the dynamic influence of α is ignored, the control strategy may misjudge the valve body cooling or heating response speed, thereby affecting the optimization path of the electromagnetic control current. Therefore, by clarifying the value range of α, the present invention enhances the physical interpretability and parameter stability of the model.

[0057] Finally, the flow coefficient C d It is a dimensionless parameter that describes the flow efficiency of gas when passing through a throttle hole or valve port, reflecting the degree of loss between the actual flow and the ideal flow. d It is usually affected by the flow channel structure, surface roughness, opening ratio and pressure difference nonlinearity, and its typical engineering range is between 0.6 and 0.9. d The value means that the flow loss is large, the vortex is significant, and the valve port throttling ability is strong; while a higher C d A value of 0.6 indicates high flow efficiency, low resistance loss, and suitability for fast-response control. By setting this parameter between 0.6 and 0.9, the present invention not only ensures the stability of the flow model but also provides an adjustable parameter window for reinforcement learning. This allows for adaptation to different valve types and operating conditions during strategy evaluation, achieving model versatility and control strategy transferability.

[0058] Example 4: The dynamic model of the pneumatic valve is expressed using the following formula:

[0059]

[0060] in, is the second derivative of x(t) with respect to time; is the first derivative of x(t) with respect to time; For p c (t) first derivative with respect to time; Q v (t) first derivative with respect to time; T v (t) The first derivative with respect to time.

[0061] Specifically, on the left side of the equation, four core state derivatives are arranged, which are the valve core mass m multiplied by the valve core acceleration The mechanical motion equation expressed as Valve flow rate change over time And the rate of change of gas temperature in the valve over time These four quantities encompass the key dynamic dimensions of pneumatic valve operation: First, the displacement and acceleration of the valve core directly determine the valve opening and response speed, and are the source of macroscopic execution. Second, changes in control chamber pressure are important intermediate variables for controlling valve core motion, continuously adjusted through the combined action of electromagnetic and fluid forces. Third, valve flow is the most critical output quantity of a pneumatic valve, determining the system's fluid delivery capacity and process regulation efficiency. Fourth, the evolution of the gas temperature within the valve reflects the influence of thermomechanical coupling on fluid properties and valve core motion. Ignoring thermodynamic behavior often leads to significant modeling errors, especially in scenarios with large temperature fluctuations or high-frequency switching. By centrally describing these four terms in vector form, the equation achieves a unified modeling of the system's core dynamic characteristics. This allows deep reinforcement learning to fully learn the complex coupling relationships of system behavior based on these state evolution trajectories during training. In the policy evaluation phase, state-value functions or action-value functions are used to globally assess control effectiveness, thereby outputting more accurate electromagnetic drive current commands.

[0062] On the right, each derivative corresponds to a carefully derived nonlinear expression that takes into account factors such as spring stiffness, damping, effective piston area, electromagnetic force, fluid force, fluid state equation, and thermodynamic equations, enabling the model to realistically reflect the dynamic characteristics of the pneumatic valve under various operating conditions. For the mechanical motion equation, the resultant force on the valve core is composed of several terms: -k s(t)x(t) represents the restoring force generated by the spring, which can exert a reaction proportional to the displacement after the valve core is pushed away from the initial position; Indicates the damping force caused by viscous friction or structural resistance, which can suppress high-frequency oscillations and stabilize the valve core movement; A e p c (t) converts the control chamber pressure into the pneumatic force driving the valve core through the piston area, which is the basic mechanism for achieving valve core displacement adjustment; -F flow (t) reflects the reaction force of gas flow on the valve core, which is a disturbance that cannot be ignored when the pneumatic valve is in the ventilation state; F mag (t) is the most important external input force source, generated by energizing the electromagnetic coil. By manipulating this value in real time through reinforcement learning, the valve core's motion can be altered, thereby achieving automated control of flow or pressure. Due to the mutual checks and balances and coupling between these forces, the equation often exhibits strong nonlinear characteristics over short periods of time. Traditional linearization or approximate simplification methods cannot accurately capture the system's rapid transient behavior. Therefore, a complete nonlinear formulation is employed to maintain high-precision predictions across a wide range of operating conditions.

[0063] In the pressure change equation, the core of the system describes the evolution of the control chamber pressure based on the thermodynamic state equation and the principle of mass conservation. When the valve core moves or the flow changes, the amount of gas entering or discharged changes the gas mass in the control chamber, causing the pressure to rise or fall. At the same time, if the valve core acceleration is large, the volume change rate of the piston area will also affect the compression or expansion process of the gas in the chamber. In addition, the real-time change of the gas temperature will be affected by the gas equation. The rules further affect the pressure value. Therefore, in the corresponding formula, By first normalizing the system constants using factors such as spool displacement velocity and gas mass flow rate, the overall relationship is then integrated, ultimately resulting in a nonlinear pressure dynamic equation that can adapt to various operating conditions. This approach provides accurate pressure predictions during reinforcement learning training, enabling the policy network to rapidly adjust its actions based on deviations between target pressure and spool position, ensuring that the desired pressure level and response speed are maintained despite load disturbances.

[0064] In the flow and temperature change equation, the present invention incorporates the energy conservation principle and the convection heat transfer effect into the modeling perspective, fully considering the coupling effect of valve core movement, flow coefficient and heat transfer coefficient on the temperature of the gas in the valve. When new gas flows into the valve cavity at temperature T1, or high-temperature gas is discharged, the temperature of the gas in the entire cavity will inevitably evolve accordingly, and the wall temperature T w (t) and gas temperature T v The difference in (t) will drive the heat transfer coefficient α and the convective heat transfer coefficient h c Factors such as α and α play a role, thereby accelerating or slowing down the change of gas temperature.c A heat The heat exchange channel between the wall and the gas is described in the form of equal products, which can not only reflect the rapid heat exchange between the gas and the valve body wall when the system is switched at high speed, but also reflect the asymptotic and consistent thermal equilibrium process in the steady state. v (t) also plays a role in bringing in and taking away sensible heat during the intake and exhaust processes at different temperatures. By comparing Q v (t)c p T1 and Q v (t)c p T v The difference between these two items (t) can be used to intuitively calculate the net temperature change caused by fluid input and output per unit time, and then combined with the constant volume specific heat capacity c v and the total mass of the gas in the cavity m gas Normalization is performed to finally obtain the gas temperature change rate The parallel modeling of this thermal evolution with fluid mechanics, the gas equation of state, and valve core dynamics not only provides accurate multi-physics field coupling prediction capabilities for the control strategy, but also enables deep reinforcement learning to perceive the energy conversion laws behind temperature changes when making decisions, thus avoiding control failure or risks such as overcurrent and overpressure under extreme high and low temperature conditions.

[0065] Overall, this multi-equation coupled model closely links the four modules of mechanical motion, electromagnetic drive, compressible gas flow, and heat exchange by layering and superimposing the valve core motion equation, pressure change equation, flow dynamic equation, and temperature evolution equation. This model reflects the multiple interaction mechanisms of pneumatic valves in real industrial environments. Deep reinforcement learning is trained in such a high-fidelity simulation environment that covers multiple working conditions, and can gradually learn how to most reasonably control the electromagnetic control current F. mag (t), balance the spring force and fluid disturbance, take into account energy consumption, pressure difference limit and response speed, and finally obtain a control strategy that can maintain stability and efficiency under different pressure difference or load conditions. Unlike traditional models that rely on approximate linearization or simply focus on the valve core position, the multi-physics coupled nonlinear equation system of the present invention significantly improves the model's adaptability to transient processes, boundary transitions and extreme temperature fluctuations, and by using this model for large-scale sample data sampling in reinforcement learning training, it allows the intelligent agent to continuously perform trial and error optimization in a virtual environment to form a control strategy that best fits the real physical process. This not only minimizes experimental risks and hardware wear, but also lays a solid theoretical and engineering practice foundation for later strategy deployment and online updates on actual production lines.

[0066] Example 5: The state space of the pneumatic valve control system is:

[0067]

[0068] Among them, H x (t) is the historical data vector of the valve core position, H x (t) = [x(t-Δt), x(t-

[0069] 2Δt), ..., x(t-nΔt)]; Δt is the time step; n is the number of historical time steps; H p (t) is the pressure history data vector, H p (t)=[p c (t-Δt), p c (t-2Δt), ..., p c (t-nΔt)];s t is the state space at time t.

[0070] Specifically, t The first part is composed of the valve core displacement x(t) and the valve core speed The two are the core indicators that describe the mechanical motion state of the pneumatic valve, and directly determine the valve opening and adjustment response speed. The larger the displacement of the valve core, the larger the flow cross-section is generally, and the changes in flow and pressure difference are more obvious; the speed reflects whether the system is in the process of acceleration, deceleration or steady state, providing deep reinforcement learning with a sense of dynamic trends when selecting actions. The third factor is the control chamber pressure p c (t), which plays the role of applying aerodynamic force to the valve core in the pneumatic valve structure, and together with the electromagnetic force, spring force and fluid interference force, determines the instantaneous state of the valve core movement. By capturing the control chamber pressure value in real time, the intelligent agent can quickly perceive the difference in aerodynamic response caused by load changes or control signal changes, so as to adjust the control current. The fourth key parameter is the upstream and downstream pressure difference p1(t)-p2(t), which directly determines the driving force of the fluid flow. Different pressure difference working conditions often require completely different action strategies to obtain the best flow regulation efficiency while ensuring the stability of the valve core movement. The present invention incorporates the real-time pressure difference into the state space, so that the deep reinforcement learning algorithm can adaptively allocate control authority according to changes in the external fluid environment, thereby balancing the steady-state and dynamic performance of the system under large and small pressure differences.

[0071] As the most intuitive output indicator of pneumatic valves, valve flow Q v(t) is included in the fifth component of the state space, providing direct performance feedback for deep reinforcement learning in terms of reward function and action selection. Since pneumatic valves ultimately need to accurately control flow or flow rate in industrial applications, the flow value is closely related to system energy consumption, response time, process requirements, etc. Therefore, it not only represents the execution effect of the current control link, but also provides a basis for the intelligent agent to predict the system output trend and optimize it at subsequent times. The sixth variable is the gas temperature T in the valve. v (t). In high-speed variable operating conditions or processes with highly compressible gases, temperature significantly affects gas density, flow characteristics, and pressure fluctuations. Ignoring this factor often leads to system deviation accumulation or control instability under high pressure differentials or frequent starts and stops. By incorporating real-time gas temperature into the state vector, the present invention enables a deep reinforcement learning network to capture the coupling relationship between thermodynamic and fluid dynamic processes, thereby accurately adjusting the valve core opening and current output to avoid error accumulation under abnormal temperature conditions.

[0072] The more distinctive feature of the present invention is the time sequence information completion module in the state space, namely H x (t) and H p (t) Two historical data vectors. The former is composed of the valve core displacement at several moments in the past, and the latter is composed of the control chamber pressure data at several moments in the past. Together, they carry the evolution trajectory of the system in the previous period, providing deep reinforcement learning with a high-dimensional understanding of the system inertia, delay, and timing correlation in the decision-making process. Traditional control methods often use the current moment data as the core input, ignoring the continuous change characteristics of the system state at the most recent moment. Therefore, it is difficult to fully capture the nonlinear behavior of the pneumatic valve in a complex dynamic process. The present invention achieves this by using the current moment data as the core input and ignoring the continuous change characteristics of the system state at the most recent moment. t The time step Δt and the number of historical time steps n are explicitly introduced in the , and a series of sampling points {x(t-Δt), x(t-2Δt), ..., x(t-nΔt)} and {p c (t-

[0073] Δt), p c (t-2Δt), ..., p c(t-nΔt)} time series information enables deep reinforcement learning to encode these time series features as historical dependencies at the network input layer. Since pneumatic valves may switch frequently in high-frequency sections or exhibit nonlinear jitter on certain pressure platforms, only by combining multiple frames of past data can we perceive the upcoming trend turning points or potential unstable factors of the system, allowing the agent to be more cautious or initiate adjustments more quickly in the next action decision. Whether it is a small fine-tuning to keep the valve core near a specific equilibrium point, or a rapid increase or decrease in current in a large pressure difference mutation scenario, the time series history memory structure designed by the present invention can release its effect in the Q network or policy network of reinforcement learning, improve the control strategy's adaptability to transient processes and its ability to compensate for hysteresis phenomena, and also provide a more complete contextual environment for reward function evaluation.

[0074] By integrating the above variables in the state space, the present invention achieves multiple benefits at the level of deep reinforcement learning. First, s t The high-dimensional information reflects the true dynamics of the multi-field coupling of the pneumatic valve, providing sufficient input richness for feature learning of deep networks in complex scenarios. Secondly, after the introduction of historical memory, the system no longer needs to rely on external observations or manual input of previous states. Instead, the policy network internally forms an adaptive modeling capability for time series evolution, thereby more accurately judging the ideal action at the next moment. Thirdly, the system can alleviate the impact of delays and sensor noise to a certain extent. By comparing the current state with the changing trends of several past moments, it corrects measurement errors or signal fluctuations, thereby improving the robustness of the overall control. Unlike the existing technology that only uses a single variable or a few observations as input, the present invention focuses more on the comprehensive characterization of multiple physical elements during the operation of the pneumatic valve in state space design, and forms a decision feedback loop based on the data-driven characteristics of deep reinforcement learning combined with historical trajectory information. It is precisely because of this design concept that the entire intelligent control system can still maintain adaptive learning capabilities and strong control stability when encountering industrial-level actual interference such as high-speed start-stop, non-uniform flow shock, and valve body fatigue, truly achieving intelligent, refined, and comprehensive dynamic regulation of pneumatic valves.

[0075] Example 6: Reward function R(s) of the deep reinforcement learning model with dual deep Q networks t , a t , s t+Δt )for:

[0076]

[0077] Among them, i(t) is the electromagnetic control current; i max is the maximum allowable current; p max is the maximum pressure difference between upstream and downstream; s t+Δt is the state space of the next time step;p +λ x +λ e +λ o =1; where λ p is a set positive number, indicating the pressure control weight; x is a set positive number, indicating the position control weight; e is a set positive number, indicating the energy consumption control weight; o is a set positive number, indicating the pressure difference control weight; a t is the action space at time t, a t =i(t)+Δi(t); where Δi(t) is the electromagnetic control current adjustment amount.

[0078] Specifically, for the control targets of pressure and position, the present invention adopts a positive excitation term that decays in an exponential form, by c (t)-p ref (t)| and |x(t)-x ref (t)| performs a negative exponential mapping on the distance, so that the incentive value decreases rapidly as the error increases, thereby prompting the network to produce accurate pressure and valve core position tracking behavior. When the pressure or position approaches the target, the corresponding exponential term approaches 1, which provides positive incentives for reinforcement learning training and also makes the network tend to keep the error at an extremely low level. At the same time, pneumatic valves often have dual requirements of energy consumption and hardware safety in industrial environments. Therefore, current penalty terms and pressure difference penalty terms are added to the reward function to suppress excessive energy consumption and reduce fluid shock and unstable behavior caused by huge upstream and downstream pressure differences. The current penalty is based on The normalized form of reflects the energy loss during the electrification of the electromagnetic coil. The larger the current, the higher the penalty. It can not only guide the network to actively reduce the current consumption while meeting the control performance, but also play a protective role in avoiding overheating and hardware damage.

[0079] The pressure difference term is The form reflects the safety risk of the working condition. Reinforcement learning urges the system to reduce the upstream and downstream pressure difference as much as possible by reducing rewards to avoid extreme flow fluctuations. On this basis, the present invention uses λ p +λ x +λ e +λ o = 1, the four weight coefficients are coordinated and adjusted, so that the control system achieves a comprehensive balance between pressure accuracy, valve core position accuracy, energy consumption level and pressure difference balance. Since the action space of the pneumatic valve is set to i(t) + Δi(t), that is, relative increase or decrease based on the current current, deep reinforcement learning can continuously make fine adjustments based on real-time pressure, valve core displacement and energy consumption feedback when exploring strategies. When the state transfers to the next moment s t+ΔtWhen the system steady-state error is small and the energy consumption is moderate, and the upstream and downstream pressure differences are also kept in a relatively safe range, the reward function will show a relatively high positive return, thereby strengthening the network's memory and re-calling of the strategy; if the pressure deviates from the target or the valve core position deviates too much from the reference, or the current consumption and pressure difference increase excessively, the corresponding negative or weak reward will force the network to correct the strategy through multiple rounds of iterative updates. By integrating the multi-objective control demand system into a reward function, the present invention completely overcomes the drawbacks of single-objective or previous separate control, allowing the strategy network to achieve comprehensive management and balance of pneumatic valve performance under different operating conditions and load interference, thereby effectively reducing energy consumption and enhancing system safety while ensuring control accuracy.

[0080] Example 7: Loss Function of Deep Reinforcement Learning Model with Dual Deep Q Networks for:

[0081]

[0082]

[0083] Where θ is the network parameter of the deep reinforcement learning model of the dual deep Q network, including the weights and biases of the hidden layer or output layer of each deep Q network; Ω(s t , p1(t)-p2(t)) is the pressure difference weight; γ is the discount factor; Q θ (·) represents the state-action value function of the deep reinforcement learning model; Indicates that in the sample (s t , a t ,r,s t+Δt ) From the experience replay buffer Find the expectation of this expression under the condition of sampling in .

[0084] Specifically, the core structure of the loss function follows the target construction method of the deep Q network, based on the principle of the Bellman equation, through the current state s t With the current action a t The error between the corresponding actual Q value and the target Q value is calculated by the mean square error, and this is used to promote the update of the network parameter θ. The target Q value is determined by the immediate reward R(s t , a t , s t+Δt ) and the Q value of the next state under the optimal action Together, γ is a discount factor used to balance the contribution of immediate rewards and future long-term rewards. By introducing greedy estimation of target actions, the present invention can effectively alleviate the problem of policy deviation caused by overestimation of Q values in traditional Q networks, while ensuring a stable and clear directionality during policy iteration. Compared with the conventional DQN structure, the present invention introduces an additional pressure difference weight term Ω(s t , p1(t)-p2(t)), which is a key adjustment mechanism set according to the physical characteristics of the pneumatic valve control system. In actual working conditions, the operating state of the pneumatic valve is significantly affected by the upstream and downstream pressure difference. Especially when the pressure difference is large, the valve core movement is more sensitive, the system stability is more easily disturbed, and the control current output is more likely to exceed the hardware limit. Therefore, the present invention introduces the pressure difference weight function to give different optimization priorities to the training sample errors under different pressure difference conditions. Specifically, when the pressure difference |p1(t)-p2(t)| is high, the weight function Ω(·) will increase accordingly, thereby amplifying the loss value of the sample at this time, so that the network pays more attention to the accuracy and robustness of the strategy output under high pressure difference conditions during training; on the contrary, when the pressure difference is small, the network focuses more on detail adjustment and energy consumption optimization to achieve steady-state fine control. This dynamic weighting mechanism not only enhances the model's ability to identify key working conditions, but also provides a mechanism to ensure the stability and security of the strategy during actual deployment.

[0085] In addition, the expected sign in the loss function Represents the experience replay buffer By randomly sampling samples from historical interaction data and introducing the average error of multiple state-action pairs in each parameter update, the present invention can improve the generalization ability of training and make the control strategy applicable to a wider range of actual working conditions. It is worth noting that the Q involved in the loss function θ(·) It is realized by a dual-depth Q network structure, that is, two networks with symmetrical structures but independent parameters are used to alternately perform strategy selection and target estimation to alleviate the problem of Q-value estimation deviation. Since the present invention introduces a cross-linking mechanism between the output layer and the hidden layer in the network structure, as well as a nonlinear compression design of the action space, the gradient calculation of the loss function is more sensitive and high-dimensional. Without a reasonable weight modulation mechanism, it is very easy to cause gradient explosion or convergence oscillation. Therefore, in the overall design, the loss function of the present invention enhances the stability of the error backpropagation path at the theoretical level by introducing a pressure difference weight term and a strategy greedy estimation term, and at the same time integrates the boundary conditions of the pneumatic valve control into the optimization target at the physical level, thereby improving the adaptability of the control strategy to the system dynamics and physical constraints.

[0086] Example 8: The pressure difference weight Ω(s, p1(t)-p2(t)) is expressed using the following formula:

[0087]

[0088] Specifically, as an actuator that relies on electromagnetic control to achieve precise regulation of airflow, the operating state of a pneumatic valve is largely limited by changes in the upstream and downstream pressure difference. When the pressure difference |p1(t)-p2(t)| is small, the system tends to be stable, the valve core is under balanced force, and the improvement of control accuracy depends on the adjustment of subtle movements; when the pressure difference is large, the impact of the airflow is significantly enhanced, which not only makes the valve core movement subject to stronger disturbance force, resulting in a more intense and sensitive response, but also easily produces complex flow field problems such as eddy currents and oscillations. If the control strategy does not have sufficient robustness and responsiveness, it is very easy to cause control failure, equipment overload, and even physical damage. Therefore, how to guide the model to pay enough attention to this high-pressure differential "dangerous working condition" during the reinforcement learning training stage becomes the key to improving the practicality and safety of the control strategy engineering. The present invention achieves this goal through the pressure differential weight function.

[0089] The weight function is simple yet highly intuitive. The constant 1 is used as the baseline weight to ensure that the loss of all samples has at least the basic training strength; the adjustable factor γ and the pressure difference normalization term The product of |p1(t)-p2(t)| determines the loss amplification factor caused by the pressure difference change. Here, γ is a discount factor, which is usually used to adjust the importance of current rewards and future expected rewards in deep reinforcement learning, but in the present invention, it is innovatively used to amplify the training weights of high-risk working condition samples, thereby improving the model's policy fitting ability under boundary conditions. A prominent advantage of this structural design is that: under low pressure difference conditions, since |p1(t)-p2(t)| is close to zero, the weight function is close to 1, and the training process focuses more on the fine control of the policy and energy consumption optimization; in high pressure difference scenarios, the normalized pressure difference term approaches 1, and the entire weight can be significantly increased to 1+γ, which is equivalent to amplifying the training error for high-risk working condition samples in the loss function, prompting the network to converge to the optimal policy in these key areas first, thereby enhancing the stability and responsiveness of the system under extreme loads or rapid start and stop conditions in actual deployment.

[0090] Furthermore, this pressure difference weighting mechanism provides a mathematical closed loop for the dynamic adjustment of the loss function, so that the deep Q network has "physical perception ability" during the training process. In the traditional framework of reinforcement learning, the loss function usually processes all samples with a unified standard, ignoring the special importance of certain working conditions in engineering applications. The Ω(·) set by the present invention is associated with the pressure difference directly measured in the state space, so that the model can perceive the risk level of the current sample in the physical space in each round of training, and dynamically adjust the learning intensity accordingly to form a "critical and non-critical" training strategy. This mechanism significantly improves the generalization ability of the strategy, especially in industrial sites, when pneumatic valves are used in complex environments with frequent start-stops, drastic working condition switching, and high load uncertainty, the pressure difference weighting mechanism of the present invention can avoid the risk of the strategy overfitting to the ideal working conditions to the greatest extent, and truly realize adaptive intelligent control of complex real-world systems.

[0091] In addition, the design of Ω(·) is also highly controllable and adjustable during the deployment phase. Engineers can adjust the maximum pressure difference p of the actual system. max The system can set γ based on the desired control sensitivity, thereby fine-tuning the strategy focus during training. For example, in application scenarios with extremely high requirements for system safety, the value of γ can be increased to enhance the learning accuracy of high-pressure differential conditions; while in systems that pay more attention to energy consumption optimization, γ can be appropriately reduced to make the training more inclined towards energy-saving control optimization under low-pressure differential conditions. This parameter adjustability facilitates the flexible deployment of the present invention in multiple industries and with multiple control objectives, and also opens a key channel for deep reinforcement learning control systems to move towards practical industrial applications.

[0092] Example 9: The real-time electromagnetic control current I(t) is calculated using the following formula:

[0093]

[0094] Specifically, the first term i(t) in the formula is the current control current baseline of the system, which can also be understood as the current value at the previous moment or under the predicted expectation, representing the basic current to maintain stable operation of the system in the current state. The second term is the action a output by the reinforcement learning agent. t , the current regulation increment obtained after a series of physical-inspired nonlinear processing. Here, γ is the regulation intensity coefficient, which controls the degree of influence of the agent output on the final current; is the key nonlinear compression function, which converts the action a t Mapping to the interval (-1, 1) avoids the problem of action explosion that may occur in the reinforcement learning model during training (that is, the output amplitude is far greater than the system can withstand), and also concentrates the action adjustment capability in the "fine-tuning" category, effectively adapting to the high-precision, small-action, and high-response control requirements in industrial systems.

[0095] It should be noted that the input of the hyperbolic tangent function is Rather than simply a t This is an important improvement in the safety design of the control strategy of the present invention. By dividing the action by the current current value, it introduces the concept of "relative regulation rate" rather than "absolute action amount". This normalization processing method can be understood as "the disturbance amplitude of the current regulation behavior relative to the current state of the system", which makes the control system more cautious in the low-current operation stage, avoiding system instability due to relatively large absolute values; in the high-current operation stage, it allows a wider range of strategy exploration, improves the dynamic response capability of the system, and thus realizes the "state-aware" self-regulation of the control strategy.

[0096] On this basis, another key factor in the product term is This is a pressure differential normalized amplitude modulation function, used to dynamically scale current regulation based on the current system's gas drive capability. When the upstream and downstream pressure differential is small, the pneumatic system is in a relatively steady state, requiring less sensitivity for system response. Therefore, the current regulation amplitude should be reduced accordingly to maintain an energy-saving and low-fluctuation control strategy. When the pressure differential rises, the airflow drive capability increases, load disturbances may intensify, and system response time requirements become faster. Increasing the regulation amplitude can improve the valve core's reaction speed, enhancing the system's tracking capability and interference suppression. Using a square root function instead of linear normalization not only smooths the gain adjustment but also ensures that the system is not overly aggressive under conditions of drastic pressure differential fluctuations, thereby avoiding problems such as valve core shock or solenoid coil overload caused by excessively strong excitation signals.

[0097] From the overall structure, the present invention combines the output action a of deep reinforcement learning tThrough a set of highly physically sensitive and nonlinearly controlled function transformations, the real-time electromagnetic control current I(t) is ultimately constructed, achieving a safe mapping and performance coordination between abstract strategies and physical execution. This mechanism has significant engineering practicality: on the one hand, it transforms reinforcement learning from a mere "black-box policy fitter" into an intelligent controller capable of real-time perception of pneumatic system conditions, dynamic adjustment of actuators, and adaptive adjustment to hardware constraints. On the other hand, the formula is highly scalable and interpretable. Engineers can adjust γ, select different normalization methods, or replace the pressure differential regulation function based on the response speed and stability requirements of different systems, thereby quickly migrating this control strategy to other types of solenoid valves, flow valves, or pressure regulating devices, thus showing broad application prospects.

[0098] The embodiments of the present invention are described in detail above. Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core ideas. At the same time, for those skilled in the art, according to the ideas of the present invention, there may be changes in the specific implementation methods and application scopes. In summary, the contents of this specification should not be understood as limiting the present invention.

Claims

1. A pneumatic valve intelligent control method based on deep reinforcement learning, characterized in that: The method comprises: Step 1: Establish a dynamic model of the pneumatic valve based on the real-time operating parameters in the multi-physics coupled behavior of the pneumatic valve; Step 2: Define the state space of the pneumatic valve control system based on the dynamic model; Step 3: Input the state space into a pre-established deep reinforcement learning model, and output the real-time electromagnetic control current for the pneumatic valve; the deep reinforcement learning model is a dual deep Q network, including two identical deep Q networks to alleviate the over-estimation problem; each deep Q network has a single hidden layer, and the weights and biases of the hidden layer are respectively equal to the weights and biases of the output layer; the output layer of one deep Q network is connected to the hidden layer of the other deep Q network; a nonlinear adjustment mechanism is added to the dual deep Q network to compress the action space of the pneumatic valve control system through a hyperbolic tangent function, and then adjust the amplitude according to the real-time pressure difference, so that the electromagnetic control current maintains the response sensitivity to the pneumatic valve control system without exceeding the working limit of the pneumatic valve control system.

2. The pneumatic valve intelligent control method based on deep reinforcement learning according to claim 1, characterized in that: Real-time operating parameters include: valve core displacement x(t); control chamber pressure p c (t); valve flow Q v (t); Gas temperature in valve T v (t); valve core mass m; spring stiffness k s (t); damping coefficient c d ;Effective piston area A e ;Electromagnetic driving force F mag (t); fluid force F flow (t); gas specific heat ratio γ; gas constant R; control chamber volume V c ; Flow coefficient C d ; Flow area A flow ; Upstream gas density ρ1; Upstream pressure p1(t); Downstream pressure p2(t); Heat transfer coefficient α; Convection heat transfer coefficient h c ;Heat exchange area A heat Specific heat at constant volume c v Specific heat capacity at constant pressure c p ; Gas mass in the cavity m gas ; Inlet temperature T1; Wall temperature T w (t); t is time.

3. The pneumatic valve intelligent control method based on deep reinforcement learning according to claim 2, characterized in that: Convective heat transfer coefficient h c The value range of is 50 to 200; the value range of heat transfer coefficient α is 30 to 200; the flow coefficient C d The value range is 0.6 to 0.

9.

4. The pneumatic valve intelligent control method based on deep reinforcement learning according to claim 3 is characterized in that: The dynamic model of the pneumatic valve is expressed using the following formula: in, is the second derivative of x(t) with respect to time; is the first derivative of x(t) with respect to time; For p c (t) first derivative with respect to time; Q v (t) first derivative with respect to time; T v (t) The first derivative with respect to time.

5. The pneumatic valve intelligent control method based on deep reinforcement learning according to claim 4 is characterized in that: The state space of the pneumatic valve control system is: Among them, H x (t) is the historical data vector of the valve core position, H x (t) = [x(t-Δt), x(t-2Δt), ..., x(t-nΔt)]; Δt is the time step; n is the number of historical time steps; H p (t) is the pressure history data vector, H p (t)=[p c (t-Δt),p c (t-2Δt),...,p c (t-nΔt)];s t is the state space at time t.

6. The pneumatic valve intelligent control method based on deep reinforcement learning according to claim 5, characterized in that: The reward function R(s) of the deep reinforcement learning model with dual deep Q networks t ,a t ,s t+Δt )for: Among them, i(t) is the electromagnetic control current; i max is the maximum allowable current; p max is the maximum pressure difference between upstream and downstream; s t+Δt is the state space of the next time step; p +λ x +λ e +λ o =1; where λ p is a set positive number, indicating the pressure control weight; x is a set positive number, indicating the position control weight; e is a set positive number, indicating the energy consumption control weight; o is a set positive number, indicating the pressure difference control weight; a t is the action space at time t, a t =i(t)+Δi(t); where Δi(t) is the electromagnetic control current adjustment amount.

7. The pneumatic valve intelligent control method based on deep reinforcement learning according to claim 6, characterized in that: Loss function of deep reinforcement learning model with dual deep Q-network for: Where θ is the network parameter of the deep reinforcement learning model of the dual deep Q network, including the weights and biases of the hidden layer or output layer of each deep Q network; Ω(s t ,p1(t)-p2(t)) is the pressure difference weight; γ is the discount factor; Q θ (·) represents the state-action value function of the deep reinforcement learning model; Indicates that in the sample (s t ,a t ,r,s t+Δt ) From the experience replay buffer Find the expectation of this expression under the condition of sampling in .

8. The pneumatic valve intelligent control method based on deep reinforcement learning according to claim 7, characterized in that: The pressure difference weight Ω(s,p1(t)-p2(t)) is expressed using the following formula:

9. The pneumatic valve intelligent control method based on deep reinforcement learning according to claim 8, characterized in that: The real-time electromagnetic control current I(t) is calculated using the following formula:

Citation Information

Cited By

  • Self-adaptive control method, device and equipment of vacuum regulating valve and storage medium

    CN121209294A

  • Intelligent control system and method for degrading composite smelly substances in water through vacuum ultraviolet catalytic oxidation

    CN121757974A

  • Semiconductor valve fluid simulation method based on deep reinforcement learning fusion

    CN121881927A