Multi-cylinder electro-hydraulic position servo system joint control method and equipment

By optimizing the coordinated control of the multi-cylinder electro-hydraulic position servo system through the DDPG algorithm and dynamic interactive network, the control problem of the multi-cylinder electro-hydraulic position servo system under complex working conditions is solved, and high-precision and fast-response multi-cylinder coordinated control is achieved.

CN120686612APending Publication Date: 2025-09-23THREE GORGES JINSHA RIVER CHUANYUN HYDROPOWER DEV CO LTD YONGSHAN XILUODU POWER PLANT +2
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510817365.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Existing multi-cylinder electro-hydraulic position servo systems face problems such as high control difficulty, large synchronization error, and slow response speed when faced with highly nonlinear dynamic characteristics and complex working conditions. Especially in multi-cylinder coordinated control, the environmental states that the intelligent agent needs to perceive and the action space it needs to perform are expanded, which increases the difficulty of reinforcement learning.

Method used

The DDPG algorithm is used to construct a cooperative control method for a multi-cylinder electro-hydraulic position servo system. By pre-training a single agent and performing joint training, combining dynamic interaction networks and attention mechanisms, optimizing the observer design, introducing target future trajectory prediction and error differential control, and designing a comprehensive reward function to improve the system's adaptability and control accuracy.

Benefits of technology

It significantly improved the convergence speed of the multi-agent system and its adaptability to complex working conditions, reduced tracking errors, improved the system's dynamic response capability and robustness, achieved high-precision control of multiple cylinders, reduced overshoot by 23.88%, and accelerated response time by 20.11%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120686612A_ABST
    Figure CN120686612A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of electro-hydraulic servo systems, in particular to a multi-cylinder electro-hydraulic position servo system joint control method and device, and the method comprises the steps: constructing a plurality of independently operating single-cylinder electro-hydraulic position servo system models; a DDPG algorithm architecture and a reward function R are configured for each model; training an intelligent agent of the DDPG algorithm; the observation data of the intelligent agent comprises a current position, a control error of the system, an accumulated error of the system and a future trajectory of the target; the trained intelligent agent is migrated to a multi-cylinder electro-hydraulic position servo system; constructing a multi-agent cooperative control framework based on a dynamic interaction network; carrying out joint training on the multiple agents; and based on the trained intelligent agent group, completing control of the multi-cylinder electro-hydraulic position servo system. Through the joint control method and equipment, the highly nonlinear dynamic characteristic of a single-cylinder electro-hydraulic position servo system can be fully dealt with, and joint actions of synchronization, following and the like of multiple cylinders are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of electro-hydraulic servo systems, and in particular to a coordinated control method and equipment for a multi-cylinder electro-hydraulic position servo system. Background Art

[0002] Multi-cylinder electro-hydraulic position servo systems offer numerous advantages, including high power-to-weight ratio, fast response, and high control precision. They are widely used in applications such as hydropower station cylinder valves and construction machinery. However, single-cylinder electro-hydraulic position servo systems inherently suffer from numerous nonlinearities, uncertainties, and external disturbances. These issues can severely hinder improvements in the dynamic and steady-state performance of single-cylinder systems. In multi-cylinder electro-hydraulic position servo systems, the nonlinearities and uncertainties of each cylinder are coupled and superimposed, increasing the difficulty of coordinated control.

[0003] Domestic and international scholars have conducted extensive research on multi-cylinder cooperative control strategies. Zhou Qiang et al. proposed a dual-fuzzy PID tracking controller based on the concept of cross-coupling. Compared with conventional fuzzy PID control, this control strategy can significantly reduce the synchronization error of a two-cylinder synchronous system under eccentric load conditions. Duan et al. proposed a control strategy combining adjacent cross-coupling and adaptive robust control. For single-cylinder control, an adaptive robust controller was used. A PID coupling controller was designed based on the adjacent cross-coupling concept. This strategy achieved high synchronization accuracy under both eccentric load and non-eccentric load conditions. Song Zhao used a feedforward-corrected integral-separated PID neural network control algorithm to control the position synchronization of two cylinders. This method can effectively suppress external interference. Wu Lianhua studied the synchronization control strategy for an isothermal forging hydraulic press, combining sliding mode control with master-slave control to improve the system's synchronization accuracy and anti-interference capability. Zhang used fuzzy PID control for a single cylinder based on a dual-cylinder parallel synchronization control structure. Experimental results showed that this strategy can reduce the synchronization error of a scissor-type dual-cylinder lifting mechanism to less than 0.66 mm, meeting practical control requirements.

[0004] In recent years, reinforcement learning has shown great potential in solving complex control problems. Compared with traditional control methods, reinforcement learning does not rely on precise system models and has strong adaptability and high control accuracy. It is particularly suitable for scenarios with strong nonlinearity and high uncertainty, such as electro-hydraulic position servo systems. Existing studies have used reinforcement learning methods to control electro-hydraulic position servo systems and have achieved certain results. For example, Xu Baochang et al. used the DQN reinforcement learning algorithm to control the proportional servo valve, thereby achieving precise control of the hydraulic throttle valve. The experimental results show that the intelligent agent's control strategy for the valve position can meet the requirements of accuracy and response speed. Yin Fengyuan used the TD3 reinforcement learning algorithm to control the electro-hydraulic servo valve single-cylinder system, verifying the advantages of this method in terms of control performance and applicability.

[0005] However, the observer dimension of the reinforcement learning methods in existing research is low, usually only containing an error signal or a combination of an error signal and an accumulated error signal, and the output actions are mostly discrete, which makes it difficult to fully cope with the highly nonlinear dynamic characteristics of the electro-hydraulic position servo system.

[0006] Moreover, as the number of cylinders increases, the system complexity increases significantly, and the environmental states that the intelligent agent needs to perceive and the action space it needs to perform also expand, which greatly increases the difficulty of reinforcement learning control. Summary of the Invention

[0007] In order to solve the above technical problems, the present invention proposes a coordinated control method and equipment for a multi-cylinder electro-hydraulic position servo system, which can fully cope with the highly nonlinear dynamic characteristics of a single-cylinder electro-hydraulic position servo system, realize the coordinated actions such as synchronization and following of multiple cylinders, and can significantly improve the convergence speed of the multi-agent system and its adaptability to complex working conditions, thereby realizing high-precision control of multiple cylinders.

[0008] The present invention is achieved by adopting the following technical solutions:

[0009] The coordinated control method of a multi-cylinder electro-hydraulic position servo system includes the following steps:

[0010] Step S1. Constructing a multi-cylinder electro-hydraulic position servo system, including constructing several independently operating single-cylinder electro-hydraulic position servo system models;

[0011] Step S2. Construct the corresponding DDPG algorithm architecture for the single-cylinder electro-hydraulic position servo system model, including defining the reward function R:

[0012] R=R e +R de +R stop ,

[0013]

[0014] R de =-|e t -e t-1 |,

[0015]

[0016] Where R e is the reward based on the error between the target position and the actual position, R de is the reward based on the error change between adjacent sampling times, R stop Rewards based on safety; t is the control error of the system at the current moment, b is a very small positive number, e t-1is the control error of the system at the last moment; emax is the maximum control error allowed, X p is the position output of the servo system, xpmax is the maximum position output allowed, and xpmin is the minimum position output allowed;

[0017] Step S3. Train the DDPG algorithm agent corresponding to the single-cylinder electro-hydraulic position servo system model; the agent's observation data includes: the current position Xp, the system's control error e t , the cumulative error of the system ∑e t and the target future trajectory S1, S2, ..., S N ;

[0018] Step S4. Migrate the agent trained in step S3 to the multi-cylinder electro-hydraulic position servo system, initialize the multi-agent cooperative training environment, and prepare for cooperative control training;

[0019] Step S5. Construct a multi-agent collaborative control framework based on a dynamic interaction network and predict the target's future trajectory S1′, S2′, ..., S N ';

[0020] Step S6. Based on the multi-agent collaborative control framework and the target future trajectory S1′, S2′, ..., S N ′, use linear and nonlinear signals to conduct cooperative training on multiple agents in a multi-cylinder system to obtain a trained group of agents;

[0021] Step S7. Based on the trained intelligent agent group, complete the control of the multi-cylinder electro-hydraulic position servo system.

[0022] The step S5 specifically includes the following steps:

[0023] Step S 51 Real-time information sharing between agents is achieved through the communication module. Each agent broadcasts its observation state and local reward signal to other agents, forming a global state information pool.

[0024] Step S 52 Using the attention mechanism, the interaction weights of each agent with other agents are dynamically calculated, prioritizing the integration of state information that is more relevant to the current working conditions.

[0025] Step S 53 Design a global synchronization optimization objective and add a negative reward term for cross-agent synchronization error to the total reward function.

[0026] The step S 52 The specific steps include:

[0027] Step S521 The input state of each agent i is defined as:

[0028]

[0029] Where x i is the current position, e i To control the error, For future trajectories;

[0030] Step S 522 For any two agents i and j, construct their relative state difference:

[0031] Δs ij =s j -s i ;

[0032] Step S 523 .Calculate the interaction weight through the attention score function:

[0033]

[0034] Among them, the attention function f(s i ,s j ) is defined as:

[0035] f(s i ,s j )=ReLU(W q s i ) Τ ReLU(W k s j ),

[0036] Where W q 、W k are the trainable parameter matrices for queries and keys respectively;

[0037] Step S 524 The fused observation input of agent i is:

[0038]

[0039] Step S 53 The total reward function in is:

[0040] R i =R local,i +λR sync,i ,

[0041]

[0042] Where R i is the total reward function, R local,iis the local reward of agent i itself, which is consistent with the reward function R defined in step S2; R sync,i is the negative reward term of synchronization error, λ is the weight coefficient of the negative reward term of synchronization error; i 、x j Represent the position outputs of the i-th and j-th agents respectively.

[0043] Target future trajectory S1, S2, ..., S N And the target future trajectory S1′, S2′, ..., S N ′ are generated by the target future trajectory prediction model based on dynamic recursive prediction and multimodal information fusion, respectively.

[0044] The target future trajectory S1, S2, ..., S N And the target future trajectory S1′, S2′, ..., S N The specific generation method of ′ includes the following steps:

[0045] Using a long short-term memory neural network, combined with historical target signals, hydraulic cylinder speed, and external disturbance data, a target future trajectory prediction model is constructed and trained.

[0046] Introducing an adaptive time window mechanism to dynamically adjust the prediction time range and step size according to real-time working conditions;

[0047] Generate the target's future trajectory based on the trained target future trajectory prediction model and the adjusted prediction time range and step size.

[0048] The construction method of the single-cylinder electro-hydraulic position servo system model includes:

[0049] Step S 11 .Construct a mathematical model of a single-cylinder electro-hydraulic position servo system and obtain the open-loop transfer function;

[0050] Step S 12 .In the simulation platform, external disturbances are added to the open-loop transfer function to construct a closed-loop transfer function model.

[0051] The open-loop transfer function G(s) is:

[0052]

[0053] Where K a is the servo amplifier proportional gain, K q Indicates the flow gain of the hydraulic cylinder slide valve, K ν Indicates the servo valve gain, K f is the displacement sensor gain, A p Indicates the piston area of ​​the hydraulic cylinder, K cIndicates the flow pressure coefficient of the slide valve, m l is the load mass, w v represents the natural frequency, ξ v Indicates the servo valve damping ratio.

[0054] During the training of the intelligent agent in step S3, a priority experience replay mechanism is also set up.

[0055] The priority experience replay mechanism dynamically adjusts the sample sampling probability according to the TD error.

[0056] The agent includes an actor network and a critic network. Step S3 of training the agent of the DDPG algorithm specifically includes the following steps:

[0057] Step S 31 .Initialize network parameters;

[0058] Step S 32 . Determine whether the iteration is complete, if not, go to step S 33 ,If so, end the training;

[0059] Step S 33 Initialize the random process, receive the initial state of the single-cylinder electro-hydraulic position servo system, and predict the target position N steps in the future;

[0060] Step S 34 . Determine whether the position control is completed. If so, go to step S 32 If not, go to step S 35 ;

[0061] Step S 35 Select action A based on strategy π and random process, and get reward R;

[0062] Step S 36 .Transfer to the next state S', observe the state of the single-cylinder electro-hydraulic position servo system and update the target position for the next N steps;

[0063] Step S 37 Store (S, A, R, S') into the priority experience replay pool and update the corresponding TD error;

[0064] Step S 38 Dynamically adjust the sample sampling probability based on the TD error and sample K samples from the priority experience replay pool;

[0065] Step S 39 . Use K samples to train the actor network and the critic network and update the network parameters; enter the next control cycle and enter step S 34 .

[0066] A multi-cylinder electro-hydraulic position servo system coordinated control device comprises at least one processor and at least one memory communicatively connected to the processor; the memory stores program instructions that can be executed by the processor; the processor calls the program instructions to execute the control method as claimed in any one of claims 1 to 9.

[0067] Compared with the prior art, the present invention has the following beneficial effects:

[0068] 1. The collaborative control method of the present invention significantly improves the convergence speed of the multi-agent system and its adaptability to complex working conditions by pre-training a single agent and adopting a joint training approach, thereby achieving high-precision control of multiple cylinders.

[0069] Specifically, the present invention adds target prediction and state information of key time nodes to the observer design to enhance the system's ability to track dynamic targets. Secondly, in terms of action output, a continuous action space design is adopted. The system output action can obtain a continuous action signal within the threshold range, thereby improving the system's control accuracy. At the same time, combined with the error differential control strategy, the DDPG algorithm's reward function introduces an error differential term as a penalty mechanism, penalizing agents that cause large changes in system error to ensure smooth action and consistency with the control objective. In addition, in the design of the DDPG algorithm's reward function, a nonlinear reward function is introduced to optimize the agent's learning efficiency and control stability through a smoothed reward mechanism.

[0070] More specifically, the reward function of the DDPG algorithm in the present invention further optimizes the control effect by comprehensively evaluating multi-dimensional performance indicators. Specifically, the reward function not only includes negative rewards for tracking error and speed change, but also introduces positive rewards for motion smoothness to reduce oscillations in the control process, combined with negative rewards for energy consumption to improve system energy efficiency, while considering the prediction error of key time nodes to enhance the system's responsiveness to dynamic targets. In addition, due to actual industrial applications, when the position output of the electro-hydraulic position servo system exceeds the allowable range, or the control error is too large, the system needs to stop urgently to prevent device damage and accidents. Therefore, this reward function also takes the safety of the system into consideration.

[0071] Experiments show that the reward function of the DDPG algorithm significantly improves the control accuracy, stability, and robustness of the intelligent agent under complex working conditions. Specifically, the reward function adopts a nonlinear design to ensure reward smoothness and show greater sensitivity to small errors, thereby avoiding reward explosion. By combining it with an error differential control strategy, it further ensures the smoothness of the action and keeps the system state consistent with the target state.

[0072] In summary, through the mutual cooperation of the above-mentioned technical solutions, the cooperative control method of the present invention can significantly suppress external disturbances, and can effectively overcome a large number of nonlinear and uncertainty problems in the system, significantly reduce tracking errors, improve the dynamic response capability and tracking accuracy of the system, and enhance the robustness of the system.

[0073] Traditional multi-cylinder control algorithms are difficult to implement and prone to severe overshoot and lag. This invention, through its multi-dimensional observation space and deep neural network fitting, effectively overcomes these challenges and achieves high-precision multi-cylinder control. Furthermore, the multi-agent DDPG algorithm proposed in this invention reduces overshoot by 23.88% and accelerates response time by 20.11% compared to the PID algorithm.

[0074] 2. In constructing a multi-agent collaborative control framework based on a dynamic interaction network, this invention forms a global state information pool and prioritizes state information with the highest relevance to the current operating conditions through interaction weights, enabling better optimization of action decisions. Furthermore, by adding a negative reward term for cross-agent synchronization error to the overall reward function, each agent is encouraged to collaboratively minimize inter-cylinder synchronization deviation while optimizing local tracking accuracy.

[0075] Specifically, the interaction weight calculation method is designed based on the attention mechanism, and according to the characteristics of the collaborative control task, it introduces state differences as the basis of similarity, thereby optimizing the effectiveness and adaptability of state fusion expression.

[0076] 3. In the present invention, the observer has a high dimension and can fully cope with the highly nonlinear dynamic characteristics of the electro-hydraulic position servo system.

[0077] 4. The present invention dynamically adjusts the prediction time range and step size according to the real-time working conditions, so as to further enhance the tracking capability of dynamic targets and make it more adaptable to the complex working environment of the electro-hydraulic position servo system.

[0078] 5. The present invention adds external disturbances to the transfer function of the electro-hydraulic position servo system, which can simulate the uncertainty and interference that occur in the control process of the electro-hydraulic position servo system in a real environment.

[0079] 6. The prioritized experience replay mechanism and the dynamic adjustment of sample sampling probability based on TD error ensure that the model focuses on learning high-value samples, improving control accuracy and stability. Combined with the reward function, the reward signal R is fed back to the DDPG network. By prioritizing samples with large learning errors, it can more effectively correct policy deviations, significantly improving the control accuracy and convergence speed of the DDPG algorithm for nonlinear systems in complex environments.

[0080] 7. Use linear and nonlinear signals to conduct collaborative training of multiple agents in a multi-cylinder system, which can comprehensively consider the characteristics of linear and nonlinear signals to meet the project's performance indicator requirements for the combination of linear and nonlinear signals such as sinusoidal signals and trapezoidal signals. BRIEF DESCRIPTION OF THE DRAWINGS

[0081] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments, wherein:

[0082] Figure 1 It is a structural schematic diagram of the present invention;

[0083] Figure 2 Schematic diagram of the training process of a single agent in the present invention;

[0084] Figure 3 Schematic diagram of the structure of the single-cylinder electro-hydraulic position servo system of the present invention;

[0085] Figure 4 is a schematic diagram of a closed-loop transfer function model in the present invention;

[0086] Figure 5 Schematic diagram of a semi-trapezoidal curve in the present invention;

[0087] Figure 6 This is a schematic diagram of the reward curve for single-agent training in the present invention;

[0088] Figure 7 Schematic diagram of the position output curve of a single agent in the present invention;

[0089] Figure 8 This is the enlarged comparison position curve of the DDPG algorithm framework and the conventional PID algorithm in the present invention;

[0090] Figure 9 Schematic diagram of the reward curves of four agents in the four-cylinder electro-hydraulic position servo system of the present invention;

[0091] Figure 10 Schematic diagram of the position output curve of the multi-agent synchronous control in the present invention;

[0092] Figure 11 Schematic diagram of the multi-intelligent follow-up control position output curve in the present invention. DETAILED DESCRIPTION

[0093] Example 1

[0094] As a basic embodiment of the present invention, the multi-cylinder electro-hydraulic position servo system coordinated control method of the present invention includes the following steps:

[0095] Step S1. Constructing a multi-cylinder electro-hydraulic position servo system, including constructing several independently running single-cylinder electro-hydraulic position servo system models.

[0096] Step S2. Construct the corresponding DDPG algorithm architecture for the single-cylinder electro-hydraulic position servo system model, including defining the reward function R:

[0097] R=R e +R de +R stop ,

[0098]

[0099] R de =-|e t -e t-1 |,

[0100]

[0101] Where R e is the reward based on the error between the target position and the actual position, R de is the reward based on the error change between adjacent sampling times, R stop Rewards based on safety; t is the control error of the system at the current moment, b is a very small positive number, e t-1 is the control error of the system at the last moment; emax is the maximum control error allowed, X p is the position output of the servo system, xpmax is the maximum position output allowed, and xpmin is the minimum position output allowed.

[0102] Step S3: Train the DDPG algorithm agent corresponding to the single-cylinder electro-hydraulic position servo system model. The agent's observation data includes: the current position Xp, the system's control error e t , the cumulative error of the system ∑e t and the target future trajectory S1, S2, ..., S N .

[0103] Step S4. Migrate the intelligent agent trained in step S3 to the multi-cylinder electro-hydraulic position servo system, initialize the multi-agent cooperative training environment, and prepare for collaborative control training.

[0104] Step S5. Construct a multi-agent collaborative control framework based on a dynamic interaction network and predict the target's future trajectory S1′, S2′, ..., S N ′.

[0105] Step S6. Based on the multi-agent collaborative control framework and the target future trajectory S1′, S2′, ..., S N′, use linear signals and nonlinear signals to conduct collaborative training on multiple agents in a multi-cylinder system to obtain a trained group of agents.

[0106] Step S7. Based on the trained intelligent agent group, complete the control of the multi-cylinder electro-hydraulic position servo system.

[0107] Example 2

[0108] As a preferred embodiment of the present invention, the multi-cylinder electro-hydraulic position servo system coordinated control method of the present invention includes the following steps:

[0109] Step S1. Constructing a multi-cylinder electro-hydraulic position servo system, including constructing several independently running single-cylinder electro-hydraulic position servo system models.

[0110] Step S2. Construct the corresponding DDPG algorithm architecture for the single-cylinder electro-hydraulic position servo system model, including defining the reward function R:

[0111] R=R e +R de +R stop ,

[0112]

[0113] R de =-|e t -e t-1 |,

[0114]

[0115] Where R e is the reward based on the error between the target position and the actual position, R de is the reward based on the error change between adjacent sampling times, R stop Rewards based on safety; t is the control error of the system at the current moment, b is a very small positive number, e t-1 is the control error of the system at the last moment; emax is the maximum control error allowed, X p is the position output of the servo system, xpmax is the maximum position output allowed, and xpmin is the minimum position output allowed.

[0116] Step S3: Train the DDPG algorithm agent corresponding to the single-cylinder electro-hydraulic position servo system model. The agent's observation data includes: the current position Xp, the system's control error e t , the cumulative error of the system ∑e t and the target future trajectory S1, S2, ..., S NAmong them, the target future trajectory S1, S2, ..., S N It is generated by a target future trajectory prediction model based on dynamic recursive prediction and multimodal information fusion.

[0117] Step S4. Migrate the intelligent agent trained in step S3 to the multi-cylinder electro-hydraulic position servo system, initialize the multi-agent cooperative training environment, and prepare for collaborative control training.

[0118] Step S5. Construct a multi-agent collaborative control framework based on a dynamic interaction network and predict the target's future trajectory S1′, S2′, ..., S N '. The construction of a multi-agent collaborative control framework based on a dynamic interaction network specifically includes the following steps:

[0119] Step S 51 .Real-time information sharing between agents is achieved through the communication module. Each agent broadcasts its observation state and local reward signal to other agents to form a global state information pool.

[0120] Step S 52 .Use the attention mechanism to dynamically calculate the interaction weights of each agent with other agents, and prioritize the integration of state information that is more relevant to the current working conditions.

[0121] Step S 53 Design a global synchronization optimization objective and add a negative reward term for cross-agent synchronization error to the total reward function.

[0122] Step S6. Based on the multi-agent collaborative control framework and the target future trajectory S1′, S2′, ..., S N ′, use linear signals and nonlinear signals to conduct collaborative training on multiple agents in a multi-cylinder system to obtain a trained group of agents.

[0123] Step S7. Based on the trained intelligent agent group, complete the control of the multi-cylinder electro-hydraulic position servo system.

[0124] Example 3

[0125] As another preferred embodiment of the present invention, the coordinated control method of the multi-cylinder electro-hydraulic position servo system of the present invention includes the following steps:

[0126] Step S1. Constructing a multi-cylinder electro-hydraulic position servo system, including constructing several independently operated single-cylinder electro-hydraulic position servo system models. The method for constructing the single-cylinder electro-hydraulic position servo system model includes:

[0127] Step S 11 .Construct the mathematical model of the single-cylinder electro-hydraulic position servo system and obtain the open-loop transfer function G(s):

[0128]

[0129] Where K a is the servo amplifier proportional gain, K q Indicates the flow gain of the hydraulic cylinder slide valve, K v Indicates the servo valve gain, K f is the displacement sensor gain, A p Indicates the piston area of ​​the hydraulic cylinder, K c Indicates the flow pressure coefficient of the slide valve, m l is the load mass, w v represents the natural frequency, ξ ν Indicates the servo valve damping ratio.

[0130] Step S 12 .In the simulation platform, external disturbances are added to the open-loop transfer function to construct a closed-loop transfer function model.

[0131] Step S2. Construct the corresponding DDPG algorithm architecture for the single-cylinder electro-hydraulic position servo system model, including defining the reward function R:

[0132] R=R e +R de +R stop ,

[0133]

[0134] R de =-|e t -e t-1 |,

[0135]

[0136] Where R e is the reward based on the error between the target position and the actual position, R de is the reward based on the error change between adjacent sampling times, R stop Rewards based on safety; t is the control error of the system at the current moment, b is a very small positive number, e t-1 is the control error of the system at the last moment; emax is the maximum control error allowed, X p is the position output of the servo system, xpmax is the maximum position output allowed, and xpmin is the minimum position output allowed.

[0137] Step S3: Train the DDPG algorithm agent corresponding to the single-cylinder electro-hydraulic position servo system model. The agent's observation data includes: the current position Xp, the system's control error e t, the cumulative error of the system ∑e t and the target future trajectory S1, S2, ..., S N . Target future trajectory S1, S2, ..., S N It is generated by a target future trajectory prediction model based on dynamic recursive prediction and multimodal information fusion. The specific generation method includes the following steps:

[0138] Using a long short-term memory neural network, combined with historical target signals, hydraulic cylinder speed, and external disturbance data, a target future trajectory prediction model is constructed and trained.

[0139] Introducing an adaptive time window mechanism to dynamically adjust the prediction time range and step size according to real-time working conditions;

[0140] Generate the target's future trajectory based on the trained target future trajectory prediction model and the adjusted prediction time range and step size.

[0141] Step S4. Migrate the intelligent agent trained in step S3 to the multi-cylinder electro-hydraulic position servo system, initialize the multi-agent cooperative training environment, and prepare for collaborative control training.

[0142] Step S5. Construct a multi-agent collaborative control framework based on a dynamic interaction network and predict the target's future trajectory S1′, S2′, ..., S N ′.

[0143] Step S6. Based on the multi-agent collaborative control framework and the target future trajectory S1′, S2′, ..., S N ′, use linear signals and nonlinear signals to conduct collaborative training on multiple agents in a multi-cylinder system to obtain a trained group of agents.

[0144] Step S7. Based on the trained intelligent agent group, complete the control of the multi-cylinder electro-hydraulic position servo system.

[0145] Example 4

[0146] As another preferred embodiment of the present invention, the multi-cylinder electro-hydraulic position servo system coordinated control method of the present invention includes the following steps:

[0147] Step S1. Constructing a multi-cylinder electro-hydraulic position servo system, including constructing several independently running single-cylinder electro-hydraulic position servo system models.

[0148] Step S2. Construct the corresponding DDPG algorithm architecture for the single-cylinder electro-hydraulic position servo system model, including defining the reward function R:

[0149] R=R e +R de +R stop ,

[0150]

[0151] R de =-|e t -e t-1 |,

[0152]

[0153] Where R e is the reward based on the error between the target position and the actual position, R de is the reward based on the error change between adjacent sampling times, R stop Rewards based on safety; t is the control error of the system at the current moment, b is a very small positive number, e t-1 is the control error of the system at the last moment; emax is the maximum control error allowed, K p is the position output of the servo system, xpmax is the maximum position output allowed, and xpmin is the minimum position output allowed.

[0154] Step S3: Train the DDPG algorithm agent corresponding to the single-cylinder electro-hydraulic position servo system model. The agent's observation data includes: the current position Xp, the system's control error e t , the cumulative error of the system ∑e t and the target future trajectory S1, S2, ..., S N The agent includes an actor network and a critic network. In step S3, training the agent of the DDPG algorithm specifically includes the following steps:

[0155] Step S 31 .Initialize network parameters;

[0156] Step S 32 . Determine whether the iteration is complete, if not, go to step S 33 ,If so, end the training;

[0157] Step S 33 Initialize the random process, receive the initial state of the single-cylinder electro-hydraulic position servo system, and predict the target position N steps in the future;

[0158] Step S 34 . Determine whether the position control is completed. If so, go to step S 32 If not, go to step S 35 ;

[0159] Step S 35 Select action A based on strategy π and random process, and get reward R;

[0160] Step S 36 .Transfer to the next state S', observe the state of the single-cylinder electro-hydraulic position servo system and update the target position for the next N steps;

[0161] Step S 37 Store (S, A, R, S') into the priority experience replay pool and update the corresponding TD error;

[0162] Step S 38 Dynamically adjust the sample sampling probability based on the TD error and sample K samples from the priority experience replay pool;

[0163] Step S 39 . Use K samples to train the actor network and the critic network and update the network parameters; enter the next control cycle and enter step S 34 .

[0164] Step S4. Migrate the intelligent agent trained in step S3 to the multi-cylinder electro-hydraulic position servo system, initialize the multi-agent cooperative training environment, and prepare for collaborative control training.

[0165] Step S5. Construct a multi-agent collaborative control framework based on a dynamic interaction network and predict the target's future trajectory S1′, S2′, ..., S N ′.

[0166] Step S6. Based on the multi-agent collaborative control framework and the target future trajectory S1′, S2′, ..., S N ′, use linear and nonlinear signals to perform cooperative training on multiple agents in a multi-cylinder system to obtain a trained agent group. Among them, the target future trajectory S1′, S2′, ..., S N The specific generation method of ′ includes the following steps:

[0167] Using a long short-term memory neural network, combined with historical target signals, hydraulic cylinder speed, and external disturbance data, a target future trajectory prediction model is constructed and trained.

[0168] Introducing an adaptive time window mechanism to dynamically adjust the prediction time range and step size according to real-time working conditions;

[0169] Generate the target's future trajectory based on the trained target future trajectory prediction model and the adjusted prediction time range and step size.

[0170] Step S7. Based on the trained intelligent agent group, complete the control of the multi-cylinder electro-hydraulic position servo system.

[0171] Example 5

[0172] As another preferred embodiment of the present invention, the present invention includes a multi-cylinder electro-hydraulic position servo system coordinated control method, refer to the attached specification Figure 1 , including the following steps:

[0173] Step S1. Construct a multi-cylinder electro-hydraulic position servo system, including constructing several independently operated single-cylinder electro-hydraulic position servo system models. Figure 3 A single-cylinder electro-hydraulic position servo system primarily consists of command, control, execution, feedback, power, and cylinder components. The command element in the system sends a command, compares it with the feedback element, and generates an error signal that is then passed to the controller for processing. Finally, a signal converter and amplifier convert this error signal into an electrical signal, which continuously controls and monitors the system. This closed-loop system ultimately enables the system to reach the desired position.

[0174] Specifically, the construction method of the single-cylinder electro-hydraulic position servo system model includes:

[0175] Step S 11 .Construct the mathematical model of the single-cylinder electro-hydraulic position servo system and obtain the open-loop transfer function G(s).

[0176] When the natural frequency of the system actuator is higher than 50Hz, the servo valve can be regarded as an oscillation link, which is generally second-order. Its mathematical model is expressed as:

[0177]

[0178] Where K ν represents the servo valve gain, ω v represents the natural frequency, ξ ν Indicates the servo valve damping ratio.

[0179] The hydraulic energy of the hydraulic cylinder oil is converted into mechanical energy of linear motion, which continuously performs linear reciprocating motion. The hydraulic output is the displacement of the piston rod, and its transfer function is:

[0180]

[0181] Where K q Indicates the flow gain of the hydraulic cylinder slide valve, A p Indicates the piston area of ​​the hydraulic cylinder, K c Indicates the flow pressure coefficient of the slide valve, m l is the load mass.

[0182] The transfer functions of the amplifier and sensor are considered as proportional links in the model, and their transfer functions are:

[0183] G3=K a ,H=K f ,

[0184] Where K a is the servo amplifier proportional gain, K f is the displacement sensor gain.

[0185] According to the above transfer function, the open-loop transfer function of the system is G(s):

[0186]

[0187] Where K a is the servo amplifier proportional gain, K q Indicates the flow gain of the hydraulic cylinder slide valve, K v Indicates the servo valve gain, K f is the displacement sensor gain, A p Indicates the piston area of ​​the hydraulic cylinder, K c Indicates the flow pressure coefficient of the slide valve, m l is the load mass, w v represents the natural frequency, ξ v Indicates the servo valve damping ratio.

[0188] The proportional gain of the servo amplifier is K a =5A / V, servo valve gain is K v =1.96×10 -3 m 3 / (A·s), displacement feedback coefficient K f =1V / m, the piston area of ​​the hydraulic cylinder is A p =1.3×10 -2 m 2 , the flow gain of the hydraulic cylinder slide valve is K q =2.01m 2 / s, the flow pressure coefficient of the slide valve is K c =4.6×10 -10 m 5 / (W·s), the natural frequency of the servo valve is w v =100Hz, the damping coefficient is ξ v =0.65.

[0189] Step S 12 In the MATLAB / Simulink simulation platform, add external disturbances to the open-loop transfer function and construct the Figure 4 The closed-loop transfer function model is shown.

[0190] Step S2: Construct a corresponding DDPG algorithm architecture for the single-cylinder electro-hydraulic position servo system model, including defining a reward function R.

[0191] The DDPG algorithm architecture is an improvement on the traditional Actor-Critic algorithm, adding a target Actor network and a target Critic network, and using the target network to soft-update the current AC network to avoid the problem of poor stability during algorithm training. The DDPG algorithm observes the current environment state S t As the input of the network, the output a of the current state is obtained by calculating the deterministic policy function μ t , and superimpose random noise N to increase the probability of exploring unknown states:

[0192] α t =μ(S t |θ μ )+N,

[0193] The Actor network uses the gradient strategy to update the neural network parameters θ in the policy function μ:

[0194]

[0195] The Critic network takes the action output by the Actor network as input and obtains the current (S t ,α t ) under the Q value, thereby guiding the Actor network parameters θ u To select high-value actions. The update formula of the critic network parameters is:

[0196] y i =r i +γQ′(S i+1 ,μ′(S i+1 |θ μ′ ))|θ Q ,

[0197]

[0198] Where y i represents the actual evaluation value calculated by the target network, S i Indicates the environmental state, a i Indicates that in S i The action selected under , μ represents the deterministic policy function, γ represents the reward decay rate, which represents the impact of the reward value of the future step on the reward value of the current step, and L represents the loss function, which is the actual value y i The sum of squared errors between the estimated value and the

[0199] Based on the DDPG algorithm architecture, a reward function is defined. The primary performance indicator of the agent-controlled electro-hydraulic position servo system is the error between the target position and the actual position. If the error is less than a certain threshold during the control process, the agent is given a positive reward. At the same time, the smaller the error, the greater the reward value given to the agent. Its reward function R e as follows:

[0200]

[0201] Where R e is the reward based on the error between the target position and the actual position, e t is the control error of the system at the current moment, and b is a very small positive number to prevent errors in logarithmic operations when the error is 0.

[0202] In the process of controlling the electro-hydraulic position servo system, the smoothness of the system output position change should be ensured as much as possible to reduce the impact and wear on the mechanical components. Therefore, if the error changes greatly between adjacent sampling times, the agent will be severely punished. Its reward function R de as follows:

[0203] R de =-|e t -e t-1 |.

[0204] Where R de is the reward based on the error change within adjacent sampling times, e t-1 is the control error of the system at the previous moment.

[0205] In addition, in industrial applications, when the position output of the electro-hydraulic servo system exceeds the allowable range or the control error is too large, the system needs to be stopped immediately to prevent damage to the device and accidents. Therefore, when designing the reward function, the safety of the system must also be considered. The reward R given based on safety stop The design is as follows:

[0206]

[0207] Where, emax is the maximum allowable control error, X p is the position output of the servo system, xpmax is the maximum position output allowed, and xpmin is the minimum position output allowed.

[0208] When R stop When the value is 50, the entire system should stop running immediately.

[0209] Therefore, the reward function R of the DDPG algorithm architecture is:

[0210] R=R e +R de +R stop .

[0211] Step S3: Training the intelligent agent of the DDPG algorithm corresponding to the single-cylinder electro-hydraulic position servo system model.

[0212] Since it is difficult to observe the corresponding intelligent agent well in conventional dimensions for the electro-hydraulic position servo system, the dimensions need to be expanded. The observation data of the intelligent agent include: the current position Xp, the control error of the system e t , the cumulative error of the system ∑e t and the target future trajectory S1, S2, ..., S N Among them, the target future trajectory S1, S2, ..., S N It is generated by a target future trajectory prediction model based on dynamic recursive prediction and multimodal information fusion. The specific generation method includes the following steps:

[0213] Step 1. Use the long short-term memory neural network to combine historical target signals, hydraulic cylinder speed, and external disturbance data to build and train a target future trajectory prediction model.

[0214] Step 2: Introduce an adaptive time window mechanism to dynamically adjust the prediction time range and step size within 2 to 5 seconds based on real-time operating conditions, such as load changes or linear disturbance intensity, to optimize prediction accuracy.

[0215] Step 3. Based on the trained target future trajectory prediction model and the adjusted prediction time range and step size, generate the target trajectory within the next 2 to 5 seconds.

[0216] During the training of the agent, a priority experience replay mechanism is set up. The priority experience replay mechanism dynamically adjusts the sample sampling probability according to the TD error. Figure 2 , the DDPG algorithm agent training specifically includes the following steps:

[0217] Step S 31 .Initialize network parameters;

[0218] Step S 32 . Determine whether the iteration is complete, if not, go to step S 33 ,If so, end the training;

[0219] Step S 33 Initialize the random process, receive the initial state of the single-cylinder electro-hydraulic position servo system, and predict the target position N steps in the future;

[0220] Step S 34. Determine whether the position control is completed. If so, go to step S 32 If not, go to step S 35 ;

[0221] Step S 35 Select action A based on strategy π and random process, and get reward R;

[0222] Step S 36 .Transfer to the next state S', observe the state of the single-cylinder electro-hydraulic position servo system and update the target position for the next N steps;

[0223] Step S 37 Store (S, A, R, S') into the priority experience replay pool and update the corresponding TD error;

[0224] Step S 38 Dynamically adjust the sample sampling probability based on the TD error and sample K samples from the priority experience replay pool;

[0225] Step S 39 . Use K samples to train the actor network and the critic network and update the network parameters; enter the next control cycle and enter step S 34 .

[0226] More specifically, in the selection of training signals, the linear and nonlinear signal characteristics are comprehensively considered to meet the project's performance indicator requirements for the combination of linear and nonlinear signals such as sinusoidal signals and trapezoidal signals.

[0227] For details, please refer to the attached manual. Figure 5 Since the trajectory of a single cylinder electro-hydraulic position servo system to complete a single stroke is a semi-trapezoidal curve, there are two key time points: t1 is the time point when the curve is about to rise, and t2 is the time point when the curve is about to flatten. Using the semi-trapezoidal curve as the input instruction, pre-train a single agent, refer to the instructions attached. Figure 6 , simulation time 35s, t1 = 5s, t2 = 15s. 117 iterations, average reward value reaches 1700.

[0228] Using the trained agent, the agent is used to control the single cylinder electro-hydraulic position servo system with disturbance, and the position output curve is obtained as shown in the attached manual. Figure 7 shown.

[0229] Refer to the instruction manual Figure 8 Compared with the traditional PID control algorithm, the DDPG control strategy of this application can greatly improve the robustness and accuracy of the algorithm. The maximum control error and average control error are reduced by 14.06% and 24.93% compared with the traditional PID algorithm; at the same time, the overshoot is also reduced by 14.06%.

[0230] Step S4. Migrate the intelligent agent trained in step S3 to the multi-cylinder electro-hydraulic position servo system, initialize the multi-agent cooperative training environment, and prepare for collaborative control training.

[0231] Step S5. Construct a multi-agent collaborative control framework based on a dynamic interaction network, and generate the target future trajectory S1′, S2′, ..., S based on the target future trajectory prediction model. N '. Among them, building a multi-agent collaborative control framework based on a dynamic interactive network specifically includes the following steps:

[0232] Step S 51 Real-time information sharing between agents is achieved through the communication module. Each agent broadcasts its observation state (including current position, control error, and predicted target future trajectory signal) and local reward signal to other agents, forming a global state information pool.

[0233] Step S 52 .Use the attention mechanism to dynamically calculate the interaction weights of each agent with other agents, and prioritize the integration of state information that is more relevant to the current working conditions to optimize action decisions.

[0234] The interaction weight calculation method is designed based on the attention mechanism. In view of the characteristics of collaborative control tasks, it introduces state differences as the basis for similarity, optimizing the effectiveness and adaptability of state fusion expression. Specifically, the attention mechanism is used to dynamically calculate the interaction weights of each agent with other agents, including the following:

[0235] The input state of each agent i is defined as:

[0236]

[0237] Among them, x i is the current position, e i To control the error, For future trajectory.

[0238] For any two agents i and j, construct their relative state difference:

[0239] Δs ij =s j -s i .

[0240] The interaction weight is calculated by the attention score function:

[0241]

[0242] Among them, the attention function f(s i,s j ) is defined as:

[0243] f(s i ,s j )=ReLU(W q s i ) Τ ReLU(W k s j ),

[0244] Among them, W q 、W k They are the trainable parameter matrices for query and key respectively; the ReLU function is used to improve nonlinear expression capabilities.

[0245] The fused observation input of agent i is:

[0246]

[0247] Step S 53 . Design a global synchronization optimization goal and add a negative reward term for cross-agent synchronization error in the total reward function; encourage each agent to optimize the local tracking accuracy while collaboratively minimizing the synchronization deviation between multiple cylinders. Among them, the total reward function R i The definition is as follows:

[0248] R i =R local,i +λR sync,i .

[0249] Among them, R local,i is the local reward of agent i itself, which is consistent with the reward function R defined in step S2. sync,i is the negative reward for the synchronization error, and λ is the weight coefficient for the negative reward for the synchronization error. It is a hyperparameter used to adjust the trade-off between synchronization and local performance. A smaller value for the negative reward for the synchronization error indicates higher synchronization between the multiple cylinders.

[0250] Specifically, the negative reward term for synchronization error is defined as:

[0251]

[0252] Among them, x i 、x j Represent the position outputs of the i-th and j-th agents respectively.

[0253] Among them, the future target displacements S1′, S2′, ..., S N ′ and the target future trajectory S1, S2, ..., S in step S3 NThe sources are the same and are all generated by the target future trajectory prediction model. The only difference is that the objects used are expanded from single agents to multiple agents, and the targets tracked by multiple agents are different.

[0254] Step S6. Based on the multi-agent collaborative control framework and the target future trajectory S1′, S2′, ..., S N ′, use linear signals and nonlinear signals to conduct collaborative training on multiple agents in a multi-cylinder system to obtain a trained group of agents.

[0255] Step S7. Based on the trained intelligent agent group, complete the control of the multi-cylinder electro-hydraulic position servo system.

[0256] The experiment was conducted using a four-cylinder system as an example.

[0257] The trained single agent was copied into four copies, each of which controlled a corresponding single-cylinder electro-hydraulic position servo system. The signal processing module was used to distribute instructions to each agent. The four agents reached a convergence state after 120 iterations. The reward curve is shown in the attached manual. Figure 9 shown.

[0258] In the synchronous control simulation experiment, multiple agents (A, B, C, D) are used to control the multi-cylinder electro-hydraulic position servo system, tracking the semi-trapezoidal curve, and obtaining the position output curve as shown in the attached manual. Figure 10 shown.

[0259] The use of multi-agent control strategies can significantly suppress external disturbances and improve the robustness of the system.

[0260] In the following control simulation experiment, by adding deviations 2, 3, and 4 to systems B, C, and D respectively, the control situation of the multi-cylinder system under different position expectations is simulated. The position output curve obtained is shown in the attached manual. Figure 11 As shown in the figure, in this scenario, since the initial states of some systems are not zero, this is equivalent to inputting a step signal into the system. This increases the control difficulty of traditional control algorithms and can easily lead to severe overshoot and lag. However, the multi-agent DDPG algorithm of the present invention, with its multidimensional observation space and deep neural network fitting, can effectively overcome this problem and achieve high-precision control of multiple cylinders.

[0261] Compared to the PID algorithm, the multi-agent DDPG algorithm reduced overshoot by 23.88% and accelerated response time by 20.11%. In the scenario of coordinated control of a multi-cylinder system, the reinforcement learning algorithm demonstrated greater adaptability and responsiveness.

[0262] Example 6

[0263] As another preferred embodiment of the present invention, the present invention includes a coordinated control device for a multi-cylinder electro-hydraulic position servo system, comprising at least one processor and at least one memory communicatively coupled to the processor. The memory stores program instructions executable by the processor. The processor, by invoking the program instructions, can execute the control method described in any one of Examples 1 to 5 above.

Claims

1. A coordinated control method for a multi-cylinder electro-hydraulic position servo system, characterized by: The following steps are involved: Step S1. Constructing a multi-cylinder electro-hydraulic position servo system, including constructing several independently operating single-cylinder electro-hydraulic position servo system models; Step S2. Construct the corresponding DDPG algorithm architecture for the single-cylinder electro-hydraulic position servo system model, including defining the reward function R: R=R e +R de +R stop , R de =-|e t -And t-1 |, Where R e is the reward based on the error between the target position and the actual position, R de is the reward based on the error change between adjacent sampling times, R stop Rewards based on safety; t is the control error of the system at the current moment, b is a very small positive number, e t-1 is the control error of the system at the last moment; emax is the maximum control error allowed, X p is the position output of the servo system, xpmax is the maximum position output allowed, and xpmin is the minimum position output allowed; Step S3. Train the DDPG algorithm agent corresponding to the single-cylinder electro-hydraulic position servo system model; the agent's observation data includes: the current position Xp, the system's control error e t , the cumulative error of the system ∑e t and the target future trajectory S1, S2, ..., S N ; Step S4. Migrate the agent trained in step S3 to the multi-cylinder electro-hydraulic position servo system, initialize the multi-agent cooperative training environment, and prepare for cooperative control training; Step S5. Construct a multi-agent collaborative control framework based on a dynamic interaction network and predict the target's future trajectory S1′, S2′, ..., S N ';; Step S6. Based on the multi-agent collaborative control framework and the target future trajectory S1′, S2′, ..., S N ′, use linear and nonlinear signals to conduct cooperative training on multiple agents in a multi-cylinder system to obtain a trained group of agents; Step S7. Based on the trained intelligent agent group, complete the control of the multi-cylinder electro-hydraulic position servo system.

2. The coordinated control method for a multi-cylinder electro-hydraulic position servo system according to claim 1, characterized in that: The step S5 specifically includes the following steps: Step S 51 Real-time information sharing between agents is achieved through the communication module. Each agent broadcasts its observation state and local reward signal to other agents, forming a global state information pool. Step S 52 Using the attention mechanism, the interaction weights of each agent with other agents are dynamically calculated, prioritizing the integration of state information that is more relevant to the current working conditions. Step S 53 Design a global synchronization optimization objective and add a negative reward term for cross-agent synchronization error to the total reward function.

3. The coordinated control method for a multi-cylinder electro-hydraulic position servo system according to claim 2, characterized in that: The step S 52 The specific steps include: Step S 521 The input state of each agent i is defined as: Where x i is the current position, e i To control the error, For future trajectories; Step S 522 For any two agents i and j, construct their relative state difference: Δs ij =s j -s i ; Step S 523 .Calculate the interaction weight through the attention score function: Among them, the attention function f(s i ,s j ) is defined as: f(s i ,s j )=ReLU(W q s i ) Τ ·ReLU(W k s j ), Where W q 、W k are the trainable parameter matrices for queries and keys respectively; Step S 524 The fused observation input of agent i is:

4. The coordinated control method for a multi-cylinder electro-hydraulic position servo system according to claim 2 or 3, characterized in that: Step S 53 The total reward function in is: R i =R local,i +λR sync,i , Where R i is the total reward function, R local,i is the local reward of agent i itself, which is consistent with the reward function R defined in step S2; R sync,i is the negative reward term of synchronization error, λ is the weight coefficient of the negative reward term of synchronization error; i 、x j Represent the position outputs of the i-th and j-th agents respectively.

5. The coordinated control method for a multi-cylinder electro-hydraulic position servo system according to claim 2, characterized in that: Target future trajectory S1, S2, ..., S N And the target future trajectory S1′, S2′, ..., S N ′ are generated by the target future trajectory prediction model based on dynamic recursive prediction and multimodal information fusion, respectively.

6. The coordinated control method for a multi-cylinder electro-hydraulic position servo system according to claim 5, characterized in that: The target future trajectory S1, S2, ..., S N And the target future trajectory S1′, S2′, ..., S N The specific generation method of ′ includes the following steps: Using a long short-term memory neural network, combined with historical target signals, hydraulic cylinder speed, and external disturbance data, a target future trajectory prediction model is constructed and trained. Introducing an adaptive time window mechanism to dynamically adjust the prediction time range and step size according to real-time working conditions; Generate the target's future trajectory based on the trained target future trajectory prediction model and the adjusted prediction time range and step size.

7. The coordinated control method for a multi-cylinder electro-hydraulic position servo system according to claim 2 or 6, characterized in that: The construction method of the single-cylinder electro-hydraulic position servo system model includes: Step S 11 .Construct a mathematical model of a single-cylinder electro-hydraulic position servo system and obtain the open-loop transfer function; Step S 12 .In the simulation platform, external disturbances are added to the open-loop transfer function to construct a closed-loop transfer function model.

8. The coordinated control method for a multi-cylinder electro-hydraulic position servo system according to claim 7, characterized in that: The open-loop transfer function G(s) is: Where K a is the servo amplifier proportional gain, K q Indicates the flow gain of the hydraulic cylinder slide valve, K v Indicates the servo valve gain, K f is the displacement sensor gain, A p Indicates the piston area of ​​the hydraulic cylinder, K c Indicates the flow pressure coefficient of the slide valve, m l is the load mass, w v represents the natural frequency, ξ v Indicates the servo valve damping ratio.

9. The coordinated control method for a multi-cylinder electro-hydraulic position servo system according to claim 2 or 6, characterized in that: During the training of the intelligent agent in step S3, a priority experience replay mechanism is also set up.

10. The coordinated control method for a multi-cylinder electro-hydraulic position servo system according to claim 9, characterized in that: The priority experience replay mechanism dynamically adjusts the sample sampling probability according to the TD error.

11. The coordinated control method for a multi-cylinder electro-hydraulic position servo system according to claim 10, characterized in that: The agent includes an actor network and a critic network. Step S3 of training the agent of the DDPG algorithm specifically includes the following steps: Step S 31 .Initialize network parameters; Step S 32 . Determine whether the iteration is complete, if not, go to step S 33 ,If so, end the training; Step S 33 Initialize the random process, receive the initial state of the single-cylinder electro-hydraulic position servo system, and predict the target position N steps in the future; Step S 34 . Determine whether the position control is completed. If so, go to step S 32 If not, go to step S 35 ; Step S 35 Select action A based on strategy π and random process, and get reward R; Step S 36 .Transfer to the next state S', observe the state of the single-cylinder electro-hydraulic position servo system and update the target position for the next N steps; Step S 37 Store (S, A, R, S') into the priority experience replay pool and update the corresponding TD error; Step S 38 Dynamically adjust the sample sampling probability based on the TD error and sample K samples from the priority experience replay pool; Step S 39 . Use K samples to train the actor network and the critic network and update the network parameters; enter the next control cycle and enter step S 34 .

12. Multi-cylinder electro-hydraulic position servo system coordinated control equipment, characterized by: The system comprises at least one processor and at least one memory in communication with the processor; the memory stores program instructions that can be executed by the processor; and the processor calls the program instructions to execute the control method according to claim 1.

Citation Information

Cited By

  • Electro-hydraulic servo building robot and transfer learning adaptive control method thereof

    CN122323218A

  • Electro-hydraulic servo building robot and migration learning adaptive control method thereof

    CN122323218B