A method for adaptive pressure control of an electrically driven support based on reinforcement learning

CN122812686APending Publication Date: 2026-09-25TIANJIN HUANING ELECTRONICS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611247735.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-18
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0005]本发明是为了克服现有技术中传统支护压力控制方法无法适应复杂工况、无法进行多目标优化、依赖人工经验的问题,提供一种基于强化学习的电驱动支架自适应压力控制方法,能够实现根据实时工况自主决策最优支护压力,实现一架一策精准支护的控制方法

Benefits of technology

(1)本发明实现了从自动控制到自主决策的跨越:使液压支架拥有了类似人类专家的决策智能,能够根据实时工况进行一架一策、一时一策的精准、自适应支护,是实现无人化工作面的关键一步。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122812686A_ABST
    Figure CN122812686A_ABST
Patent Text Reader

Abstract

The application discloses a kind of electric drive support adaptive pressure control methods based on reinforcement learning, it is related to artificial intelligence and advanced control engineering field, including based on hydraulic support and its working environment Construction multi-physics digital twin model;In offline state, through multi-physics digital twin model let deep reinforcement learning controller interact with virtual hydraulic support, generate neural network model;Neural network model is deployed to the edge controller of the hydraulic support;When working online, edge controller real-time acquisition hydraulic support's state vector, and neural network model calculates optimal target pressure value through state vector and sends to frequency conversion station controller execution.The application provides a kind of electric drive support adaptive pressure control methods based on reinforcement learning, can realize optimal support pressure according to real-time working condition autonomous decision, realize the control method of one strategy precision support.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence and advanced control engineering, specifically to an adaptive pressure control method for an electrically driven support based on reinforcement learning. Background Technology

[0002] The core mission of hydraulic supports is to effectively support the roof, and precise control of their support force is the cornerstone of ensuring safe and efficient mining at the working face. While the introduction of electrically driven integrated supports has enabled independent control of a single support, setting the optimal support pressure remains a significant challenge. Traditional support pressure control methods, such as "constant pressure control" (maintaining a fixed high pressure at all times) or simple "follow-the-machine control" (setting a fixed pressure after the coal mining machine has passed), are essentially passive, statically rule-based controls.

[0003] The existing technology has the following problems: (1) Inability to adapt to complex and variable geological conditions: The pressure patterns of the underground roof, rock strata structure, and fault distribution are extremely complex, variable, and highly nonlinear. Fixed pressure settings cannot adapt to these changes, leading to widespread over-support or under-support phenomena. Over-support wastes energy and may even damage the roof; under-support leaves safety hazards. (2) Lack of multi-objective collaborative optimization capability: The optimal support not only needs to ensure safety, but also needs to take into account multiple objectives such as energy saving, equipment life and coal mining efficiency. Traditional methods cannot achieve a dynamic balance among these conflicting objectives.

[0004] (3) Relying on human experience, true autonomy cannot be achieved: At present, the adjustment of support parameters still largely depends on the experience judgment of on-site engineers. This method has a slow response, poor consistency, and cannot be promoted in large-scale support groups, and is far from reaching the level of "autonomous intelligence". Summary of the Invention

[0005] This invention aims to overcome the problems of traditional support pressure control methods in the prior art, such as inability to adapt to complex working conditions, inability to perform multi-objective optimization, and reliance on human experience. It provides an adaptive pressure control method for electrically driven supports based on reinforcement learning, which can realize the autonomous decision-making of the optimal support pressure according to real-time working conditions, and achieve a precise support control method with one policy for each support.

[0006] This invention provides an adaptive pressure control method for an electrically driven support based on reinforcement learning, comprising: A multi-physics digital twin model was constructed based on the hydraulic support and its working environment; In offline mode, a deep reinforcement learning controller based on a deep deterministic policy gradient algorithm interacts with a virtual hydraulic support through a multiphysics digital twin model. The deep deterministic policy gradient algorithm, based on the Actor-Critic framework, outputs a continuous target pressure adjustment ΔP in a continuous action space. The deep reinforcement learning controller takes a multi-dimensional state vector S, including the hydraulic support column pressure, adjacent support column pressure, roof separation instrument reading, microseismic sensor signal strength, relative distance between the hydraulic support and the coal mining machine, motor power, and oil temperature, as input. It uses a reward function R, which is a weighted sum of a safety reward function r1, an energy-saving reward function r2, and a stability reward function r3, as the optimization objective, generating a neural network model including a policy network U and a value network Q. The policy network U is used to decide "what to do," including deterministic actions A, and the value network Q is used to evaluate "how well it is done," including the reward function R. Deploy the neural network model into the edge controller of the hydraulic support; When working online, the edge controller collects the state vector of the hydraulic support in real time. The neural network model calculates the optimal target pressure value through the state vector and sends it to the frequency converter controller for execution.

[0007] The Deep Deterministic Policy Gradient (DDPG) algorithm is a classic algorithm for solving reinforcement learning problems in continuous action spaces. It combines deep neural networks with the concept of deterministic policy gradients and is widely used in fields such as robot control, autonomous driving, and industrial automation.

[0008] The algorithm is based on the Actor-Critic framework and contains two core networks: Policy network (Actor network): Outputs deterministic actions that directly determine the agent's behavior.

[0009] Value network (Critic network): Evaluates the Q-value of the current state and action, guiding the optimization direction of the Actor network.

[0010] During offline training in a digital twin environment, the DDPG agent interacts with a virtual scaffold millions of times. In each interaction, the agent selects an action based on the state, receives a reward from the environment, and is given the next state. The DDPG agent continuously updates the weights of its Actor and Critic networks based on this experience. Through massive training, the Actor network is able to output the optimal action policy in any state.

[0011] When working online, the controller collects the state vector of the support in real time and inputs it into the Actor network. The network calculates the current optimal target pressure value instantly (in milliseconds) and sends it to the underlying variable frequency pump station controller for execution.

[0012] The adaptive pressure control method for an electrically driven support based on reinforcement learning, as a preferred embodiment, includes the following steps: S1. Construction of a multiphysics digital twin model: A multiphysics digital twin model is constructed based on the hydraulic support and its working environment; S2. Construction of the neural network model: In offline mode, a deep reinforcement learning controller with deep deterministic policy gradient algorithm as its core interacts with the virtual hydraulic support through a multi-physics digital twin model to generate a neural network model. S3. Deep reinforcement learning of neural network models: In each interaction, the deep reinforcement learning controller selects a state vector as input to the neural network model, updates the weights of the policy network and the value network, and obtains the neural network model after training. S4. Deployment of the neural network model: Deploy the trained neural network model to the edge controller of each hydraulic support downhole; S5. Online control: The edge controller collects state vectors, and the trained neural network model calculates the optimal target pressure value and sends it to the frequency converter controller for execution.

[0013] The construction method of multiphysics digital twin model is as follows: (1) Modeling of the mechanical field of the top plate Based on geological exploration data from the working face, a three-dimensional elastoplastic mechanical model of the roof strata was established using the finite element method (FEM). The roof strata in the model were divided into three layers: the immediate roof, the main roof, and the overlying load layer. The rock mechanical parameters (elastic modulus E, Poisson's ratio ν, and uniaxial compressive strength UCS) of each layer were obtained from borehole core test reports. The model boundary conditions were set as follows: the goaf side was a free surface, the solid coal side of the two roadways was a fixed constraint, and the overlying strata were subjected to an equivalent uniformly distributed load. As the coal mining machine advanced, the model dynamically updated the goaf extent, simulated the bending deformation and periodic fracture process of the roof cantilever beam in real time, and output the predicted values ​​of roof subsidence and incoming pressure load at each support location.

[0014] (2) Hydraulic fluid field modeling Based on the hydraulic system schematic, a fluid-mechanical coupling simulation model of the hydraulic support column was established using AMESim software. The model includes: a variable frequency motor-hydraulic pump unit (input is motor speed N, output is flow rate Q and pressure P), the column hydraulic cylinder (piston area, sealing friction, leakage coefficient), safety valves, and hydraulically controlled check valves, among other hydraulic components. The model parameters were calibrated using bench test data, accurately simulating the dynamic response of the column pressure under different roof loads. The simulation step size was set to 10ms, consistent with the sampling period of the actual controller.

[0015] (3) Modeling of the kinematic field of the coal mining machine Based on the actual operating trajectory data of the coal mining machine, a three-dimensional kinematic model of the coal mining machine is established to describe the changes in the spatial position, cutting speed, and cutting depth of the coal mining machine on the working face over time. The model takes the real-time position data uploaded by the coal mining machine's encoder as input and outputs the relative distance between the coal mining machine and each hydraulic support as one of the input components of the state vector S.

[0016] The three-field coupling method involves using the pressure load output from the roof mechanics field as the external excitation input to the hydraulic fluid field, and the column pressure output from the hydraulic fluid field as the boundary condition for the support reaction force of the roof mechanics field. These two fields form a bidirectional coupled iterative solution. The position information output from the coal mining machine kinematics field synchronously updates the goaf boundary of the roof mechanics field, achieving collaborative simulation of the three fields. The digital twin model's simulation speed is 100 times faster than real-time, capable of completing 1000 hours of operational simulation within 10 minutes, providing ample offline training data for the DDPG algorithm.

[0017] The adaptive pressure control method for an electrically driven support based on reinforcement learning described in this invention, as a preferred embodiment, includes a working environment comprising the roof, floor, and coal mining machine in step S1.

[0018] The present invention discloses an adaptive pressure control method for an electrically driven support based on reinforcement learning. In a preferred embodiment, the multiphysics digital twin model in step S1 is used to simulate the periodicity and randomness of roof pressure, the mechanical behavior of the interaction between the hydraulic support and the surrounding rock, and the dynamic response of the hydraulic system.

[0019] The adaptive pressure control method for an electrically driven support based on reinforcement learning described in this invention, as a preferred embodiment, involves a neural network model M in step S2 modeled using a Markov decision process, including: Where M is the neural network model, S is the state vector, A is the deterministic action, and R is the reward function.

[0020] The present invention discloses an adaptive pressure control method for electrically driven supports based on reinforcement learning. In a preferred embodiment, the state vector S is a set of multi-dimensional vectors describing the current situation, including the column pressure of the hydraulic support, the column pressure of adjacent hydraulic supports, the reading of the roof separation instrument, the signal strength of the micro-vibration sensor, the relative distance between the hydraulic support and the coal mining machine, the motor power, and the oil temperature.

[0021] In the present invention, an adaptive pressure control method for an electrically driven support based on reinforcement learning is preferred in which the deterministic action A is the target pressure adjustment amount ΔP output by the edge controller.

[0022] The adaptive pressure control method for an electrically driven support based on reinforcement learning described in this invention, as a preferred embodiment, has the following reward function R: Where R is the reward function, r1 is the safety reward function, r2 is the energy-saving reward function, r3 is the stability reward function, and a, b, and c are weighting coefficients; When the column pressure P of the hydraulic support is within the safe range [P1, P2], the safety reward function r1=C1, C1>0, where C1 is the first safety reward value; When the column pressure of the hydraulic support is P < P1 or P > P2, the safety reward function is r1 = -C2, C2 > C1, where C2 is the second safety reward value. Energy saving reward function r2=-C3×P t / P0, where C3 is the energy-saving bonus value, P t P0 is the instantaneous power of the pumping station, and P1 is the standard power of the pumping station. The stable reward function is r3 = -C4 × |P(t) - P(t-1)| / P(t-1), where C4 is the stable reward value, P(t) is the column pressure of the hydraulic support at time t, and P(t-1) is the column pressure of the hydraulic support at time t-1.

[0023] The adaptive pressure control method for an electrically driven support based on reinforcement learning described in this invention, as a preferred embodiment, has a policy network U whose input is a state vector S and whose output is a deterministic action A. The input to the value network Q is the state vector S and the deterministic action A, and the output is the expected value q.

[0024] The adaptive pressure control method for an electrically driven support based on reinforcement learning described in this invention, as a preferred embodiment, further includes an updated policy network and an updated value network in the neural network model. The policy network is updated by policy gradient, with the goal of enabling the deterministic action A output by the policy network to obtain a higher expected value q. The value network is updated by minimizing the loss function, with the goal of making the output of the value network approximate the true expected value q.

[0025] The present invention has the following beneficial effects: (1) This invention realizes the leap from automatic control to autonomous decision-making: it enables hydraulic supports to have decision-making intelligence similar to human experts, and can carry out precise and adaptive support according to real-time working conditions, which is a key step in realizing unmanned working face.

[0026] (2) Multi-objective optimization of the present invention under the premise of ensuring absolute safety: Through a carefully designed reward function, the method can automatically find the best balance point between multiple objectives such as energy consumption, equipment stability, and support efficiency while ensuring support effect (safety). The comprehensive benefits far exceed those of traditional methods, and the energy saving effect can reach more than 15%.

[0027] (3) This invention solves the core problem of industrial AI implementation: through the innovative paradigm of "digital twin offline training + online deployment", it effectively solves the core pain points of reinforcement learning in real industrial scenarios, such as difficulty in obtaining massive interactive data, high trial and error costs, and inability to guarantee security, and provides a feasible engineering path for the application of AI in the field of heavy equipment. Attached Figure Description

[0028] Figure 1 This is an adaptive pressure control method for an electrically driven support based on reinforcement learning. Detailed Implementation

[0029] Example 1 like Figure 1 As shown, an adaptive pressure control method for an electrically driven support based on reinforcement learning, as a preferred embodiment, includes the following steps: S1. Construction of Multiphysics Digital Twin Model: A multiphysics digital twin model is constructed based on the hydraulic support and its working environment; the working environment includes the roof, floor, and coal mining machine; the multiphysics digital twin model is used to simulate the periodicity and randomness of roof pressure, the mechanical behavior of the interaction between the hydraulic support and the surrounding rock, and the dynamic response of the hydraulic system; The construction method of multiphysics digital twin model is as follows: (1) Modeling of the mechanical field of the top plate Based on geological exploration data from the working face, a three-dimensional elastoplastic mechanical model of the roof strata was established using the finite element method (FEM). The roof strata in the model were divided into three layers: the immediate roof, the main roof, and the overlying load layer. The rock mechanical parameters (elastic modulus E, Poisson's ratio ν, and uniaxial compressive strength UCS) of each layer were obtained from borehole core test reports. The model boundary conditions were set as follows: the goaf side was a free surface, the solid coal side of the two roadways was a fixed constraint, and the overlying strata were subjected to an equivalent uniformly distributed load. As the coal mining machine advanced, the model dynamically updated the goaf extent, simulated the bending deformation and periodic fracture process of the roof cantilever beam in real time, and output the predicted values ​​of roof subsidence and incoming pressure load at each support location.

[0030] (2) Hydraulic fluid field modeling Based on the hydraulic system schematic, a fluid-mechanical coupling simulation model of the hydraulic support column was established using AMESim software. The model includes: a variable frequency motor-hydraulic pump unit (input is motor speed N, output is flow rate Q and pressure P), the column hydraulic cylinder (piston area, sealing friction, leakage coefficient), safety valves, and hydraulically controlled check valves, among other hydraulic components. The model parameters were calibrated using bench test data, accurately simulating the dynamic response of the column pressure under different roof loads. The simulation step size was set to 10ms, consistent with the sampling period of the actual controller.

[0031] (3) Modeling of the kinematic field of the coal mining machine Based on the actual operating trajectory data of the coal mining machine, a three-dimensional kinematic model of the coal mining machine is established to describe the changes in the spatial position, cutting speed, and cutting depth of the coal mining machine on the working face over time. The model takes the real-time position data uploaded by the coal mining machine's encoder as input and outputs the relative distance between the coal mining machine and each hydraulic support as one of the input components of the state vector S.

[0032] The three-field coupling method involves using the pressure load output from the roof mechanics field as the external excitation input to the hydraulic fluid field, and the column pressure output from the hydraulic fluid field as the boundary condition for the support reaction force of the roof mechanics field. These two fields form a bidirectional coupled iterative solution. The position information output from the coal mining machine kinematics field synchronously updates the goaf boundary of the roof mechanics field, achieving collaborative simulation of the three fields. The digital twin model's simulation speed is 100 times faster than real-time, capable of completing 1000 hours of operational simulation within 10 minutes, providing ample offline training data for the DDPG algorithm. S2. Construction of the Neural Network Model: In offline mode, a deep reinforcement learning controller based on a deep deterministic policy gradient algorithm interacts with a virtual hydraulic support through a multiphysics digital twin model to generate a neural network model; the neural network model M is modeled using a Markov decision process, including: Where M is the neural network model, S is the state vector, A is the deterministic action, and R is the reward function; The state vector S is a set of multi-dimensional vectors describing the current situation, including the column pressure of the hydraulic support, the column pressure of the adjacent hydraulic support, the reading of the roof separation instrument, the signal strength of the micro-seismic sensor, the relative distance between the hydraulic support and the coal mining machine, the motor power and the oil temperature. The deterministic action A is the target pressure adjustment amount ΔP output by the edge controller; The reward function R is as follows: Where R is the reward function, r1 is the safety reward function, r2 is the energy-saving reward function, r3 is the stability reward function, and a, b, and c are weighting coefficients; When the column pressure P of the hydraulic support is within the safe range [P1, P2], the safety reward function r1=C1, C1>0, where C1 is the first safety reward value; When the column pressure of the hydraulic support is P < P1 or P > P2, the safety reward function is r1 = -C2, C2 > C1, where C2 is the second safety reward value. Energy saving reward function r2=-C3×P t / P0, where C3 is the energy-saving bonus value, P t P0 is the instantaneous power of the pumping station, and P1 is the standard power of the pumping station. The stable reward function is r3 = -C4 × |P(t) - P(t-1)| / P(t-1), where C4 is the stable reward value, P(t) is the column pressure of the hydraulic support at time t, and P(t-1) is the column pressure of the hydraulic support at time t-1. S3. Deep reinforcement learning of neural network models: In each interaction, the deep reinforcement learning controller selects a state vector as input to the neural network model, updates the weights of the policy network and the value network, and obtains the neural network model after training. S4. Deployment of the Neural Network Model: Deploy the trained neural network model to the edge controller of each hydraulic support in the well. The neural network model includes a policy network U and a value network Q. The input of the policy network U is the state vector S, and the output is the deterministic action A. The input of the value network Q is the state vector S and the deterministic action A, and the output is the expected value q. S5. Online control: The edge controller collects state vectors, and the trained neural network model calculates the optimal target pressure value and sends it to the frequency converter controller for execution.

[0033] Example 2 Compared to Example 1, the neural network model in this example further includes an updated policy network and an updated value network; The policy network is updated by policy gradient, with the goal of enabling the deterministic action A output by the policy network to obtain a higher expected value q. The value network is updated by minimizing the loss function, with the goal of making the output of the value network approximate the true expected value q.

[0034] Example 3 To verify the effectiveness of the method of the present invention, a 30-day industrial test was conducted on a fully mechanized mining face in a certain mine (average roof pressure step distance 15m, pressure intensity 25~42MPa) using the method of Example 1, and a quantitative comparison was made with the following two comparative schemes: Comparison with Option A: Traditional constant pressure PID control (target pressure is fixed at 35MPa); Comparison Scheme B: Discrete Action Space Reinforcement Learning Control Based on DQN Algorithm (Action Space Discrete in 5 Levels: 28 / 30 / 32 / 35 / 38MPa); Scheme C of the present invention: Adaptive pressure control of continuous motion space based on DDPG algorithm.

[0035] The experimental results are shown in the table below:

[0036] As shown in the table above, the pressure control accuracy of Scheme C in this invention is improved by approximately 81% compared to the constant-pressure PID scheme and by approximately 67% compared to the DQN scheme; the energy saving rate reaches 15.1%, which is better than the 12.9% of the DQN scheme; and the online inference time is only 3.1ms, meeting the requirements of real-time control. These results demonstrate that the continuous action output characteristics of the DDPG algorithm, compared to the discrete action of DQN, can more precisely track the optimal support pressure, achieving significant improvements in safety, energy saving, and stability.

[0037] The above description is illustrative only and not restrictive of the present invention. Those skilled in the art will understand that any modifications, variations or equivalents that can be made without departing from the spirit and scope defined by the claims will fall within the protection scope of the present invention.

Claims

1. An adaptive pressure control method for an electrically driven support based on reinforcement learning, characterized in that: include: A multiphysics digital twin model is constructed based on the hydraulic support and its working environment. The working environment includes the roof, floor, and coal mining machine. The multiphysics digital twin model includes a two-way coupled simulation model of three physical fields: the roof mechanical field, the hydraulic fluid field, and the coal mining machine kinematic field. In offline mode, the deep reinforcement learning controller, based on the deep deterministic policy gradient algorithm, interacts with the virtual hydraulic support through the multi-physics digital twin model. The deep deterministic policy gradient algorithm is based on the Actor-Critic framework and outputs a continuous target pressure adjustment ΔP in the continuous action space. The deep reinforcement learning controller takes a multi-dimensional state vector S, which includes the hydraulic support column pressure, the pressure of adjacent support columns, the roof separation instrument reading, the signal strength of the microseismic sensor, the relative distance between the hydraulic support and the coal mining machine, the motor power, and the oil temperature, as input. It takes a reward function R, which is formed by the weighted sum of the safety reward function r1, the energy-saving reward function r2, and the stability reward function r3, as the optimization objective, and generates a neural network model including a policy network U and a value network Q. The neural network model is deployed to the edge controller of the hydraulic support; When operating online, the edge controller collects the state vector S of the hydraulic support in real time. The neural network model calculates the optimal target pressure value through the state vector S and sends it to the frequency converter controller for execution.

2. The adaptive pressure control method for an electrically driven support based on reinforcement learning according to claim 1, characterized in that: Includes the following steps: S1. Construction of the multiphysics digital twin model: The multiphysics digital twin model is constructed based on the hydraulic support and its working environment; S2. Construction of the neural network model: In offline mode, the deep reinforcement learning controller with deep deterministic policy gradient algorithm as its core interacts with the virtual hydraulic support through the multi-physics digital twin model to generate the neural network model; S3. Deep reinforcement learning of neural network model: In each interaction, the deep reinforcement learning controller selects one of the state vectors and inputs it into the neural network model to update the weights of the policy network and the value network, thereby obtaining the neural network model after training. S4. Deployment of the neural network model: Deploy the trained neural network model to the edge controller of each hydraulic support in the well. S5. Online control: The edge controller collects the state vector, and the trained neural network model calculates the optimal target pressure value and sends it to the frequency converter controller for execution.

3. The adaptive pressure control method for an electrically driven support based on reinforcement learning according to claim 2, characterized in that: The working environment described in step S1 includes the roof, floor, and coal mining machine.

4. The adaptive pressure control method for an electrically driven support based on reinforcement learning according to claim 3, characterized in that: The multiphysics digital twin model mentioned in step S1 is used to simulate the periodicity and randomness of the roof pressure, the mechanical behavior of the interaction between the hydraulic support and the surrounding rock, and the dynamic response of the hydraulic system.

5. The adaptive pressure control method for an electrically driven support based on reinforcement learning according to claim 2, characterized in that: The neural network model M described in step S2 is modeled using a Markov decision process, including: Where M is the neural network model, S is the state vector, A is the deterministic action, and R is the reward function.

6. The adaptive pressure control method for an electrically driven support based on reinforcement learning according to claim 5, characterized in that: The state vector S is a set of multi-dimensional vectors describing the current situation, including the column pressure of the hydraulic support, the column pressure of the adjacent hydraulic support, the reading of the roof separation instrument, the signal strength of the micro-vibration sensor, the relative distance between the hydraulic support and the coal mining machine, the motor power, and the oil temperature.

7. The adaptive pressure control method for an electrically driven support based on reinforcement learning according to claim 5, characterized in that: The deterministic action A is the target pressure adjustment amount ΔP output by the edge controller.

8. The adaptive pressure control method for an electrically driven support based on reinforcement learning according to claim 5, characterized in that: The reward function R is as follows: Where R is the reward function, r1 is the safety reward function, r2 is the energy-saving reward function, r3 is the stability reward function, and a, b, and c are weighting coefficients; When the column pressure P of the hydraulic support is within the safe range [P1, P2], the safety reward function r1 = C1, C1 > 0, where C1 is the first safety reward value; When the column pressure of the hydraulic support is P < P1 or P > P2, the safety reward function r1 = -C2, C2 > C1, where C2 is the second safety reward value; The energy-saving reward function r2=-C3×P t / P0, where C3 is the energy-saving bonus value, P t P0 is the instantaneous power of the pumping station, and P0 is the standard power of the pumping station. The stable reward function r3 = -C4 × |P(t) - P(t-1)| / P(t-1) is given by the function r3 = -C4 × |P(t) - P(t-1)| / P(t-1), where C4 is the stable reward value, P(t) is the column pressure of the hydraulic support at time t, and P(t-1) is the column pressure of the hydraulic support at time t-1.

9. The adaptive pressure control method for an electrically driven support based on reinforcement learning according to claim 1, characterized in that: The policy network U takes a state vector S as input and outputs a deterministic action A. The value network Q takes the state vector S and the deterministic action A as inputs and outputs the expected value q.

10. The adaptive pressure control method for an electrically driven support based on reinforcement learning according to claim 9, characterized in that: The neural network model also includes an update policy network and an update value network; The updated policy network updates the policy network through policy gradients, with the goal of enabling the deterministic action A output by the policy network to obtain a higher expected value q. The updated value network is updated by minimizing the loss function, with the goal of making the output of the value network approximate the true expected value q.