Dredged mud pipeline flow velocity hierarchical control method fusing model predictive control and deep reinforcement learning

By adopting a hierarchical MPC and DRL method in the dredged mud pipeline system, combined with multi-source sensor data and reward mechanism, the control problem of traditional control methods in nonlinear, large time delay, and large inertia systems is solved, and efficient and accurate flow rate adjustment and stability are achieved.

CN120295124APending Publication Date: 2025-07-11HOHAI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510432208.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

Traditional control methods are difficult to achieve precise control and optimized adjustment in nonlinear, large time delay, large inertia and susceptible to disturbance dredged mud pipeline systems. The existing MPC methods fail to fully consider long-term system performance in long-term operations, and the DRL method has poor control effect under hard constraints.

Method used

Adopting a hierarchical architecture, the MPC module is responsible for global planning and long-term optimization, and the DRL module is responsible for real-time adjustment, and is optimized and controlled online through the DDPG network model, combining multi-source sensor data and reward mechanism to achieve efficient and accurate adjustment of pipeline flow rate.

Benefits of technology

It realizes efficient and stable control in complex dynamic environments, the DRL module quickly adapts to real-time changes, and the MPC module ensures long-term stability, reduces system error rate and failure risk, and meets precise control requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120295124A_ABST
    Figure CN120295124A_ABST
Patent Text Reader

Abstract

The invention discloses a dredged mud pipeline flow velocity hierarchical control method fusing model predictive control and deep reinforcement learning, and relates to the technical field of intelligent control of dredging equipment. Comprising the steps of obtaining dredging slurry pipeline operation data to make a historical operation data set; preprocessing the historical operation data set, and dividing the historical operation data set into a training set and a verification set according to a time sequence in a ratio of 8: 2; a DDPG network model is constructed, and a simulation training environment is established; training and verifying the DDPG network model through the training set and verification set data; real-time operation data of the dredging slurry pipeline are obtained, the expected flow speed is set, through the verified DDPG network model, the MPC module triggers global planning with the execution period Tc, the DRL module conducts real-time adjustment with the execution period Td at the bottom layer, and online optimization control over the flow speed of the dredging slurry pipeline is achieved. According to the invention, quick response and real-time optimization can be realized in a dynamic environment, and efficient and accurate adjustment of the flow velocity of the pipeline is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent control of dredging equipment, and particularly to a method for hierarchical control of the flow velocity of dredging mud pipelines that combines model predictive control and deep reinforcement learning. Background Art

[0002] As an important part of marine engineering, dredging projects are widely used in projects such as port construction and waterway maintenance. In dredging projects, the excavated sediment needs to be efficiently transported through a pipeline system. In this process, precise control of the pipeline flow velocity is crucial.

[0003] The mud pipeline transportation system itself has significant nonlinear, large time-delay, and large inertia characteristics, and is vulnerable to multi-factor disturbances and other problems, making traditional control methods face many challenges. The traditional PID control method relies on a feedback regulation mechanism. Although it can maintain system stability to a certain extent, it is difficult for the PID adjustment parameters to accurately match the dynamic environmental changes. As a result, when facing external disturbances, it is difficult to achieve precise control and optimal regulation, thereby affecting the operation efficiency and energy consumption management. The MPC control method uses the dynamic model of the system to predict future behavior for decision-making control. However, when making decisions within a limited time horizon, it often only obtains the optimal solution and fails to fully consider the impact of a certain action on the overall system performance over a long period of time. This limitation is particularly prominent in long-cycle operations and resource optimization processes, and may lead to an unsatisfactory overall control effect.

[0004] In recent years, deep reinforcement learning (DRL) has been widely applied to the field of unmanned system control due to its powerful control ability. Deep reinforcement learning learns dynamic control strategies by interacting with the environment and has strong long-term decision-making and the ability to adapt to complex disturbances. However, when DRL is applied alone, due to low sample efficiency, slow convergence speed, and deficiencies in dealing with hard constraint problems, its control strategy is difficult to achieve a high-precision control effect in a short time.

[0005] Therefore, proposing a method for hierarchical control of the flow velocity of dredging mud pipelines that combines model predictive control and deep reinforcement learning to solve the difficulties existing in the prior art is an urgent problem that needs to be solved by those skilled in the art. Summary of the Invention

[0006] In view of this, the present invention provides a method for hierarchical control of the flow velocity of dredging mud pipelines that combines model predictive control and deep reinforcement learning. Through the design of a hierarchical architecture, the MPC module serves as a high-level planning module, responsible for global planning and long-term optimization, and the DRL module serves as a low-level optimization module, responsible for real-time adjustment and local optimization. It can quickly respond and perform real-time optimization in a dynamic environment, achieving efficient and precise adjustment of the pipeline flow velocity and ensuring the efficient and stable operation of the pipeline transportation system.

[0007] To achieve the above object, the present invention adopts the following technical solutions:

[0008] A dredging mud pipeline flow velocity stratified control method integrating model predictive control and deep reinforcement learning, comprising:

[0009] S1. Obtain the operation data of the dredging mud pipeline and make it into a historical operation data set;

[0010] S2. Preprocess the historical operation data set and divide it into a training set and a validation set according to a ratio of 8:2 based on the time series;

[0011] S3. Construct a DDPG network model and build a simulation training environment;

[0012] S4. Train and validate the DDPG network model with the training set and validation set data;

[0013] S5. Obtain the real-time operation data of the dredging mud pipeline, set the desired flow velocity, and through the validated DDPG network model, the MPC module triggers global planning with an execution period of T c and the DRL module makes real-time adjustments at the bottom layer with an execution period of T d to achieve online optimal control of the flow velocity of the dredging mud pipeline.

[0014] In the above method, optionally, the operation data of the dredging mud pipeline obtained in S1 is specifically: collecting real-time data of multiple sensors on the dredging mud pipeline transportation experimental platform, including: pipeline flow velocity V t , pressure P t , mud concentration C t , motor torque τ t , motor speed n t .

[0015] In the above method, optionally, the preprocessing of the historical operation data set in S2 specifically includes:

[0016] S201. Smooth the data by using a filtering method;

[0017] S202. Normalize the filtered data to unify the numerical range;

[0018] S203. Use the sliding window method for the normalized data to segment the preprocessed time series data according to a fixed window length.

[0019] In the above method, optionally, constructing the DDPG network model in S3 specifically includes:

[0020] Constructing the DDPG network model in S3 specifically includes:

[0021] S301. Actor Network Design:

[0022] Define the six-dimensional state space of the Actor network input:

[0023] S = {v t , P t , C t , τ t , n t , v ref}

[0024] where v t is the pipeline flow velocity, with the unit of m / s; P t is the pipeline pressure, with the unit of Kpa; C t is the mud concentration, with the unit of kg / m 3 ; τ t is the motor torque, with the unit of N·m; n t is the motor speed, with the unit of rmp / min; v ref is the desired flow velocity, with the unit of m / s;

[0025] Define the action space A t of the Actor network output, and the control action is the motor drive frequency control quantity

[0026]

[0027] where is the motor drive frequency control quantity at the current time t, with the unit of Hz; u max is the maximum control quantity of the motor drive frequency for the current flow velocity, with the unit of Hz;

[0028] The Actor network adopts a multi-layer fully connected structure, with the input being the six-dimensional state space S. Feature extraction is performed through 3 hidden layers. The output layer uses the tanh function to constrain the action value in the interval [-1, 1], and a linear transformation is used to map it to the motor frequency control range;

[0029] The Actor network is updated using the gradient descent method, and the optimization problem is transformed into minimizing the loss function:

[0030]

[0031] where N is the sampling batch size, and Q(s i , a i ) is the Q value corresponding to the current state s i - action a i ;

[0032] S302. Design of the Critic Network:

[0033] The Critic network concatenates the 1D action space A output by the Actor network with the six-dimensional state space as the input, extracts the non-linear combined features of the state and action through three fully connected hidden layers, and outputs the Q value for the current state-action pair; t μ′

[0034] S303. Target network update:

[0035] The target Actor network is the same as the main Actor network, and the target Critic network is the same as the main Critic network. Both the target Actor network and the target Critic network adopt a soft update strategy, and the formula is as follows:

[0036]

[0037] where τ is the soft update coefficient, θ μ′ is the parameter of the target Actor network, θ μ is the parameter of the main Actor network, θ Q′ is the parameter of the target Critic network, θ Q is the parameter of the main Critic network;

[0038] S304. Set up the experience replay mechanism:

[0039] According to the historical operation data, store the samples (s t , u rl , r, s t+1 ) in the experience replay pool, and assign priorities to the samples according to the absolute value of the temporal difference error. The formula is as follows.

[0040]

[0041] where δ i is the TD error of the i-th sample, r i is the immediate reward, γ is the discount factor, Q(s i , a i ) is the Q value corresponding to the current state s i -action a i , s′ i is the next state, p i is the priority of the i-th sample, and ε prevents zero priority;

[0042] Combined with the prioritized experience replay mechanism, the loss function of the Critic network is as follows.

[0043]

[0044] where N is the sampling batch size, λi Assign weight parameters to the samples;

[0045] S305. Design a reward function, including: a flow velocity tracking error term, an action penalty term, and a steady-state reward term;

[0046] r(t) = -ω1|v t -v ref | 2 -ω2|u t -u t-1 | 2 +ω3·I(|v t -v ref | < α)

[0047] where v t is the pipeline flow velocity, ω1 is used to control the importance of flow velocity tracking accuracy, v ref is the desired flow velocity, ω2 is used to control the system smoothness, ω3 is used to control the excitation degree after the system reaches the steady state, u t is the control quantity at time t, u t-1 is the control quantity at time t - 1, I(·) is an index function, and α is the threshold of the velocity deviation.

[0048] For the above method, optionally, in S4, the DDPG network model is trained, specifically including:

[0049] S401. Receive the preprocessed training set data;

[0050] S402. Sample a batch of data from the experience replay pool according to the priority, and use the Critic network to calculate the temporal difference error;

[0051] S403. Use the gradient descent method to minimize the loss function, clip the gradient, and update the Critic network parameter θ Q ;

[0052] S404. Update the Actor network every 10 steps, optimize the policy function, the Critic network evaluates the Q value in real time, the Critic network calculates the policy gradient, and updates the Actor network parameter θ along the policy gradient direction μ ;

[0053] S405. The target network adopts a soft update strategy, and the target network is updated by slow synchronization to stabilize the training process until the policy network can stably achieve a control effect with a flow velocity deviation less than 5% in the simulation environment.

[0054] For the above method, optionally, in S5, the online optimization control of the dredging mud pipeline flow velocity is implemented, specifically including:

[0055] S501. Set the desired flow velocity vref Collect the real-time data of the pipeline through the multi-source sensors of the dredging mud pipeline transportation experimental platform;

[0056] S502. Taking the set desired flow velocity v ref as the control target, according to the real-time data of the pipeline, the MPC module triggers global planning at a high level with an execution period T c to generate the corresponding reference control sequence

[0057] S503. Deploy the trained DDPG model to generate the control quantity u of the DRL module drl ; The DRL module receives the collected real-time data of the pipeline and the set desired flow velocity v ref ;

[0058] S504. The DRL module runs at a high frequency within the time range of T c in the underlying control system, and its running period is T d , T c >T d , and infers the control quantity u corresponding to that moment drl ;

[0059] S505. Design the comprehensive evaluation function Φ(u), and the formula is as follows,

[0060]

[0061] where u is the control quantity output by the system at the current moment, λ∈[0, 1] is the dynamic weight, v is the pipeline flow velocity at the current moment, α = 0.5, β = 3%, indicating that when the error is less than 3%, it biases towards MPC, otherwise it biases towards DRL;

[0062] The output by the MPC module in step S502 and the u drl output by the DRL module in step S504 are jointly evaluated through the comprehensive evaluation function Φ(u) to synthesize the best control quantity u t under the current state S opt ;

[0063] S506. Send the obtained best control quantity u opt to the physical system and execute it;

[0064] S507. After executing the best control quantity u opt , according to the feedback of the real-time data, judge whether the current flow velocity meets the flow velocity stability judgment;

[0065] If the flow velocity stability is satisfied, keep the control quantity unchanged and stabilize the current flow velocity;

[0066] If the flow velocity stability is not satisfied, the optimal control quantity u will be executed opt The new state S after that t+1 As the input for the next closed-loop cycle, update the real-time pipeline data obtained in step S501. In step S502, the MPC module starts rolling optimization to update the reference control sequence, and repeat steps S503 to S506 until the obtained optimal control quantity satisfies the flow velocity stability.

[0067] For the above method, optionally, the design of the MPC module includes:

[0068] During the operation period, the prediction horizon of the MPC control method is N p = T c , and the control horizon is N c = T d , T d is the control frequency of the DRL module, and the formula is as follows:

[0069]

[0070] where Q is the output error weight matrix, R is the control increment penalty coefficient, v k+i|k is the predicted value of the velocity at the future k+i moment at time k, u k+i is the control input at the k+i moment, u k+i-1 is the control input at the k+i-1 moment;

[0071] Set the corresponding constraint conditions, including: motor frequency control quantity constraint, flow velocity safety constraint, specifically as follows:

[0072]

[0073] where u min and u max are the minimum and maximum control quantities of the motor drive frequency control quantity respectively, v min and v max are the safety lower limit and upper limit of the flow velocity.

[0074] For the above method, optionally, the flow velocity stability judgment specifically includes:

[0075] Judge whether the flow velocity meets the stable control requirements through joint judgment of multi-dimensional indicators, and define the flow velocity stability index as follows:

[0076] Design instantaneous dynamic error judgment:

[0077]

[0078] where ε dyn (t) is the dynamic allowable deviation, ε base is the basic allowable deviation, Tdelay is the time delay, k p is the proportionality coefficient, and e(τ) is the error function, representing the error between the actual flow velocity and the reference flow velocity at time τ;

[0079] Design trend stability judgment:

[0080]

[0081] Among them, is the change rate of the pipeline flow velocity, reflecting how fast the velocity changes with time, and |·| MA represents the moving average processing of the change rate of the pipeline flow velocity, and v thres is the set change rate threshold, T dyn is the dynamic time parameter, σ v (τ) is the standard deviation of the flow velocity at time τ, and σ max is the maximum allowable standard deviation.

[0082] As can be seen from the above technical solutions, compared with the prior art, the present invention provides a dredging mud pipeline flow velocity hierarchical control method that combines model predictive control and deep reinforcement learning, and has the following beneficial effects: 1) By effectively combining two control methods, namely model predictive control (MPC) and deep reinforcement learning (DRL), the present invention overcomes the limitations of traditional PID control algorithms in the face of nonlinear, large time-delay, and large inertia systems. Although the PID control algorithm can ensure the stability of the system, it is difficult to achieve precise control and optimal adjustment. Moreover, compared with the traditional MPC control algorithm, the use of the DRL method in the present invention can perform optimal control on the entire prediction time domain scale, fully considering the long-term benefits of the system, rather than being limited to short-term decisions, thus making up for the problem that the traditional MPC method may fall into a suboptimal solution in the face of complex dynamic systems; 2) Through the global planning of the MPC module, the present invention ensures the long-term stability of the system. At the same time, after the DRL module is offline trained with historical pipeline operation data and then deployed online, it can quickly adapt according to the real-time feedback of each sensor on the experimental platform. In the short term, the DRL module generates control quantities considering the long-term benefits within the prediction time domain to ensure that the flow velocity can quickly reach the desired range; 3) The present invention adopts a closed-loop feedback mechanism to generate and update control instructions according to real-time feedback and perform adaptive adjustment to ensure the accuracy of the control quantity and the stability of the flow velocity. Considering the performance parameters of the physical system and setting corresponding safety constraints can effectively reduce the error rate and failure risk of the system. In terms of the control strategy, the present invention gives full play to the advantages of the two control methods of MPC and DRL to meet the precise control requirements in a complex dynamic environment. Description of the Drawings

[0083] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to the provided drawings.

[0084] Figure 1 It is a flowchart of a dredging mud pipeline flow velocity stratified control method that combines model predictive control and deep reinforcement learning provided by the present invention;

[0085] Figure 2 It is a schematic diagram of an experimental platform for obtaining real-time data of pipeline operation provided by the present invention;

[0086] Figure 3 It is a schematic diagram of the change of the hierarchical execution time scale and control quantity of a dredging mud pipeline flow velocity stratified control method that combines model predictive control and deep reinforcement learning provided by the present invention;

[0087] Figure 4 It is a schematic diagram of the expected flow velocity curve of the flow velocity stable control of a dredging mud pipeline flow velocity stratified control method that combines model predictive control and deep reinforcement learning provided by the present invention. Detailed implementation manners

[0088] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0089] In this application, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. The terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including an..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.

[0090] Refer to Figure 1As shown in the figure, the present invention discloses a method for controlling the velocity stratification of dredging mud pipelines by integrating model predictive control and deep reinforcement learning, including:

[0091] S1. Obtain the operation data of the dredging mud pipeline and make it into a historical operation data set;

[0092] S2. Preprocess the historical operation data set and divide it into a training set and a validation set according to a ratio of 8:2 based on the time series;

[0093] S3. Construct a DDPG network model and build a simulation training environment;

[0094] S4. Train and validate the DDPG network model with the data in the training set and the validation set;

[0095] S5. Obtain the real-time operation data of the dredging mud pipeline, set the desired flow velocity, and through the validated DDPG network model, the MPC module triggers the global planning with an execution period of T c and the DRL module makes real-time adjustments at the bottom layer with an execution period of T d to achieve the online optimal control of the flow velocity of the dredging mud pipeline.

[0096] Further, the obtaining of the operation data of the dredging mud pipeline in S1 is specifically: referring to Figure 2 as shown in the figure, collect the real-time data of multiple sensors on the dredging mud pipeline transportation experimental platform, including: pipeline flow velocity V t , pressure P t , mud concentration C t , motor torque τ t , motor speed n t .

[0097] Further, the preprocessing of the historical operation data set in S2 specifically includes:

[0098] S201. Use a filtering method to smooth the data and reduce the influence of sensor noise and external disturbances;

[0099] S202. Normalize the data after filtering to unify the numerical range;

[0100] S203. Use the sliding window method for the normalized data, and cut the preprocessed time series data according to a fixed window length. The window length should cover the main dynamic response period of the system, and the step size can be appropriately overlapped to ensure the temporal continuity of the data.

[0101] Further, the construction of the DDPG network model in S3 specifically includes:

[0102] The construction of the DDPG network model in S3 specifically includes:

[0103] S301. Actor Network Design:

[0104] Define the six - dimensional state space of the Actor network input:

[0105] S = {v t , P t , C t , τ t , n t , v ref}

[0106] Among them, v t is the pipeline flow velocity, with the unit of m / s; P t is the pipeline pressure, with the unit of Kpa; C t is the mud concentration, with the unit of kg / m 3 ; τ t is the motor torque, with the unit of N·m; b t is the motor speed, with the unit of rmp / min; v ref is the desired flow velocity, with the unit of m / s;

[0107] Define the action space A of the Actor network output t , and the control action is the motor drive frequency control quantity

[0108]

[0109] Among them, is the motor drive frequency control quantity at the current time t, with the unit of Hz; u max is the maximum control quantity of the motor drive frequency for the current flow velocity, with the unit of Hz;

[0110] The Actor network adopts a multi - layer fully - connected structure. The input is the six - dimensional state space S. Feature extraction is performed through 3 hidden layers. The output layer uses the tanh function to constrain the action value in the interval [-1, 1], and uses a linear transformation to map it to the motor frequency control range;

[0111] The Actor network is updated using the gradient descent method, and the optimization problem is transformed into minimizing the loss function:

[0112]

[0113] Among them, N is the sampling batch size, Q(s i , a i ) is the Q - value corresponding to the current state s i - action a i ;

[0114] S302. Design of the Critic Network:

[0115] The Critic network concatenates the 1D action space A output by the Actor network t with the six-dimensional state space as the input, and extracts the non-linear combined features of the state and the action through three fully connected hidden layers (the activation function is the ReLU function), and outputs the Q value under the current state-action pair;

[0116] S303. Target network update:

[0117] The target Actor network is the same as the main Actor network, the target Critic network is the same as the main Critic network, and both the target Actor network and the target Critic network adopt the soft update strategy. The formula is as follows:

[0118]

[0119] where τ is the soft update coefficient, θ μ is the parameter of the target Actor network, θ μ is the parameter of the main Actor network, θ Q′ is the parameter of the target Critic network, θ Q is the parameter of the main Critic network;

[0120] S304. Set up the experience replay mechanism:

[0121] According to the historical operation data, the samples (s t , u rl , r, s t+1 ) are stored in the experience replay pool, and the priority of the samples is assigned according to the absolute value of the temporal difference error. The formula is as follows.

[0122]

[0123] where δ i is the TD error of the i-th sample, r i is the immediate reward, γ is the discount factor, Q(s i , a i ) is the Q value corresponding to the current state s i -action a i , s′ i is the next state, p i is the priority of the i-th sample, and ε prevents zero priority;

[0124] Combined with the prioritized experience replay mechanism, the loss function of the Critic network is as follows.

[0125]

[0126] where N is the sampling batch size, and λ i is the weight parameter assigned to the samples (to prevent sampling bias);

[0127] S305. Design a reward function, including: a flow velocity tracking error term, an action penalty term, and a steady-state reward term;

[0128] r(t) = -ω1|v t -v ref | 2 -ω2|u t -u t-1 | 2 +ω3·I(|v t -v ref | < α)

[0129] where v t is the pipeline flow velocity, ω1 is used to control the importance of flow velocity tracking accuracy, v ref is the desired flow velocity, ω2 is used to control the system smoothness, ω3 is used to control the excitation degree after the system reaches the steady state, u t is the control quantity at time t, and u t-1 is the control quantity at time t-1, I(·) is an indicator function, and α is the threshold of the velocity deviation.

[0130] Furthermore, in S4, the DDPG network model is trained, specifically including:

[0131] S401. Receive the preprocessed training set data;

[0132] S402. Sample a batch of data from the experience replay pool according to the priority, and use the Critic network to calculate the temporal difference error;

[0133] S403. Use the gradient descent method to minimize the loss function, clip the gradient, and update the Critic network parameters θ Q , and during the update process, clip the gradient to prevent gradient explosion;

[0134] S404. Update the Actor network every 10 steps, optimize the policy function, the Critic network evaluates the Q value in real time and suppresses overfitting through gradient clipping, and the Critic network calculates the policy gradient and updates the Actor network parameters θ μ ;

[0135] S405. The target network adopts a soft update strategy, and the target network is updated by slow synchronization to stabilize the training process until the policy network can stably achieve a control effect with a flow velocity deviation less than 5% in the simulation environment.

[0136] Further, in S5, the online optimization control of the dredging mud pipeline flow rate is realized, specifically including:

[0137] S501. Set the desired flow rate v ref , and collect the real-time pipeline data through the multi-source sensors of the dredging mud pipeline transportation experimental platform;

[0138] S502. Taking the set desired flow rate v ref as the control target, according to the real-time pipeline data, the MPC module triggers global planning at the high level with the execution period T c , and the time scale and control quantity change of the MPC module execution refer to Figure 3 as shown, and generate the corresponding reference control sequence

[0139] Transmit the first value of the generated reference control sequence to the comprehensive evaluation function for subsequent synthesis of the optimal control quantity u opt ;

[0140] S503. Deploy the trained DDPG model to generate the DRL module control quantity u drl ; The DRL module receives the collected real-time pipeline data and the set desired flow rate v ref ;

[0141] S504. The DRL module runs at a high frequency within the time range of T c in the underlying control system, and its running period is T d , T c >T d , and infer the control quantity u drl corresponding to this moment;

[0142] S505. Design the comprehensive evaluation function Φ(u), and the formula is as follows,

[0143]

[0144] where u is the control quantity output by the system at the current moment, λ∈[0,1] is the dynamic weight, v is the pipeline flow rate at the current moment, α = 0.5, β = 3%, indicating that when the error is less than 3%, it biases towards MPC, otherwise it biases towards DRL;

[0145] Combine the output by the MPC module in step S502 and the u drl output by the DRL module in step S504 through the comprehensive evaluation function Φ(u) for joint evaluation, and synthesize the optimal control quantity u t at the current state S opt ;

[0146] S506. Transmit the obtained optimal control quantity uopt Send it to the physical system (dredging slurry pipeline transportation experimental platform) and execute it;

[0147] S507. Execute the optimal control quantity u opt After that, according to the feedback of real-time data (considering the time delay in the change of system flow velocity after executing the control quantity), judge whether the current flow velocity meets the flow velocity stability judgment;

[0148] If it meets the requirement, keep the control quantity unchanged and stabilize the current flow velocity. For the schematic diagram of the flow velocity control curve, please refer to Figure 4 as shown;

[0149] If the flow velocity stability is not met, then execute the optimal control quantity u opt The new state S after t+1 is used as the input of the next closed-loop cycle, update the pipeline real-time data obtained in step S501, and the MPC module in step S502 starts rolling optimization to update the reference control sequence. Repeat steps S503 to S506 until the obtained optimal control quantity meets the flow velocity stability.

[0150] Furthermore, the design of the MPC module includes:

[0151] Model Predictive Control (MPC), as an online control method, is used for high-level global planning in the present invention, and its operation period is T c , within the operation period, the prediction horizon of the MPC control method is N p =T c , the control horizon is N c =T d , T d , is the control frequency of the DRL module. When designing the objective function, fully consider minimizing the flow velocity error (the difference between the actual flow velocity and the desired flow velocity) and the smoothness constraint of the control increment (to avoid drastic fluctuations in the control input). The formula is as follows:

[0152]

[0153] where, Q is the output error weight matrix, R is the control increment penalty coefficient, v k+i|k is the predicted value of the velocity at the k-th moment for the (k + i)-th moment in the future, u k+i is the control input at the (k + i)-th moment, and u k+i-1 is the control input at the (k + i - 1)-th moment;

[0154] Fully consider the basic performance of the system and set corresponding constraint conditions, including: motor frequency control quantity constraint, flow velocity safety constraint, as follows:

[0155]

[0156] where, u min and u max are the minimum and maximum control quantities of the motor drive frequency control quantity respectively, and v min and v max are the safety lower limit and upper limit of the flow rate.

[0157] Specifically, the steps of the high-level control process of the MPC module include:

[0158] Initialize the relevant parameters of the MPC module according to the received desired flow rate and the real-time operation data of the acquisition pipeline;

[0159] Based on the Kalman filter to observe the current state of the system, update the model parameters online;

[0160] Solve the constrained quadratic programming optimization problem through the objective function J(ω), and use the system dynamic model to infer the reference control sequence p within the prediction time domain N

[0161] Furthermore, the judgment of flow rate stability specifically includes:

[0162] Judge whether the flow rate meets the requirements of stable control through the joint judgment of multi-dimensional indicators to avoid misjudgment of a single indicator. Define the flow rate stability indicator as follows:

[0163] Design the instantaneous dynamic error judgment to reflect the flow rate control accuracy of the system at a certain moment:

[0164]

[0165] where, ε dyn (t) is the dynamic allowable deviation, ε base is the basic allowable deviation, T delay is the time delay, k p is the proportional coefficient, and e(τ) is the error function, representing the error between the actual flow rate and the reference flow rate at time τ;

[0166] Design the trend stability judgment, which is used to represent the change rate of the flow rate after moving average filtering. The filtered differential estimate smooths out the instantaneous noise and more accurately reflects the true change trend. It is mainly divided into the judgment of the flow rate change rate and the flow rate fluctuation degree:

[0167]

[0168] where, is the change rate of the pipeline flow rate, reflecting how fast the speed changes with time, |·| MA represents the moving average processing of the pipeline flow rate change rate, and v thres is the set change rate threshold, and T dynis a dynamic time parameter, σ v (τ) is the standard deviation of the flow velocity at time τ, σ max is the maximum allowable standard deviation.

[0169] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. The key point of each embodiment is to illustrate the differences from other embodiments. In particular, for a system or system embodiment, since it is basically similar to the method embodiment, the description is relatively simple. For the relevant parts, reference can be made to the partial description of the method embodiment. The systems and system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative efforts.

[0170] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for stratified control of the flow velocity of dredging slurry pipelines that combines model predictive control and deep reinforcement learning, characterized in that, Including: S1. Obtain the operation data of the dredging mud pipeline and make it into a historical operation dataset; S2. Preprocess the historical operation dataset and divide it into a training set and a validation set according to the time series at a ratio of 8:2; S3. Construct a DDPG network model and build a simulation training environment; S4. Train and validate the DDPG network model with the training set and validation set data; S5. Obtain the real-time operation data of the dredging mud pipeline, set the desired flow rate, and through the verified DDPG network model, the MPC module executes with the execution period T c to trigger global planning, and the DRL module performs real-time adjustment at the bottom layer with the execution period T d to achieve online optimal control of the flow rate of the dredging mud pipeline.

2. A method for stratified control of the flow velocity of a dredging mud pipeline that combines model predictive control and deep reinforcement learning according to claim 1, characterized in that The specific operation data of the dredging mud pipeline obtained in S1 are as follows: collecting the real-time data of multi-source sensors on the dredging mud pipeline transportation experimental platform, including: pipeline flow velocity V t , pressure P t , mud concentration C t , motor torque τ t , motor speed n t .

3. A method for stratified control of the flow velocity of a dredging mud pipeline that combines model predictive control and deep reinforcement learning according to claim 1, characterized in that In S2, the preprocessing of the historical operation dataset specifically includes: S201. Smooth the data using a filtering method; S202. Normalize the data after filtering to unify the numerical range; S203. Use a sliding window method for the data after normalization processing, and divide the preprocessed time series data according to a fixed window length.

4. A method for stratified control of the flow velocity of a dredging mud pipeline that combines model predictive control and deep reinforcement learning according to claim 1, characterized in that In S3, the construction of the DDPG network model specifically includes: S301. Actor network design: Define a six-dimensional state space for the input of the Actor network: S = {v t , P t , C t , τ t , n t , v ref} Among them, v t is the pipeline flow velocity, with the unit of m / s; P t is the pipeline pressure, with the unit of Kpa; C t is the mud concentration, with the unit of kg / m 3 ; τ t is the motor torque, with the unit of N·m; n t is the motor speed, with the unit of rmp / min; v ref is the desired flow velocity, with the unit of m / s; Define the action space A output by the Actor network t , and the control action is the motor drive frequency control quantity Among them, is the control quantity of the motor drive frequency at the current moment t, with the unit of Hz; u max is the maximum control quantity of the motor drive frequency of the current flow rate, with the unit of Hz; The Actor network adopts a multi-layer fully connected structure. The input is a six-dimensional state space S. Feature extraction is performed through 3 hidden layers. The output layer uses the tanh function to constrain the action value in the range of [-1, 1], and is mapped to the motor frequency control range through a linear transformation; The Actor network uses the gradient descent method for update, and transforms the optimization problem into minimizing the loss function: where N is the sampling batch size, Q(s i , a i ) is the Q-value corresponding to the current state s i - action a i ; S302. Design of the Critic network: The Critic network concatenates the 1D action space A output by the Actor network with the six-dimensional state space as the input, extracts the non-linear combined features of the state and the action through three fully connected hidden layers, and outputs the Q value for the current state-action pair; t After concatenating with the six-dimensional state space as the input, the non-linear combined features of the state and the action are extracted through three fully connected hidden layers, and the Q value for the current state-action pair is output; S303. Target network update: The target Actor network is the same as the main Actor network, the target Critic network is the same as the main Critic network, and both the target Actor network and the target Critic network adopt a soft update strategy. The formula is as follows: where τ is the soft update coefficient, θ μ′ is the target Actor network parameter, θ μ is the main Actor network parameter, θ Q′ is the target Critic network parameter, θ Q is the main Critic network parameter; S304. Set up an experience replay mechanism: According to historical operation data, store the sample (s t , u rl , r, s t+1 ) in the experience replay pool, and assign priorities to the samples according to the absolute value of the temporal difference error. The formula is as follows: where δ i is the TD error of the i-th sample, r i is the immediate reward, γ is the discount factor, Q(s i , a i ) is the Q-value corresponding to the current state s i - action a i , S′ i is the next state, p i is the priority of the i-th sample, and ε prevents zero priority; Combined with the prioritized experience replay mechanism, the loss function of the Critic network is as follows, where N is the sampling batch size, and λ i is the weight parameter assigned to the sample; S305. Design a reward function, including: a flow velocity tracking error term, an action penalty term, and a steady-state reward term; r(t) = -ω1|v t -v ref | 2 -ω2|u t -u t-1 | 2 +ω3·I(|v t -v ref |<α) Among them, v t is the pipeline flow velocity, ω1 is the importance for controlling the flow velocity tracking accuracy, v ref is the desired flow velocity, ω2 is used for the system smoothness, ω3 is used for the excitation degree after the system reaches the steady state, u t is the control quantity at time t, u t-1 is the control quantity at time t-1, I(·) is the index function, and α is the threshold of the velocity deviation.

5. A method for stratified control of the flow velocity of a dredging mud pipeline that combines model predictive control and deep reinforcement learning according to claim 4, characterized in that In S4, the training of the DDPG network model specifically includes: S401. Receive the preprocessed training set data; S402. Sample a batch of data from the experience replay pool according to the priority, and use the Critic network to calculate the temporal difference error; S403. Minimize the loss function using the gradient descent method, clip the gradient, and update the Critic network parameters θ Q ; S404. Update the Actor network every 10 steps to optimize the policy function. The Critic network evaluates the Q value in real time, calculates the policy gradient, and updates the Actor network parameter θ along the direction of the policy gradient. μ ; S405. The target network adopts a soft update strategy, and the target network is updated by slow synchronization to stabilize the training process until the policy network can stably achieve a control effect with a flow velocity deviation less than 5% in the simulation environment.

6. A method for stratified control of the flow velocity of dredging slurry pipelines by integrating model predictive control and deep reinforcement learning, characterized in that In S5, online optimal control of the flow velocity of the dredging slurry pipeline is realized, specifically including: S501. Set the desired flow velocity v ref , and collect the real-time data of the pipeline through the multi-source sensors of the dredging mud pipeline transportation experimental platform; S502. With the set desired flow rate v ref as the control target, according to the real-time pipeline data, the MPC module triggers global planning at a high level with an execution period T c to generate the corresponding reference control sequence S503. Deploy the trained DDPG model to generate the control quantity u of the DRL module drl ; The DRL module receives the real-time data of the pipeline collected and the set desired flow velocity v ref ; S504. The DRL module operates at a high frequency within the underlying control system within the time range of T c with its operating cycle being T d , where T c >T d , and infers the control quantity u corresponding to that moment drl ; S505. Design a comprehensive evaluation function Φ(u), and the formula is as follows where u is the control quantity output by the system at the current moment, λ ∈ [0, 1] is the dynamic weight, v is the pipeline flow velocity at the current moment, α = 0.5, β = 3%, indicating that when the error is less than 3%, it biases towards MPC, otherwise it biases towards DRL; The output of the MPC module in step S502 and the output \(u\) of the DRL module in step S504 drl are jointly evaluated through the comprehensive evaluation function \(\varPhi(u)\) to synthesize the optimal control quantity \(u\) under the current state \(S\) t ; opt ; S506. Send the obtained optimal control quantity u opt to the physical system for execution; S507. Execute the optimal control quantity u opt After that, based on the feedback of real-time data, determine whether the current flow rate meets the flow rate stability judgment; If the flow velocity stability is satisfied, the control quantity remains unchanged to stabilize the current flow velocity; If the flow velocity stability is not satisfied, the optimal control quantity u will be executed opt The new state S after t+1 As the input for the next closed-loop cycle, update the real-time pipeline data obtained in step S501. In step S502, the MPC module starts rolling optimization to update the reference control sequence, and repeat steps S503 to S506 until the obtained optimal control quantity satisfies the flow velocity stability.

7. A method for stratified control of the flow velocity of dredging slurry pipelines by integrating model predictive control and deep reinforcement learning, characterized in that The design of the MPC module includes: During the operation cycle, the prediction horizon of the MPC control method is N p = T c , and the control horizon is N c = T d , where T d is the control frequency of the DRL module, and the formula is as follows: where Q is the output error weight matrix, R is the control increment penalty coefficient, v k+i|k is the predicted value of the velocity at time k+i in the future, u k+i is the control input at time k+i, u k+i-1 is the control input at time k+i-1; Set corresponding constraint conditions, including: motor frequency control quantity constraint, flow velocity safety constraint, specifically as follows: where u min and u max are the minimum and maximum control amounts of the motor drive frequency control amount respectively, and v min and v max are the safety lower limit and upper limit of the flow rate.

8. A method for stratified control of the flow velocity of dredging slurry pipelines by integrating model predictive control and deep reinforcement learning, characterized in that The judgment of flow velocity stability specifically includes: Jointly judge whether the flow velocity meets the requirements of stable control through multi-dimensional indicators, and define the flow velocity stability index as follows: Design instantaneous dynamic error judgment: Among them, ε dyn (t) is the dynamic allowable deviation, ε base is the basic allowable deviation, T delay is the time delay, k p is the proportionality coefficient, and e(τ) is the error function, representing the error between the actual flow rate and the reference flow rate at time τ; Design trend stability judgment: Among them, is the change rate of the pipeline flow velocity, reflecting how fast the velocity changes with time, |·| MA represents the moving average processing of the change rate of the pipeline flow velocity, v thres is the set change rate threshold, T dyn is the dynamic time parameter, σ v (τ) is the standard deviation of the flow velocity at time τ, σ max is the maximum allowable standard deviation.