Ore pulp conveying pipe network optimization method based on historical data and agent learning
By transforming the slurry delivery pipeline model into a directed graph and constructing Markov decision-making process, combined with deep reinforcement learning to optimize the operating strategy of pumps, the problems of lack of accurate dynamic hydraulic models, high computational burden and excessive pump start and stop are solved in the existing technology, and a more efficient and economical slurry delivery pipeline management is achieved.
Patent Information
- Application Number
- CN202510210438.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-27
AI Technical Summary
The prior art lacks accurate dynamic hydraulic model in the slurry conveying pipeline network, the calculation burden is heavy, and the excessive start-stop problem of high-pressure diaphragm pumps is difficult to solve.
The slurry delivery pipeline network optimization method based on historical data and deep reinforcement learning is adopted. By converting the slurry delivery pipeline network model into a directed graph, the Markov decision-making process is constructed, and the deep neural network is used for agent learning to optimize the pump operation strategy.
It reduces the computational complexity, avoids excessive start-stop of high-pressure diaphragm pumps, improves the service life of hydraulic components, and adapts to the time-varying demand for ore.
Smart Images

Figure CN120046691A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of slurry pipeline transportation, and particularly relates to an optimization method for a slurry transportation pipe network based on historical data and agent learning. Background Art
[0002] The slurry transportation pipe network is a large-scale key infrastructure. The continuously developing slurry transportation pipe network poses challenges to related enterprises to ensure the stability and safety of the transportation pipe network and minimize economic costs. The most fundamental decision-making problem in the slurry transportation pipe network is the pressure management problem. The modeling and optimal scheduling of the pressure at distributed connection nodes are the primary issues of concern in the field of slurry pipeline transportation. The pressure at the connection nodes is not allowed to be too high, otherwise there is a high probability of pipeline rupture. At the same time, the pressure at the connection nodes is not allowed to be too low. Otherwise, the production goals of the enterprise cannot be met. The hydraulic models of head, pressure, pump flow rate, and required ore quantity are non-linear and non-convex, and it is challenging to find the optimal pump scheduling strategy in a computationally efficient manner. Currently, there are many studies on the scheduling optimization of high-pressure diaphragm pumps in slurry transportation pipe networks, mainly based on accurate hydraulic models of slurry transportation pipe networks.
[0003] However, these model-based methods all require obtaining an accurate steady-state hydraulic model in advance. In actual engineering, it is difficult to obtain the dynamic model of the slurry transportation pipe network, which limits the application of these methods. In addition, the steady-state model lacks the ability to handle stochastic required ore quantities and thus cannot adapt to the increasingly complex slurry transportation pipe network. In the field of modern control, model predictive control (MPC) has been widely applied to the optimal scheduling of high-pressure diaphragm pumps in slurry transportation pipe networks. However, MPC-based methods all require knowledge of the system dynamics before optimization. In addition, since MPC itself needs to solve an optimization problem at each step, the complexity of the optimization problem is related to the length of the receding horizon, resulting in a heavy computational burden. To reduce the online computational burden, policy search methods have been proposed. Such methods can effectively reduce the computational complexity. In an actual slurry transportation pipe network, the required ore quantity at the connection nodes changes rapidly over time. If the pressure at the connection nodes is controlled at a fixed value, adjustable hydraulic components such as pumps and valves need to be frequently adjusted. Excessive operation of pumps and valves will shorten the service life of these hydraulic components and increase additional energy consumption. In summary, the existing research has the following problems: lack of accurate hydraulic models, heavy computational burden, and excessive start-stop of pumps. With the continuous development of deep learning and reinforcement learning, it is possible to consider using deep reinforcement learning for the optimal scheduling of high-pressure diaphragm pumps in slurry transportation pipe networks to overcome the above-mentioned technical problems. Summary of the Invention
[0004] Aiming at the above-mentioned shortcomings of the prior art, the present invention provides an optimization method for a slurry transportation pipe network based on historical data and agent learning.
[0005] To achieve the above object, the technical solution adopted by the present invention is: an optimization method for a pulp transportation pipe network based on historical data and agent learning, comprising the following steps:
[0006] S1. Convert the pulp transportation pipe network model into a directed graph;
[0007] Represent the directed graph as G(V, ε), where V is the set of nodes and ε is the set of pipes;
[0008] V is divided into V = J ∪ T ∪ R, J represents the set of connection points n j of, T represents the set of agitation tanks n t of, R represents the set of smelters n r of, j, t, r ∈ n, and j ≠ t ≠ r, n is the number of nodes;
[0009] And define ε = P ∪ M ∪ W, where P, M, and W respectively represent the pipes n p , pumps n m and valves n w of, p, m, w ∈ n, p ≠ m ≠ w, n is the number of nodes;
[0010] S2. Construct an optimization scheduling model for the pulp transportation pipe network through the directed graph to obtain an optimal performance function, the steps are as follows:
[0011] S2.1. The optimal control problem of the optimization scheduling model of the pulp transportation pipe network is described by the following expression:
[0012] x t+1 = F(x t , m t ), t ∈ (0, ∞)
[0013] In the formula, x t is the head of the intersection node, m t is the pump speed ratio, F(x t , m t ) is the non-linear system function of the pulp transportation pipe network; among them, x t ∈ (x t,min , x t,max ), x t,min and x t,max respectively represent the minimum head and the maximum head of the intersection node, m t ∈ (m t,min , m t,max ), m t,min and m t,max respectively represent the minimum pump speed ratio and the maximum pump speed ratio;
[0014] S2.2. Through the control strategy u, make the performance index function Q(x0 , u) is minimized, and the specific expression of the performance index function is as follows:
[0015]
[0016] In the formula, R(x t , a t ) is the utility function, where a t is the action taken, and R(x t , a t ) satisfies R(0, 0) = 0 and R(x i , a t ) ≥ 0.
[0017] The control strategy u not only ensures the optimality of the result of the optimal control problem expression but also ensures the finiteness of the performance index function.
[0018] S2.3. Transform the original performance function to obtain the optimal performance function. The expression of the original performance function is as follows:
[0019]
[0020] In the formula, ψ(Ω u ) represents the set (pump speed, flow rate, pressure) containing all acceptable controls.
[0021] According to the Bellman equation, the original performance function Q * (x t ) can be described as:
[0022]
[0023] Then the optimal control strategy can be expressed as:
[0024]
[0025] Since a t = u(x t ), and u * (x t ) is obtained from Q * (x t ), if the traditional dynamic programming algorithm is used to calculate Q * (x t ), the "curse of dimensionality" will occur.
[0026] Therefore, the optimal performance function Q * (x t ) can be expressed as:
[0027]
[0028] Wherein, a t * is the optimal action taken;
[0029] S3. Convert the optimized scheduling model of the pulp transportation pipeline network into a Markov decision process;
[0030] The Markov decision process is expressed as: <s t , a t , R t , s t+1 >, where s t represents the state at time t, a t represents the action taken, R t represents the reward obtained, s t+1 represents the state at t + 1;
[0031] Specifically, the state s t is expressed as:
[0032]
[0033] Wherein, h i is the head of the connecting node, h max is the maximum node head of the connecting nodes in the entire pulp transportation pipeline network, v i is the relative speed of the i-th pump, v max represents the maximum number of pumps, d i is the ore demand at the intersection point i, d sum is the total ore demand;
[0034] The head of the connecting node, the relative speed of the pump, and the ore demand need to be normalized to avoid the influence of different units;
[0035] Further, since the speed at time t has no influence on the state at t + 1, the real-time pump scheduling in the pulp transportation pipeline network belongs to an online learning problem. When the action space is continuous, it is difficult to model it as a Markov decision process; therefore, the present invention constructs the operation of the high-pressure diaphragm pump as a discrete action space [0, 1, 2], which respectively represents that the speed ratio of the i-th pump increases by 0.05, the speed ratio decreases by 0.05, and no action is taken; wherein, i ∈ (1, ∞);
[0036] The reward R of the Markov decision process t plays a crucial role in the reinforcement learning algorithm process and can affect the effectiveness and results of the training process. In order to solve the sequential decision problem according to the node order, the expression of the reward R t is as follows:
[0037]
[0038] State st The instantaneous value V st , which is used to represent the reward of the Markov decision process. In traditional reinforcement learning algorithms, the state s t is calculated by discounted cumulative rewards; in the present invention, the state s t is calculated by the constructed optimal performance metric function;
[0039] To obtain the state value V st , an evaluation algorithm for the pulp transportation pipe network is designed, including three parts: V head , V eff and V reser , and the expressions are as follows:
[0040]
[0041] In the formula, V head represents the head ratio controlling the entry into the smelter; n satisfaction is the number of nodes within the acceptable range restricted by the minimum node head h min and the maximum node head h max in the entire pulp transportation pipe network; n total is the total number of nodes; the higher V head , the better the scheduling strategy; V eff represents the efficiency ratio; η t is the efficiency of the pump at time step t; η max is the maximum efficiency of the pump; V reser represents the flow ratio; f source represents the flow rate of the pulp produced by the concentrator; f reservoir / tank represents the flow rate of the pulp in the agitation tank;
[0042] Combining the three expressions gives the expression for the state V st as follows:
[0043]
[0044] In the formula, α 1 represents the weight of V head , α 1 ∈(0,1); α 2 represents the weight of V eff , α 2 ∈(0,1); α 3 represents the weight of V reser , α 3 ∈(0,1); α 1 , α 2 , α 3 can be equal or not equal; The maximum value of
[0045] S4. Construct a deep reinforcement learning optimization process to improve the learning efficiency of the Markov decision process; due to the long convergence time, to accelerate the training process, a deep reinforcement learning framework integrating historical data is constructed to improve the convergence efficiency, as follows:
[0046] Because the Supervisory Control And Data Acquisition (SCADA) system is widely used in the pulp transportation pipeline network, information can be collected from historical data. Therefore, the historical data is integrated into the predictor of deep reinforcement learning to assist learning when predicting the maximum evaluation state value based on historical data in the case of uncertain pulp demand at nodes;
[0047] The predictor is a preprocessing step before the agent interacts with the pulp transportation pipeline network environment; the agent consists of a policy network and a value network in a deep neural network;
[0048] To assist the agent, precise models of the predictor in two different scenarios (actual and simulated) are defined, as follows:
[0049] For the actual pulp transportation pipeline network, the predictor V predictor is the evaluation state value (expectation), and the predictor is defined as:
[0050] V predictor = max V st
[0051] where V st ∈(1, N), and N is the total number of historical operation data;
[0052] For the pulp transportation pipeline network in the simulated scenario, each is calculated by the local search method of nonlinear optimization. At this time, the expression of V predictor is as follows:
[0053] V predictor = E(V st )
[0054] where V st ∈(1, N), and N is the total number of historical operation data;
[0055] V predictor is calculated only once at the beginning of the training process. Then, the agent and the pulp transportation pipeline network environment interact at the time step t, specifically: the agent obtains the state S t of the pulp transportation pipeline network environment, and based on this state, the agent selects an action a t for the pulp transportation pipeline network environment at the time step t.; After that, at time step t+1, the agent obtains S t+1 and the reward R generated by the pulp transportation pipe network environment t ;
[0056] At the same time, the proximal policy optimization method is used to optimize the model. The proximal policy optimization uses a clipping term to replace the objective function instead of directly using the original objective function because it is difficult to estimate the gradient of the policy function with the original objective function;
[0057] The objective function of the proximal policy optimization method is:
[0058]
[0059] In the formula, ∈ is the clipping range; is the clipping probability ratio; r inside min(.) t (θ)A t is the original objective function; the second term modifies the surrogate objective function through the clipping probability ratio to eliminate the incentive to move r t outside the interval ; Finally, by taking the minimum of the clipped objective and the unclipped objective, the final objective is the lower bound of the unclipped objective;
[0060] The gradient of the proximal policy optimization method can be easily estimated through past trajectories (experiences, optimal control policies), and the parameters of the network (learning rate, clipping range, action) are updated using the gradient. The clipping function ensures that the update step size is within the clipping range;
[0061] In the present invention, the probability ratio is ignored only when the change in the probability ratio can improve the objective; otherwise, when it makes the objective worse, the probability ratio is considered; by using the clipped surrogate objective function, the gradient can be constrained to prevent the policy update from being too large;
[0062] A t is the advantage function, which is used to replace the performance metric function Q(x 0 ,u) to reduce the variance of the state value function. The expression of A t is as follows:
[0063] A t =Q π (s,a)-V π (s)
[0064] In the formula, Q π (s, a) and V π (s) are the state-action value function and the value function respectively, and their expressions are as follows:
[0065]
[0066] In the formula, γ is the learning rate; r is the reward for each iteration; l is the number of iterations
[0067] The advantage of Proximal Policy Optimization is that it is not sensitive to hyperparameters because only a few parameters need to be adjusted, including: the learning rate γ, the clipping range
[0068] Therefore, the deep reinforcement learning optimization process is as follows:
[0069] After obtaining the pulp demand of each node, to simplify the problem, it is assumed that the pulp demand remains unchanged during optimization; the output of the deep reinforcement learning optimization process controller is W, that is, the control signal of the pump. By introducing historical data into the deep reinforcement learning algorithm, the predictor and As input; initialize the state-action value network Q(s,a), the state value network V(s) and the policy network π(s) by using a deep neural network; run the policy π θ (s) for experience replay for time step T to improve the efficiency of the proposed algorithm;
[0070] S5. Cumulatively reward through total discount to determine that the optimal effect of the pulp transportation pipe network optimization scheduling model;
[0071] The purpose of constructing the optimal performance function of the pulp transportation pipe network is to maximize the total discounted cumulative reward in an unknown dynamic environment that satisfies the following conditions:
[0072] The unknown dynamic environment includes:
[0073] (1) When the required ore quantity changes with time, keep the node liquid level within a specific range;
[0074] (2) Minimize the number of pump operations;
[0075] (3) Reduce the water age of the pulp transportation pipe network; the water age refers to the time it takes for a fixed fluid point to be transported from the inlet to a specified point, and usually the water age at the inlet is set to 0;
[0076] The optimization problem of the pulp transportation pipe network considering the total discounted cumulative reward can be expressed as:
[0077]
[0078] In the formula, τ represents the discount factor.
[0079] The beneficial effects of the present invention
[0080] (1) The present invention is more practical in addressing the scheduling optimization problem of high-pressure diaphragm pumps considering the time-varying demand in slurry transportation pipe networks. The scheduling optimization problem of high-pressure diaphragm pumps is formulated as a Markov decision process, facilitating the handling through reinforcement learning methods.
[0081] (2) To reduce the computational complexity, the present invention designs a data-driven agent learning framework that does not require the precise dynamics of the slurry transportation pipe network as prior knowledge.
[0082] (3) The present invention constructs a key objective function as the performance index function for the pump optimization problem. By adjusting the weighting matrix, trade-offs can be made between different objectives to meet various requirements in practical applications, avoiding excessive start-stop processes of high-pressure diaphragm pumps and extending the service life of hydraulic components. BRIEF DESCRIPTION OF THE DRAWINGS
[0083] Figure 1 is a flowchart of the present invention;
[0084] Figure 2 is a schematic diagram of the deep reinforcement learning training process of the present invention;
[0085] Figure 3 is a topological diagram of the slurry transportation pipe network in an embodiment of the present invention;
[0086] Figure 4 is the training result of proximal policy optimization in an embodiment of the present invention;
[0087] Figure 5 is the state evaluation value within 24 hours in an embodiment of the present invention;
[0088] Figure 6 is the speed value within 24 hours in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0089] To better illustrate the purpose, technical solution, and advantages of the present invention, the present invention will be further described below in conjunction with specific embodiments.
[0090] In this embodiment, the topological diagram of the ore pulp transportation pipe network is selected as shown in Figure 3 . Among them, the number of high-pressure diaphragm pumps in a single pumping station is two, and both pumps operate at the same stepless speed. There are a total of 41 pipelines and 1 pumping station (equipped with two high-pressure diaphragm pumps). The ore pulp transportation pipe network environment can randomize the demands of the connection nodes and the speed ratios of the pumps.
[0091] As shown in Figure 1 , the optimization method for the ore pulp transportation pipe network based on historical data and agent learning includes the following steps:
[0092] S1. Convert the ore pulp transportation pipe network model into a directed graph;
[0093] The directed graph is represented as G(V, ε), where V is the set of nodes and ε is the set of pipelines;
[0094] V is partitioned into V = J ∪ T ∪ R, where J represents the set of connection nodes n j The set of, T represents the set of agitation tanks n t The set of, R represents the set of smelters n r The set of, j, t, r ∈ n, and j ≠ t ≠ r, where n is the number of nodes;
[0095] And define ε = P ∪ M ∪ W, where P, M, and W represent pipelines n p , pumps n m And valves n w The set of, p, m, w ∈ n, p ≠ m ≠ w, where n is the number of nodes;
[0096] S2. Build an optimal scheduling model for the pulp transportation pipeline network through the directed graph to obtain the optimal performance function. The steps are as follows:
[0097] S2.1. The optimal control problem of the pulp transportation pipeline network optimization scheduling model is described by the following expression:
[0098] x t+1 = F(x t , m t ), t ∈ (0, ∞)
[0099] In the formula, x t Is the head of the intersection node, m t Is the pump speed ratio, and F(·, ·) is the nonlinear system function of the pulp transportation pipeline network; among them, x t ∈ (x t,min , x t,max ), x t,min And x t,max Respectively represent the minimum head and maximum head of the intersection node, m t ∈ (m t,min , x t,max ), m t,min And m t,max Respectively represent the minimum pump speed ratio and the maximum pump speed ratio;
[0100] S2.2. Through the control strategy u, minimize the performance index function Q(x 0 , u) within an infinite time range. The specific expression of the performance index function is as follows:
[0101]
[0102] In the formula, R(x t , a t ) is the utility function, where at represents the action taken, R(x t , a t ) satisfies R(0, 0) = 0 and R(x t , a t ) ≥ 0
[0103] The control strategy u not only ensures the optimality of the optimal control problem expression but also guarantees the finiteness of the performance index function;
[0104] S2.3. Transform the original performance function to obtain the optimal performance function. The expression of the original performance function is as follows:
[0105]
[0106] In the formula, Ψ(Ω u ) represents the set containing all acceptable controls (pump speed, flow rate, pressure);
[0107] According to the Bellman equation, the original performance function Q * (x t ) can be described as:
[0108]
[0109] Then the optimal control strategy can be expressed as:
[0110]
[0111] Because a t = u(x t ), u * (x t ) is obtained from Q * (x t ), if the traditional dynamic programming algorithm is used to calculate Q * (x t ), the "curse of dimensionality" will occur;
[0112] Therefore, the optimal performance function Q * (x t ) can be expressed as:
[0113]
[0114] In the formula, a t * is the optimal action taken;
[0115] S3. Convert the optimized scheduling model of the pulp transportation pipe network into a Markov decision process;
[0116] The Markov decision process is expressed as: <s t , at , R t , s t+1 >, where s t represents the state at time t, a t represents the action taken, R t represents the reward obtained, s t+1 represents the state at t+1;
[0117] The state s t has the expression:
[0118]
[0119] In the formula, h i is the head connecting the nodes, h max is the maximum node head connecting the nodes in the entire pulp transportation pipe network, v i is the relative speed of the i-th pump, v max represents the maximum number of pumps, d i is the required pulp volume at the intersection point i, d sum is the total required pulp volume;
[0120] The head connecting the nodes, the relative speed of the pump, and the required pulp volume need to be normalized. In this embodiment, Min-Max normalization is adopted to avoid the influence of different units;
[0121] Since the speed at time t has no influence on the state at t+1, real-time pump scheduling in the pulp transportation pipe network belongs to an online learning problem. When the action space is continuous, it is difficult to model it as a Markov decision process. Therefore, in the present invention, the operation of the high-pressure diaphragm pump is constructed as a discrete action space [0, 1, 2], which respectively represents that the speed ratio of the i-th pump increases by 0.05, the speed ratio decreases by 0.05, and no action is taken; where i ∈ (1, ∞), and i takes 1 in this embodiment;
[0122] The reward R of the Markov decision process t plays a crucial role in the process of the reinforcement learning algorithm and can affect the effectiveness and results of the training process. In order to solve the sequential decision problem according to the node order, the reward R t has the following expression:
[0123]
[0124] The instantaneous value V of the state s t , which is used to represent the reward of the Markov decision process. In the traditional reinforcement learning algorithm, the state s st is calculated through discounted cumulative rewards; in the present invention, the state s t is calculated through the constructed optimal performance index function; t
[0125] To obtain the state value V st , an evaluation algorithm for the pulp transportation pipeline network is designed, which includes three parts: V head , V eff and V reser , and the expressions are as follows:
[0126]
[0127] In the formula, V head represents the ratio of the head controlling the entry into the smelter; n satisfaction is the number of nodes within the acceptable range restricted by the minimum node head h min and the maximum node head h max in the entire pulp transportation pipeline network; n total is the total number of nodes; the higher V head , the better the scheduling strategy; V eff represents the efficiency ratio; η t is the efficiency of the pump at the time step t; η max is the maximum efficiency of the pump; V reser represents the flow ratio; f source represents the flow rate of the pulp produced by the concentrator; f reservoir / tank represents the flow rate of the pulp in the agitation tank;
[0128] Combining the three expressions to obtain the expression of the state V st is as follows:
[0129]
[0130] In the formula, α 1 represents the weight of V head , α 1 ∈(0,1); α 2 represents the weight of V eff , α 2 ∈(0,1); α 3 represents the weight of V reser , α 3 ∈(0,1); α 1 , α 2 , α 3 can be equal or not equal; The maximum value of is 1;
[0131] S4. Construct a deep reinforcement learning optimization process to improve the learning efficiency of the Markov decision process; Since the convergence time is long, therefore, to accelerate the training process, a deep reinforcement learning framework integrating historical data is constructed to improve the convergence efficiency, specifically as follows:
[0132] As Figure 2As shown in the figure, the wide use of the Supervisory Control And Data Acquisition (SCADA) system in the pulp transportation pipeline network can collect information from historical data. Therefore, the historical data is integrated into the predictor of deep reinforcement learning, which aims to predict the maximum evaluation state value based on the historical data and assist in learning when the pulp demand at the node is uncertain.
[0133] The predictor is a preprocessing step before the agent interacts with the pulp transportation pipeline network environment. The agent consists of a policy network and a value network of a deep neural network.
[0134] To assist the agent, precise models of the predictor in two different scenarios (actual and simulated) are defined.
[0135] For the actual pulp transportation pipeline network, the predictor V predictor is the evaluation state value (expectation), and the predictor is defined as:
[0136] V predictor = max V st
[0137] where V st ∈(1, N), and N is the total number of historical operation data.
[0138] For the pulp transportation pipeline network in the simulated scenario, in this embodiment, N = 100 is selected, and each is calculated by the local search method of nonlinear optimization. V predictor is the evaluation state value (expectation), and the expression of V predictor is as follows:
[0139] V predictor = E(V st )
[0140] V predictor is calculated only once at the beginning of the training process. Then, the agent and the pulp transportation pipeline network environment interact at the time step t, specifically: the agent obtains the state S t of the pulp transportation pipeline network environment, and based on this state, the agent selects the action a t for the pulp transportation pipeline network environment at the time step t; then, at the time step t + 1, the agent obtains S t+1 and the reward R t generated by the pulp transportation pipeline network environment.
[0141] Meanwhile, the proximal policy optimization method is used for model optimization. The main idea of proximal policy optimization is to use a clipped surrogate objective function instead of the original objective function, because it is difficult to estimate the gradient of the policy function;
[0142] The objective function of the proximal policy optimization method is:
[0143]
[0144] In the formula, ∈ is the clipping range; is the clipping probability ratio; r inside min(.) t (θ)A t is the original objective function; the second term modifies the surrogate objective function through the clipping probability ratio, thereby eliminating the incentive to move r t outside the interval Finally, by taking the minimum of the clipped objective and the unclipped objective, the final objective is the lower bound of the unclipped objective;
[0145] In the present invention, the probability ratio is ignored only when the change in the probability ratio will improve the objective; otherwise, when it makes the objective worse, the probability ratio will be considered; by using the clipped surrogate objective function, the gradient can be constrained to prevent the policy update from being too large; A t is the advantage function, which is used to replace the performance metric function Q(x 0 ,u) to reduce the variance of the state value function, A t The expression is as follows:
[0146] A t = Q π (s,a)-V π (s)
[0147] In the formula, Q π (s, a) and V π (s) are the state-action value function and the value function respectively, and their expressions are as follows:
[0148]
[0149] In the formula, γ is the learning rate; r is the reward for each iteration; l is the number of iterations
[0150] The advantage of proximal policy optimization is that proximal policy optimization is not sensitive to hyperparameters because only a few parameters need to be adjusted, including: learning rate γ, clipping range
[0151] Therefore, the deep reinforcement learning optimization process is as follows:
[0152] After obtaining the pulp demand of each node, to simplify the problem, it is assumed that the pulp demand remains unchanged during optimization; the output of the deep reinforcement learning optimization process controller is W, that is, the control signal of the pump. By introducing historical data into the deep reinforcement learning algorithm, the predictor and are used as inputs; by using a deep neural network to initialize the state-action value network Q(s,a), state value network V(s), and policy network π(s); running the policy π θ (s) for experience replay for time step T to improve the efficiency of the proposed algorithm;
[0153] S5. Determine that the optimal effect of the pulp transportation pipeline network optimization scheduling model is achieved by accumulating the total discounted reward;
[0154] The purpose of constructing the performance index function of the pulp transportation pipeline network is to maximize the total discounted cumulative reward in an unknown dynamic environment that satisfies the following conditions:
[0155] The unknown dynamic environment includes:
[0156] (1) When the ore demand changes with time, keep the node liquid level within a specific range;
[0157] (2) Minimize the number of pump operations;
[0158] (3) Reduce the water age of the pulp transportation pipeline network; the water age refers to the time it takes for a fixed fluid point to travel from the inlet to a specified point. Usually, the water age at the inlet is set to 0;
[0159] The optimization problem of the pulp transportation pipeline network considering the total discounted cumulative reward can be expressed as:
[0160]
[0161] In the formula, τ represents the discount factor.
[0162] Figure 4 Shows the training results of Proximal Policy Optimization running 20 times with different random seeds. Figure 4 Shows the average length of the training process; in the first 30K time steps, the average episode length rapidly drops from 9.5 to 6.5; in the last 20K time steps, the average length gradually becomes 6.0, and the lower the average length, the lower the operating cost; therefore, Proximal Policy Optimization has better training efficiency and reduces energy consumption;
[0163] During the 20 training processes using different random seeds, the average computing time required to train the agent is 178.3 seconds, which can be used in practical engineering applications. Since the agent only needs to be trained once, the agent can directly perform inference; as for evaluation, after training the agent and obtaining the optimal policy, the agent can be used online. At this stage, the computing resources can be ignored because the agent is only used to generate the optimal control signal without training and updating.
[0164] To further illustrate the effectiveness of the present invention, it is assumed that the demand remains unchanged during the control period. Figure 5 、 Figure 6 are the state evaluation values and control signals (relative speed ratio of the pump) at t = 1,..., 24. The results of the optimization values and speeds are as Figure 5 、 Figure 6 shown. It can be seen that the optimal values calculated by the present invention are superior to the Nelder-Mead method and the traditional optimization method, which demonstrates the applicability and robustness of the present invention.
[0165] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit the protection scope of the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the essence and scope of the technical solutions of the present invention.
Claims
1. A method for optimizing a slurry transportation network based on historical data and agent learning, characterized in that: The following steps are involved: S1. Convert the slurry transportation network model into a directed graph; S2. Constructing a slurry transportation pipeline network optimization scheduling model through a directed graph to obtain the optimal performance function; S3, converting the slurry transportation pipeline network optimization scheduling model into a Markov decision process; S4. Improving the learning efficiency of Markov decision process through deep reinforcement learning optimization process; S5. The total discount is used to accumulate rewards, so as to determine whether the slurry transportation pipeline network optimization scheduling model is the most effective.
2. The method for optimizing a slurry transportation network based on historical data and agent learning according to claim 1, characterized in that: The directed graph is represented as G(V, ε), where V is a node set and ε is a pipeline set; V is divided into V = J ∪ T ∪ R, J represents the connection point n j A collection of, T represents the stirring tank n t A collection of R, where R represents the smelter n r The set of , j, t, r ∈ n, and j ≠ t ≠ r, n is the number of nodes; And define ε=P∪M∪W, where P, M and W represent pipeline n respectively. p , Pump m And valve w A set of , p, m, w∈n, p≠m≠w, n is the number of nodes.
3. The method for optimizing a slurry transportation network based on historical data and agent learning according to claim 1, characterized in that: The optimal performance function expression is as follows: In the formula, x t is the water head at the intersection node; a t * is the optimal action to take; R(x t ,a t *) is the optimal utility function.
4. The method for optimizing a slurry transportation network based on historical data and agent learning according to claim 1, characterized in that: The Markov decision process is expressed as: t , a t , R t ,s t+1 >, where represents the state at time t, a t Indicates the action taken, R t Indicates that you have received a reward, s t+1 Indicates the state at t+1; Among them, the state s t The expression is: In the formula, h i is the water head at the connection node, h max is the maximum node head of the connecting node in the entire slurry transportation network, v i is the relative speed of the ith pump, v max Indicates the maximum number of pumps, d i is the demand for ore at intersection point i, d sum is the total mineral demand; Reward R t The expression is as follows: Where V st For state s t The instantaneous value of V St-1 For state S t-1 The instantaneous value of Status V st The expression is as follows: In the formula, α1 represents V head The weight of V eff The weight of α2∈(0,1); α3 represents V reser The weight of , α3∈(0,1); α1, α2, α3 are equal; The maximum value of V head It represents the head ratio that controls the water entering the smelter; V eff Indicates efficiency ratio; V reser Indicates the flow rate ratio.
5. The method for optimizing a slurry transportation network based on historical data and agent learning according to claim 1, characterized in that: The method for improving the learning efficiency of the Markov decision process through the deep reinforcement learning optimization process is a proximal strategy optimization method, which is expressed as follows: In the formula, is the cropping range; is the clipping probability ratio; r within min(.) t (θ)A t is the initial objective function; the second The surrogate objective function is modified by clipping the probability ratio, thereby eliminating the t Move to interval Finally, by taking the minimum value between the cropped target and the uncropped target, the final target is the lower bound of the uncropped target.
6. The method for optimizing a slurry transportation network based on historical data and agent learning according to claim 1, characterized in that: The expression for judging the optimal effect of the slurry transportation pipeline network optimization scheduling model by using the total discount cumulative reward is as follows: Where τ represents the discount factor.
Citation Information
Cited By
Universe water supply scheduling method and system
CN120317634A