A method for optimizing a mill control loop setpoint with experience-driven adaptive reinforcement decision making

By employing an experience-driven adaptive reinforcement decision-making method to optimize the setpoint of the mill control loop, combined with an Actor-Critic model and a case-based reasoning algorithm, the feed rate and water replenishment are optimized in real time. This solves the problem of precise control of grinding particle size and circulating load during the grinding process, thereby improving the production efficiency and product quality of the grinding process.

CN120630716BActive Publication Date: 2026-07-21CHINA UNIV OF MINING & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA UNIV OF MINING & TECH
Filing Date
2025-07-17
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing grinding process control methods struggle to achieve precise control of grinding particle size and circulating load under complex and variable operating conditions. Traditional reinforcement learning methods suffer from slow learning convergence and insufficient adaptability when ore properties change frequently, resulting in control lag.

Method used

An experience-driven adaptive reinforcement decision-making method for optimizing the setpoint of the mill control loop is adopted. By identifying the operating indicators of the mill grinding process, constructing a performance index function, building an Actor-Critic model, and using case reasoning algorithms and reinforcement learning techniques, the setpoints of feed rate and water replenishment are optimized in real time to achieve convergence of grinding particle size and mill load.

Benefits of technology

It effectively solves the control lag problem of traditional methods under complex working conditions, realizes dynamic and precise control of the grinding process setpoint, improves production efficiency and product quality, and provides a new method for intelligent control of the grinding process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120630716B_ABST
    Figure CN120630716B_ABST
Patent Text Reader

Abstract

The application provides a mill control loop setting value optimization method for experience-driven adaptive reinforcement decision. The method comprises the following steps: identifying the operation index of the mill grinding process; constructing a performance index function based on the identified grinding granularity and mill load; constructing a causal diagram of the historical operation data of the mill grinding process; calculating the similarity of the current working condition causal diagram and each causal diagram of the historical operation data; building an Actor-Critic model and training the model; setting the similarity threshold of the current working condition causal diagram and all causal diagrams of the historical operation data as η; based on the calculated similarity and the similarity threshold η, screening the pre-trained Actor-Critic model that is most matched with the current working condition causal diagram, and updating the strategy parameter θ of the pre-trained Actor-Critic model in real time k ; and delivering the real-time generated setting value to the controller of the mill control loop. The application solves the control lag problem caused by the sudden change of working condition in the traditional method, realizes the dynamic and accurate regulation and control of the setting value of the grinding process, and improves the production efficiency of the grinding process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the technical field of optimized control of grinding processes, and more specifically, to an experience-driven adaptive reinforcement decision-making method for optimizing the setpoint of a mill control loop. Background Technology

[0002] As a pillar industry of the national economy, the mining sector is at a critical juncture of deep transformation towards intelligent and efficient operations. In the mineral processing field, the grinding process, as the core of the entire beneficiation process, directly impacts the overall efficiency of downstream beneficiation operations and plays a decisive role in the quality of the final product. As key indicators for measuring the performance of the grinding process, grinding particle size and circulating load not only affect the energy consumption level of the grinding process but also have a profound impact on core production indicators such as concentrate grade and metal recovery rate. Since grinding particle size and circulating load are mainly controlled by loop setpoints such as feed rate and water replenishment, how to dynamically optimize these setpoints based on real-time operating conditions has become a key research focus.

[0003] In the field of grinding process control, existing methods can be broadly categorized into two main types: optimization methods based on mechanistic models and data-driven methods based on artificial intelligence. Globally, some overseas concentrators, due to their high ore grades, uniform particle size distribution, relatively stable mineral composition, and the ability to homogenize ore properties through reasonable blending, exhibit good stability in their grinding processes. Based on this, precise mathematical models can be established, and mechanistic model-based methods such as real-time optimization (RTO), model predictive control (MPC), and multivariate decoupling control can be employed to optimize loop setpoints and achieve precise control of grinding particle size and circulating load. However, Chinese concentrators primarily process complex ores such as hematite, whose raw ore grades are generally low, with fine and uneven particle sizes and drastic variations in composition, resulting in grinding processes exhibiting significant nonlinearity, strong coupling, and dynamic uncertainty. This complexity significantly limits the practical application of mathematical model-based optimization control methods in Chinese concentrators, making it difficult to meet the optimization requirements of setpoints under complex environments.

[0004] For grinding processes where precise mechanistic models are difficult to establish, researchers have been exploring artificial intelligence-based optimization strategies for grinding loop setpoints in recent years. For example, some studies employ multivariate fuzzy supervised control methods, using fuzzy inference rules to dynamically adjust grinding loop setpoints to achieve adaptive regulation of cyclic load. Other studies combine expert systems with fuzzy control technology, constructing intelligent decision-making systems based on knowledge bases, databases, and fuzzy logic inference engines to determine optimal loop setpoints, thereby achieving stable control of grinding particle size. Industrial applications show that these methods have improved mill production efficiency to some extent and effectively reduced unit energy consumption in concentrators. However, such intelligent control methods still heavily rely on the experience and knowledge of domain experts during the design process, and their system decision-making capabilities are significantly affected by human factors, making it difficult to effectively adapt to complex and changing operating conditions. Furthermore, the adjustment strategies of these controllers lack adaptive adjustment capabilities; when faced with new changes in operating conditions, manual adjustments by experts are usually required. Human subjective judgment can easily lead to instability in the control strategy, thus affecting the overall optimization capability of the system.

[0005] In recent years, reinforcement learning (RL) technology has demonstrated broad application potential in solving complex process control problems with unknown mechanisms and strong nonlinearity due to its adaptability and self-learning capabilities. However, in actual grinding production environments, the frequent and significant changes in ore properties lead to highly dynamic and unstable operating conditions in the grinding process. This causes traditional RL methods to suffer from slow learning convergence speed and insufficient adaptability. Under drastically changing operating conditions, traditional RL struggles to achieve rapid environmental perception and real-time adjustment, making it difficult to adapt to the optimization control requirements under all operating conditions, thus limiting its application effectiveness in grinding setpoint optimization tasks. Summary of the Invention

[0006] This invention aims to address the problems of existing setpoint optimization methods in the grinding process, such as difficulty in model building and poor adaptability when facing changes in ore properties and dynamic operating conditions. It proposes an experience-driven adaptive reinforcement decision-making setpoint optimization method for the mill control loop.

[0007] Therefore, the purpose of this invention is to propose an experience-driven adaptive reinforcement decision-making method for optimizing the setpoint of the mill control loop.

[0008] To achieve the above objectives, the present invention provides a method for optimizing the setpoint of a mill control loop based on experience-driven adaptive reinforcement decision-making. This optimization method includes:

[0009] Step S1: Identify the operating indicators of the mill grinding process, including grinding particle size r1(k) and mill load r2(k); where mill load refers to the total mass of ore and media in the mill.

[0010] Step S2: Based on the grinding particle size r1(k) and mill load r2(k), construct a performance index function for the grinding process setpoint optimization method based on experience-driven adaptive reinforcement decision-making, so as to optimize the feed rate setpoint w1(k) and mill inlet water supply setpoint w2(k) in the subsequent grinding process, thereby achieving convergence of the grinding particle size r1(k) and mill load r2(k); the mathematical expression corresponding to the performance index function is:

[0011]

[0012] In equation (1), k is the simulation time, and i is the start time; r(k) = [r1(k), r2(k)] T w(k) = [w1(k), w2(k)] T w1(k) represents the feed rate setpoint, w2(k) represents the mill inlet water makeup rate setpoint; R is a positive definite matrix; γ k-i r*(k) represents the weighting factor at time (ki); r*(k) represents the expected value of r(k); w min and w max These are the upper and lower limits for the feed rate setpoint and the mill inlet water supply setpoint.

[0013] Step S3: Based on the case-based reasoning algorithm, extract the causal dependencies between the operating variables v1, v2, and v3 and the operating indices r1(k) and r2(k) from the historical operating data of the mill grinding process, in order to construct a causal graph of the historical operating data of the mill grinding process. Where V is the set of operating condition variable nodes, and V = {v1, v2, v3}; v1 represents the expected value of mill load; v2 represents the expected value of grinding particle size; v3 represents the raw ore particle size; E is the set of directed causal edges, which represents the set of directed causal edges between variables v1, v2, v3 in the set of mill grinding operating condition variable nodes V and operating indicators r1(k) and r2(k);

[0014] Step S4: Calculate the similarity between the current causal graph of the mill grinding process and the i-th causal graph of the historical operating data; wherein, the formula for calculating the similarity is:

[0015]

[0016] In equation (2), This is a cause-and-effect diagram showing the current operating conditions of the grinding mill process. This is the i-th causal graph in the causal graph of the historical operating data of the mill grinding process; ε c and ε i P represents the edge set of the current causal graph and the edge set of the i-th causal graph in the causal graph of the historical running data, respectively; c P represents the probability distribution of the current case, that is, the joint probability distribution of the variables in the causal graph under the current operating conditions of the grinding process; i KL(P) represents the probability distribution of the i-th historical case in the causal graph of historical operating data of the mill grinding process, that is, the joint probability distribution of each variable in the causal graph of historical operating data corresponding to the historical operating conditions of the mill grinding process; c ||P i ) represents the KL divergence between the conditional probability distributions of the causal mechanism; β is the decay coefficient;

[0017] Step S5: Construct an Actor-Critic model and train it based on historical operating condition variables v1, v2, and v3, historical operating indices r1(k) and r2(k), and historical setpoints w1(k) and w2(k). The Actor-Critic model includes a Critic network model and an Actor network model. The input to the Critic network is the current operating condition s of the mill. k s k = [v1,v2,v3,r1(k),r2(k)]; the output of the Critic network is the expected cumulative reward V(s) under the current operating condition of the mill. k The input to the Actor network is the current operating state s of the mill. k The output of the Actor network is a setpoint vector w(k) = [w1(k), w2(k)]. T ;

[0018] The architecture of the Critic network model is as follows: V(s) k ) = MLP(s k MLP(·) is a multilayer perceptron network. The input dataset of MLP(·) is [s1,…,s] constructed from historical data. k The output dataset of MLP(·) is [V(s1),…,V(s)]. k )]; and according to formula (2) and using the gradient descent algorithm to pre-train MLP(·), and save the parameters of MLP(·);

[0019] The architecture of the Actor network model is as follows: For s k mean For s k Variance, θ k Let θ be the policy parameter of the Actor network model. k The control strategy used to evaluate and optimize the state-to-setpoint performance of the grinding process of the mill; the input dataset of the Actor network model is [s1,...,s] constructed from historical data. k The output dataset of the Actor network model is [w(1),…,w(k)]; and the Actor network model is pre-trained according to formula (1) using the gradient descent algorithm, and the policy parameters θ of the Actor network model are saved. k ;

[0020] Step S6: Set the cause-effect graph of the current operating conditions of the mill grinding process. The similarity threshold between the causal graphs and all causal graphs in the historical operational data is η; where the formula for calculating the similarity threshold η is:

[0021] η = μ ε -ασ ε (3)

[0022] In equation (3), μ ε It is the set of edges ε of the causal graph of historical running data. i The mean; σ ε It is the set of edges ε of the causal graph of historical running data. i The standard deviation; α is the adjustment factor;

[0023] Step S7: Based on the calculated similarity Using a similarity threshold η, a pre-trained Actor-Critic model is selected to best match the causal graph of the current working condition of the grinding process; where, when a similarity to a historical working condition is determined, the model is selected. If the Actor-Critic model at this time is selected as the pre-trained Actor-Critic model that best matches the causal graph of the current working condition of the grinding process, the parameters of the best-matching pre-trained Actor-Critic model are saved; when it is determined that there is similarity between multiple historical working conditions... Then, the Actor-Critic model corresponding to the historical working condition with the highest similarity is selected as the pre-trained Actor-Critic model that best matches the causal graph of the current working condition of the grinding process, and the parameters of the best-matching pre-trained Actor-Critic model are saved; if all When the gradient correction strategy is applied, the parameter θ is directly adjusted. k Perform online correction;

[0024] Step S8: Utilize the natural policy gradient to update the policy parameters θ of the pre-trained Actor-Critic model that best matches the current causal graph of the grinding process in real time. k Wherein, the strategy parameter θ k The update formula is:

[0025] θ k+1 =θ k +αG -1 (θ k )▽ θ J(θ k (4)

[0026] In equation (4), α is the learning rate; Fisher's information matrix; Let J(θ) be the standard policy gradient, representing the policy performance. k Relative parameter θ k The gradient;

[0027] in, Let be the strategy function, which represents the mill's state s during grinding operations. k The probability of selecting the setpoint vector w(k); policy function The update formula is:

[0028]

[0029] In equation (5), Q(s) k w(k) is the action value function, which is used to evaluate the mill's state s during grinding operation. k The long-term expected reward of the setpoint vector w(k); the action value function Q(s) k The update formula for w(k) is:

[0030]

[0031] Step S9: Calculate the setpoint vector based on formula (4) The setpoints w1(k) and w2(k) are generated in real time through the Actor network model. These setpoints are then transmitted to the controller in the mill control loop. The controller adjusts the frequency of the feed motor and the opening of the water supply valve in real time based on the setpoints w1(k) and w2(k) until the operating indicators, grinding particle size r1(k) and mill load r2(k), converge to the target range. The convergence of grinding particle size r1(k) and mill load r2(k) to the target range means that the grinding particle size r1(k) and mill load r2(k) satisfy the performance index function constructed in step S2.

[0032] Preferably, the adjustment coefficient α is 0.9.

[0033] Preferably, the similarity threshold η ranges from [0,1].

[0034] The beneficial effects of this invention are:

[0035] This invention provides an experience-driven adaptive reinforcement decision-making method for optimizing the setpoint of a mill control loop, adapting to dynamic operating conditions characterized by frequent changes in ore properties and strong nonlinearity. Specifically, this optimization method quickly locks historical experience through a case matching mechanism and, combined with the online optimization characteristics of reinforcement learning, effectively solves the control lag problem caused by sudden changes in operating conditions in traditional methods. It achieves dynamic and precise control of the mill setpoint, improving production efficiency and product quality, and providing a novel approach for intelligent control of the mill process.

[0036] Additional aspects and advantages of the invention will become apparent from the description which follows, or may be learned by practice of the invention. Attached Figure Description

[0037] Figure 1 A schematic flowchart of an experience-driven adaptive reinforcement decision-making mill control loop setpoint optimization method according to an embodiment of the present invention is shown.

[0038] Figure 2 A process flow diagram of the grinding process according to an embodiment of the present invention is shown;

[0039] Figure 3 A schematic block diagram of an experience-driven adaptive reinforcement decision-making mill control loop setpoint optimization method according to an embodiment of the present invention is shown. Detailed Implementation

[0040] To better understand the above-mentioned objects, features, and advantages of the present invention, such as Figures 1 to 3 As shown in the accompanying drawings and specific embodiments, the present invention will be further described in detail below. It should be noted that, unless otherwise specified, the embodiments and features described in these embodiments can be combined with each other.

[0041] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and therefore the scope of protection of the invention is not limited to the specific embodiments disclosed below.

[0042] Figure 1 A schematic flowchart of an experience-driven adaptive reinforcement decision-making method for mill control loop setpoint optimization, according to an embodiment of the present invention, is shown. Figure 1As shown, the mill control loop setpoint optimization method driven by experience and adaptive reinforcement decision-making includes:

[0043] Step S1: Identify the operating indicators of the mill grinding process, including grinding particle size r1(k) and mill load r2(k); where mill load refers to the total mass of ore and media in the mill.

[0044] Step S2: Based on the grinding particle size r1(k) and mill load r2(k), a performance index function is constructed for the grinding process setpoint optimization method based on experience-driven adaptive reinforcement decision-making. This function is used to optimize the feed rate setpoint w1(k) and mill inlet water supply setpoint w2(k) in the subsequent grinding process, thereby achieving convergence of the grinding particle size r1(k) and mill load r2(k). The mathematical expression corresponding to the performance index function is:

[0045]

[0046] In equation (1), k is the simulation time, and i is the start time; r(k) = [r1(k), r2(k)] T w(k) = [w1(k), w2(k)] T w1(k) represents the feed rate setpoint, w2(k) represents the mill inlet water makeup rate setpoint; R is a positive definite matrix; γ k-i r*(k) represents the weighting factor at time (ki); r*(k) represents the expected value of r(k); w min and w max These are the upper and lower limits for the feed rate setpoint and the mill inlet water supply setpoint.

[0047] Step S3: Based on the case-based reasoning algorithm, extract the causal dependencies between the operating variables v1, v2, and v3 and the operating indices r1(k) and r2(k) from the historical operating data of the mill grinding process, in order to construct a causal graph of the historical operating data of the mill grinding process. Where V is the set of operating condition variable nodes, and V = {v1, v2, v3}; v1 represents the expected value of mill load; v2 represents the expected value of grinding particle size; v3 represents the raw ore particle size; E is the set of directed causal edges, which represents the set of directed causal edges between variables v1, v2, v3 in the set of mill grinding operating condition variable nodes V and operating indicators r1(k) and r2(k);

[0048] Step S4: Calculate the similarity between the current causal graph of the mill grinding process and the i-th causal graph of the historical operating data; wherein, the formula for calculating the similarity is:

[0049]

[0050] In equation (2), This is a cause-and-effect diagram showing the current operating conditions of the grinding mill process. This is the i-th causal graph in the causal graph of the historical operating data of the mill grinding process; ε c and ε i P represents the edge set of the current causal graph and the edge set of the i-th causal graph in the causal graph of the historical running data, respectively; c P represents the probability distribution of the current case, that is, the joint probability distribution of the variables in the causal graph under the current operating conditions of the grinding process; i KL(P) represents the probability distribution of the i-th historical case in the causal graph of historical operating data of the mill grinding process, that is, the joint probability distribution of each variable in the causal graph of historical operating data corresponding to the historical operating conditions of the mill grinding process; c ||P i ) represents the KL divergence between the conditional probability distributions of the causal mechanism; β is the decay coefficient;

[0051] Step S5: Construct an Actor-Critic model and train it based on historical operating condition variables v1, v2, and v3, historical operating indices r1(k) and r2(k), and historical setpoints w1(k) and w2(k). The Actor-Critic model includes a Critic network model and an Actor network model. The input to the Critic network is the current operating condition s of the mill. k s k = [v1,v2,v3,r1(k),r2(k)]; the output of the Critic network is the expected cumulative reward V(s) under the current operating condition of the mill. k The input to the Actor network is the current operating state s of the mill. k The output of the Actor network is a setpoint vector w(k) = [w1(k), w2(k)]. T ;

[0052] The architecture of the Critic network model is as follows: V(s) k ) = MLP(s k MLP(·) is a multilayer perceptron network. The input dataset of MLP(·) is [s1,…,s] constructed from historical data. k The output dataset of MLP(·) is [V(s1),…,V(s)]. k )]; and according to formula (2) and using the gradient descent algorithm to pre-train MLP(·), and save the parameters of MLP(·);

[0053] The architecture of the Actor network model is as follows: For s k mean For s k Variance, θ k Let θ be the policy parameter of the Actor network model. k The control strategy used to evaluate and optimize the state-to-setpoint performance of the grinding process of the mill; the input dataset of the Actor network model is [s1,…,s] constructed from historical data. k The output dataset of the Actor network model is [w(1),…,w(k)]; and the Actor network model is pre-trained according to formula (1) using the gradient descent algorithm, and the policy parameters θ of the Actor network model are saved. k ;

[0054] Step S6: Set the cause-effect graph of the current operating conditions of the mill grinding process. The similarity threshold between the causal graphs and all causal graphs in the historical operational data is η; where the formula for calculating the similarity threshold η is:

[0055] η = μ ε -ασ ε (3)

[0056] In equation (3), μ ε It is the set of edges ε of the causal graph of historical running data. i The mean; σ ε It is the set of edges ε of the causal graph of historical running data. i The standard deviation; α is the adjustment factor;

[0057] Step S7: Based on the calculated similarity Using a similarity threshold η, a pre-trained Actor-Critic model is selected to best match the causal graph of the current working condition of the grinding process; where, when a similarity to a historical working condition is determined, the model is selected. If the Actor-Critic model at this time is selected as the pre-trained Actor-Critic model that best matches the causal graph of the current working condition of the grinding process, the parameters of the best-matching pre-trained Actor-Critic model are saved; when it is determined that there is similarity between multiple historical working conditions... Then, the Actor-Critic model corresponding to the historical working condition with the highest similarity is selected as the pre-trained Actor-Critic model that best matches the causal graph of the current working condition of the grinding process, and the parameters of the best-matching pre-trained Actor-Critic model are saved; if all When the gradient correction strategy is applied, the parameter θ is directly adjusted. k Perform online correction;

[0058] Step S8: Utilize the natural policy gradient to update the policy parameters θ of the pre-trained Actor-Critic model that best matches the current causal graph of the grinding process in real time. k Wherein, the strategy parameter θ k The update formula is:

[0059] θ k+1 =θ k +αG -1 (θ k )▽ θ J(θ k (4)

[0060] In equation (4), α is the learning rate; Fisher's information matrix; Let J(θ) be the standard policy gradient, representing the policy performance. k Relative parameter θ k The gradient;

[0061] in, Let be the strategy function, which represents the mill's state s during grinding operations. k The probability of selecting the setpoint vector w(k); policy function The update formula is:

[0062]

[0063] In equation (5), Q(s) k w(k) is the action value function, which is used to evaluate the mill's state s during grinding operation. k The long-term expected reward of the setpoint vector w(k); the action value function Q(s) k The update formula for w(k) is:

[0064]

[0065] Step S9: Calculate the setpoint vector based on formula (4) The setpoints w1(k) and w2(k) are generated in real time using the Actor network model. These setpoints are then transmitted to the controller in the mill control loop. The controller adjusts the frequency of the feed motor and the opening of the water supply valve in real time based on the setpoints w1(k) and w2(k) until the operating indicators, grinding particle size r1(k) and mill load r2(k), converge to the target range. The convergence of grinding particle size r1(k) and mill load r2(k) to the target range means that the grinding particle size r1(k) and mill load r2(k) satisfy the performance index function constructed in step S2.

[0066] In this embodiment, the experience-driven adaptive reinforcement decision-making mill control loop setpoint optimization method combines case reasoning and Actor-Critic reinforcement learning techniques. It selects an Actor-Critic model suitable for the current working conditions through case reasoning and optimizes the loop setpoint using reinforcement learning algorithms, thereby realizing online optimization control of the grinding process.

[0067] In one embodiment of the present invention, the adjustment coefficient α is 0.9.

[0068] In one embodiment of the present invention, the similarity threshold η ranges from [0,1].

[0069] like Figure 2 and Figure 3 As shown below, a specific embodiment will be used to illustrate the technical solution of the present invention. This experience-driven adaptive reinforcement decision-making method for optimizing the setpoint of a mill control loop is implemented through the following steps:

[0070] S1. Characterization of Grinding Process Problems and Analysis of Key Operating Indicators: This section delves into the dynamic characteristics of the grinding system, focusing on two core operating indicators: grinding particle size and mill load. By analyzing their relationship with key variables such as feed rate and mill inlet water volume, the main factors influencing grinding efficiency and energy consumption are explored. Based on this, and considering process constraints and optimization objectives, the section investigates how to rationally adjust loop setpoints to improve the stability and efficiency of the grinding process, achieving precise control of grinding particle size and mill load.

[0071] (1) Description of grinding problems:

[0072] like Figure 2 In the typical closed-circuit grinding process shown, the ball mill serves as the control core of the operation layer. Its output indicators include grinding particle size r1(k) (reflecting product quality) and mill load r2(k) (reflecting operating efficiency). The inputs of the operation layer originate from the setpoints of the bottom loop, including the feed rate setpoint w1(k) and the mill inlet makeup water setpoint w2(k).

[0073] Grinding particle size r1(k) and mill load r2(k) not only reflect the energy consumption level of grinding, but also affect important production indicators such as concentrate grade and metal recovery rate in the concentrator. Therefore, they are considered two key operating indicators for grinding production. Figure 2As can be seen, excessively large or small operating parameters will lead to suboptimal operation of the grinding process. Therefore, to improve the quality and efficiency of the grinding process, it is essential to control the grinding particle size r1(k) and mill load r2(k) within a certain range and as close as possible to their expected values ​​when the operating environment changes, that is, to minimize the following performance indicators:

[0074]

[0075] In equation (1), k is the simulation time, and i is the start time; r(k) = [r1(k), r2(k)] T w(k) = [w1(k), w2(k)] T w1(k) represents the feed rate setpoint, w2(k) represents the mill inlet water makeup rate setpoint; R is a positive definite matrix; γ k-i r*(k) represents the weighting factor at time (ki); r*(k) represents the expected value of r(k); w min and w max These are the upper and lower limits for the feed rate setpoint and the mill inlet water supply setpoint.

[0076] Since the grinding particle size r1(k) and mill load r2(k) are closely related to the mill feed rate y1(k) and mill inlet water supply y2(k), in order to achieve the above objectives, it is necessary to optimize the loop setpoints w1(k) and w2(k) and track the setpoints by adjusting the feeder frequency u1(k) and mill inlet water supply valve opening u2(k) through the bottom loop.

[0077] (2) Dynamic characteristics analysis of the grinding process:

[0078] In the grinding process, coarse ore slurry is first ground to a specific discharge particle size r by a ball mill. m (k) then enters the classifier, and finally the classifier screens out the grinding particle size r1(k) that meets the standard.

[0079] Experiments show that for a constant discharge particle size r m For slurry of particle size (k), as the water addition at the classifier inlet increases, the overflow concentration gradually decreases, and the grinding particle size r1(k) first increases to its maximum value (critical point) and then decreases. For a constant overflow concentration, different discharge particle sizes and different return sand amounts result in different grinding particle sizes r1(k) and mill load r2(k). Therefore, the operating parameters can be expressed as:

[0080] X(k+1)=X(r(k),y(k),r m (k))

[0081] Where X is a nonlinear function, y(k) = [y1(k), y2(k)] TFor a closed-circuit grinding process, the mill discharge particle size r m (k) is not only related to the feed rate y1(k) and the mill inlet makeup water flow rate y2(k), but also depends on the return sand concentration v d (k), flow rate v q (k), and is also affected by the return sand particle size v of the classifier. r The influence of (k). Therefore, the discharge particle size in the grinding process is

[0082] r m (k)=g m (y1(k),y2(k),v r (k),v d (k))

[0083] Among them, g m This is a nonlinear function. Due to the frequent changes in the particle size of the returned sand from the classifier and the raw ore, it is difficult to describe using a mathematical model, resulting in a nonlinear function, g. m Difficult to describe using a mathematical model, the grinding operation index can be expressed as follows, as shown in the above formula:

[0084] X(k+1)=X(r(k),y(k),v(k))

[0085] Where: X is an unknown nonlinear function; y(k) = [y1(k), y2(k)] T v(k) = [v d (k),v q (k),v r (k)] T .

[0086] In summary, grinding operation indicators have complex and comprehensive characteristics that are highly nonlinear and difficult to describe using mathematical models.

[0087] S2. Using a greedy equivalence search algorithm, extract the causal dependencies between operating variables from historical operating data, construct a causal graph of the grinding process, and during online operation, based on the matching results of real-time operating features and the causal graph, select the Actor-Critic model (the Actor-Critic model is a reinforcement learning model) that best matches the current causal structure, and generate an optimized sequence of loop setpoints through a reinforcement learning algorithm.

[0088] When ore properties are stable and operating conditions change little, reinforcement learning can be used to calculate setpoints and continuously adjust them to adapt to environmental changes. However, in most domestic ore processing plants, the grinding process does not involve blending, leading to frequent fluctuations in ore properties and large dynamic changes in operating conditions. Reinforcement learning alone is insufficient for quickly adjusting setpoints. Therefore, this paper establishes multiple Actor-Critic reinforcement learning models for different operating conditions. During online execution, a case-based reasoning algorithm is used to select the most suitable Actor-Critic reinforcement learning model, thereby combining reinforcement learning algorithms to adjust the setpoints.

[0089] Specifically, the Actor network outputs the probability distribution of the setpoint adjustment action, the Critic network evaluates the long-term value function of the state action, and achieves setpoint optimization by maximizing the cumulative reward function.

[0090]

[0091] Where: θ k The parameters represent the policy parameters, where τ is the trajectory and γ is the tactical parameter. t ∈(0,1) is the discount factor, r(s) t ,a t ) represents the instantaneous reward at time t.

[0092] To improve the stability and efficiency of policy updates, natural policy gradients are used to update the Actor-Critic network parameters. Natural policy gradients are achieved by incorporating the Fisher information matrix G(θ). k The update rule for handling the curvature of the parameter space is as follows:

[0093] θ k+1 =θ k +αG -1 (θ k )▽ θ J(θ k (4)

[0094] In equation (4), α is the learning rate; Fisher's information matrix; Let J(θ) be the standard policy gradient, representing the policy performance. k Relative parameter θ k The gradient;

[0095] The causal reasoning module uses a greedy equivalence search algorithm to identify causal dependencies between operating condition variables from historical operating data and constructs a causal graph. Where V = {v1, v2, v3} is the set of nodes representing operating condition variables, v1 represents the expected value of mill load, v2 represents the expected value of grinding particle size, and v3 represents the raw ore particle size. E is the set of directed edges, representing the direction of causal interactions between variables. The similarity of causal graph matching is defined as:

[0096]

[0097] In equation (2), This is a cause-and-effect diagram showing the current operating conditions of the grinding mill process. This is the i-th causal graph in the causal graph of the historical operating data of the mill grinding process; ε c and ε i P represents the edge set of the current causal graph and the edge set of the i-th causal graph in the causal graph of the historical running data, respectively; c P represents the probability distribution of the current case, that is, the joint probability distribution of the variables in the causal graph under the current operating conditions of the grinding process; i KL(P) represents the probability distribution of the i-th historical case in the causal graph of historical operating data of the mill grinding process, that is, the joint probability distribution of each variable in the causal graph of historical operating data corresponding to the historical operating conditions of the mill grinding process; c ||P i ) represents the KL divergence between the conditional probability distributions of the causal mechanism; β is the decay coefficient;

[0098] During the model reuse phase, a causal similarity threshold η∈[0,1] is set. If there exists... If the parameters are selected for initialization, the corresponding Actor-Critic model parameters are chosen; otherwise, new policy network parameters are generated based on causal intervention counterfactual inference and fine-tuned using online natural policy gradients. During parameter updates, the Critic network uses temporal difference error optimization to estimate the value function.

[0099] δ k =r k +γV(s k+1 )-V(s k )

[0100] In the above formula, γ is the discount factor; r k For immediate reward; V(s) k V(s) is the state-value function estimate at time k; k+1 ) is the state value function estimate for time (k+1);

[0101] Step 1: Initialize the cause-effect graph set and the Actor-Critic model library;

[0102] Step 2: Collect real-time operating data and perform a greedy equivalence search to update the causal graph.

[0103] Step 3: Calculate the similarity between the current cause-effect graph and the case library, and select or generate a suitable Actor-Critic model.

[0104] Step 4: Update network parameters based on natural policy gradient and TD-error to achieve closed-loop optimization of set values ​​and model self-learning.

[0105] S3. To address the variable operating conditions in the grinding process, multiple Actor-Critic network models adapted to different operating conditions are constructed. Using a case reasoning mechanism, the optimal Actor-Critic model is matched from the case library based on the characteristics of the current operating conditions, ensuring the dynamic adaptability of the optimization strategy.

[0106] The goal of grinding setpoint optimization is to find the optimal loop setpoint y*(t) for a nonlinear system with an unknown model, such that the operating index tracks its expected value, i.e., minimizing the performance index function. Clearly, y*(k) is difficult to obtain by solving the Bellman equation; therefore, this paper employs a strategy iterative algorithm to find the solution to the Bellman equation.

[0107] 1) Policy Iteration Algorithm

[0108] The strategy iteration algorithm uses the Bellman equations to evaluate the current setpoint and updates the setpoint in the form of the optimal control solution to find an improved control strategy.

[0109] Algorithm 1: Iterative Algorithm for Indicator Control

[0110] Step 1: First, conduct a strategy evaluation and select a control strategy that will stabilize the system;

[0111]

[0112] Step 2: Next, implement the strategy improvement steps, using the following methods to determine the controls for improvement:

[0113]

[0114] Step 3: Stop when convergence is close enough.

[0115] 2) Behavioral evaluation structure of optimal solution for industrial process

[0116] The value function and setpoint function are approximated by two independent three-layer perceptrons, forming a structure called the behavior evaluation network. The evaluation network estimates the value function. The behavior network represents the control policy. Algorithm 2 provides a data-driven optimization algorithm for real-time performance-based control.

[0117] Let n be the number of hidden layer neurons in the evaluation network and the behavior network, respectively. c and na This indicates that the weights between the input and hidden layers of the two neural networks are represented by V. c and V a This indicates that the weights between the hidden layers and output layers of the two neural networks are represented by W. c and W a express.

[0118] The output of the Critic-NN network is given by the following formula:

[0119]

[0120] The output of the Actor-NN network is given by the following formula:

[0121]

[0122] Where φ c (·)=tanh(·) and φ a (·)=tanh(·) are the activation functions of Critic-NN and Actor-NN, respectively. The Bellman equation for the nonlinear optimal tracking problem is given by the following equation.

[0123]

[0124] To update the weights of the Critic-NN, the error of the Bellman equation is defined as:

[0125]

[0126] This is called the time difference error. If the nonlinear Bellman equation holds, the time difference error becomes zero. Therefore, the weights of the Critic-NN are adjusted to minimize the square of the Bellman error, which is given by the following equation:

[0127]

[0128] To update the weights of the Actor-NN, we define the Actor-NN error as follows using supervised learning:

[0129]

[0130] in It is the output of the Actor-NN, with X(k-1) as input and V as output. a (k),W a (k) weights, The target value is obtained from the above:

[0131]

[0132] Therefore, the weights of the Actor-NN can be adjusted to minimize the squared error of the Actor.

[0133]

[0134] The approximation applies to a two-layer neural network, where the input layer V a and hidden layer V c The weights between them are randomly selected, while the hidden layer weights W c and output layer weights W a It is adjustable.

[0135] Algorithm 2: A data-driven optimization method for real-time indicator control

[0136] Initialization: Select the initial stable settings for the system.

[0137] Step 1: First, conduct a strategy evaluation and determine the set values;

[0138] Step 2: Next, improve the strategy and determine the improved control strategy;

[0139] Step 3: Stop when convergence is close enough.

[0140] S4. Dynamic adjustment and intelligent control of setpoints. This dynamic adjustment is based on optimized loop setpoints, which consist of mill feed rate setpoints and mill inlet water flow rate setpoints. The joint adjustment of these two setpoints is used to precisely control the core operating indicators of the grinding process—grind particle size and mill load. This precise control relies on real-time acquired operating data, including multiple key variables such as mill load, grinding particle size, and raw ore particle size. Fluctuations in these key variables are corrected online through a feedback adjustment strategy. The adjustment strategy combines a multi-objective optimization method, dynamically calculates the optimal setpoint adjustment amount based on process constraints and operational stability requirements, and applies this setpoint adjustment amount to the bottom-level actuators through a closed-loop control system. These bottom-level actuators include a feed motor frequency adjustment device and a mill inlet water supply valve control system. The control system adjusts the working state of the actuators in real time to keep the grinding particle size and mill load within the set optimization range. The determination of this set optimization range comprehensively considers process requirements, energy consumption levels, and system dynamic characteristics, achieving adaptive optimization and intelligent adjustment of the grinding process.

[0141] In this specific embodiment, the controller of the mill control loop includes: a controller for the feed motor frequency and a controller for the mill water supply valve. The control loop for the mill feed rate and the control loop for the mill inlet water supply respectively act on the controller for the feed motor frequency and the controller for the mill water supply valve; these are used to apply the optimized setpoints w1 and w2 to the actuators to achieve precise control.

[0142] In this specific embodiment, the update mechanism of the causal graph of historical operating data includes: encapsulating the current operating condition features, optimized model weights and control effects into a new causal graph of historical operating data, and updating it to the case library (i.e., the causal graph storing historical operating data) according to similarity priority.

[0143] In this specific embodiment, such as Figure 2 As shown, the grinding process setpoint optimization control system of this invention, through a case selection and matching module, retrieves the Actor-Critic network model that best matches the current operating conditions from the case library based on real-time collected operating conditions (including key variables such as grinding particle size, mill load, and raw ore particle size), and performs case correction and storage to update model parameters. Subsequently, the controller calculates setpoints based on the selected Actor-Critic model, where the Actor network is used to generate the feed rate setpoints and mill inlet water supply setpoints for the equipment loop, and the Critic network is used to evaluate the cost function and expected reward value of the current control strategy. The generated setpoint vector acts on the underlying actuators, including the feed motor frequency controller and the water supply valve controller, through feedback closed-loop adjustment, to achieve real-time and accurate control of grinding process operating indicators (grinding particle size and mill load). The control results and new operating condition data are encapsulated into new cases and updated to the case library, forming a dynamic adaptive optimization closed loop, thereby improving the production efficiency and control stability of the grinding process.

[0144] In summary, this specific embodiment discloses an experience-driven adaptive reinforcement decision-making method for optimizing the setpoint of the grinding process. Through a synergistic control strategy of causal reasoning matching mechanism and natural policy gradient optimization, it solves the control lag problem caused by ore property fluctuations.

[0145] Specifically, the experience-driven adaptive reinforcement decision-making grinding process setpoint optimization method disclosed in this embodiment first extracts the causal dependencies between operating condition variables from historical operating data based on a greedy equivalence search algorithm, constructs a causal graph of the grinding process, and matches the optimal pre-trained Actor-Critic network parameters through similarity calculation. Then, the Actor network is updated using natural policy gradients to generate continuous adjustment sequences for the feed rate setpoint and the mill inlet water supply setpoint. The value function estimation of the Critic network is optimized using time-difference (TD) error optimization, achieving collaborative iterative optimization of the policy network and the value network. Further, the optimized setpoints are transmitted to the underlying closed-loop control system, and dynamic and precise control of key operating indicators such as grinding particle size and mill load is achieved by adjusting the feed motor frequency and the water supply valve opening. Simultaneously, the system dynamically updates the case library based on the control effect and new operating condition characteristics, and combines multi-objective optimization methods to correct the setpoint adjustment amount in real time, thereby improving the system's adaptive and steady-state performance under conditions of frequent fluctuations in ore properties.

[0146] Therefore, this invention significantly improves the control accuracy and dynamic response speed of the grinding process, providing a reliable solution for the intelligent control of mineral processing plants.

[0147] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. An experience-driven adaptive reinforcement decision-making method for optimizing the setpoint of a mill control loop, comprising: Step S1: Identify the operating indicators of the mill grinding process, including grinding particle size r1(k) and mill load r2(k); where mill load refers to the total mass of ore and media in the mill. Step S2: Based on the grinding particle size r1(k) and mill load r2(k), construct a performance index function for the grinding process setpoint optimization method based on experience-driven adaptive reinforcement decision-making, so as to optimize the feed rate setpoint w1(k) and mill inlet water supply setpoint w2(k) in the subsequent grinding process, thereby achieving convergence of the grinding particle size r1(k) and mill load r2(k); the mathematical expression corresponding to the performance index function is: In equation (1), k is the simulation time, and i is the start time; r(k) = [r1(k), r2(k)] T w(k) = [w1(k), w2(k)] T w1(k) represents the feed rate setpoint, w2(k) represents the mill inlet water makeup rate setpoint; R is a positive definite matrix; γ k-i r*(k) represents the weighting factor at time (ki); r*(k) represents the expected value of r(k); w min and w max These are the upper and lower limits for the feed rate setpoint and the mill inlet water supply setpoint. Step S3: Based on the case-based reasoning algorithm, extract the causal dependencies between the operating variables v1, v2, and v3 and the operating indices r1(k) and r2(k) from the historical operating data of the mill grinding process, in order to construct a causal graph of the historical operating data of the mill grinding process. Where V is the set of operating condition variable nodes, and V = {v1, v2, v3}; v1 represents the expected value of mill load; v2 represents the expected value of grinding particle size; v3 represents the raw ore particle size; E is the set of directed causal edges, which represents the set of directed causal edges between variables v1, v2, v3 in the set of mill grinding operating condition variable nodes V and operating indicators r1(k) and r2(k); Step S4: Calculate the similarity between the current causal graph of the mill grinding process and the i-th causal graph of the historical operating data; wherein, the formula for calculating the similarity is: In equation (2), This is a cause-and-effect diagram showing the current operating conditions of the grinding mill process. This is the i-th causal graph in the causal graph of the historical operating data of the mill grinding process; ε c and ε i P represents the edge set of the current causal graph and the edge set of the i-th causal graph in the causal graph of the historical running data, respectively; c P represents the probability distribution of the current case, that is, the joint probability distribution of the variables in the causal graph under the current operating conditions of the grinding process; i KL(P) represents the probability distribution of the i-th historical case in the causal graph of historical operating data of the mill grinding process, that is, the joint probability distribution of each variable in the causal graph of historical operating data corresponding to the historical operating conditions of the mill grinding process; c ||P i ) represents the KL divergence between the conditional probability distributions of the causal mechanism; β is the decay coefficient; Step S5: Construct an Actor-Critic model and train it based on historical operating condition variables v1, v2, and v3, historical operating indices r1(k) and r2(k), and historical setpoints w1(k) and w2(k). The Actor-Critic model includes a Critic network model and an Actor network model. The input to the Critic network is the current operating condition s of the mill. k s k = [v1,v2,v3,r1(k),r2(k)]; the output of the Critic network is the expected cumulative reward V(s) under the current operating condition of the mill. k The input to the Actor network is the current operating state s of the mill. k The output of the Actor network is a setpoint vector w(k) = [w1(k), w2(k)]. T ; The architecture of the Critic network model is as follows: V(s) k ) = MLP(s k MLP(·) is a multilayer perceptron network. The input dataset of MLP(·) is [s1,...,s] constructed from historical data. k The output dataset of MLP(·) is [V(s1),...,V(s)]. k )]; and according to formula (2) and using the gradient descent algorithm to pre-train MLP(·), and save the parameters of MLP(·); The architecture of the Actor network model is as follows: For s k mean For s k Variance, θ k Let θ be the policy parameter of the Actor network model. k The control strategy used to evaluate and optimize the state-to-setpoint performance of the grinding process of the mill; the input dataset of the Actor network model is [s1,...,s] constructed from historical data. k The output dataset of the Actor network model is [w(1),…,w(k)]; and the Actor network model is pre-trained according to formula (1) using the gradient descent algorithm, and the policy parameters θ of the Actor network model are saved. k ; Step S6: Set the cause-effect graph of the current operating conditions of the mill grinding process. The similarity threshold between the causal graphs and all causal graphs in the historical operational data is η; where the formula for calculating the similarity threshold η is: h=m ε -as ε (3) In equation (3), μ ε It is the set of edges ε of the causal graph of historical running data. i The mean; σ ε It is the set of edges ε of the causal graph of historical running data. i The standard deviation; α is the adjustment factor; Step S7: Based on the calculated similarity Using a similarity threshold η, a pre-trained Actor-Critic model is selected to best match the causal graph of the current working condition of the grinding process; where, when a similarity to a historical working condition is determined, the model is selected. If the Actor-Critic model at this time is selected as the pre-trained Actor-Critic model that best matches the causal graph of the current working condition of the grinding process, the parameters of the best-matching pre-trained Actor-Critic model are saved; when it is determined that there is similarity between multiple historical working conditions... Then, the Actor-Critic model corresponding to the historical working condition with the highest similarity is selected as the pre-trained Actor-Critic model that best matches the causal graph of the current working condition of the grinding process, and the parameters of the best-matching pre-trained Actor-Critic model are saved; if all When the gradient correction strategy is applied, the parameter θ is directly adjusted. k Perform online correction; Step S8: Utilize the natural policy gradient to update the policy parameters θ of the pre-trained Actor-Critic model that best matches the current causal graph of the grinding process in real time. k Wherein, the strategy parameter θ k The update formula is: In equation (4), α is the learning rate; Fisher's information matrix; Let J(θ) be the standard policy gradient, representing the policy performance. k Relative parameter θ k The gradient; in, Let be the strategy function, which represents the mill's state s during grinding operations. k The probability of selecting the setpoint vector w(k); policy function The update formula is: In equation (5), Q(s) k w(k) is the action value function, which is used to evaluate the mill's state s during grinding operation. k The long-term expected reward of the setpoint vector w(k); the action value function Q(s) k The update formula for w(k) is: Step S9: Calculate the setpoint vector based on formula (4) The setpoints w1(k) and w2(k) are generated in real time through the Actor network model. These setpoints are then transmitted to the controller in the mill control loop. The controller adjusts the frequency of the feed motor and the opening of the water supply valve in real time based on the setpoints w1(k) and w2(k) until the operating indicators, grinding particle size r1(k) and mill load r2(k), converge to the target range. The convergence of grinding particle size r1(k) and mill load r2(k) to the target range means that the grinding particle size r1(k) and mill load r2(k) satisfy the performance index function constructed in step S2.

2. The mill control loop setpoint optimization method based on experience-driven adaptive reinforcement decision-making according to claim 1, characterized in that, The adjustment factor α is 0.

9.

3. The mill control loop setpoint optimization method based on experience-driven adaptive reinforcement decision-making according to claim 1 or 2, characterized in that, The similarity threshold η ranges from [0,1].