Mill control loop set value optimization method based on empirical-driven adaptive enhancement decision
By using an experience-driven adaptive reinforcement decision-making method to optimize the mill control loop set values, combined with the Actor-Critic model and case-based reasoning algorithm, the feed rate and water replenishment set values are optimized in real time, solving the control lag problem under complex working conditions during the grinding process, achieving dynamic and precise control of the grinding process, and improving production efficiency and product quality.
Patent Information
- Application Number
- CN202510987459.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-17
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-07-17
AI Technical Summary
Existing grinding process control methods find it difficult to achieve dynamic and precise control of grinding set values under complex and changeable working conditions. Traditional reinforcement learning methods have slow learning convergence speed and insufficient adaptability in an environment where ore properties change frequently, resulting in control lag.
A mill control loop set value optimization method based on experience-driven adaptive reinforcement decision-making is adopted. By identifying the operating indicators of the mill grinding process, constructing a performance indicator function, and building an Actor-Critic model, the feed rate and water replenishment set values are optimized in real time using case-based reasoning algorithm and reinforcement learning technology to achieve convergence of grinding particle size and mill load.
It effectively solves the control lag problem of traditional methods under complex working conditions, realizes dynamic and precise regulation of the set value of the grinding process, improves production efficiency and product quality, and provides a new method for the intelligent control of the grinding process.
Smart Images

Figure CN120630716A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of optimized control of a grinding process, and in particular to a method for optimizing a setting value of a grinding mill control loop using experience-driven adaptive enhanced decision-making. Background Art
[0002] As a pillar industry of the national economy, the mining industry is at a critical juncture of deep transformation towards intelligence and high efficiency. In the field of mineral processing, the grinding link is the core of the entire mineral processing process. Its operating status not only directly affects the overall efficiency of downstream mineral processing operations, but also plays a decisive role in the quality of the final product. As key indicators for measuring the performance of the grinding process, the grinding particle size and circulation load not only affect the energy consumption level of the grinding process, but also have a profound impact on core production indicators such as the grade of the mineral processing concentrate and the metal recovery rate. Since the grinding particle size and circulation load are mainly controlled by the circuit set values such as the feed rate and the amount of added water, how to dynamically optimize these set values based on real-time working conditions has become a key issue in the current research field.
[0003] In the field of grinding process control, existing methods can be roughly categorized into two main categories: optimization methods based on mechanism models and data-driven methods based on artificial intelligence. Globally, some foreign mineral processing plants exhibit relatively stable grinding processes due to their high ore grade, uniform particle size distribution, relatively stable mineral composition, and the ability to achieve uniform ore properties through rational ore blending. Based on this, precise mathematical models can be established and mechanism-based methods such as real-time optimization (RTO), model predictive control (MPC), and multivariable decoupling control can be employed to optimize loop setpoints and achieve precise control of grinding particle size and circulating load. However, Chinese mineral processing plants primarily process complex ore resources such as hematite. These ore grades are generally low, with fine and uneven particle size distribution, and their compositional properties vary dramatically. This results in significant nonlinearity, strong coupling, and dynamic uncertainty in the grinding process. This complexity significantly limits the practical application of mathematical model-based optimization control methods in Chinese mineral processing plants, making it difficult to meet the setpoint optimization requirements in complex environments.
[0004] In recent years, researchers have been exploring artificial intelligence-based optimization strategies for grinding circuit setpoints, particularly those challenging to establish precise mechanistic models of the grinding process. For example, some studies employ multivariable fuzzy supervisory control methods, leveraging fuzzy inference rules to dynamically adjust grinding circuit setpoints to achieve adaptive regulation of the circulating load. Other studies have combined expert systems with fuzzy control technology, constructing intelligent decision-making systems based on knowledge bases, databases, and fuzzy logic inference engines to determine optimal circuit setpoints and thus achieve stable control of grinding particle size. Industrial applications have demonstrated that these methods have significantly improved mill production efficiency and effectively reduced specific energy consumption in mineral processing plants. However, these intelligent control methods still rely heavily on the experience and knowledge of domain experts during the design process, making their decision-making capabilities susceptible to human factors and unable to effectively adapt to complex and changing operating conditions. Furthermore, the adjustment strategies of these controllers lack adaptive capabilities, often requiring manual adjustments by experts in response to new operating conditions. This subjective judgment can lead to instability in the control strategy, thus compromising the overall optimization capability of the system.
[0005] In recent years, reinforcement learning (RL) technology has shown broad application potential in solving complex process control problems with unknown mechanisms and strong nonlinearity due to its adaptability and self-learning capabilities. However, in actual grinding production environments, the operating conditions of the grinding process are highly dynamically unstable due to the frequent and large changes in ore properties. This makes traditional RL methods face the problems of slow learning convergence and insufficient adaptability. Under drastically changing operating conditions, traditional RL has difficulty in achieving rapid environmental perception and real-time adjustment, making it difficult to adapt to the optimization control requirements under all working conditions, thereby limiting its application effect in the task of optimizing grinding set points. Summary of the Invention
[0006] The present invention aims to address the problems of existing set value optimization methods in the grinding process, such as difficulty in model establishment and poor adaptability when facing changes in ore properties and dynamic working conditions. A mill control loop set value optimization method with experience-driven adaptive enhanced decision-making is proposed.
[0007] To this end, the purpose of the present invention is to propose a mill control loop setpoint optimization method with experience-driven adaptive enhanced decision-making.
[0008] To achieve the above objectives, the technical solution of the present invention provides a method for optimizing the setpoints of a mill control loop using experience-driven adaptive reinforcement decision-making. The optimization method comprises:
[0009] Step S1: Identifying operating indicators of the mill grinding process, to identify operating indicators of the mill grinding process including grinding particle size r1(k) and mill load r2(k); wherein the mill load refers to the total mass of the ore and the medium in the mill;
[0010] Step S2: Based on the grinding particle size r1(k) and the mill load r2(k), a performance index function of the grinding process setting value optimization method based on experience-driven adaptive intensified decision-making is constructed to achieve subsequent optimization of the feed rate setting value w1(k) and the mill inlet water replenishment setting value w2(k) of the mill grinding process to achieve convergence of the grinding particle size r1(k) and the mill load r2(k); the mathematical expression corresponding to the performance index function is:
[0011]
[0012] In formula (1), k is the simulation time, i is the starting time; r(k) = [r1(k), r2(k)] T ; w(k)=[w1(k),w2(k)] T , w1(k) represents the set value of feed rate, w2(k) represents the set value of water supply at mill inlet; R is a positive definite matrix; γ k-i is the weighting factor at time (ki); r*(k) represents the expected value of r(k); w min and w max The upper and lower limits of the feed rate setting value and the mill inlet water supply setting value;
[0013] Step S3: Based on the case-based reasoning algorithm, the causal dependency relationship between the operating variables v1, v2, and v3 and the operating indicators r1(k) and r2(k) is extracted from the historical operating data of the mill grinding process to construct a causal graph of the historical operating data of the mill grinding process. Where V is the set of working condition variable nodes, and V = {v1, v2, v3}; v1 represents the expected value of mill load; v2 represents the expected value of grinding particle size; v3 represents the original ore particle size; E is a directed causal edge set, representing the directed causal edge set between variables v1, v2, v3 in the mill grinding condition variable node set V and the operating indicators r1(k) and r2(k);
[0014] Step S4: Calculate the similarity between the current operating condition causal graph of the grinding process of the mill and the ith causal graph in the causal graph of the historical operating data; wherein the calculation formula corresponding to the similarity is:
[0015]
[0016] In formula (2), The cause-effect diagram of the current working condition of the grinding process of the mill; is the i-th causal graph in the causal graph of the historical operation data of the grinding process of the mill; ε c and ε i P represents the edge set of the current causal graph and the edge set of the i-th causal graph in the causal graph of the historical operation data respectively; c represents the probability distribution of the current case, that is, the joint probability distribution of each variable in the causal diagram under the current working conditions of the grinding process; P i The probability distribution of the i-th historical case in the causal graph of the historical operating data of the grinding process of the mill is the joint probability distribution of each variable in the i-th causal graph of the historical operating data corresponding to the historical working conditions of the grinding process; KL(P c ||P i ) is the KL divergence between the conditional probability distributions of the causal mechanism; β is the decay coefficient;
[0017] Step S5: Build an Actor-Critic model and train it based on historical operating variables v1, v2, and v3, historical operating indicators r1(k) and r2(k), and historical set values w1(k) and w2(k); wherein the Actor-Critic model includes: a Critic network model and an Actor network model; the input of the Critic network is the current operating state s of the mill. k , s k =[v1,v2,v3,r1(k),r2(k)]; the output of the critic network is the expected cumulative reward V(s) under the current working state of the mill. k ); The input of the Actor network is the current working state of the mill s k The output of the Actor network is the set value vector w(k)=[w1(k),w2(k)] T ;
[0018] The architecture of the Critic network model is: V(s k )=MLP(s k ); MLP(·) is a multi-layer perceptron network. The input data set of MLP(·) is [s1,…,s k ], the output dataset of MLP(·) is [V(s1),…,V(s k )]; and pre-training the MLP(·) using the gradient descent algorithm according to formula (2), and saving the parameters of the MLP(·);
[0019] Among them, the architecture of the Actor network model is: For s k mean, For s k Variance, θ k is the strategy parameter of the Actor network model, strategy parameter θ k It is used to evaluate and optimize the performance of the control strategy from the state to the set value of the grinding process of the mill; the input data set of the Actor network model is [s1,...,s k ], the output data set of the Actor network model is [w(1),…,w(k)]; and the Actor network model is pre-trained using the gradient descent algorithm according to formula (1), and the strategy parameter θ of the Actor network model is saved k ;
[0020] Step S6: Setting the current operating condition cause-effect diagram of the grinding process The similarity threshold with all causal graphs in the historical operation data is η; the calculation formula of the similarity threshold η is:
[0021] η=μ ε -ασ ε (3)
[0022] In formula (3), μ ε is the edge set ε of the causal graph of historical operation data i The mean of ε is the edge set ε of the causal graph of historical operation data i The standard deviation of ;α is the adjustment coefficient;
[0023] Step S7: Based on the calculated similarity With the similarity threshold η, the pre-trained Actor-Critic model that best matches the causal graph of the current working condition of the grinding process is selected; wherein, when it is determined that there is a similarity of the historical working condition When the Actor-Critic model at this time is selected as the pre-trained Actor-Critic model that best matches the causal diagram of the current working condition of the grinding process, and the parameters of the pre-trained Actor-Critic model that best matches are saved; when it is determined that there are multiple similarities in the historical working conditions The Actor-Critic model corresponding to the historical working condition with the greatest similarity is selected as the pre-trained Actor-Critic model that best matches the causal graph of the current working condition of the grinding process, and the parameters of the pre-trained Actor-Critic model that best matches are saved; if all When , the parameter θ is directly corrected by the gradient correction strategy k Perform online correction;
[0024] Step S8: Using natural policy gradient to update in real time the policy parameter θ of the pre-trained Actor-Critic model that best matches the causal graph of the current working conditions of the grinding process k ; Wherein, the strategy parameter θ k The update formula is:
[0025] θ k+1 =θ k +αG -1 (θ k )▽ θ J(θ k )(4)
[0026] In formula (4), α is the learning rate; is the Fisher information matrix; is the standard policy gradient, which represents the policy performance J(θ k ) relative parameter θ k gradient;
[0027] in, is the strategy function, which represents the grinding mill in the grinding state s k The probability of selecting the set value vector w(k); the policy function The update formula is:
[0028]
[0029] In formula (5), Q(s k ,w(k)) is the action value function, which is used to evaluate the grinding mill in the grinding condition s k The long-term expected reward of executing the set value vector w(k); the action value function Q(s k , the update formula of w(k)) is:
[0030]
[0031] Step S9: Calculate the set value vector based on formula (4) The set values w1(k) and w2(k) are generated in real time through the Actor network model, and then the real-time generated set values w1(k) and w2(k) are transmitted to the controller of the mill control loop, so that the controller can adjust the frequency of the feeding motor and the opening of the water supply valve in real time according to the real-time generated set values w1(k) and w2(k), until the operating indicators of the grinding particle size r1(k) and the mill load r2(k) converge to the target range; wherein, the grinding particle size r1(k) and the mill load r2(k) converge to the target range, that is, the grinding particle size r1(k) and the mill load r2(k) meet the performance indicator function constructed in step S2.
[0032] Preferably, the adjustment coefficient α is 0.9.
[0033] Preferably, the value range of the similarity threshold η is [0, 1].
[0034] Beneficial effects of the present invention:
[0035] The present invention provides an experience-driven, adaptive, and reinforced decision-making method for optimizing mill control loop setpoints, adapting to dynamic operating conditions characterized by frequent changes in ore properties and strong nonlinearity. Specifically, this optimization method rapidly locks in historical experience through a case-matching mechanism. Combined with the online optimization characteristics of reinforcement learning, it effectively addresses the control lag caused by sudden changes in operating conditions in traditional methods. This method enables dynamic and precise control of grinding process setpoints, improving both production efficiency and product quality, and providing a novel approach to intelligent grinding process control.
[0036] Additional aspects and advantages of the invention will become apparent from the description which follows, or may be learned by practice of the invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 A schematic flow chart showing a method for optimizing a mill control loop setpoint using experience-driven adaptive enhanced decision-making according to an embodiment of the present invention is provided;
[0038] Figure 2 A process flow chart of a grinding process according to an embodiment of the present invention is shown;
[0039] Figure 3 A schematic structural block diagram of a method for optimizing a mill control loop setting value with experience-driven adaptive enhanced decision-making according to an embodiment of the present invention is shown. DETAILED DESCRIPTION
[0040] In order to more clearly understand the above-mentioned objects, features and advantages of the present invention, Figures 1 to 3 As shown, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that, in the absence of conflict, the embodiments of the present application and the features in the embodiments can be combined with each other.
[0041] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Therefore, the scope of protection of the present invention is not limited to the specific embodiments disclosed below.
[0042] Figure 1 FIG. 1 is a schematic flow chart showing a method for optimizing a mill control loop setting value using experience-driven adaptive reinforcement decision making according to an embodiment of the present invention. Figure 1As shown in FIG, the mill control loop set value optimization method based on experience-driven adaptive enhanced decision-making includes:
[0043] Step S1: Identifying operating indicators of the mill grinding process, to identify operating indicators of the mill grinding process including grinding particle size r1(k) and mill load r2(k); wherein the mill load refers to the total mass of the ore and the medium in the mill;
[0044] Step S2: Based on the grinding particle size r1(k) and the mill load r2(k), a performance index function of the grinding process setting value optimization method based on experience-driven adaptive intensified decision-making is constructed to achieve subsequent optimization of the feed rate setting value w1(k) and the mill inlet water replenishment setting value w2(k) of the mill grinding process to achieve convergence of the grinding particle size r1(k) and the mill load r2(k); the mathematical expression corresponding to the performance index function is:
[0045]
[0046] In formula (1), k is the simulation time, i is the starting time; r(k) = [r1(k), r2(k)] T ; w(k)=[w1(k),w2(k)] T , w1(k) represents the set value of feed rate, w2(k) represents the set value of water supply at mill inlet; R is a positive definite matrix; γ k-i is the weighting factor at time (ki); r*(k) represents the expected value of r(k); w min and w max The upper and lower limits of the feed rate setting value and the mill inlet water supply setting value;
[0047] Step S3: Based on the case-based reasoning algorithm, the causal dependency relationship between the operating variables v1, v2, and v3 and the operating indicators r1(k) and r2(k) is extracted from the historical operating data of the mill grinding process to construct a causal graph of the historical operating data of the mill grinding process. Where V is the set of working condition variable nodes, and V = {v1, v2, v3}; v1 represents the expected value of mill load; v2 represents the expected value of grinding particle size; v3 represents the original ore particle size; E is a directed causal edge set, representing the directed causal edge set between variables v1, v2, v3 in the mill grinding condition variable node set V and the operating indicators r1(k) and r2(k);
[0048] Step S4: Calculate the similarity between the current operating condition causal graph of the grinding process of the mill and the ith causal graph in the causal graph of the historical operating data; wherein the calculation formula corresponding to the similarity is:
[0049]
[0050] In formula (2), The cause-effect diagram of the current working condition of the grinding process of the mill; is the i-th causal graph in the causal graph of the historical operation data of the grinding process of the mill; ε c and ε i P represents the edge set of the current causal graph and the edge set of the i-th causal graph in the causal graph of the historical operation data respectively; c represents the probability distribution of the current case, that is, the joint probability distribution of each variable in the causal diagram under the current working conditions of the grinding process; P i The probability distribution of the i-th historical case in the causal graph of the historical operating data of the grinding process of the mill is the joint probability distribution of each variable in the i-th causal graph of the historical operating data corresponding to the historical working conditions of the grinding process; KL(P c ||P i ) is the KL divergence between the conditional probability distributions of the causal mechanism; β is the decay coefficient;
[0051] Step S5: Build an Actor-Critic model and train it based on historical operating variables v1, v2, and v3, historical operating indicators r1(k) and r2(k), and historical set values w1(k) and w2(k); wherein the Actor-Critic model includes: a Critic network model and an Actor network model; the input of the Critic network is the current operating state s of the mill. k , s k =[v1,v2,v3,r1(k),r2(k)]; the output of the critic network is the expected cumulative reward V(s) under the current working state of the mill. k ); The input of the Actor network is the current working state of the mill s k The output of the Actor network is the set value vector w(k)=[w1(k),w2(k)] T ;
[0052] Among them, the architecture of the Critic network model is: V(s k )=MLP(s k ); MLP(·) is a multi-layer perceptron network. The input data set of MLP(·) is [s1,…,s k ], the output dataset of MLP(·) is [V(s1),…,V(s k )]; and pre-training the MLP(·) using the gradient descent algorithm according to formula (2), and saving the parameters of the MLP(·);
[0053] Among them, the architecture of the Actor network model is: For s k mean, For s k Variance, θ k is the strategy parameter of the Actor network model, strategy parameter θ k It is used to evaluate and optimize the performance of the control strategy from the state to the set value of the grinding process of the mill; the input data set of the Actor network model is [s1,…,s k ], the output data set of the Actor network model is [w(1),…,w(k)]; and the Actor network model is pre-trained using the gradient descent algorithm according to formula (1), and the strategy parameter θ of the Actor network model is saved k ;
[0054] Step S6: Setting the current operating condition cause-effect diagram of the grinding process The similarity threshold with all causal graphs in the historical operation data is η; the calculation formula of the similarity threshold η is:
[0055] η=μ ε -ασ ε (3)
[0056] In formula (3), μ ε is the edge set ε of the causal graph of historical operation data i The mean of ε is the edge set ε of the causal graph of historical operation data i The standard deviation of ;α is the adjustment coefficient;
[0057] Step S7: Based on the calculated similarity With the similarity threshold η, the pre-trained Actor-Critic model that best matches the causal graph of the current working condition of the grinding process is selected; wherein, when it is determined that there is a similarity of the historical working condition When the Actor-Critic model at this time is selected as the pre-trained Actor-Critic model that best matches the causal diagram of the current working condition of the grinding process, and the parameters of the pre-trained Actor-Critic model that best matches are saved; when it is determined that there are multiple similarities in the historical working conditions The Actor-Critic model corresponding to the historical working condition with the greatest similarity is selected as the pre-trained Actor-Critic model that best matches the causal graph of the current working condition of the grinding process, and the parameters of the pre-trained Actor-Critic model that best matches are saved; if all When , the parameter θ is directly corrected by the gradient correction strategy k Perform online correction;
[0058] Step S8: Use natural policy gradient to update the policy parameter θ of the pre-trained Actor-Critic model that best matches the causal graph of the current working conditions of the grinding process in real time k ; Wherein, the strategy parameter θ k The update formula is:
[0059] θ k+1 =θ k +αG -1 (θ k )▽ θ J(θ k )(4)
[0060] In formula (4), α is the learning rate; is the Fisher information matrix; is the standard policy gradient, which represents the policy performance J(θ k ) relative parameter θ k gradient;
[0061] in, is the strategy function, which represents the grinding mill in the grinding state s k The probability of selecting the set value vector w(k); the policy function The update formula is:
[0062]
[0063] In formula (5), Q(s k ,w(k)) is the action value function, which is used to evaluate the grinding mill in the grinding condition s k The long-term expected reward of executing the set value vector w(k); the action value function Q(s k , the update formula of w(k)) is:
[0064]
[0065] Step S9: Calculate the set value vector based on formula (4) The set values w1(k) and w2(k) are generated in real time through the Actor network model, and then the real-time generated set values w1(k) and w2(k) are transmitted to the controller of the mill control loop, so that the controller can adjust the frequency of the feeding motor and the opening of the water supply valve in real time according to the real-time generated set values w1(k) and w2(k), until the operating indicators of the grinding particle size r1(k) and the mill load r2(k) converge to the target range; wherein, the grinding particle size r1(k) and the mill load r2(k) converge to the target range, that is, the grinding particle size r1(k) and the mill load r2(k) meet the performance indicator function constructed in step S2.
[0066] In this embodiment, the experience-driven adaptive reinforcement decision-making method for optimizing the mill control loop set values combines case-based reasoning and actor-critic reinforcement learning technology. Through case-based reasoning, the actor-critic model suitable for the current working conditions is selected, and the reinforcement learning algorithm is used to optimize the loop set values, thereby realizing online optimization control of the grinding process.
[0067] In one embodiment of the present invention, the adjustment coefficient α is 0.9.
[0068] In one embodiment of the present invention, the value range of the similarity threshold η is [0, 1].
[0069] like Figure 2 and Figure 3 As shown, the technical solution of the present invention will be demonstrated below with a specific embodiment. The experience-driven adaptive strengthening decision-making mill control loop set value optimization method is specifically implemented by the following steps:
[0070] S1. Characterizing the Grinding Process Problems and Analyzing Key Operating Indicators. This study focuses on the dynamic characteristics of the grinding system, focusing on two core operating indicators: grinding particle size and mill load. By analyzing their relationships with key variables such as feed rate and mill inlet water replenishment, the primary factors influencing grinding performance and energy consumption are explored. Based on this, and incorporating process constraints and optimization objectives, the authors investigate how to rationally adjust circuit setpoints to improve grinding process stability and efficiency, achieving precise control of grinding particle size and mill load.
[0071] (1) Description of grinding problem:
[0072] like Figure 2 In the typical single-stage closed-circuit grinding process shown, the ball mill serves as the core control unit of the operating layer. Its output indicators include grinding particle size r1(k) (reflecting product quality) and mill load r2(k) (reflecting operating efficiency). The operating layer's inputs are derived from the setpoints of the underlying circuits, including the feed rate setpoint w1(k) and the mill inlet water setpoint w2(k).
[0073] Grinding particle size r1(k) and mill load r2(k) not only reflect the energy consumption level of grinding, but also affect important production indicators such as concentrate grade and metal recovery rate of the dressing plant. Therefore, they are regarded as two key operating indicators of grinding production. Figure 2As can be seen from the figure, operating indicators that are too large or too small will lead to suboptimal operation of the grinding process. Therefore, if we want to improve the quality and efficiency of the grinding process, it is necessary to control the operating indicators of grinding particle size r1(k) and mill load r2(k) within a certain range when the working environment changes, and keep them as close to their expected values as possible, that is, to achieve the minimum of the following performance indicators:
[0074]
[0075] In formula (1), k is the simulation time, i is the starting time; r(k) = [r1(k), r2(k)] T ; w(k)=[w1(k),w2(k)] T , w1(k) represents the set value of feed rate, w2(k) represents the set value of water supply at mill inlet; R is a positive definite matrix; γ k-i is the weighting factor at time (ki); r*(k) represents the expected value of r(k); w min and w max The upper and lower limits of the feed rate setting value and the mill inlet water supply setting value;
[0076] Since the grinding particle size r1(k) and mill load r2(k) are closely related to the mill feed rate y1(k) and the mill inlet water supply y2(k), in order to achieve the above goals, it is necessary to optimize the loop set values w1(k) and w2(k), and adjust the feeder frequency u1(k) and the mill inlet water supply flow valve opening u2(k) through the underlying loop to track the set values.
[0077] (2) Analysis of dynamic characteristics of grinding process:
[0078] In the grinding operation, the coarse ore pulp is first ground into a specific discharge particle size by a ball mill. m (k), and then enters the classifier, and finally the classifier screens out the ore pulp r1(k) that meets the grinding particle size standards.
[0079] The experiment shows that for a constant discharge particle size r m For a slurry with a specific flow rate (k), as the amount of water added at the classifier inlet increases, the overflow concentration gradually decreases, and the grinding particle size r1(k) first increases to a maximum value (critical point) and then decreases. For a constant overflow concentration, different discharge particle sizes, different return sand amounts, different grinding particle sizes r1(k) and mill load r2(k) also vary. Therefore, the operating indicators can be expressed as:
[0080] X(k+1)=X(r(k),y(k),r m (k))
[0081] Where: X is a nonlinear function, y(k) = [y1(k), y2(k)] TFor a closed-circuit grinding process, the mill discharge particle size r m (k) is not only related to the feed rate y1(k) and the water flow rate y2(k) at the mill inlet, but also depends on the return sand concentration v d (k), flow v q (k), and is also affected by the return sand particle size v of the classifier r (k). Therefore, the discharge particle size of the grinding process is
[0082] r m (k) = g m (y1(k),y2(k),v r (k),v d (k))
[0083] Among them, g m It is a nonlinear function. Since the return sand of the classifier and the particle size of the raw ore change frequently, it is difficult to describe it with a mathematical model, resulting in a nonlinear function, g m It is difficult to describe it with a mathematical model. From the above formula, it can be seen that the grinding operation index can be expressed as
[0084] X(k+1)=X(r(k),y(k),v(k))
[0085] Where: X is an unknown nonlinear function; y(k) = [y1(k), y2(k)] T ; v(k)=[v d (k),v q (k),v r (k)] T .
[0086] In summary, grinding operation indicators have complex comprehensive characteristics that are strongly nonlinear and difficult to describe with mathematical models.
[0087] S2. Use the greedy equivalent search algorithm to extract the causal dependency between operating variables from historical operating data and construct a causal graph for the grinding process. During online operation, based on the matching results between the real-time operating condition characteristics and the causal graph, select the Actor-Critic model (the Actor-Critic model is a reinforcement learning model) that best fits the current causal structure. Generate an optimized sequence of loop setting values through the reinforcement learning algorithm.
[0088] When ore properties are stable and operating conditions vary slightly, using reinforcement learning to calculate setpoints allows for continuous adjustment to environmental changes. However, most domestic mineral processing plants lack blending in their grinding processes, resulting in frequent fluctuations in ore properties and significant dynamic changes in operating conditions. This makes it difficult to quickly adjust setpoints using reinforcement learning alone. Therefore, this paper establishes multiple actor-critic reinforcement learning models for different operating conditions. During online operation, a case-based reasoning algorithm is used to select the most suitable actor-critic reinforcement learning model, allowing the reinforcement learning algorithm to be used to adjust the setpoints.
[0089] Specifically, the Actor network outputs the probability distribution of the set value adjustment action, and the Critic network evaluates the long-term value function of the state action, achieving set value optimization by maximizing the cumulative reward function:
[0090]
[0091] Where: θ k represents the policy parameters, τ is the trajectory, γ t ∈(0,1) is the discount factor, r(s t ,a t ) is the immediate reward at time t.
[0092] In order to improve the stability and efficiency of policy updates, the natural policy gradient is used to update the Actor-Critic network parameters. The natural policy gradient introduces the Fisher information matrix G(θ k ) processes the parameter space curvature, and its update rule is:
[0093] θ k+1 =θ k +αG -1 (θ k )▽ θ J(θ k )(4)
[0094] In formula (4), α is the learning rate; is the Fisher information matrix; is the standard policy gradient, which represents the policy performance J(θ k ) relative parameter θ k gradient;
[0095] The causal reasoning module uses the greedy equivalent search algorithm to identify the causal dependency between operating variables from historical operating data and construct a causal graph. Where V = {v1, v2, v3} is the set of operating condition variable nodes, v1 represents the expected value of mill load, v2 represents the expected value of grinding particle size, and v3 represents the raw ore particle size. E is a set of directed edges, representing the direction of causal interactions between variables. The causal graph matching similarity is defined as:
[0096]
[0097] In formula (2), The cause-effect diagram of the current working condition of the grinding process of the mill; is the i-th causal graph in the causal graph of the historical operation data of the grinding process of the mill; ε c and ε i P represents the edge set of the current causal graph and the edge set of the i-th causal graph in the causal graph of the historical operation data respectively; c represents the probability distribution of the current case, that is, the joint probability distribution of each variable in the causal diagram under the current working conditions of the grinding process; P i The probability distribution of the i-th historical case in the causal graph of the historical operating data of the grinding process of the mill is the joint probability distribution of each variable in the i-th causal graph of the historical operating data corresponding to the historical working conditions of the grinding process; KL(P c ||P i ) is the KL divergence between the conditional probability distributions of the causal mechanism; β is the decay coefficient;
[0098] In the model reuse phase, set the causal similarity threshold η∈[0,1]. If there is The corresponding Actor-Critic model parameters are selected for initialization; otherwise, new policy network parameters are generated based on causal intervention counterfactual reasoning and fine-tuned through online natural policy gradient. During the parameter update process, the Critic network uses the temporal difference error to optimize the value function estimation:
[0099] δ k =r k +γV(s k+1 )-V(s k )
[0100] In the above formula, γ is the discount factor; r k is the immediate reward; V(s k ) is the state value function estimate at time k; V(s k+1 ) is the state value function estimate at time (k+1);
[0101] Step 1: Initialize the causal graph set and the Actor-Critic model library;
[0102] Step 2: Collect real-time operating data and perform greedy equivalent search to update the causal graph
[0103] Step 3: Calculate the similarity between the current causal graph and the case library, and select or generate an adapted Actor-Critic model
[0104] Step 4: Update network parameters based on natural policy gradient and TD-error to achieve closed-loop optimization of set values and model self-learning.
[0105] S3. In response to the variable working conditions of the grinding process, multiple Actor-Critic network models that adapt to different operating conditions are constructed. The case-based reasoning mechanism is used to match the optimal Actor-Critic model from the case library based on the current working condition characteristics to ensure the dynamic adaptability of the optimization strategy.
[0106] The goal of grinding setpoint optimization is to find the optimal loop setpoint y*(t) for a nonlinear system with an unknown model, such that the operating indicators track their expected values, i.e., minimize the performance indicator function. Obviously, y*(k) is difficult to obtain by solving the Bellman equation. Therefore, this paper uses a policy iteration algorithm to find a solution to the Bellman equation.
[0107] 1) Policy Iteration Algorithm
[0108] The policy iteration algorithm utilizes the Bellman equation to evaluate the current set point and updates the set point in the form of an optimal control solution to find an improved control policy.
[0109] Algorithm 1: Policy iteration algorithm for indicator control
[0110] Step 1: First, conduct strategy evaluation and select a control strategy that stabilizes the system.
[0111]
[0112] Step 2: Next, proceed to the strategy improvement step and use the following methods to determine the improved control:
[0113]
[0114] Step 3: Stop when the convergence is close enough.
[0115] 2) Behavioral evaluation structure for optimal solutions of industrial processes
[0116] The value function and setpoint function are approximated by two independent three-layer perceptrons, forming a structure called a behavioral evaluation network. The evaluation network estimates the value function. The behavioral network represents the control policy. Algorithm 2 provides a data-driven optimization algorithm for real-time indicator control.
[0117] Assume that the number of hidden layer neurons of the evaluation network and the behavior network are n c and na Indicates that the weights between the input layer and the hidden layer of the two NNs are represented by V c and V a Indicates that the weights between the hidden layer and the output layer of the two NNs are represented by W c and W a express.
[0118] The output of the Critic-NN network is given by:
[0119]
[0120] The output of the Actor-NN network is given by:
[0121]
[0122] where φ c (·)=tanh(·) and φ a (·) = tanh(·) are the activation functions of Critic-NN and Actor-NN respectively. The Bellman equation for the nonlinear optimal tracking problem is given by the following equation.
[0123]
[0124] To update the weights of Critic-NN, the error of the Bellman equation is defined as:
[0125]
[0126] This is called the temporal difference error. If the nonlinear Bellman equation holds, the temporal difference error becomes 0. Therefore, the weights of the Critic-NN are adjusted to minimize the square of the Bellman error, which is given by:
[0127]
[0128] To update the weights of Actor-NN, we define the Actor-NN error in a supervised learning manner as
[0129]
[0130] in is the output of Actor-NN, X(k-1) as input and V a (k),W a (k) weight, is the target value, from the above we can get:
[0131]
[0132] Therefore, the weights of the Actor-NN can be adjusted to minimize the squared error of the Actor
[0133]
[0134] The approximation is applicable to a two-layer neural network where the input layer V a and hidden layer V c The weights between are randomly selected, and the hidden layer weights W c and the output layer weight W a It is adjustable.
[0135] Algorithm 2: Data-driven optimization method for real-time indicator control
[0136] Initialization: Select the system initial stable setting value
[0137] Step 1: First, conduct a strategy evaluation and determine the set value;
[0138] Step 2: Secondly, improve the strategy and determine the improved control strategy;
[0139] Step 3: Stop when the convergence is close enough.
[0140] S4. Dynamic adjustment and intelligent control of set values. The dynamic adjustment is based on the optimized loop set value, which consists of the mill feed set value and the mill inlet water flow set value. The joint adjustment of the mill feed set value and the mill inlet water flow set value is used to accurately control the core operating indicators of the grinding process - grinding particle size and mill load. The precise control relies on the real-time collected working condition data, which includes multiple key variables such as mill load, grinding particle size, and raw ore particle size. The fluctuation of the key variable is corrected online through the feedback adjustment strategy. The regulation strategy combines a multi-objective optimization method and dynamically calculates the optimal set value adjustment according to process constraints and operational stability requirements. The set value adjustment acts on the underlying actuator through a closed-loop control system. The underlying actuator includes a feed motor frequency adjustment device and a mill inlet water supply valve control system. The control system adjusts the working state of the actuator in real time to ensure that the grinding particle size and mill load are always maintained within the set optimization range. The determination of the set optimization range comprehensively considers process requirements, energy consumption levels and system dynamic characteristics, thereby realizing adaptive optimization and intelligent regulation of the grinding process.
[0141] In this specific embodiment, the mill control circuit controllers include a feed motor frequency controller and a mill water supply valve controller. The mill feed rate control circuit and the mill inlet water supply control circuit act on the feed motor frequency controller and the mill water supply valve controller, respectively, to apply the optimized setpoints w1 and w2 to the actuators for precise control.
[0142] In this specific embodiment, the update mechanism of the causal graph of historical operating data includes: encapsulating the current operating condition characteristics, optimized model weights and control effects into a new causal graph of historical operating data, and updating it to the case library (i.e., storing the causal graph of historical operating data) according to the similarity priority.
[0143] In this specific embodiment, Figure 2 As shown, the grinding process set value optimization control system of the present invention uses a case selection and matching module to retrieve the Actor-Critic network model that best matches the current working condition in the case library based on the working condition characteristics (including key variables such as grinding particle size, mill load, and raw ore particle size) collected in real time, and performs case correction and storage to update the model parameters. Subsequently, the controller calculates the set value based on the selected Actor-Critic model, where the Actor network is used to generate the feed rate set value and the mill inlet water supply set value of the equipment loop, and the Critic network is used to evaluate the cost function and expected reward value of the current control strategy. The generated set value vector acts on the underlying actuator through feedback closed-loop regulation, including the feed motor frequency controller and the water supply valve controller, to achieve real-time and precise control of the grinding process operating indicators (grinding particle size and mill load). The control results and the new working condition data are packaged together as a new case and updated to the case library to form a dynamic adaptive optimization closed loop, thereby improving the production efficiency and control stability of the grinding process.
[0144] In summary, this specific embodiment discloses a method for optimizing grinding process setpoints with experience-driven adaptive intensified decision-making, which solves the control lag problem caused by fluctuations in ore properties through a collaborative control strategy combining a causal reasoning matching mechanism with natural strategy gradient optimization.
[0145] Specifically, the experience-driven adaptive reinforcement decision-making grinding process set value optimization method disclosed in this specific embodiment first extracts the causal dependency between operating variables from historical operating data based on the greedy equivalent search algorithm, constructs a causal graph of the grinding process, and matches the optimal pre-trained Actor-Critic network parameters through similarity calculation; then, the natural policy gradient is used to update the Actor network to generate a continuous adjustment sequence of the feed set value and the mill inlet water set value, and the value function estimation of the Critic network is optimized in combination with the time difference (TD) error to achieve collaborative iterative optimization of the policy network and the value network. Furthermore, the optimized set value is transmitted to the underlying closed-loop control system, and by adjusting the feed motor frequency and the water valve opening, dynamic and precise control of key operating indicators such as grinding particle size and mill load is achieved. At the same time, the system dynamically updates the case library based on the control effect and the new operating condition characteristics, and combines the multi-objective optimization method to correct the set value adjustment amount in real time, thereby improving the system's adaptive and steady-state performance under conditions of frequent fluctuations in ore properties.
[0146] Therefore, the present invention significantly improves the control accuracy and dynamic response speed of the grinding process, and provides a reliable solution for the intelligent control of the mineral processing plant.
[0147] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
Claims
1. A method for optimizing mill control loop setpoints using experience-driven adaptive reinforcement decision-making, comprising: Step S1: Identifying operating indicators of the mill grinding process, to identify operating indicators of the mill grinding process including grinding particle size r1(k) and mill load r2(k); wherein the mill load refers to the total mass of the ore and the medium in the mill; Step S2: Based on the grinding particle size r1(k) and the mill load r2(k), a performance index function of the grinding process setting value optimization method based on experience-driven adaptive intensified decision-making is constructed to achieve subsequent optimization of the feed rate setting value w1(k) and the mill inlet water replenishment setting value w2(k) of the mill grinding process to achieve convergence of the grinding particle size r1(k) and the mill load r2(k); the mathematical expression corresponding to the performance index function is: In formula (1), k is the simulation time, i is the starting time; r(k) = [r1(k), r2(k)] T ; w(k)=[w1(k),w2(k)] T , w1(k) represents the set value of feed rate, w2(k) represents the set value of water supply at mill inlet; R is a positive definite matrix; γ k-i is the weighting factor at time (ki); r*(k) represents the expected value of r(k); w min and w max The upper and lower limits of the feed rate setting value and the mill inlet water supply setting value; Step S3: Based on the case-based reasoning algorithm, the causal dependency relationship between the operating variables v1, v2, and v3 and the operating indicators r1(k) and r2(k) is extracted from the historical operating data of the mill grinding process to construct a causal graph of the historical operating data of the mill grinding process. Where V is the set of working condition variable nodes, and V = {v1, v2, v3}; v1 represents the expected value of mill load; v2 represents the expected value of grinding particle size; v3 represents the original ore particle size; E is a directed causal edge set, representing the directed causal edge set between variables v1, v2, v3 in the mill grinding condition variable node set V and the operating indicators r1(k) and r2(k); Step S4: Calculate the similarity between the current operating condition causal graph of the grinding process of the mill and the ith causal graph in the causal graph of the historical operating data; wherein the calculation formula corresponding to the similarity is: In formula (2), The cause-effect diagram of the current working condition of the grinding process of the mill; is the i-th causal graph in the causal graph of the historical operation data of the grinding process of the mill; ε c and ε i P represents the edge set of the current causal graph and the edge set of the i-th causal graph in the causal graph of the historical operation data respectively; c represents the probability distribution of the current case, that is, the joint probability distribution of each variable in the causal diagram under the current working conditions of the grinding process; P i The probability distribution of the i-th historical case in the causal graph of the historical operating data of the grinding process of the mill is the joint probability distribution of each variable in the i-th causal graph of the historical operating data corresponding to the historical working conditions of the grinding process; KL(P c ||P i ) is the KL divergence between the conditional probability distributions of the causal mechanism; β is the decay coefficient; Step S5: Build an Actor-Critic model and train it based on historical operating variables v1, v2, and v3, historical operating indicators r1(k) and r2(k), and historical set values w1(k) and w2(k); wherein the Actor-Critic model includes: a Critic network model and an Actor network model; the input of the Critic network is the current operating state s of the mill. k , s k =[v1,v2,v3,r1(k),r2(k)]; the output of the critic network is the expected cumulative reward V(s) under the current working state of the mill. k ); The input of the Actor network is the current working state of the mill s k The output of the Actor network is the set value vector w(k)=[w1(k),w2(k)] T ; The architecture of the Critic network model is: V(s k )=MLP(s k ); MLP(·) is a multi-layer perceptron network. The input data set of MLP(·) is [s1,...,s k ], the output dataset of MLP(·) is [V(s1),...,V(s k )]; and pre-training the MLP(·) using the gradient descent algorithm according to formula (2), and saving the parameters of the MLP(·); Among them, the architecture of the Actor network model is: For s k mean, For s k Variance, θ k is the strategy parameter of the Actor network model, strategy parameter θ k It is used to evaluate and optimize the performance of the control strategy from the state to the set value of the grinding process of the mill; the input data set of the Actor network model is [s1,...,s k ], the output data set of the Actor network model is [w(1),…,w(k)]; and the Actor network model is pre-trained using the gradient descent algorithm according to formula (1), and the strategy parameter θ of the Actor network model is saved k ; Step S6: Setting the current operating condition cause-effect diagram of the grinding process The similarity threshold with all causal graphs in the historical operation data is η; the calculation formula of the similarity threshold η is: h=m ε -as ε (3) In formula (3), μ ε is the edge set ε of the causal graph of historical operation data i The mean of ε is the edge set ε of the causal graph of historical operation data i The standard deviation of ;α is the adjustment coefficient; Step S7: Based on the calculated similarity With the similarity threshold η, the pre-trained Actor-Critic model that best matches the causal graph of the current working condition of the grinding process is selected; wherein, when it is determined that there is a similarity of the historical working condition When the Actor-Critic model at this time is selected as the pre-trained Actor-Critic model that best matches the causal diagram of the current working condition of the grinding process, and the parameters of the pre-trained Actor-Critic model that best matches are saved; when it is determined that there are multiple similarities in the historical working conditions The Actor-Critic model corresponding to the historical working condition with the greatest similarity is selected as the pre-trained Actor-Critic model that best matches the causal graph of the current working condition of the grinding process, and the parameters of the pre-trained Actor-Critic model that best matches are saved; if all When , the parameter θ is directly corrected by the gradient correction strategy k Perform online correction; Step S8: Using natural policy gradient to update in real time the policy parameter θ of the pre-trained Actor-Critic model that best matches the causal graph of the current working conditions of the grinding process k ; Wherein, the strategy parameter θ k The update formula is: In formula (4), α is the learning rate; is the Fisher information matrix; is the standard policy gradient, which represents the policy performance J(θ k ) relative parameter θ k gradient; in, is the strategy function, which represents the grinding mill in the grinding state s k The probability of selecting the set value vector w(k); the policy function The update formula is: In formula (5), Q(s k ,w(k)) is the action value function, which is used to evaluate the grinding mill in the grinding condition s k The long-term expected reward of executing the set value vector w(k); the action value function Q(s k , the update formula of w(k)) is: Step S9: Calculate the set value vector based on formula (4) The set values w1(k) and w2(k) are generated in real time through the Actor network model, and then the real-time generated set values w1(k) and w2(k) are transmitted to the controller of the mill control loop, so that the controller can adjust the frequency of the feeding motor and the opening of the water supply valve in real time according to the real-time generated set values w1(k) and w2(k), until the operating indicators of the grinding particle size r1(k) and the mill load r2(k) converge to the target range; wherein, the grinding particle size r1(k) and the mill load r2(k) converge to the target range, that is, the grinding particle size r1(k) and the mill load r2(k) meet the performance indicator function constructed in step S2.
2. The mill control loop setting value optimization method based on experience-driven adaptive reinforcement decision-making according to claim 1 is characterized in that: The adjustment coefficient α is 0.
9.
3. The mill control loop setting value optimization method with experience-driven adaptive enhanced decision-making according to claim 1 or 2, characterized in that: The value range of the similarity threshold η is [0,1].
Citation Information
Patent Citations
Ore grinding process modeling method based on neural network and evolutionary computation
CN108469797A
Closed-loop modeling optimization control method and equipment for multiple groups of vertical mills
CN116550459A
Ore grinding granularity optimization control method under variable working conditions
CN118768073A
Gold ore flotation full-process intelligent monitoring and optimal control method and system
CN120276349A
Method for controlling mill for pulverized coal burning boiler and device therefor
JP1997000969A
Cited By
Intelligent servo control method of glass edge grinding machine guided by causal priori
CN121348966A