Flotation process automatic control method for intelligently optimizing reinforcement learning algorithm

Through intelligent optimization of reinforcement learning algorithms, combined with deep reinforcement learning and PID control strategies, precise control of flotation agents and process parameters is solved, and the problem of poor flotation indicators is achieved, the maximum utilization of mineral resources and system stability is improved.

CN119926674AActive Publication Date: 2025-05-06CENT SOUTH UNIV

Patent Information

Application Number
CN202510248562.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-05-06
Estimated Expiration
2045-03-04

AI Technical Summary

Technical Problem

The existing technology is difficult to achieve precise control of the flotation agent system, resulting in poor flotation indicators and affecting the maximum utilization of mineral resources.

Method used

The intelligent optimization reinforcement learning algorithm is adopted to accurately control process parameters such as flotation agent addition, slurry pH, flotation temperature and agent reaction time through deep reinforcement learning technology. Combined with PID control strategy and improved FPA algorithm, the PPO algorithm is optimized to automatically adjust the PID parameters.

Benefits of technology

The optimization of flotation indicators has been achieved, the utilization efficiency of mineral resources has been improved, and the stability and automation level of the system have been enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119926674A_ABST
    Figure CN119926674A_ABST
Patent Text Reader

Abstract

The invention relates to a flotation process automatic control method for intelligently optimizing a reinforcement learning algorithm. In the mineral flotation operation process, a flotation reagent system needs to be precisely controlled, for example, the dosage of reagents such as a collecting agent, an inhibitor and a foaming agent is adjusted, or process reaction parameters such as the pH value of ore pulp, the flotation temperature and the reagent reaction time are adjusted, so that the flotation index reaches the best, and maximum utilization of mineral resources is achieved. Main equipment for adding flotation reagents is a variable frequency pump, the reagent amount is adjusted through PID, meanwhile, process index data of a flotation tank are collected through a sensor, the whole control algorithm is optimized through improved FPA reinforcement learning, flotation process index data parameters serve as environment state observation values of reinforcement learning intelligent agents, and the flotation process index data parameters are calculated. Meanwhile, the design environment is rewarded according to the conditions meeting the flotation requirements, and the intelligent agent is trained through repeated environment interaction. And finally, during operation, the intelligent agent continuously gives an optimal action value parameter of a flotation cell actuator in real time, and unmanned reinforcement learning optimization control of mineral flotation is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of algorithm learning for mineral flotation operations, and in particular to an automatic control method for a flotation process using an intelligently optimized reinforcement learning algorithm. Background Art

[0002] During the mineral flotation operation, the flotation reagent system needs to be precisely controlled, such as adjusting the dosage of collectors, depressants, frothers, slurry pH, flotation temperature, reagent reaction time, etc., so as to achieve the best flotation indicators and maximize the utilization of mineral resources.

[0003] The present invention is the application of intelligent optimization deep reinforcement learning in flotation process control. Combined with the characteristics of the flotation process, an efficient and stable deep reinforcement learning control algorithm is designed. With the rapid development of artificial intelligence technology, deep reinforcement learning (DRL) provides a new idea for solving such complex control problems in flotation processes. The main equipment for adding flotation reagents is a variable frequency pump. The dosage of the reagents is adjusted by PID. The PID control algorithm of deep reinforcement learning PPO is optimized by the improved FPA algorithm, which has achieved remarkable results in improving flotation efficiency, optimizing resource utilization and enhancing system stability. The reagent addition control of the flotation process collects the process index parameters of the flotation tank through sensors, and then uses the IFPA-PPO-PID control algorithm for automatic adjustment. The flotation process index data parameters are used as the environmental state observation values ​​of the reinforcement learning agent. At the same time, the environment reward is designed according to the conditions that meet the flotation requirements. After repeated environmental interaction training, the agent is trained. Finally, the agent continuously gives the best action value parameters of the flotation tank control and adjustment actuator in real time during operation. The metal mineral flotation reinforcement learning optimization control realizes refined automatic control and realizes the unmanned reinforcement learning optimization control of mineral flotation. Summary of the invention

[0004] The technical problem to be solved by the present invention is to provide an automatic flotation process control method for precisely controlling the flotation reagent system to achieve the best flotation index in view of the shortcomings of the prior art.

[0005] In order to solve the above technical problems, the technical solution provided by the present invention is: an automatic control method for flotation process of intelligent optimization reinforcement learning algorithm, in the mineral flotation operation process in the field of lead and zinc metal mineral resources processing, by accurately controlling the flotation reagent dosage (such as collector, inhibitor, frother, etc.) and the pH value of the ore pulp, the flotation temperature, the reagent reaction time and other process parameters, the flotation index reaches the optimal state, and the maximum utilization of mineral resources is achieved. The specific implementation content is as follows:

[0006] A. Preliminary analysis and technology introduction: First, we deeply analyze the characteristics and control difficulties of the flotation process, and then apply deep reinforcement learning technology to flotation process control;

[0007] B. Algorithm selection and architecture determination: PID control strategy is adopted in the flotation system process control, and the hybrid reinforcement learning PPO method of strategy rules and execution evaluation is selected. In view of the characteristics of strong coupling, complex nonlinearity and time-varying parameters of the flotation process control system, the Actor-Critic (AC) network architecture suitable for outputting continuous action space is selected;

[0008] C. Algorithm improvement and parameter optimization: A variety of intelligent optimization algorithms are used to improve the deep reinforcement learning PPO algorithm. The optimized algorithm is used to automatically optimize the PID control parameters to improve the control effect of the flotation process, enhance the automation level and control accuracy.

[0009] Furthermore, the algorithm selection is specifically that, in the corresponding flotation system process control, a PID control strategy is adopted, and an evaluation strategy network that gives rewards to the fitting environment is continuously trained and iteratively optimized with the agent action strategy rule network, and a hybrid reinforcement learning PPO method of strategy rules and execution evaluation is selected.

[0010] Furthermore, the algorithm selection is specifically for the flotation process control system, whose characteristics such as strong coupling, complex nonlinearity and time-varying parameters constitute significant operational control difficulties, and a reinforcement learning strategy suitable for outputting continuous action space, namely the Actor-Critic (AC) network architecture, is selected.

[0011] Furthermore, in the reinforcement learning agent of the AC network, the training iterative update gradient strategy of the action strategy function and the state value function mainly includes three methods:

[0012] The first is the update of the policy network (Actor): through the policy gradient algorithm, such as the REINFORCE algorithm or its improved version, the parameters of the policy network are updated according to the evaluation signal provided by the value function network (Critic). The policy gradient algorithm aims to maximize the expectation of the cumulative reward and adjust the policy parameters through the gradient ascent method;

[0013] The second is the update of the value function network (Critic): using temporal difference learning or Monte Carlo method, the parameters of the value function deep learning network are updated according to the data generated by the interaction between the agent and the environment. The goal of the Critic network is to accurately estimate the value of a given state or state-action pair to guide the update of the Actor network;

[0014] The third is joint update: In the AC framework, the Actor and Critic networks are usually updated alternately and simultaneously, that is, the Actor network adjusts its strategy according to the evaluation results of the Critic network, and the Critic network updates its value function estimate based on the interaction data under the new strategy generated by the Actor network. This joint update mechanism helps the agent find a balance between exploration and exploitation, thereby learning the optimal strategy more effectively.

[0015] Furthermore, in order to obtain the optimal strategy to maximize the reward, according to the training of the action strategy network in the previous section, a gradient algorithm is required to update the strategy parameters by optimizing the performance index gradient to maximize the reward index obtained by the agent;

[0016] The PPO algorithm combines the advantages of the policy gradient method and importance sampling, and ensures the stability of policy updates by limiting the difference between the new and old policies. The PPO algorithm includes an Actor network and a Critic network, which are used to generate actions and evaluate the value of actions respectively.

[0017] During the training process, PPO updates the policy parameters by optimizing the alternative objective function and uses the Clip truncation mechanism to control the magnitude of the policy update.

[0018] The PPO algorithm allows the same data to be used for multiple strategy updates in each iteration, which improves data utilization efficiency and makes the PID parameter optimization process more efficient;

[0019] At the same time, the probability density function of Gaussian distribution is used as the Actor action strategy function, so that the control system can better adapt to complex and changeable environments and working conditions.

[0020] Furthermore, the PPO algorithm is applied to the flotation process control system framework for adaptively adjusting PID parameters. The PPO algorithm, as a parameter adjustment mechanism for PID control, obtains environmental state information of the flotation control process. The action information is the output of the action strategy, which is a three-dimensional feature vector array, including kp, ki, and kd composed of the three PID parameters. The action space is:

[0021] A={k p ,k i ,k d}

[0022] Then it continuously interacts with the environment of the flotation control process, collects data and updates network parameters, implements the learning and training of PPOagent, and continuously optimizes the strategy network and evaluation network of PPO until the preset number of training rounds or reward threshold. After the training is completed, the PPO-PID algorithm model will provide the best adjustment parameters for flotation optimization control.

[0023] The entire algorithm model framework of the autonomous unmanned system with optimal control established by reinforcement learning uses various sensors to collect perception data of the environment, inputs it into the constructed reinforcement learning agent, makes decisions and controls through continuous action strategies, and calculates rewards with the value function network, thereby iteratively updating the parameters of the reinforcement learning training to achieve the highest reward of the learning control strategy, and finally uses the trained agent to realize intelligent real-time control in interaction with the environment.

[0024] Furthermore, the algorithm improvement and parameter optimization are specifically to improve the chaotic optimization FPA algorithm, use piecewise chaotic mapping to improve the FPA algorithm (IPFA) initialization, and improve the parameter optimization effect and fitness curve convergence speed;

[0025] The IFPA algorithm is introduced to optimize the reinforcement learning PPO (IFPA-PPO-PID). The good parameter iteration optimization characteristics of IFPA are utilized to select the global optimal parameters in the PPO hyperparameter adjustment link, thereby improving the training reward value, accelerating the convergence speed, and reducing the training time consumption.

[0026] Furthermore, the flotation process indicator data parameters are used as the environmental state observation values ​​of the reinforcement learning agent. At the same time, the environmental rewards are designed according to the conditions that meet the flotation requirements. The agent is trained through repeated environmental interactions, and the improved IFPA algorithm is used to continuously improve and adjust the hyperparameters of the PPO reinforcement learning. Finally, during operation, the agent continuously gives the optimal action value parameters of the flotation tank actuator in real time, realizing unmanned reinforcement learning optimization control of mineral flotation and solving the problem of poor control effect under nonlinear time-varying mineral flotation process.

[0027] After adopting the above algorithm, the present invention has the following advantages: in the field of lead and zinc metal mineral resource processing, the flotation process is a key link in extracting valuable minerals, and its control effect directly affects the quality and recovery rate of metal resource products. The present invention proposes a method for the mineral flotation process, through precise algorithm control of the flotation reagent system, such as adjusting the dosage of collectors, inhibitors and frothers, the pH value of the slurry, the flotation temperature, the reagent reaction time, etc., so as to optimize the flotation index and maximize the utilization of mineral resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 Schematic diagram of the model for the interaction between the agent and the environment;

[0029] Figure 2 This is a schematic diagram of the critic evaluation algorithm network structure;

[0030] Figure 3 This is a schematic diagram of the actor action strategy algorithm network structure;

[0031] Figure 4 This is a logical diagram of the deep reinforcement learning PPO strategy algorithm;

[0032] Figure 5 Schematic diagram of the algorithm model framework for deep reinforcement learning optimization;

[0033] Figure 6 Schematic diagram of the framework of the control algorithm optimized for deep reinforcement learning;

[0034] Figure 7 This is a schematic diagram of the deep reinforcement learning PPO adaptive optimization PID control system framework;

[0035] Figure 8 This is a schematic diagram of the piecewise chaotic mapping distribution;

[0036] Fig. 9 This is a schematic diagram of the distribution frequency of piecewise chaotic mapping;

[0037] Fig.10 Schematic diagram of the F1(x) function and fitness value iteration curve.

[0038] Fig.11 It is a schematic diagram of the F2(x) function and the fitness value iteration curve;

[0039] Fig.12 It is a schematic diagram of the running results of the F3(x) function and the fitness value iteration curve;

[0040] Fig.13 It is a schematic diagram of the F4(x) function and the fitness value iteration curve;

[0041] Fig.14 This is a schematic diagram of the deep reinforcement learning PPO adaptive optimization PID control system framework;

[0042] Fig.15 Schematic diagram of the fitness curve for improving the deep reinforcement learning IFPA-PPO-PID control algorithm;

[0043] Fig.16 Schematic diagram of the fitness curve for improving the deep reinforcement learning IFPA-PPO-PID control algorithm;

[0044] Fig.17 Schematic diagram of the fitness curve for improving the deep reinforcement learning IFPA-PPO-PID control algorithm;

[0045] Fig.18 Schematic diagram of IFPA-PPO-PID control parameter output changes. DETAILED DESCRIPTION

[0046] The present invention is further described in detail below in conjunction with the accompanying drawings.

[0047] Combined with Figure 1-18 The object of the present invention is achieved by the following manner: the flotation process of metal minerals is regulated, especially the flotation of lead and zinc ores, and reagents are added in the flotation tank for separation and purification. It is necessary to adjust the flotation reaction reagents such as collectors, depressants and frothers, as well as the flotation reaction conditions such as pH and temperature, so as to achieve the best flotation effect, that is, to improve the indicators such as the lead and zinc recovery rate.

[0048] The deep reinforcement learning in this invention is designed based on the Markov decision process. Reinforcement learning allows the agent to learn by trial and error in the environment and continuously optimize its behavior strategy to maximize the cumulative reward. The agent and the environment interact continuously through state and action. The state of the environment is random, and the action strategy of the agent is also random. The state and action time t, the state of the environment s t , the action output at of the reinforcement learning agent actuator, the action set A made by the agent, the reward value rt given by the environment to the reinforcement learning after the agent executes the action at time t, the reward evaluation function R, E is the target expectation value of the reinforcement learning training process, the value function Vy of the state at this time t, the action strategy y decided by the agent state actuator, the network parameter of the strategy function y is θ, the value function qy of the permanent reward after the state at time t executes the action, p is the possibility of state change after executing the action, the best strategy y*, the best value function q*, and the evaluation attenuation factor γ of the execution effect.

[0049] According to the probability density function distribution of whether the environmental state is discrete or continuous, the formula for obtaining the environmental state is expressed as follows

[0050] p(s t |s t-1 ,a t-1 ) → s t (0-1)

[0051] The reinforcement learning actuator action output and the actuator action strategy probability function are obtained by random sampling

[0052] y(a t |s t )→a t (0-2)

[0053] The rewards set by the reinforcement learning action are obtained by random sampling of the state sampling and execution dynamics of the environment.

[0054] R(rt|st,at)→rt (0-3)

[0055] The state value function at time t is the expectation of the actuator reward and the decaying prospective reward q obtained by the subsequent continuous actions in this state, and is obtained using the Bellman expectation equation.

[0056]

[0057] The action value function after executing the action at this time t is obtained by the expectation of the actuator reward and the decayed prospective return Ut

[0058]

[0059] The best actuator action strategy is selected in any state s obtained through the final reinforcement learning training.

[0060]

[0061] At this time, the corresponding optimal value function will also be obtained by any state and any action of the intelligent agent.

[0062]

[0063] Reinforcement learning is generally carried out around the iteration of action value function and state value function. The critic, as the judge of reinforcement learning, uses the value function network to evaluate the state and action at the same time. This uses the deep learning network to evaluate and obtain q y (s t ,a)≈q(s t ,a;w), w is the weight parameter of the deep learning network, such as Figure 2 The evaluation value function of reinforcement learning is shown in Figure 2.

[0064] The reinforcement learning algorithm for training action strategy improvement continuously improves the actor strategy, that is, through complete output actions, in a complete interaction with the environment, the cumulative reward is obtained to achieve the best value. Then the goal of training reinforcement learning is the optimal action strategy function distribution, which can be obtained

[0065]

[0066] Therefore, the best action for reinforcement learning under the optimal action strategy function is:

[0067]

[0068] From the above, we can see that the value learning of the evaluation function network requires iterative calculation of the gradient training of the value function, where the error is the loss function That is, the estimated value given by the value network of deep learning, so the gradient of the iteratively updated parameters is

[0069] Then in order to reduce the error of the evaluation function, according to the principle of gradient descent, the learning rate, that is, the step size, is α, so the iterative update of the evaluation function network parameter w is

[0070] A typical actor network for reinforcement learning is as follows Figure 3 shown.

[0071] Using the powerful fitting calculation ability of the deep learning network, and performing approximate calculation of the action strategy based on the policy network parameters, we can obtain

[0072] y(a|s t )≈y(a|s t ;θ)(0-10)

[0073] And the probability density of the action policy network

[0074]

[0075] Then, the state value function is obtained from equation (0-4) as follows

[0076]

[0077] However, the strategy gradient is obtained randomly, as shown below:

[0078]

[0079] According to the result of the policy gradient, we need to maximize the reward after the action is executed, so the update can be obtained as follows: The above constitutes the most basic AC algorithm framework and training method for deep reinforcement learning.

[0080] For the flotation process control system, it is necessary to select a reinforcement learning strategy suitable for outputting continuous action space, namely the Actor-Critic (AC) network architecture. The PPO algorithm includes the Actor network and the Critic network, which are used to generate actions and evaluate the value of actions respectively.

[0081] PPO algorithm optimization process architecture Figure 4 As shown in the figure, it is divided into two parts: one is the part where the agent interacts with the environment, and the other is the iterative improvement part of the training network parameters. When training reinforcement learning, the state value function V estimated by the Critic evaluation network is discounted and added with the reward rt to obtain the reward function R. The advantage function is estimated by calculating the time difference error between the reward function R and the state value function V, and then performing the generalized advantage estimation gae calculation for:

[0082]

[0083] Then, the back-propagation gradient descent updates the Critic evaluation network parameter w, and the loss function of the updated Critic evaluation network parameter w is:

[0084]

[0085] At the same time, the probability of the distribution of new and old actions is sampled by the importance weight δ, and then the first-order optimization constraint and truncation limit function operation coefficient ε are used. ε is generally taken as 0.2, and the loss function is obtained as follows:

[0086]

[0087] The loss function is used as the objective function for gradient update, and the truncation operation better limits the update of the new strategy. is positive, the reward estimate generated by the current action is greater than the return estimate of the old action strategy, and the new action strategy needs to be updated, and the probability of the action should be larger in distribution, but it needs to be restricted and truncated within the range of (1-ε,1+ε) multiples of the old strategy. Similarly, If is positive, the action probability is smaller in distribution and the update is truncated within the range of (1-ε,1+ε) multiples of the old policy.

[0088] The design uses the deep reinforcement learning PPO algorithm to adaptively adjust the PID control model (PPO-PID). Taking the flotation dosing pH control system as an example, the control target is to achieve the control process index evaluation of the reaction parameters expected by the flotation system based on multiple requirements. The environmental reward function of reinforcement learning is very important. For the pH adjustment automatic process control of the flotation system, the expected pH value is TpHq, and the actual pH is TpH. Then the discretization is

[0089] e(n)=[T pHq (n)-T ph (n)] (0-17)

[0090] Then the state observation space S is calculated according to the error e and the change rate between the actual pH value TpHq and the expected set pH of the flotation system. To design:

[0091]

[0092] And the termination signal design of the reinforcement learning agent is used to stop in advance, that is, to stop the sign according to the control boundary conditions to prevent exceeding the safety range, and to prevent errors and poor training data from affecting the learning update of the agent. According to the process conditions of flotation dosing to control pH, the calculation of the stop sign is:

[0093]

[0094] The safety margin reward according to the stop sign is rd:

[0095] r d =-20I done (0-20)

[0096] The control error reward re is:

[0097]

[0098] The training ends at tq, and the response reward rt at time t when the control response reaches the target value within 50% error is:

[0099]

[0100] The reward rq at the end of training is:

[0101]

[0102] The faster the response speed, the better, and the smaller the steady-state fluctuation and overshoot, the better. Then the reward function of the entire control process is

[0103] r=r d +r e +r t +r q (0-24)

[0104] The flotation control system uses a PID controller for basic control operations. The PID control formula has three parameters: kp, ki, and kd. After discretization, it is

[0105]

[0106] Then the formula for the pH PID controller is

[0107]

[0108] The flotation process control system framework using the PPO algorithm to adaptively adjust PID parameters is shown in Figure 7 As shown in the figure, the PPO algorithm is used as a parameter adjustment mechanism for PID control to obtain the environmental state information of the flotation control process. The action information is the output of the action strategy. a is a three-dimensional feature vector array, which includes kp, ki, and kd composed of the three PID parameters. The action space is:

[0109] A={k p ,k i ,k d}(0-27)

[0110] Then it continuously interacts with the environment of the flotation control process, collects data and updates network parameters to realize the learning and training of PPOagent, and continuously optimizes the strategy network and evaluation network of PPO until the preset number of training rounds or reward threshold. After the training is completed, the PPO-PID algorithm model will provide the best adjustment parameters for flotation optimization control.

[0111] In the PPO algorithm, the PID parameters are adaptively adjusted. Its parameter space is continuous and a continuous Gaussian probability density function is used for sampling. Therefore, the function distribution of the action strategy y is:

[0112]

[0113] Where a is the action output. In the policy-based reinforcement learning algorithm, the policy is often expressed as a distribution. By combining the distribution transformation with the reversible compression function, the PPO algorithm is applied to the flotation process control. The detailed process steps of the algorithm pseudo code are shown in the algorithm table (I), which realizes the adaptive adjustment of PID parameters, improves the control accuracy, and maintains the continuity of the distribution.

[0114] Table (I) Reinforcement learning flotation control detailed process steps algorithm pseudo code

[0115]

[0116] The Flower pollenation algorithm (FPA) introduces a parameter p by slightly biasing the switching probability towards local pollination, with an initial value of 0.5, to adjust the frequency of switching between local pollination and global pollination and improve the optimization efficiency. The global and flower specificity of biological pollination can be described mathematically, and the update formula is expressed as:

[0117]

[0118] Here, Represents the solution vector The update during the tth iteration, x g is the optimal solution for the current fitness function evaluation, and λ is the proportional coefficient for adjusting the step length. Function T() represents the step length in the pollination process, and the random step length sampled based on Levy flight reflects the intensity of pollination activity.

[0119] The formula for the Levy flight distribution is as follows:

[0120] [T~Levy(ξ),(ξ>0)] (0-30)

[0121] The calculation formula of Levy flight is as follows:

[0122]

[0123] This distribution is valid for ξ>0, where Γ() represents the standard Gamma function.

[0124] Another equation for the local self-pollination process uses the switching probability parameter p, and the self-pollination process formula is:

[0125]

[0126] in, and represents two solutions drawn randomly, Follow the random sampling of values ​​between the uniform distribution [0,1]. Introduce the parameter p, and set the initial value to 0.5. FPA is described by 4 limiting conditions based on the laws of nature, which are written as the above four equations to limit it. In addition, the initialization position distribution of the general pollen population will not be uniformly distributed if the coefficients randomly generated by [0,1] are extracted within the spatial range. Therefore, piecewise chaotic mapping is used to improve the initialization of the FPA algorithm (Improved Flowerpollenation algorithm, IPFA), so that the IPFA algorithm can improve the parameter optimization effect and accelerate the convergence speed of the fitness curve. The derivation formula of piecewise chaotic mapping is as follows:

[0127]

[0128] Where n is the number of pollen populations, and 1≤i<n, then the initial distribution of the population is

[0129]

[0130] Write the code and run it according to the above formula. The piecewise chaotic map obtains the [0,1] population distribution according to the population size of 30 as follows Figure 8 and Fig. 9 , the best parameter solution space segment can be found as quickly as possible in the initial operation stage of the algorithm, increasing the probability of solving the global optimal parameters.

[0131] To verify the solving performance of the improved IFPA algorithm, the experiment selected four representative standard performance test functions, namely Generalized Rastrigin (F1(x)), Ackley (F2(x)), Step (F3(x)), and Griewank (F4(x)). Among them, F1(x) and F2(x) are single-peak functions, which can test the optimization algorithm's ability to find the best solution in the local solution space, and F3(x) and F4(x) are multi-peak characteristic functions, which can evaluate the optimization algorithm's performance in jumping out of the local extreme point to obtain the global optimal parameters. At the same time, the ability of IFPA to solve the optimal parameters is tested, and then compared with the other three algorithms in the standard test function. The experimental parameter settings are shown in Table (II).

[0132] Table (II) Algorithm test comparison experiment settings

[0133]

[0134] According to the unified experimental parameter settings, the algorithm experimental function and running results are as follows Figure 10 to Figure 13 As shown in the running curve results, it can be seen that the improved IFPA algorithm can iteratively solve the optimal solution of the function. Compared with the unimproved FPA algorithm, the optimization ability is significantly improved, and it can quickly iterate and converge to the optimal solution, achieving better algorithm optimization performance for solving the optimal parameters.

[0135] In the process of training the IFPA-PPO-PID algorithm, the learning rate of PPO-PID using IFPA is first selected, and the population number of IFPA is set to n, the search range is [LB, UB], the maximum number of iterations IterMax, the conversion probability p and other parameters are set. Then the piecewise chaotic mapping is used to initialize the distribution of the population, and the initialized fitness function value is calculated. Here, the reward of reinforcement learning is used as the fitness function, and PPO-PID is called to initialize the fitness calculation to obtain the fitness value of each particle in the population, and the global optimal solution is obtained. Then, FPA is used to update the particle distribution position, and the fitness function is calculated again, so as to iteratively update the particle distribution position, and continuously iteratively update the particle distribution position and the fitness function value until the maximum number of iterations IterMax is reached to stop updating the iteration. At this time, the optimal PPO-PID learning rate parameter is output, such as Fig.14 .

[0136] To verify the effectiveness of the IFPA-PPO-PID algorithm, the controlled object selects the commonly used first-order time-delay model. First, the parameters of PPO-PID with and without IFPA are set as shown in Table (III), ensuring that the flotation process control experiment is carried out under the same initial condition setting and experimental parameters. The learning rate range is [0.05, 0.0001], and the learning rate without IFPA is generally set to 0.01.

[0137] Table (III) Reinforcement learning algorithm experimental settings

[0138]

[0139]

[0140] The algorithm fitness function curve using IFPA is as follows: Fig.15 As shown in the figure, the iteration runs 20 times. The fitness value can be reduced quickly at the beginning of the iteration. At the same time, the search continues after the local optimal parameters are found, and converges after the global optimal parameters are obtained. A good balance is obtained between local search and global search. The reward values ​​of PPO-PID with and without IFPA are -134.5 and -130.7, respectively, which improves the performance by 2.9%.

[0141] After further training the IFPA-PPO-PID algorithm 30,000 times, Fig.16 The reward value has converged after 12,000 times, indicating that the improved deep reinforcement learning control algorithm can achieve good control effects in flotation process control. After training the saved IFPA-PPO-PID algorithm model, the flotation process control simulation is performed and compared with the PID control after manual tuning. The results are as follows: Fig.17 As shown, the control effect is greatly improved.

[0142] In the flotation process control simulation of the IFPA-PPO-PID algorithm model, the effect of the action space, i.e., the PID parameter output, is as follows: Fig.18 As shown in the figure, compared with the PID control algorithm, the parameters are adjusted adaptively and in real time, achieving superior results compared to manual adjustment.

[0143] The IFPA-PPO-PID improved reinforcement learning optimization control algorithm for mineral flotation proposed in the present invention uses the flotation process index data parameters as the environmental state observation values ​​of the reinforcement learning agent, and designs the environment reward according to the conditions that meet the flotation requirements. The agent is trained through repeated environmental interactions, and the improved IFPA algorithm is used to continuously improve and adjust the hyperparameters of the PPO reinforcement learning. Finally, during operation, the agent continuously gives the optimal action value parameters of the flotation tank actuator in real time, realizing unmanned reinforcement learning optimization control of mineral flotation, and solving the problem of poor control effect under nonlinear time-varying mineral flotation process. The reagent addition control of the flotation process collects the process index parameters of the flotation tank through sensors, and then uses the IFPA-PPO-PID control algorithm for automatic adjustment, and the metal mineral flotation reinforcement learning optimization control realizes refined automatic control.

[0144] The present invention and its implementation methods are described above, and such description is not restrictive, and the actual structure is not limited thereto. In short, if a person skilled in the art is inspired by it, and does not deviate from the purpose of the invention, and does not creatively design a structure and implementation method similar to the technical solution, they should all fall within the protection scope of the present invention.

Claims

1. A flotation process automatic control method based on intelligent optimization reinforcement learning algorithm, characterized in that: In the mineral flotation process in the field of lead and zinc metal mineral resource processing, the flotation index is optimized by accurately controlling the flotation reagent dosage (such as collector, inhibitor, frother, etc.) and the process parameters such as pulp pH, flotation temperature, and reagent reaction time, so as to maximize the utilization of mineral resources. The specific implementation contents are as follows: A. Preliminary analysis and technology introduction: First, we deeply analyze the characteristics and control difficulties of the flotation process, and then apply deep reinforcement learning technology to flotation process control; B. Algorithm selection and architecture determination: PID control strategy is adopted in the flotation system process control, and the hybrid reinforcement learning PPO method of strategy rules and execution evaluation is selected. In view of the characteristics of strong coupling, complex nonlinearity and time-varying parameters of the flotation process control system, the Actor-Critic (AC) network architecture suitable for outputting continuous action space is selected; C. Algorithm improvement and parameter optimization: A variety of intelligent optimization algorithms are used to improve the deep reinforcement learning PPO algorithm. The optimized algorithm is used to automatically optimize the PID control parameters to improve the control effect of the flotation process, enhance the automation level and control accuracy.

2. The flotation process automatic control method of an intelligent optimization reinforcement learning algorithm according to claim 1 is characterized in that: The algorithm selection is specifically that, in the corresponding flotation system process control, a PID control strategy is adopted, an evaluation strategy network that gives rewards to the fitting environment is continuously trained and iteratively optimized with the agent action strategy rule network, and a hybrid reinforcement learning PPO method of strategy rules and execution evaluation is selected.

3. The flotation process automatic control method of an intelligent optimization reinforcement learning algorithm according to claim 2 is characterized in that: The algorithm selection is specifically, for the flotation process control system, its strong coupling, complex nonlinearity and time-varying parameters and other characteristics constitute significant operational control difficulties, and a reinforcement learning strategy suitable for outputting continuous action space, namely, the Actor-Critic (AC) network architecture, is selected.

4. The flotation process automatic control method of the intelligent optimization reinforcement learning algorithm according to claim 3 is characterized by: In the reinforcement learning agent of the AC network, the training iterative update gradient strategy of the action strategy function and the state value function mainly includes three methods: The first is the update of the policy network (Actor): through the policy gradient algorithm, such as the REINFORCE algorithm or its improved version, the parameters of the policy network are updated according to the evaluation signal provided by the value function network (Critic). The policy gradient algorithm aims to maximize the expectation of the cumulative reward and adjust the policy parameters through the gradient ascent method; The second is the update of the value function network (Critic): using temporal difference learning or Monte Carlo method, the parameters of the value function deep learning network are updated according to the data generated by the interaction between the agent and the environment. The goal of the Critic network is to accurately estimate the value of a given state or state-action pair to guide the update of the Actor network; The third is joint update: In the AC framework, the Actor and Critic networks are usually updated alternately and simultaneously, that is, the Actor network adjusts its strategy according to the evaluation results of the Critic network, and the Critic network updates its value function estimate based on the interaction data under the new strategy generated by the Actor network. This joint update mechanism helps the agent find a balance between exploration and exploitation, thereby learning the optimal strategy more effectively.

5. The flotation process automatic control method of the intelligent optimization reinforcement learning algorithm according to claim 4 is characterized by: In order to obtain the best strategy to maximize the reward, according to the training of the action strategy network in the previous section, a gradient algorithm is required to update the strategy parameters by optimizing the performance index gradient to maximize the reward index obtained by the agent; The PPO algorithm combines the advantages of the policy gradient method and importance sampling, and ensures the stability of policy updates by limiting the difference between the new and old policies. The PPO algorithm includes an Actor network and a Critic network, which are used to generate actions and evaluate the value of actions respectively. During the training process, PPO updates the policy parameters by optimizing the alternative objective function and uses the Clip truncation mechanism to control the magnitude of the policy update. The PPO algorithm allows the same data to be used for multiple strategy updates in each iteration, which improves data utilization efficiency and makes the PID parameter optimization process more efficient; At the same time, the probability density function of Gaussian distribution is used as the Actor action strategy function, so that the control system can better adapt to complex and changeable environments and working conditions.

6. The flotation process automatic control method of the intelligent optimization reinforcement learning algorithm according to claim 5 is characterized by: The PPO algorithm is applied to the flotation process control system framework of adaptively adjusting PID parameters. The PPO algorithm is used as a parameter adjustment mechanism for PID control to obtain the environmental state information of the flotation control process. The action information is the output of the action strategy. a is a three-dimensional feature vector array, which is composed of three PID parameters kp, ki, and kd. The action space is: A={k p ,k i ,k d } Then it continuously interacts with the environment of the flotation control process, collects data and updates network parameters, implements the learning and training of PPOagent, and continuously optimizes the strategy network and evaluation network of PPO until the preset number of training rounds or reward threshold. After the training is completed, the PPO-PID algorithm model will provide the best adjustment parameters for flotation optimization control. The entire algorithm model framework of the autonomous unmanned system with optimal control established by reinforcement learning uses various sensors to collect perception data of the environment, inputs it into the constructed reinforcement learning agent, makes decisions and controls through continuous action strategies, and calculates rewards with the value function network, thereby iteratively updating the parameters of the reinforcement learning training to achieve the highest reward of the learning control strategy, and finally uses the trained agent to realize intelligent real-time control in interaction with the environment.

7. The flotation process automatic control method of an intelligent optimization reinforcement learning algorithm according to claim 1 is characterized in that: The algorithm improvement and parameter optimization are specifically to improve the chaotic optimization FPA algorithm, use piecewise chaotic mapping to improve the FPA algorithm (IPFA) initialization, and improve the parameter optimization effect and fitness curve convergence speed; The IFPA algorithm is introduced to optimize the reinforcement learning PPO (IFPA-PPO-PID). The good parameter iteration optimization characteristics of IFPA are utilized to select the global optimal parameters in the PPO hyperparameter adjustment link, thereby improving the training reward value, accelerating the convergence speed, and reducing the training time consumption.

8. A flotation process automatic control method based on an intelligent optimization reinforcement learning algorithm according to claims 1-7, characterized in that: The flotation process indicator data parameters are used as the environmental state observation values ​​of the reinforcement learning agent. At the same time, the environmental rewards are designed according to the conditions that meet the flotation requirements. The agent is trained through repeated environmental interactions, and the improved IFPA algorithm is used to continuously improve and adjust the hyperparameters of PPO reinforcement learning. Finally, during operation, the agent continuously gives the optimal action value parameters of the flotation tank actuator in real time, realizing unmanned reinforcement learning optimization control of mineral flotation and solving the problem of poor control effect under nonlinear time-varying mineral flotation process.

Citation Information

Patent Citations

  • Segmented interpretable intelligent chemical adding method for coal flotation

    CN115390450A

  • Flotation intelligent dosing control system and method for coal preparation plant

    CN116037324A

  • Determination method, system and equipment for flotation reagent dosage and medium

    CN116140074A

  • Method for automatically regulating explicit congestion notification of data center network based on multi-agent reinforcement learning

    US20240080270A1

Cited By

  • Model prediction parameter dynamic change method in flotation pH adjusting process

    CN121122456A

  • Mine safety interlocking control method, device and equipment based on reinforcement learning and medium

    CN122172603A