A multivariate hierarchical adaptive filling regulation method based on reinforcement learning

By employing a multivariate hierarchical adaptive filling control method in the mine filling system, and utilizing reinforcement learning and hierarchical control mechanisms, the adaptive coordination problem among multiple variables was solved, thereby improving the stability and accuracy of the filling system, avoiding control strategy oscillations, and enhancing the system's adaptive capability and response stability.

CN120506267BActive Publication Date: 2026-02-06INNER MONGOLIA YULONG MINING IND CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510948591.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-10
Publication Date
2026-02-06
Estimated Expiration
2045-07-10

AI Technical Summary

Technical Problem

Existing reinforcement learning in mine filling systems focuses only on a single operating parameter, failing to achieve adaptive dynamic coordination among multiple key variables. The control structure design is mainly based on end-to-end strategies, resulting in low accuracy and stability of filling control.

Method used

A multivariate hierarchical adaptive filling control method based on reinforcement learning is adopted. The initial operating parameters of the filling system are periodically collected to construct a state vector. The target action parameters are output by the target reinforcement learning policy network model, and a multidimensional reward function is constructed. Combined with the hierarchical control mechanism, the parameters are adjusted and the policy is updated to achieve coordinated control of multiple variables.

Benefits of technology

The stability and accuracy of the filling system have been improved. Comprehensive optimization control of the filling system has been achieved through multi-dimensional state characteristic functions and reward functions, which enhances the system's adaptability and response stability and avoids the direct impact of policy oscillations on the equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120506267B_ABST
    Figure CN120506267B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of mine filling control, and particularly relates to a multivariate hierarchical adaptive filling regulation method based on reinforcement learning, comprising: constructing a state vector; inputting the state vector into a target reinforcement learning strategy network model to obtain a target action parameter and a state value estimation; determining whether the state value estimation meets an expected standard, and if so, adjusting a filling control parameter of a filling system based on the target action parameter; in response to completion of adjustment of the filling control parameter of the filling system, constructing a state characteristic function corresponding to the filling system in a current control cycle, and constructing a multidimensional reward function based on each state characteristic function; determining whether a strategy update is triggered based on a control cycle number and a current cumulative reward value, and if so, updating the target reinforcement learning strategy network based on the multidimensional reward function. The present application can realize adaptive coordination optimization of multivariate targets, and effectively improve the precision and stability of filling control.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of mine filling control, and particularly relates to a multivariate hierarchical self-adaptive filling regulation method based on reinforcement learning. BACKGROUND

[0002] As a key means of safe stoping and stability control of goaf in underground mines, filling mining technology has been widely used in non-ferrous, ferrous and non-metallic deposits. In recent years, with the increase of ore body depth and the complexity of ground pressure environment, filling operation faces higher reliability, continuity and adaptability requirements. Various processes such as slurry filling, paste filling and tailings cementation filling are continuously developing, and the filling system gradually moves from traditional manual control to automation and intelligent stage.

[0003] The filling system operation process covers material proportioning, conveying and pumping, pipeline regulation, stope filling and other links, involving concentration, water-cement ratio, pumping pressure, cement mixing ratio, temperature and humidity environment and other multi-dimensional control parameters. These parameters have strong coupling characteristics, and fluctuate significantly with changes in construction environment, material properties and operation tasks. Traditional processes rely on operator experience to set parameters, which are easily affected by subjective judgment and reaction lag, and the regulation process lacks systematicness, making it difficult to adapt to complex and variable mine operation environment.

[0004] To improve the efficiency and stability of the filling system, researchers have gradually introduced model predictive control, genetic algorithm, multi-objective optimization and other methods to coordinate and optimize between objective functions. However, most of these methods are designed under static data conditions, and the control strategy lacks adaptability, making it difficult to respond quickly to abnormal conditions. With the gradual improvement of industrial data acquisition and edge computing platforms, data-driven intelligent control models have become a research focus. Reinforcement learning, as an interactive learning method for dynamic decision-making problems, has shown good application prospects in industrial process control. It can continuously modify the control strategy in actual operation without relying on explicit mathematical models, and can improve the adaptive ability and long-term performance of the system. Current research attempts to apply reinforcement learning to mine beneficiation, grinding, filling and other aspects, and has achieved preliminary results in nonlinear system modeling and control decision-making. Although reinforcement learning control has made some breakthroughs in beneficiation automation, its application in mine filling systems is still in the exploratory stage. Existing research focuses on single variable control objectives such as concentration control or bottom flow concentration stability, and lacks joint control mechanisms between water-cement ratio, pump speed, material usage and other multivariables. At the same time, most control structures still use end-to-end strategies, with reinforcement learning output directly acting on the device execution layer. In the filling system, parameter control involves dynamic coordination of multiple variables, and control actions have continuity, coupling and irreversibility. The existing reinforcement learning control structure cannot meet the engineering-level stability and precision requirements. SUMMARY

[0005] To this end, the present application provides a multivariate hierarchical adaptive filling regulation method based on reinforcement learning, to overcome the problem that the prior art only focuses on a single operating parameter, fails to achieve adaptive dynamic coordination between multiple key variables, and the control structure design mainly uses an end-to-end strategy, so that the reinforcement learning output directly acts on the device execution layer, resulting in relatively low precision and stability of filling control.

[0006] To achieve the above object, the present application provides a multivariate hierarchical adaptive filling regulation method based on reinforcement learning, comprising:

[0007] Periodically collecting initial operating parameters of the filling system to construct a state vector, wherein the operating parameters include slurry concentration, slurry flow rate, delivery pump outlet pressure, water-cement ratio, backwater rate, and strength estimate value;

[0008] Inputting the state vector into a target reinforcement learning strategy network model to obtain target action parameters and state value estimates, wherein the action parameters include slurry concentration, water-cement ratio, and pump speed;

[0009] Determining whether the state value estimates meet the expected standard, and if so, adjusting filling control parameters of the filling system based on the target action parameters, wherein the filling control parameters include cement content, clean water dosage, and delivery pump frequency;

[0010] In response to completion of adjustment of the filling control parameters of the filling system, constructing state feature functions corresponding to the filling system in the current control period, and constructing a multi-dimensional reward function based on the state feature functions, wherein the state feature functions include slurry concentration feature function, water-cement ratio feature function, pumping energy consumption feature function, and strength feature function;

[0011] Obtaining the number of control cycles in the current control period, and calculating a current cumulative reward value corresponding to the current control period based on the multi-dimensional reward function;

[0012] Determining whether to trigger strategy update based on the number of control cycles and the current cumulative reward value, and if so, updating the target reinforcement learning strategy network model based on the multi-dimensional reward function.

[0013] Further, adjusting the filling control parameters of the filling system based on the target action parameters comprises:

[0014] Determining a concentration control deviation based on a comparison result of the target slurry concentration and the initial slurry concentration, and determining cement content control amount and water-cement ratio control amount based on the concentration control deviation;

[0015] determine a water-cement ratio control deviation based on a comparison result of the target water-cement ratio and the initial water-cement ratio, and determine a net water dosage control amount based on the water-cement ratio control deviation;

[0016] determine a pump speed control deviation based on a comparison result of the target pump speed and the initial pump speed, and determine a conveying pump frequency control amount based on the pump speed control deviation.

[0017] Further, a multi-dimensional reward function is constructed, including:

[0018] determine a concentration deviation penalty based on the slurry concentration characteristic function and a slurry concentration reference value;

[0019] determine a water-cement ratio deviation penalty based on the water-cement ratio characteristic function and a water-cement ratio reference value;

[0020] determine a pumping energy consumption penalty based on the pumping energy consumption characteristic function;

[0021] determine a strength compliance index based on the strength characteristic function;

[0022] construct a multi-dimensional reward function based on the concentration deviation penalty, the water-cement ratio deviation penalty, the pumping energy consumption penalty, the strength compliance index, a concentration weight coefficient, a water-cement ratio weight coefficient, an energy consumption weight coefficient, and a strength weight coefficient.

[0023] Further, determine whether the expected standard is met based on the state value estimate, including:

[0024] determine an estimate comparison value based on a comparison result of the state value estimate and a preset state value estimate;

[0025] determine whether the expected standard is met based on a comparison result of the estimate comparison value and a preset comparison value.

[0026] Further, determine whether to trigger a policy update, including:

[0027] determine whether to trigger a policy update based on a comparison result of the control cycle number and a preset cycle number, and a comparison result of the current cumulative reward value and a preset cumulative reward value.

[0028] Further, construct a state vector, including:

[0029] periodically collect initial operating parameters of the filling system, and determine a parameter change representation value corresponding to each operating parameter based on the initial operating parameters within a target time period;

[0030] determine whether to mark the target time period based on a comparison result of each parameter change representation value and a preset change representation value;

[0031] If the flag is marked, the state input variable corresponding to each operation parameter is generated based on the initial operation parameter in the target time period;

[0032] The state vector is constructed based on each state input variable.

[0033] Further, if the condition of not triggering the strategy update is determined, the control cycle number in the current control period is updated, and the initial operation parameter of the filling system is re-collected to update the state vector.

[0034] Further, if the current cumulative reward value is greater than the preset cumulative reward value, the cumulative reward change in a plurality of continuous control periods is calculated, wherein if the control cycle number is greater than the preset cycle number, or the cumulative reward change meets the preset change condition, it is determined that the strategy update is triggered.

[0035] Further, the filling control parameter of the filling system is adjusted based on the target action parameter, and the method further comprises:

[0036] Based on the comparison result of the concentration control deviation and the concentration safety index, the comparison result of the target water-cement ratio and the water-cement ratio safety index, and the comparison result of the target pump speed and the pump speed safety index, it is determined whether the target freezing logic is triggered, if not, the cement content control amount is determined based on the concentration control deviation, the water-cement ratio control amount is determined based on the water-cement ratio control deviation, and the conveying pump frequency control amount is determined based on the pump speed control deviation.

[0037] Further, if it is determined that the target freezing logic is triggered, the filling control parameter of the filling system is adjusted based on the action parameter in the control period before the current control period.

[0038] Compared with the prior art, the present application has the beneficial effects that the present application can comprehensively and accurately reflect the current actual operation state of the filling system by periodically collecting the initial operation parameters of the filling system to construct a state vector, providing a rich data basis for subsequent precise regulation and control. Through the target reinforcement learning strategy network model, the target action parameters are output according to the state vector, which provides precise decision support for the regulation and control of the filling system, can determine the optimal action parameters according to the current complex system state, and at the same time outputs the state value estimation, which helps to quantitatively evaluate the advantages and disadvantages of the current state, and provides a basis for judging whether the state vector meets the expected standard. According to the target action parameters, the filling control parameters such as cement content, water addition amount and conveying pump frequency of the filling system are precisely adjusted, instead of directly controlling the control system with the target action parameters, which avoids abnormal system execution caused by strategy shock or fluctuation, and can improve the stability of the filling system. The state characteristic functions including the slurry concentration characteristic function, the water-cement ratio characteristic function, the pumping energy consumption characteristic function and the strength characteristic function are constructed, which can finely depict the running state of the filling system in the current control period from multiple dimensions. Based on the state characteristic functions, a multi-dimensional reward function is constructed, which comprehensively considers the concentration stability, water-cement ratio accuracy, pumping energy consumption and strength compliance rate and other key performance indicators, and realizes the comprehensive optimization control of the filling system, thereby improving the precision of the filling system. Based on the control cycle number and the current cumulative reward value, it is judged whether the strategy update is triggered, an adaptive update mechanism based on state feedback and performance fluctuation is introduced, involving two strategy update triggering modes, fixed cycle triggering and reward trend triggering, which further improves the robustness of the update, avoids the interference of frequent restart or control mutation on the field filling operation, and can improve the adaptive coordination ability of the filling system. Based on the multi-dimensional reward function, the target reinforcement learning strategy network is updated, which can make the control system continuously improve the control performance of the filling system on the basis of continuous learning and experience accumulation, realize the intelligent and adaptive operation of the filling system, and improve the stability and precision of the filling system.

[0039] Further, the present application introduces a hierarchical control mechanism, uses reinforcement learning for upper target value setting, decouples the high-level reinforcement learning strategy output from the lower execution control, accurately determines the concentration control deviation, water-cement ratio control deviation and pump speed control deviation, and reasonably adjusts the filling control parameters, effectively avoiding the risk of strategy shock directly acting on the equipment, and improving the response stability of the control system.

[0040] Further, the multi-dimensional reward function of the present application comprehensively considers the concentration, accuracy deviation of water-cement ratio, system energy consumption and standard compliance of the filling body mechanical performance, forms a reward structure facing multi-objective coordination, and adjusts through weight coefficients, effectively realizes the coordination and unity among the control targets, can realize adaptive trade-off and dynamic adjustment of the control targets under complex working conditions, improves the robustness and adaptability of the strategy, and thus further improves the stability and precision of the filling system.

[0041] Further, the present application can continuously and comprehensively obtain the running state information of the filling system in the target time period by periodically collecting initial running parameters, can accurately quantify the change trend of each parameter in the target time period by determining the parameter change representation value corresponding to each running parameter, can timely and accurately judge whether the system running in the target time period is abnormal or deviates from the normal range by comparing the parameter change representation value with the preset change representation value, and can form a comprehensive running state description of the filling system in the target time period by generating the state input variable corresponding to each running parameter to build a state vector, fully reflects the multi-dimensional information of the system, and thus improves the accuracy of the strategy output and realizes accurate regulation and control of the filling system. BRIEF DESCRIPTION OF DRAWINGS

[0042] Figure 1 Fig. 1 is a flowchart of the multi-variable hierarchical adaptive filling regulation and control method based on reinforcement learning of the embodiment of the present application;

[0043] Figure 2 Fig. 2 is a flowchart of the reinforcement learning strategy of the embodiment of the present application;

[0044] Figure 3 Fig. 3 is a flowchart of constructing a multi-dimensional reward function of the embodiment of the present application;

[0045] Figure 4 Fig. 4 is a structural diagram of the hierarchical control system of the embodiment of the present application. DETAILED DESCRIPTION

[0046] In order to make the purpose and advantages of the present application clearer and more apparent, the present application will be further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application.

[0047] The preferred embodiments of the present application will be described below with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are only used to explain the technical principles of the present application and do not limit the protection scope of the present application.

[0048] It should be noted that in the description of the present application, the terms indicating the direction or positional relationship of "upper", "lower", "left", "right", "inner", "outer" and the like are based on the direction or positional relationship shown in the drawings, which is only for the convenience of description, and does not indicate or imply that the device or element must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation on the present application.

[0049] In addition, it should be noted that in the description of the present application, unless otherwise explicitly specified and limited, the terms "mounting", "connection", "connection" should be understood broadly, for example, it can be fixedly connected, or it can be detachably connected, or integrally connected, it can be mechanically connected, or it can be electrically connected, it can be directly connected, or it can be indirectly connected through an intermediate medium, it can be the communication inside two elements. For those skilled in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.

[0050] It can be understood that the filling system includes a batching plant, a pump station, an underground conveying pipeline and a goaf backfill unit, the batching plant is used to prepare the filling slurry, including the cement content and the water-cement ratio, the main pump station is provided with a conveying pump and a clean water pump, the clean water pump is used to control the clean water flow rate, the conveying pump is used to control the slurry flow rate, the underground conveying pipeline is used to convey the filling slurry to the filling area, and the goaf backfill unit is used to backfill the goaf.

[0051] It can be understood that the overall structure of the present application is divided into three main functional levels according to the "perception - decision - execution - feedback" logic: data acquisition layer, reinforcement learning strategy layer and execution control layer, and data transmission and instruction interaction are realized between the layers through industrial Ethernet or Mdbus TCP protocol. The data acquisition layer is composed of multiple types of sensors and field acquisition devices, which is used to acquire the running parameters of the filling system in real time; the reinforcement learning strategy layer is the core, which is mainly responsible for multi-objective strategy optimization output through the target reinforcement learning strategy network model, the input is the state variable, and the output is a set of optimal target set values, which is used to guide the lower controller to execute; the execution control layer is mainly composed of a group of tracking type controllers (such as PID controller or MPC controller), which receives the optimal target set value output by the target reinforcement learning strategy network model, and finely adjusts the slurry proportioning and pump station.

[0052] Please refer to Figures 1-4 It is a flowchart of the multivariate hierarchical adaptive filling regulation method based on reinforcement learning of the embodiment of the present application; Figure 2 It is a flowchart of the reinforcement learning strategy of the embodiment of the present application; Figure 3 It is a flowchart of constructing a multi-dimensional reward function of the embodiment of the present application; Figure 4The figure is a structure schematic diagram of a layered control system of an embodiment of the present application. The embodiment of the present application provides a multivariable layered self-adaptive filling regulation and control method based on reinforcement learning, comprising:

[0053] In step S1, initial operation parameters of the filling system are periodically collected to construct a state vector, wherein the operation parameters include pulp concentration, pulp flow rate, delivery pump outlet pressure, water-cement ratio, backwater rate and strength estimation value;

[0054] Specifically, in the step S1, a state vector is constructed, comprising:

[0055] In step S11, initial operation parameters of the filling system are periodically collected, and parameter change representation values corresponding to each operation parameter are determined based on the initial operation parameters in a target time period;

[0056] In step S12, whether to mark the target time period is determined based on a comparison result of each parameter change representation value and a preset change representation value;

[0057] In step S13, if marked, state input variables corresponding to each operation parameter are generated based on the initial operation parameters in the target time period;

[0058] In step S14, a state vector is constructed based on each state input variable.

[0059] In implementation, the device or method for collecting initial operation parameters of the filling system is not specifically limited, for example, a pulp concentration sensor is arranged to collect the pulp concentration (unit: %) of the filling pulp; a pulp flow rate sensor is arranged to collect the pulp flow rate (unit: m 3 / h) of the filling pulp; a delivery pump outlet pressure sensor is arranged to collect the delivery pump outlet pressure (unit: MPa); a water-cement ratio estimation unit is arranged to determine the water-cement ratio; a backwater rate sensor is arranged to collect the backwater rate (unit: %); and an early strength prediction module is arranged to output the future compressive strength estimation value (unit: MPa) based on early test blocks or data-driven models, preferably, the collection period is set to 0.2s-1s.

[0060] It can be understood that the parameter change representation values corresponding to each initial operation parameter are determined based on the maximum parameter change rate (the ratio of the operation parameter collected at the adjacent collection time to the adjacent collection time difference) of each initial operation parameter in the target time period. The change comparison values corresponding to each initial operation parameter are determined based on the ratio of each parameter change representation value to the corresponding preset change representation value, and if each change comparison value is less than 1, it is determined that the target time period is marked. The actual implementer can set the preset change representation value corresponding to each operation parameter based on the average value of the parameter change representation values corresponding to the filling system operation parameters that pass the qualification test in the historical data.

[0061] It can be understood that if marked, the initial operation parameters are taken as dependent variables, and time is taken as independent variable to generate the state input variables corresponding to each operation parameter.

[0062] The present application can continuously and comprehensively obtain the running state information of the filling system in the target time period by periodically collecting the initial operation parameters, accurately quantify the change trend of each parameter in the target time period by determining the parameter change representation value corresponding to each operation parameter, timely and accurately judge whether the system running in the target time period is abnormal or deviates from the normal range by comparing the parameter change representation value with the preset change representation value, and form a comprehensive running state description of the filling system in the target time period by generating the state input variables corresponding to each operation parameter to construct the state vector, which can comprehensively reflect the multi-dimensional information of the system, thereby improving the accuracy of the strategy output and realizing the precise regulation of the filling system.

[0063] In step S2, the state vector is input into a target reinforcement learning strategy network model to obtain a target action parameter and a state value estimation, wherein the action parameter includes slurry concentration, water-cement ratio and pump speed.

[0064] In implementation, the state vector is wherein, is the slurry concentration, is the slurry flow rate, is the delivery pump outlet pressure, is the water-cement ratio, is the backwater rate, is the strength estimation value, and the output target action parameter is wherein, is the target slurry concentration, is the target water-cement ratio, is the target pump speed, and the target action parameter is a continuous variable.

[0065] In step S3, whether the state value estimation meets the expected standard is determined, and if yes, the filling control parameters of the filling system are adjusted based on the target action parameter, wherein the filling control parameters include cement content, clean water dosage and delivery pump frequency.

[0066] Specifically, in step S3, whether the state value estimation meets the expected standard is determined, including:

[0067] In step S301, an estimation comparison value is determined based on the comparison result of the state value estimation and the preset state value estimation.

[0068] In step S302, whether the estimation comparison value meets the expected standard is determined based on the comparison result of the estimation comparison value and the preset comparison value.

[0069] In implementation, the ratio of the state value estimate to the preset state value estimate is determined as an estimate ratio value, the greater the state value estimate, the more cumulative rewards the agent expects to obtain in the state, if the estimate ratio value is greater than the preset ratio value, it is determined that the expected standard is met, if the estimate ratio value is less than or equal to the preset ratio value, it is not determined that the expected standard is met. The actual implementer can set the preset state value estimate based on the maximum value of the state value estimate passing the eligibility test in the historical data, the greater the preset ratio value, the higher the requirement for the closeness of the state value estimate to the preset state value, preferably, the preset ratio value is in the range of 0.8-0.9.

[0070] Specifically, in the step S3, the filling control parameters of the filling system are adjusted based on the target action parameters, including:

[0071] Step S31, determining a concentration control deviation based on the comparison result of the target slurry concentration and the initial slurry concentration, and determining a cement content control amount and a water-cement ratio control amount based on the concentration control deviation;

[0072] Step S32, determining a water-cement ratio control deviation based on the comparison result of the target water-cement ratio and the initial water-cement ratio, and determining a clean water addition amount control amount based on the water-cement ratio control deviation;

[0073] Step S33, determining a pump speed control deviation based on the comparison result of the target pump speed and the initial pump speed, and determining a conveying pump frequency control amount based on the pump speed control deviation.

[0074] In implementation, the execution control layer includes a group of tracking type controllers (such as PID controllers or MPC controllers) and receives target action parameters to adjust the filling control parameters, and the adjustment logic is: adjusting the cement content and the water-cement ratio according to the deviation of the target cement content and the initial cement content; adjusting the cement content and the water-cement ratio according to the deviation of the target water-cement ratio and the initial water-cement ratio; controlling the clean water addition rate to adjust the clean water addition amount according to the deviation of the target clean water addition rate and the initial clean water addition rate; adjusting the conveying pump frequency according to the deviation of the target conveying pump frequency and the initial conveying pump frequency, and those skilled in the art know that the PID controller control formula is: wherein, is a controller output instruction signal, is a proportional gain coefficient, is an integral gain coefficient, is a differential gain coefficient, , is the deviation of the target action parameter and the initial operation parameter. The actual implementer can control the feeding motor, regulating valve, conveying pump and other execution equipment based on the standard I interface or communication protocol (such as Mdbus RTU, Prfinet).

[0075] Specifically, in the step S3, adjusting the filling control parameters of the filling system based on the target action parameters further comprises:

[0076] Based on the comparison results of the concentration control deviation and the concentration safety index, the comparison results of the target water-cement ratio and the water-cement ratio safety index, and the comparison results of the target pump speed and the pump speed safety index, it is determined whether to trigger the target freezing logic. If not, the cement content control amount is determined based on the concentration control deviation, the water-cement ratio control amount is determined based on the water-cement ratio control deviation, and the conveying pump frequency control amount is determined based on the pump speed control deviation.

[0077] In implementation, in order to ensure stable response capability under sudden disturbance, an upper limit protection mechanism is set. When the deviation exceeds the allowable range (for example, the absolute value of the concentration error is greater than 5%), the target freezing logic is triggered. For example, the concentration safety index is the maximum allowable value of the slurry concentration, the difference between the concentration control deviation and the concentration safety index is determined as the concentration difference, and the ratio of the concentration deviation to the concentration safety index is determined as the concentration deviation amount. The water-cement ratio safety index is the maximum allowable value of the water-cement ratio, the difference between the water-cement ratio control deviation and the water-cement ratio safety index is determined as the water-cement ratio difference, and the ratio of the water-cement ratio deviation to the water-cement ratio safety index is determined as the water-cement ratio deviation amount. The pump speed safety index is the maximum allowable value of the pump speed, the difference between the pump speed control deviation and the pump speed safety index is determined as the pump speed difference, and the ratio of the pump speed deviation to the pump speed safety index is determined as the pump speed deviation amount. It can be understood that if the absolute value of the concentration deviation amount is greater than 5% to 10%, and / or the absolute value of the water-cement ratio deviation amount is greater than 5% to 10%, and / or the absolute value of the pump speed deviation amount is greater than 5% to 10%, it is determined that the target freezing logic is triggered.

[0078] It can be understood that, in the system initialization or insufficient training stage, in order to prevent the reinforcement learning strategy from outputting unfeasible instructions, a strategy soft limit boundary can be set: Instructions outside the range will be clipped to the boundary value for execution.

[0079] Specifically, in the step S3, under the condition that it is determined that the target freezing logic is triggered, the filling control parameters of the filling system are adjusted based on the action parameters in the control period before the current control period. The high-level target update is suspended until the deviation returns to the allowable range and the adjustment process is restarted.

[0080] The present application introduces a hierarchical control mechanism, uses reinforcement learning for upper-level target value setting, decouples the output of the high-level reinforcement learning strategy from the lower-level execution control, accurately determines the concentration control deviation, the water-cement ratio control deviation and the pump speed control deviation, reasonably adjusts the filling control parameters, effectively avoids the risk that strategy shock directly acts on the equipment, and improves the response stability of the control system.

[0081] Step S4, in response to the completion of the adjustment of the filling control parameters of the filling system, constructing the state characteristic functions corresponding to the filling system in the current control period, and constructing a multi-dimensional reward function based on the state characteristic functions, wherein the state characteristic functions include a slurry concentration characteristic function, a water-cement ratio characteristic function, a pumping energy consumption characteristic function, and a strength characteristic function;

[0082] Specifically, in the step S4, the multi-dimensional reward function is constructed, including:

[0083] Step S41, determining a concentration deviation penalty based on the slurry concentration characteristic function and a slurry concentration reference value;

[0084] Step S42, determining a water-cement ratio deviation penalty based on the water-cement ratio characteristic function and a water-cement ratio reference value;

[0085] Step S43, determining a pumping energy consumption penalty based on the pumping energy consumption characteristic function;

[0086] Step S44, determining a strength compliance index based on the strength characteristic function;

[0087] Step S45, constructing a multi-dimensional reward function based on the concentration deviation penalty, the water-cement ratio deviation penalty, the pumping energy consumption penalty, the strength compliance index, a concentration weight coefficient, a water-cement ratio weight coefficient, an energy consumption weight coefficient, and a strength weight coefficient.

[0088] In implementation, the multi-dimensional reward function , wherein, is the concentration deviation penalty, is the slurry concentration reference value; is the water-cement ratio deviation penalty, is the water-cement ratio reference value; is the pumping energy consumption penalty; is the strength compliance index, which is 1 when the strength estimate value is greater than the minimum allowed value (such as 3.5 MPa), and 0 otherwise, is the concentration weight coefficient, is the water-cement ratio weight coefficient, is the energy consumption weight coefficient, is the strength weight coefficient, satisfying In actual application, the weight coefficients can be dynamically adjusted according to different working condition priorities, for example, when working in deep high stress areas, the can be appropriately increased to strengthen the control of strength; while in the areas where material cost control is strict or power load is sensitive, the can be increased to improve the control of pumping energy consumption.The weight is to optimize the pumping energy consumption preferentially. In addition, considering that there are order-of-magnitude differences in the numerical scales of various control targets, all reward items are standardized by a normalization function before being aggregated by weight, to enhance the balance of policy gradient in the training process. For example, the normalization function is: For each penalty term, The lower limit estimate of the target, The upper limit estimate of the target, The lower limit estimate of the target, The upper limit estimate of the target, and the window length includes 30-50 cycle times.

[0089] The multi-dimensional reward function comprehensively considers the accuracy deviation of concentration and water-cement ratio, system energy consumption, and compliance of the mechanical properties of the filling body, forms a reward structure oriented to multi-target coordination, and is adjusted by a weight coefficient, so as to effectively realize the coordination and unity among control targets, realize adaptive trade-off and dynamic adjustment of control targets under complex working conditions, improve the robustness and adaptability of the strategy, and thus further improve the stability and precision of the filling system.

[0090] Step S5: obtaining the control cycle number in the current control period, and calculating a current cumulative reward value corresponding to the current control period based on the multi-dimensional reward function;

[0091] Step S6: determining whether to trigger policy update based on the control cycle number and the current cumulative reward value, and if so, updating the target reinforcement learning policy network model based on the multi-dimensional reward function.

[0092] Specifically, in step S6, determining whether to trigger policy update includes:

[0093] Determining whether to trigger policy update based on the comparison result of the control cycle number and the preset cycle number and the comparison result of the current cumulative reward value and the preset cumulative reward value.

[0094] Specifically, in step S6, if the current cumulative reward value is greater than the preset cumulative reward value, the cumulative reward change in a plurality of consecutive control periods is calculated, wherein if the control cycle number is greater than the preset cycle number, or the cumulative reward change meets a preset change condition, it is determined that policy update is triggered.

[0095] In implementation, the accumulated reward ratio value M in the control cycle is determined based on the accumulated reward value A in the control cycle and the accumulated reward value B in the previous cycle, M=(B-A) / B, if the control cycle number is greater than the preset cycle number, or the accumulated reward ratio value in the continuous H control cycles is greater than the preset accumulated reward ratio value, the strategy update is determined. The actual implementer can set the preset cycle number and the preset accumulated reward ratio value based on the actual situation, preferably, the preset cycle number is set to 200-250, the preset accumulated reward ratio value is set to 15%-20%, and H is set to 10-15.

[0096] Specifically, in the step S6, under the condition that it is determined that the strategy update is not triggered, the control cycle number in the current control cycle is updated, and the initial operation parameters of the filling system are re-collected to update the state vector.

[0097] It can be understood that the normalized multi-dimensional reward function can be used as the instant reward input of the target reinforcement learning strategy network model, so as to guide the model to adaptively optimize the strategy among multiple targets, and the soft coupling mechanism is used in the implementation process, and there is no forced priority among the sub-targets, but the weighted contribution is used to jointly constitute the basis for strategy update, which can avoid the excessive optimization of a target and cause the overall degradation of system performance. In the strategy execution stage, the logical relationship between the target setting values is constrained by the rule base and the boundary limit, for example, if , then cannot be higher than 76%, so as to avoid the abnormality of conveying caused by the decrease of the fluidity of the slurry, and further enhance the physical rationality and industrial implementability of the multi-target control output. Through the introduction of the above-mentioned coordination mechanism, the dynamic balance among the control targets can be realized under complex working conditions, so as to take into account the filling system running efficiency, material consumption, equipment load and mechanical properties and multiple performance indicators, and the practicality and robustness of the control strategy are significantly improved.

[0098] It can be understood that the target reinforcement learning strategy network model structure is as follows: input layer (6 dimensions); 1st hidden layer (128 nodes, ReLU activation); 2nd hidden layer (64 nodes, ReLU activation); output layer: action space is 3 dimensions, using Tanh activation function; and outputting state value estimation at the same time. The training adopts the Actor-Critic architecture, and the truncated ratio optimization strategy is introduced, and the loss function is as follows: , wherein is the policy truncation threshold, which is set to 0.2 by default. The model can output stable strategy after about 100,000 steps of training in the simulation environment, and has strong generalization ability and strategy execution stability.

[0099] It can be understood that the policy network adopts a feedforward neural network structure with two hidden layers, each layer has 128 nodes, and the ReLU activation function is used as the activation function. The output is mapped to the corresponding action interval after Tanh transformation. The policy training adopts a gradient descent optimization method based on time difference, and the objective function is: wherein, represents the current policy, is the old policy, is the advantage function approximation value, and the Clip mechanism is used to stabilize the training process. After the completion of the training of this module, the inference model is solidified and deployed on the industrial edge computing device. Through the input of the field, the optimal target setting value is calculated in real time, and the control command is updated at a set period (such as 10 seconds), so as to realize the adaptive high-level scheduling of the filling process multivariable.

[0100] The initial running parameters of the filling system are periodically collected to construct a state vector, which can comprehensively and accurately reflect the current actual running state of the filling system, providing a rich data basis for subsequent precise regulation and control. Through the target reinforcement learning policy network model, the target action parameters are output according to the state vector, providing precise decision support for the regulation and control of the filling system. The optimal action parameters can be determined according to the current complex system state, and the state value estimation is output at the same time, which helps to quantitatively evaluate the advantages and disadvantages of the current state, and provides a basis for judging whether the state vector meets the expected standard. The cement content, clean water dosage and conveying pump frequency of the filling system are precisely adjusted according to the target action parameters, rather than directly controlling the control system with the target action parameters, which avoids abnormal system execution caused by policy shock or fluctuation, and improves the stability of the filling system. The state characteristic functions including the slurry concentration characteristic function, water-cement ratio characteristic function, pumping energy consumption characteristic function and strength characteristic function are constructed, which can finely describe the running state of the filling system in the current control cycle from multiple dimensions. Based on the state characteristic functions, a multi-dimensional reward function is constructed, which comprehensively considers the concentration stability, water-cement ratio accuracy, pumping energy consumption and strength compliance rate and other key performance indicators, and realizes the comprehensive optimization control of the filling system, thereby improving the precision of the filling system. Based on the control cycle number and the current cumulative reward value, it is determined whether to trigger policy update, and an adaptive update mechanism based on state feedback and performance fluctuation is introduced, involving two policy update triggering modes, fixed period triggering and reward trend triggering, which further improves the robustness of the update, avoids interference to the filling operation caused by frequent restart or control mutation, and improves the adaptive coordination ability of the filling system. Based on the multi-dimensional reward function, the target reinforcement learning policy network is updated, which can continuously improve the control performance of the filling system based on the continuous learning and accumulation of experience, realize the intelligent and adaptive operation of the filling system, and improve the stability and precision of the filling system. Embodiment

[0101] The actual operation duration is 110 minutes. The target concentration is set to 75.5% and the target water-cement ratio is 0.34 at the initial stage of system operation. The target pump speed is 44 Hz. The cement content control amount, water-cement ratio control amount, clean water addition control amount, and conveying pump frequency control amount are adjusted according to the concentration control deviation, water-cement ratio control deviation, and pump speed control deviation. After the water-cement ratio control is stabilized, the 28-day predicted compressive strength fluctuation is maintained between 3.85-4.12 MPa, meeting the engineering requirements. The energy consumption module monitoring data during the process show that under the reinforcement learning strategy, the average power consumption of the unit volume paste pump is 4.72 kWh / m³, which is about 6.5% lower than that of the conventional control mode in the same mining area. Due to the dynamic adjustment of the strategy dosage ratio, the unit cement consumption is reduced by 3.2%. The changes of key indicators during the control period are shown in Table 1.

[0102] Table 1 Changes of key indicators during the control period

[0103]

[0104] During the entire operation cycle of the system, there are no abnormal situations such as pipeline blockage and pressure fluctuation alarm. The control effect is stable, the concentration target is maintained under low water-cement ratio conditions, and the collaborative control purpose of "saving materials and increasing strength" is achieved. Based on the operation data during the control process, the reinforcement learning model will be fine-tuned in the next cycle to adapt to the new material batch and pipeline state, improving the stability and precision of the filling system.

[0105] So far, the technical solutions of the present application have been described in combination with the preferred embodiments shown in the drawings, but those skilled in the art can easily understand that the protection scope of the present application is obviously not limited to these specific embodiments. Those skilled in the art can make equivalent changes or replacements to related technical features without deviating from the principles of the present application. The technical solutions after these changes or replacements will all fall within the protection scope of the present application.

Claims

1. A multivariate hierarchical adaptive filling control method based on reinforcement learning, characterized in that, include: The initial operating parameters of the filling system are periodically collected to construct a state vector, wherein the operating parameters include slurry concentration, slurry flow rate, delivery pump outlet pressure, water-cement ratio, water return rate, and strength estimate. The state vector is input into the target reinforcement learning policy network model to obtain the target action parameters and state value estimate, wherein the action parameters include slurry concentration, water-cement ratio and pump speed. Based on the state value estimation, it is determined whether it meets the expected standard. If it does, the filling control parameters of the filling system are adjusted based on the target action parameters. The filling control parameters include cement dosage, water dosage and delivery pump frequency. In response to the completion of the filling control parameter adjustment of the filling system, a state characteristic function corresponding to the filling system in the current control cycle is constructed, and a multi-dimensional reward function is constructed based on each state characteristic function. The state characteristic function includes a slurry concentration characteristic function, a water-cement ratio characteristic function, a pumping energy consumption characteristic function, and an intensity characteristic function. Obtain the number of control cycles within the current control period, and calculate the current cumulative reward value corresponding to the current control period based on the multidimensional reward function; Based on the number of control loops and the current cumulative reward value, it is determined whether to trigger a policy update. If triggered, the target reinforcement learning policy network model is updated based on the multidimensional reward function. Adjusting the filling control parameters of the filling system based on the target action parameters includes: The concentration control deviation is determined based on the comparison between the target slurry concentration and the initial slurry concentration, and the cement dosage control amount and water-cement ratio control amount are determined based on the concentration control deviation. The water-cement ratio control deviation is determined based on the comparison between the target water-cement ratio and the initial water-cement ratio, and the net water dosage control amount is determined based on the water-cement ratio control deviation. The pump speed control deviation is determined based on the comparison between the target pump speed and the initial pump speed, and the delivery pump frequency control amount is determined based on the pump speed control deviation. Construct a multidimensional reward function, including: The concentration deviation penalty is determined based on the slurry concentration characteristic function and the slurry concentration reference value; The water-cement ratio deviation penalty is determined based on the water-cement ratio characteristic function and the water-cement ratio reference value. The pumping energy consumption penalty is determined based on the pumping energy consumption characteristic function; The strength compliance index is determined based on the strength characteristic function; A multidimensional reward function is constructed based on the concentration deviation penalty, the water-cement ratio deviation penalty, the pumping energy consumption penalty, the intensity compliance index, the concentration weight coefficient, the water-cement ratio weight coefficient, the energy consumption weight coefficient, and the intensity weight coefficient.

2. The multivariate hierarchical adaptive filling control method based on reinforcement learning according to claim 1, characterized in that, Determining whether the state value estimate meets the expected criteria based on the above includes: The estimated comparison value is determined based on the comparison result between the state value estimate and the preset state value estimate; The determination of whether the estimated comparison value meets the expected standard is based on the comparison result between the estimated comparison value and the preset comparison value.

3. The multivariate hierarchical adaptive filling control method based on reinforcement learning according to claim 2, characterized in that, Determining whether a policy update has been triggered includes: Based on the comparison between the number of control loops and the preset number of loops, and the comparison between the current cumulative reward value and the preset cumulative reward value, it is determined whether to trigger a strategy update.

4. The multivariate hierarchical adaptive filling control method based on reinforcement learning according to claim 3, characterized in that, Constructing the state vector includes: The initial operating parameters of the filling system are periodically collected, and the parameter change characterization values ​​corresponding to each operating parameter are determined based on the initial operating parameters within the target time period. The determination of whether to mark the target time period is based on the comparison results between the change characterization values ​​of each parameter and the preset change characterization values; If marked, then the state input variables corresponding to each operating parameter are generated based on the initial operating parameters within the target time period; A state vector is constructed based on each of the aforementioned state input variables.

5. The multivariate hierarchical adaptive filling control method based on reinforcement learning according to claim 4, characterized in that, If it is determined that the strategy update will not be triggered, the number of control cycles in the current control cycle is updated, and the initial operating parameters of the filling system are re-acquired to update the state vector.

6. The multivariate hierarchical adaptive filling control method based on reinforcement learning according to claim 5, characterized in that, If the current cumulative reward value is greater than the preset cumulative reward value, the cumulative reward changes within several consecutive control cycles are calculated. If the number of control cycles is greater than the preset number of cycles, or if the cumulative reward changes meet the preset changes, a strategy update is triggered.

7. The multivariate hierarchical adaptive filling control method based on reinforcement learning according to claim 6, characterized in that, Adjusting the filling control parameters of the filling system based on the target action parameters further includes: Based on the comparison results of the concentration control deviation and the concentration safety index, the comparison results of the target water-cement ratio and the water-cement ratio safety index, and the comparison results of the target pump speed and the pump speed safety index, it is determined whether the target freezing logic is triggered. If it is not triggered, the cement dosage control amount and the water-cement ratio control amount are determined based on the concentration control deviation, the net water addition amount control amount is determined based on the water-cement ratio control deviation, and the delivery pump frequency control amount is determined based on the pump speed control deviation.

8. The multivariate hierarchical adaptive filling control method based on reinforcement learning according to claim 7, characterized in that, Under the condition that the target freeze logic is triggered, the filling control parameters of the filling system are adjusted based on the action parameters in the control cycles prior to the current control cycle.

Citation Information

Patent Citations

  • High-performance foamed mortar filling method for mining sites

    CN102261262A

  • Method for optimizing filling material ratio

    CN106746946A