Multivariable layered adaptive filling regulation and control method based on reinforcement learning
By adopting multivariate hierarchical adaptive filling control method in the mine filling system, and using reinforcement learning to construct state vectors and multidimensional reward functions, the problem of insufficient multivariate coordination in the existing technology is solved, and the high accuracy and stability of the filling system are achieved.
Patent Information
- Application Number
- CN202510948591.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-10
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2045-07-10
AI Technical Summary
The existing reinforcement learning only focuses on a single operating parameter in the mine filling system, and fails to achieve adaptive dynamic coordination between multiple key variables. The control structure design is mainly end-to-end strategy, resulting in low accuracy and stability of filling control.
Using a multivariate hierarchical adaptive filling control method based on reinforcement learning, the state vector is constructed by periodically collecting the initial operating parameters of the filling system, inputting the target reinforcement learning strategy network model, adjusting the filling control parameters, and constructing a multi-dimensional reward function to achieve multi-dimensional state feedback and policy updates.
The stability and accuracy of the filling system are improved, and the comprehensive optimization control of multiple key performance indicators such as slurry concentration, water-cement ratio, and pumping energy consumption is achieved, which enhances the adaptability and robustness of the system.
Smart Images

Figure CN120506267A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of mine filling control, and in particular to a multi-variable hierarchical adaptive filling control method based on reinforcement learning. Background Art
[0002] Backfill mining technology, a key means of ensuring safe recovery and stable goaf control in underground mines, has been widely adopted in nonferrous, ferrous, and non-metallic deposits. In recent years, with the increasing depth of ore bodies and the increasing complexity of ground pressure environments, backfill operations face higher demands for reliability, continuity, and adaptability. With the continuous development of various processes such as slurry backfilling, paste backfilling, and tailings cemented backfilling, backfill systems have gradually evolved from traditional manual control to automated and intelligent systems.
[0003] The operation of the filling system encompasses multiple steps, including material proportioning, conveying and pumping, pipeline regulation, and stope filling. These parameters involve multi-dimensional control parameters such as concentration, water-cement ratio, pumping pressure, ash content, and temperature and humidity. These parameters are strongly coupled and fluctuate significantly with the construction environment, material properties, and operational tasks. Traditional processes rely on operational experience to set parameters, which are susceptible to subjective judgment and delayed response. The control process lacks systematicity and is difficult to adapt to the complex and changing mining environment.
[0004] To improve the efficiency and stability of filling systems, researchers have gradually introduced methods such as model predictive control, genetic algorithms, and multi-objective optimization to achieve coordinated optimization between objective functions. However, most of these methods are designed based on static data, resulting in control strategies that lack adaptability and are unable to quickly respond to abnormal conditions. With the gradual improvement of industrial data acquisition and edge computing platforms, data-driven intelligent control models have become a research focus. Reinforcement learning, an interactive learning method for dynamic decision-making problems, shows promising application prospects in industrial process control. By continuously interacting with the environment and not relying on explicit mathematical models, it can continuously modify control strategies during actual operation, improving the system's adaptability and long-term performance. Current research has attempted to apply reinforcement learning to processes such as mineral processing, grinding, and filling batching, achieving initial results in nonlinear system modeling and control decision-making. Although reinforcement learning control has achieved certain breakthroughs in mineral processing automation, its application in mine filling systems is still in the exploratory stage. Existing research mostly focuses on single-variable control objectives, such as concentration control or underflow concentration stability, and lacks a joint control mechanism between multiple variables such as water-cement ratio, pump speed, and material consumption. At the same time, most control structures still adopt end-to-end strategies, and reinforcement learning outputs act directly on the equipment execution layer. In the filling system, parameter control involves the dynamic coordination of multiple variables, and the control action is continuous, coupled, and irreversible. The existing reinforcement learning control structure is difficult to meet engineering-level stability and accuracy requirements. Summary of the Invention
[0005] To this end, the present invention provides a multi-variable hierarchical adaptive filling control method based on reinforcement learning to overcome the problems in the existing technology that only focuses on a single operating parameter and fails to achieve adaptive dynamic coordination between multiple key variables. The control structure design is mainly based on end-to-end strategy, and the reinforcement learning output directly acts on the equipment execution layer, resulting in relatively low filling control accuracy and stability.
[0006] To achieve the above objectives, the present invention provides a multivariable hierarchical adaptive filling control method based on reinforcement learning, comprising:
[0007] Periodically collecting initial operating parameters of the filling system to construct a state vector, wherein the operating parameters include slurry concentration, slurry flow rate, delivery pump outlet pressure, water-cement ratio, return water rate, and strength estimation value;
[0008] Inputting the state vector into a target reinforcement learning policy network model to obtain target action parameters and state value estimates, wherein the action parameters include slurry concentration, water-cement ratio, and pump speed;
[0009] Determining whether the state value estimate meets the expected standard based on the state value estimate, and if so, adjusting the filling control parameters of the filling system based on the target action parameters, wherein the filling control parameters include cement content, clean water dosage, and delivery pump frequency;
[0010] In response to completion of the filling control parameter adjustment of the filling system, constructing a state characteristic function corresponding to the filling system in the current control period, and constructing a multidimensional reward function based on each state characteristic function, wherein the state characteristic function includes a slurry concentration characteristic function, a water-cement ratio characteristic function, a pumping energy consumption characteristic function, and an intensity characteristic function;
[0011] Obtaining the number of control cycles in the current control period, and calculating a current cumulative reward value corresponding to the current control period based on the multidimensional reward function;
[0012] Whether to trigger a policy update is determined based on the number of control cycles and the current accumulated reward value; if triggered, the target reinforcement learning policy network model is updated based on the multidimensional reward function.
[0013] Furthermore, adjusting the filling control parameters of the filling system based on the target action parameters includes:
[0014] Determining a concentration control deviation based on a comparison result of the target slurry concentration and the initial slurry concentration, and determining a cement content control amount and a water-cement ratio control amount based on the concentration control deviation;
[0015] Determining a water-cement ratio control deviation based on a comparison result of the target water-cement ratio and the initial water-cement ratio, and determining a controlled amount of net water addition based on the water-cement ratio control deviation;
[0016] A pump speed control deviation is determined based on a comparison result between the target pump speed and the initial pump speed, and a delivery pump frequency control amount is determined based on the pump speed control deviation.
[0017] Furthermore, a multi-dimensional reward function is constructed, including:
[0018] determining a concentration deviation penalty based on the slurry concentration characteristic function and a slurry concentration reference value;
[0019] determining a water-cement ratio deviation penalty based on the water-cement ratio characteristic function and a water-cement ratio reference value;
[0020] determining a pumping energy consumption penalty based on the pumping energy consumption characteristic function;
[0021] determining a strength compliance index based on the strength characteristic function;
[0022] A multidimensional reward function is constructed based on the concentration deviation penalty, the water-cement ratio deviation penalty, the pumping penalty, the strength compliance index, the concentration weight coefficient, the water-cement ratio weight coefficient, the energy consumption weight coefficient, and the strength weight coefficient.
[0023] Furthermore, determining whether the estimated state value meets the expected standard includes:
[0024] determining an estimated comparison value based on a comparison result of the state value estimate and a preset state value estimate;
[0025] The comparison result between the estimated comparison value and the preset comparison value determines whether it meets the expected standards.
[0026] Furthermore, determining whether to trigger a policy update includes:
[0027] Based on the comparison result of the control cycle number and the preset cycle number and the comparison result of the current cumulative reward value and the preset cumulative reward value, it is determined whether to trigger a policy update.
[0028] Furthermore, a state vector is constructed, including:
[0029] Periodically collecting initial operating parameters of the filling system, and determining parameter change representation values corresponding to each operating parameter based on the initial operating parameters within a target time period;
[0030] Determining whether to mark the target time period based on a comparison result of each parameter change representation value with a preset change representation value;
[0031] If marked, the state input variables corresponding to each operating parameter are generated based on the initial operating parameters in the target time period;
[0032] A state vector is constructed based on each of the state input variables.
[0033] Furthermore, under the condition that it is determined that the strategy update is not triggered, the number of control cycles in the current control period is updated, and the initial operating parameters of the filling system are re-collected to update the state vector.
[0034] Furthermore, if the current cumulative reward value is greater than the preset cumulative reward value, the cumulative reward change within several consecutive control cycles is calculated, wherein, if the number of control cycles is greater than the preset number of cycles, or the cumulative reward change meets the preset change, it is determined to trigger a strategy update.
[0035] Furthermore, adjusting the filling control parameters of the filling system based on the target action parameters further includes:
[0036] Based on the comparison results of the concentration control deviation and the concentration safety index, the comparison results of the target water-cement ratio and the water-cement ratio safety index, and the comparison results of the target pump speed and the pump speed safety index, it is determined whether the target freezing logic is triggered. If not triggered, the cement content control amount and the water-cement ratio control amount are determined based on the concentration control deviation, the net water addition amount control amount is determined based on the water-cement ratio control deviation, and the delivery pump frequency control amount is determined based on the pump speed control deviation.
[0037] Furthermore, under the condition of determining that the target freezing logic is triggered, the filling control parameters of the filling system are adjusted based on the action parameters in the control cycle before the current control cycle.
[0038] Compared with the prior art, the beneficial effect of the present invention is that the present invention can comprehensively and accurately reflect the current actual operating status of the filling system by periodically collecting the initial operating parameters of the filling system to construct a state vector, providing a rich data basis for subsequent precise control. Through the target reinforcement learning strategy network model, the target action parameters are output according to the state vector, providing accurate decision support for the control of the filling system, and the optimal action parameters can be determined according to the current complex system state. At the same time, the state value estimate is output, which helps to quantitatively evaluate the pros and cons of the current state and provide a basis for judging whether the state vector meets the expected standards. The filling control parameters such as the cement content, clean water addition amount and delivery pump frequency of the filling system are accurately adjusted according to the target action parameters, rather than directly controlling the control system with the target action parameters, so as to avoid abnormal system execution due to strategy shock or fluctuation, and improve the stability of the filling system. The construction of state characteristic functions, including those for slurry concentration, water-cement ratio, pumping energy consumption, and strength, allows for a detailed characterization of the filling system's operating state within the current control cycle from multiple dimensions. A multidimensional reward function is constructed based on each state characteristic function, comprehensively considering multiple key performance indicators, such as concentration stability, water-cement ratio accuracy, pumping energy consumption, and strength compliance rate. This achieves comprehensive optimal control of the filling system, thereby improving its accuracy. The triggering of policy updates is determined based on the number of control cycles and the current accumulated reward value. An adaptive update mechanism based on state feedback and performance fluctuations is introduced, involving two policy update triggering methods: fixed-period triggering and reward trend triggering. This further enhances the robustness of the update, prevents frequent restarts or control mutations from interfering with on-site filling operations, and improves the adaptive coordination capabilities of the filling system. Updating the target reinforcement learning policy network based on the multidimensional reward function enables the control system to continuously improve the control performance of the filling system based on continuous learning and experience accumulation, achieving intelligent and adaptive operation of the filling system and improving its stability and accuracy.
[0039] Furthermore, the present invention introduces a hierarchical control mechanism and uses reinforcement learning for upper-level target value setting. By decoupling the high-level reinforcement learning strategy output from the lower-level execution control, the concentration control deviation, water-cement ratio control deviation and pump speed control deviation are accurately determined, and the filling control parameters are reasonably adjusted, effectively avoiding the risk of strategy oscillations directly affecting the equipment and improving the response stability of the control system.
[0040] Furthermore, the multidimensional reward function of the present invention comprehensively considers the accuracy deviation of concentration and water-cement ratio, system energy consumption and compliance with the mechanical properties of the filling body, forming a reward structure for multi-objective collaboration, and adjusting it through weight coefficients to effectively achieve coordination and unification between control objectives. It can realize adaptive weighing and dynamic adjustment of control objectives under complex working conditions, improve the robustness and adaptability of the strategy, and thus further improve the stability and accuracy of the filling system.
[0041] Furthermore, the present invention can continuously and comprehensively obtain the operating status information of the filling system within the target time period by periodically collecting initial operating parameters. By determining the parameter change characterization value corresponding to each operating parameter, the changing trend of each parameter within the target time period can be accurately quantified. By comparing the parameter change characterization value with the preset change characterization value, it can timely and accurately judge whether the system operation within the target time period is abnormal or deviates from the normal range. When the target time period is marked, the state input variables corresponding to each operating parameter are generated, and the state vector is constructed to form a comprehensive operating status description of the filling system within the target time period, which can comprehensively reflect the multi-dimensional information of the system, thereby improving the accuracy of the strategy output and realizing precise regulation of the filling system. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 Schematic diagram of the process of a multi-variable hierarchical adaptive filling control method based on reinforcement learning according to an embodiment of the present invention;
[0043] Figure 2 A schematic diagram of a process for a reinforcement learning strategy according to an embodiment of the present invention;
[0044] Figure 3 A schematic diagram of a process for constructing a multi-dimensional reward function according to an embodiment of the present invention;
[0045] Figure 4 Schematic diagram of the structure of a hierarchical control system according to an embodiment of the present invention. DETAILED DESCRIPTION
[0046] In order to make the objects and advantages of the present invention more clearly understood, the present invention is further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are merely used to explain the present invention and are not intended to limit the present invention.
[0047] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood by those skilled in the art that these embodiments are only used to explain the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.
[0048] It should be noted that, in the description of the present invention, terms such as "up", "down", "left", "right", "inside", and "outside" indicating directions or positional relationships are based on the directions or positional relationships shown in the accompanying drawings. This is only for the convenience of description and does not indicate or imply that the device or element must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it cannot be understood as a limitation on the present invention.
[0049] Furthermore, it should be noted that, in the description of the present invention, unless otherwise expressly specified or limited, the terms "mounted," "connected," and "connected" should be understood in a broad sense. For example, they may refer to fixed connections, detachable connections, or integral connections; mechanical connections or electrical connections; direct connections or indirect connections through an intermediate medium; and internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on specific circumstances.
[0050] It can be understood that the filling system of the present invention includes a batching workshop, a pumping station, an underground transmission pipeline and a goaf backfill unit. The batching workshop is used to prepare the filling slurry, including the cement content and the water-cement ratio. The main pumping station is equipped with a delivery pump and a clean water pump. The clean water pump is used to control the clean water flow rate, and the delivery pump is used to control the slurry flow rate. The underground transmission pipeline is used to transport the filling slurry to the filling area, and the goaf backfill unit is used to backfill the goaf.
[0051] It can be understood that the overall structure of the present invention is divided into three main functional levels according to the "perception-decision-execution-feedback" logic: data acquisition layer, reinforcement learning strategy layer, and execution control layer. Data transmission and instruction interaction are achieved between each layer through industrial Ethernet or Mdbus TCP protocol. The data acquisition layer is composed of multiple types of sensors and field acquisition devices, which are used to obtain the operating parameters of the filling system in real time; the reinforcement learning strategy layer is the core, which is mainly responsible for multi-objective strategy optimization output through the target reinforcement learning strategy network model. Its input is state variables, and its output is a set of optimal target setting values, which are used to guide the execution of the lower-level controller; the execution control layer is mainly composed of a group of tracking controllers (such as PID controllers or MPC controllers), which receive the optimal target setting values output by the target reinforcement learning strategy network model and make fine adjustments to the slurry ratio and pump station.
[0052] See also Figure 1-Figure 4 As shown, it is a flow chart of a multi-variable hierarchical adaptive filling control method based on reinforcement learning according to an embodiment of the present invention; Figure 2 A schematic diagram of a process for a reinforcement learning strategy according to an embodiment of the present invention;
[0053] Figure 3 A schematic diagram of a process for constructing a multi-dimensional reward function according to an embodiment of the present invention; Figure 4 Schematic diagram of the structure of a hierarchical control system according to an embodiment of the present invention. The embodiment of the present invention provides a multivariable hierarchical adaptive filling control method based on reinforcement learning, including:
[0054] Step S1, periodically collecting initial operating parameters of the filling system to construct a state vector, wherein the operating parameters include slurry concentration, slurry flow rate, delivery pump outlet pressure, water-cement ratio, return water rate, and strength estimation value;
[0055] Specifically, in step S1, constructing a state vector includes:
[0056] Step S11, periodically collecting initial operating parameters of the filling system, and determining parameter change characterization values corresponding to each operating parameter based on the initial operating parameters within a target time period;
[0057] Step S12, determining whether to mark the target time period based on a comparison result of each parameter change representation value with a preset change representation value;
[0058] Step S13: if marked, then generate state input variables corresponding to each operating parameter based on the initial operating parameters within the target time period;
[0059] Step S14: constructing a state vector based on each of the state input variables.
[0060] In practice, there is no specific limitation on the device or method for collecting the initial operating parameters of the filling system. For example, a slurry concentration sensor is set to collect the slurry concentration (unit: %) of the filling slurry; a slurry flow rate sensor is set to collect the slurry flow rate (unit: m 3 / h); a delivery pump outlet pressure sensor is set to collect the delivery pump outlet pressure (unit: MPa); the water-cement ratio is determined by a water-cement ratio estimation unit; a return water rate sensor is set to collect the return water rate (unit: %); a strength prediction module is set to output a future compressive strength estimate (unit: MPa) based on early test block tests or data-driven models. Preferably, the collection period is set to 0.2s to 1s.
[0061] It is understood that the parameter change representation value corresponding to each initial operating parameter is determined based on the maximum parameter change rate of each initial operating parameter within the target time period (the ratio of the operating parameters collected at adjacent collection times to the difference between adjacent collection times). The change comparison value corresponding to each initial operating parameter is determined based on the ratio of each parameter change representation value to the corresponding preset change representation value. If each change comparison value is less than 1, the target time period is determined and marked. In practice, the implementer can set the preset change representation value for each operating parameter based on the average parameter change representation value corresponding to the filling system operating parameters that have passed the qualification test in historical data.
[0062] It can be understood that if marked, the initial operating parameters are used as dependent variables and time is used as independent variables to generate state input variables corresponding to each operating parameter.
[0063] The present invention can continuously and comprehensively obtain the operating status information of the filling system within the target time period by periodically collecting initial operating parameters. By determining the parameter change characterization value corresponding to each operating parameter, the changing trend of each parameter within the target time period can be accurately quantified. By comparing the parameter change characterization value with the preset change characterization value, it can timely and accurately judge whether the system operation within the target time period is abnormal or deviates from the normal range. When the target time period is marked, the state input variables corresponding to each operating parameter are generated, and the state vector is constructed to form a comprehensive operating status description of the filling system within the target time period, which can comprehensively reflect the multi-dimensional information of the system, thereby improving the accuracy of the strategy output and realizing precise regulation of the filling system.
[0064] Step S2: inputting the state vector into a target reinforcement learning strategy network model to obtain target action parameters and state value estimation, wherein the action parameters include slurry concentration, water-cement ratio, and pump speed;
[0065] In the implementation, the state vector is S t ={c t ,q t ,p t ,w t ,r t ,s t}, where c t is the slurry concentration, q t is the slurry flow rate, p t is the outlet pressure of the delivery pump, w t is the water-cement ratio, r t is the return water rate, s t is the strength estimation value, and the output target action parameter is in, is the target slurry concentration, is the target water-cement ratio, is the target pump speed, and the target action parameters are all continuous variables.
[0066] Step S3, determining whether the estimated state value meets the expected standard based on the estimated state value; if so, adjusting the filling control parameters of the filling system based on the target action parameters, wherein the filling control parameters include cement content, clean water dosage, and delivery pump frequency;
[0067] Specifically, in step S3, determining whether the estimated state value meets the expected standard includes:
[0068] Step S301, determining an estimated comparison value based on a comparison result between the state value estimate and a preset state value estimate;
[0069] Step S302 : determining whether the estimated comparison value meets the expected standard based on the comparison result with the preset comparison value.
[0070] In implementation, the ratio of the state value estimate to the preset state value estimate is determined as the estimated comparison value. The larger the state value estimate, the more cumulative rewards the agent expects to receive in that state. If the estimated comparison value is greater than the preset comparison value, the agent is judged to have met the expected standard. If the estimated comparison value is less than or equal to the preset comparison value, the agent is not judged to have met the expected standard. Actual implementers can set the preset state value estimate based on the maximum state value estimate that has passed the qualification test in historical data. The larger the preset comparison value, the higher the requirement for the closeness of the state value estimate to the preset state value. Preferably, the preset comparison value ranges from 0.8 to 0.9.
[0071] Specifically, in step S3, adjusting the filling control parameters of the filling system based on the target action parameters includes:
[0072] Step S31, determining a concentration control deviation based on a comparison result between the target slurry concentration and the initial slurry concentration, and determining a cement content control amount and a water-cement ratio control amount based on the concentration control deviation;
[0073] Step S32, determining a water-cement ratio control deviation based on a comparison result between the target water-cement ratio and the initial water-cement ratio, and determining a controlled amount of clean water addition based on the water-cement ratio control deviation;
[0074] Step S33 : determining a pump speed control deviation based on a comparison result between the target pump speed and the initial pump speed, and determining a delivery pump frequency control amount based on the pump speed control deviation.
[0075] In practice, the execution control layer consists of a set of tracking controllers (such as PID controllers or MPC controllers), which receive target action parameters and adjust the filling control parameters. The adjustment logic is as follows: Adjust cement content and water-cement ratio; value, control the water addition acceleration rate to adjust the water addition amount; Value, adjust the delivery pump frequency, those skilled in the art know that the PID controller control formula is: Among them, u(t) is the controller output command signal, K p is the proportional gain coefficient, K i is the integral gain coefficient, K d is the differential gain coefficient, e(t) is the deviation between the target action parameters and the initial operating parameters. In practice, implementation personnel can control feeding motors, regulating valves, delivery pumps, and other actuators based on standard I interfaces or communication protocols (such as Mdbus RTU and Prfinet).
[0076] Specifically, in step S3, adjusting the filling control parameters of the filling system based on the target action parameters further includes:
[0077] Based on the comparison results of the concentration control deviation and the concentration safety index, the comparison results of the target water-cement ratio and the water-cement ratio safety index, and the comparison results of the target pump speed and the pump speed safety index, it is determined whether the target freezing logic is triggered. If not triggered, the cement content control amount and the water-cement ratio control amount are determined based on the concentration control deviation, the net water addition amount control amount is determined based on the water-cement ratio control deviation, and the delivery pump frequency control amount is determined based on the pump speed control deviation.
[0078] During implementation, in order to ensure stable response capability under sudden disturbances, an upper limit protection mechanism is set. When the deviation exceeds the allowable range (such as the absolute value of the concentration error is greater than 5%), the target freezing logic will be triggered. For example, the concentration safety index is the maximum allowable value of the slurry concentration, and the difference between the concentration control deviation and the concentration safety index is determined as the concentration difference, and the ratio of the concentration deviation to the concentration safety index is determined as the concentration deviation; the water-cement ratio safety index is the maximum allowable value of the water-cement ratio, and the difference between the water-cement ratio control deviation and the water-cement ratio safety index is determined as the water-cement ratio difference, and the ratio of the water-cement ratio deviation to the water-cement ratio safety index is determined as the water-cement ratio deviation; the pump speed safety index is the maximum allowable value of the pump speed, and the difference between the pump speed control deviation and the pump speed safety index is determined as the pump speed difference, and the ratio of the pump speed deviation to the pump speed safety index is determined as the pump speed deviation. It is understandable that if the absolute value of the concentration deviation is greater than 5% to 10%, and / or the absolute value of the water-cement ratio deviation is greater than 5% to 10%, and / or the absolute value of the pump speed deviation is greater than 5% to 10%, it is determined that the target freezing logic is triggered.
[0079] It is understandable that during system initialization or insufficient training, in order to prevent the reinforcement learning strategy from outputting infeasible instructions, the strategy soft limit can be set: Instructions outside the range will be pruned to the boundary value for execution.
[0080] Specifically, in step S3, if the target freeze logic is determined to be triggered, the filling control parameters of the filling system are adjusted based on the action parameters in the control cycle before the current control cycle. The high-level target update is suspended until the deviation returns to within the allowable range, and then the adjustment process is restarted.
[0081] The present invention introduces a hierarchical control mechanism and uses reinforcement learning for upper-level target value setting. By decoupling the high-level reinforcement learning strategy output from the lower-level execution control, the concentration control deviation, water-cement ratio control deviation and pump speed control deviation are accurately determined, and the filling control parameters are reasonably adjusted. This effectively avoids the risk of strategy oscillations directly affecting the equipment and improves the response stability of the control system.
[0082] Step S4, in response to the completion of the filling control parameter adjustment of the filling system, constructing a state characteristic function corresponding to the filling system in the current control period, and constructing a multidimensional reward function based on each state characteristic function, wherein the state characteristic function includes a slurry concentration characteristic function, a water-cement ratio characteristic function, a pumping energy consumption characteristic function, and a strength characteristic function;
[0083] Specifically, in step S4, a multi-dimensional reward function is constructed, including:
[0084] Step S41, determining a concentration deviation penalty based on the slurry concentration characteristic function and a slurry concentration reference value;
[0085] Step S42, determining a water-cement ratio deviation penalty based on the water-cement ratio characteristic function and a water-cement ratio reference value;
[0086] Step S43, determining a pumping energy consumption penalty based on the pumping energy consumption characteristic function;
[0087] Step S44, determining a strength compliance index based on the strength characteristic function;
[0088] Step S45: constructing a multidimensional reward function based on the concentration deviation penalty, the water-cement ratio deviation penalty, the pumping penalty, the strength compliance index, the concentration weight coefficient, the water-cement ratio weight coefficient, the energy consumption weight coefficient, and the strength weight coefficient.
[0089] In implementation, the multidimensional reward function Among them, -|c t -c ref | is the concentration deviation penalty, c ref is the reference value of slurry concentration; -|w t -w ref | is the water-cement ratio deviation penalty, w ref is the reference value of water-cement ratio; -E t Penalties for pumping energy consumption; is the strength compliance indicator. When the estimated strength value is greater than the minimum allowable value s min (such as 3.5MPa), take 1, otherwise take 0, α1 is the concentration weight coefficient, α2 is the water-cement ratio weight coefficient, α3 is the energy consumption weight coefficient, α4 is the strength weight coefficient, satisfy In practical applications, the weight coefficient can be dynamically adjusted based on the priority of different working conditions. For example, when operating in deep, high-stress areas, α4 can be appropriately increased to strengthen strength control. In areas with strict material cost control or sensitive power loads, the weight of α3 can be increased to prioritize pumping energy consumption. In addition, considering the orders of magnitude differences in the numerical scales of various control objectives, all reward items are normalized using a normalization function before weighted aggregation to enhance the balance of policy gradients during training. For example, the normalization function: f i (t) is each penalty item, minf i is the lower limit estimate of the i-th target in the empirical data window, maxf i is the upper limit estimate of the i-th target in the empirical data window, and the window length includes 30 to 50 cycles.
[0090] The multidimensional reward function of the present invention comprehensively considers the accuracy deviation of concentration and water-cement ratio, system energy consumption and compliance with the mechanical properties of the filling body, forming a reward structure for multi-objective collaboration, and adjusting it through weight coefficients to effectively achieve coordination and unification between control objectives. It can realize adaptive weighing and dynamic adjustment of control objectives under complex working conditions, improve the robustness and adaptability of the strategy, and thus further improve the stability and accuracy of the filling system.
[0091] Step S5, obtaining the number of control cycles in the current control period, and calculating a current cumulative reward value corresponding to the current control period based on the multidimensional reward function;
[0092] Step S6: determining whether to trigger a policy update based on the number of control cycles and the current accumulated reward value; if triggered, updating the target reinforcement learning policy network model based on the multidimensional reward function.
[0093] Specifically, in step S6, determining whether to trigger a policy update includes:
[0094] Based on the comparison result of the control cycle number and the preset cycle number and the comparison result of the current cumulative reward value and the preset cumulative reward value, it is determined whether to trigger a policy update.
[0095] Specifically, in step S6, if the current cumulative reward value is greater than the preset cumulative reward value, the cumulative reward change within several consecutive control cycles is calculated, wherein, if the number of control cycles is greater than the preset number of cycles, or the cumulative reward change meets the preset change, it is determined to trigger a strategy update.
[0096] In implementation, the cumulative reward comparison value M for any control cycle is determined based on the cumulative reward value A within the previous cycle and the cumulative reward value B within the previous cycle, where M = (B-A) / B. If the number of control cycles is greater than the preset number of cycles, or if the cumulative reward comparison value for H consecutive control cycles is greater than the preset cumulative reward comparison value, then a trigger strategy update is determined. The actual implementation personnel can set the preset number of cycles and the preset cumulative reward comparison value based on actual conditions. Preferably, the preset number of cycles is set to a value range of 200-250, the preset cumulative reward comparison value is set to a value range of 15%-20%, and the value range of H is set to 10-15.
[0097] Specifically, in step S6, if it is determined that the strategy update is not triggered, the number of control cycles in the current control period is updated, and the initial operating parameters of the filling system are re-collected to update the state vector.
[0098] It is understandable that the normalized multidimensional reward function can be used as the immediate reward input of the target reinforcement learning strategy network model, thereby guiding the model to perform adaptive strategy optimization among multiple targets. The coordination mechanism adopts a soft coupling mechanism in the implementation process. There is no mandatory priority between the sub-targets, and the weighted contribution constitutes the basis for strategy update. This can avoid over-optimization of a certain target and lead to overall degradation of system performance. In the strategy execution stage, the logical relationship between the target setting values is constrained by the rule base and boundary restrictions. For example, if it is set: but The maximum pressure drop (P / R) of the filling system should not exceed 76% to avoid transport anomalies caused by decreased slurry fluidity, further enhancing the physical rationality and industrial feasibility of the multi-objective control output. By introducing this coordination mechanism, the present invention can achieve a dynamic balance between control objectives under complex working conditions, thereby taking into account multiple performance indicators such as filling system operating efficiency, material consumption, equipment load, and mechanical properties, significantly improving the practicality and robustness of the control strategy.
[0099] It can be understood that the target reinforcement learning strategy network model structure is as follows: input layer (6 dimensions); first hidden layer (128 nodes, ReLU activation); second hidden layer (64 nodes, ReLU activation); output layer: action space is 3-dimensional, using Tanh activation function; and output state value estimate V (S t ). The training adopts the Actor-Critic architecture and introduces the truncation ratio optimization strategy. Its loss function is as follows: in ∈ is the policy cutoff threshold, which is set to 0.2 by default. This model can output a stable policy after completing about 100,000 steps of training in a simulation environment, and has strong generalization ability and policy execution stability.
[0100] It can be understood that the policy network adopts a feedforward neural network structure with two hidden layers. The number of nodes in each layer is set to 128. The activation function uses the ReLU activation function. The output is mapped to the corresponding action interval after Tanh transformation. The policy training adopts the gradient descent optimization method based on time difference. The objective function is: Among them, π θ Indicates the current strategy, For the old strategy, To approximate the advantage function, a Clip mechanism is used to stabilize the training process. After training, the module is solidified into an inference model and deployed on industrial edge computing devices. It calculates and outputs the optimal target setpoint in real time based on field input, and updates control instructions at a set interval (e.g., 10 seconds), thereby achieving adaptive high-level scheduling of multiple variables in the filling process.
[0101] The present invention periodically collects the initial operating parameters of the filling system to construct a state vector, which can comprehensively and accurately reflect the current actual operating status of the filling system and provide a rich data basis for subsequent precise control. Through the target reinforcement learning strategy network model, the target action parameters are output according to the state vector, providing accurate decision support for the control of the filling system. It can determine the optimal action parameters according to the current complex system state, and output the state value estimate at the same time, which helps to quantitatively evaluate the pros and cons of the current state and provide a basis for judging whether the state vector meets the expected standards. The filling control parameters such as the cement content, clean water addition amount and delivery pump frequency of the filling system are accurately adjusted according to the target action parameters, rather than directly controlling the control system with the target action parameters, so as to avoid abnormal system execution due to strategy shock or fluctuation, and improve the stability of the filling system. The construction of state characteristic functions, including those for slurry concentration, water-cement ratio, pumping energy consumption, and strength, allows for a detailed characterization of the filling system's operating state within the current control cycle from multiple dimensions. A multidimensional reward function is constructed based on each state characteristic function, comprehensively considering multiple key performance indicators, such as concentration stability, water-cement ratio accuracy, pumping energy consumption, and strength compliance rate. This achieves comprehensive optimal control of the filling system, thereby improving its accuracy. The triggering of policy updates is determined based on the number of control cycles and the current accumulated reward value. An adaptive update mechanism based on state feedback and performance fluctuations is introduced, involving two policy update triggering methods: fixed-period triggering and reward trend triggering. This further enhances the robustness of the update, prevents frequent restarts or control mutations from interfering with on-site filling operations, and improves the adaptive coordination capabilities of the filling system. Updating the target reinforcement learning policy network based on the multidimensional reward function enables the control system to continuously improve the control performance of the filling system based on continuous learning and experience accumulation, achieving intelligent and adaptive operation of the filling system and improving its stability and accuracy.
[0102] Example 1:
[0103] The actual operation lasted 110 minutes. The initial strategy layer of the system operation set the target concentration to 75.5%, the target water-cement ratio to 0.34, and the target pump speed to 44Hz. The cement content control amount, water-cement ratio control amount, clean water addition amount control amount, and delivery pump frequency control amount were adjusted according to the concentration control deviation, water-cement ratio control deviation, and pump speed control deviation. After the water-cement ratio control was stabilized, the 28-day predicted compressive strength fluctuation remained between 3.85 and 4.12 MPa, meeting the project requirements. The energy consumption module monitoring data during the process showed that under the reinforcement learning strategy, the average power consumption per unit volume of paste pumping was 4.72 kWh / m 3, which is about 6.5% lower than the conventional control mode in the same mining area; due to the dynamic adjustment of the dosage ratio by the strategy, the unit cement consumption was reduced by 3.2%. The changes in key indicators during the control period are shown in Table 1.
[0104] Table 1 Changes in key indicators during the control period
[0105]
[0106] During the entire operation cycle, the system did not experience any abnormal situations such as pipeline blockage and pressure fluctuation alarms. The control effect was stable and the concentration target was maintained under low water-cement ratio conditions, achieving the collaborative control purpose of "material saving and enhancement". Based on the operating data during the control process, the reinforcement learning model will fine-tune the weights in the next cycle to adapt to the new material batches and pipeline status, thereby improving the stability and accuracy of the filling system.
[0107] Thus far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art may make equivalent changes or substitutions to the relevant technical features, and the technical solutions after such changes or substitutions will fall within the scope of protection of the present invention.
Claims
1. A multivariable hierarchical adaptive filling control method based on reinforcement learning, characterized in that: include: Periodically collecting initial operating parameters of the filling system to construct a state vector, wherein the operating parameters include slurry concentration, slurry flow rate, delivery pump outlet pressure, water-cement ratio, return water rate, and strength estimation value; Inputting the state vector into a target reinforcement learning policy network model to obtain target action parameters and state value estimates, wherein the action parameters include slurry concentration, water-cement ratio, and pump speed; Determining whether the state value estimate meets the expected standard based on the state value estimate, and if so, adjusting the filling control parameters of the filling system based on the target action parameters, wherein the filling control parameters include cement content, clean water dosage, and delivery pump frequency; In response to completion of the filling control parameter adjustment of the filling system, constructing a state characteristic function corresponding to the filling system in the current control period, and constructing a multidimensional reward function based on each state characteristic function, wherein the state characteristic function includes a slurry concentration characteristic function, a water-cement ratio characteristic function, a pumping energy consumption characteristic function, and an intensity characteristic function; Obtaining the number of control cycles in the current control period, and calculating a current cumulative reward value corresponding to the current control period based on the multidimensional reward function; Whether to trigger a policy update is determined based on the number of control cycles and the current accumulated reward value; if triggered, the target reinforcement learning policy network model is updated based on the multidimensional reward function.
2. The multivariable hierarchical adaptive filling control method based on reinforcement learning according to claim 1 is characterized in that: Adjusting a filling control parameter of the filling system based on the target action parameter includes: Determining a concentration control deviation based on a comparison result of the target slurry concentration and the initial slurry concentration, and determining a cement content control amount and a water-cement ratio control amount based on the concentration control deviation; Determining a water-cement ratio control deviation based on a comparison result of the target water-cement ratio and the initial water-cement ratio, and determining a controlled amount of net water addition based on the water-cement ratio control deviation; A pump speed control deviation is determined based on a comparison result between the target pump speed and the initial pump speed, and a delivery pump frequency control amount is determined based on the pump speed control deviation.
3. The multivariable hierarchical adaptive filling control method based on reinforcement learning according to claim 2, characterized in that: Construct a multi-dimensional reward function, including: determining a concentration deviation penalty based on the slurry concentration characteristic function and a slurry concentration reference value; determining a water-cement ratio deviation penalty based on the water-cement ratio characteristic function and a water-cement ratio reference value; determining a pumping energy consumption penalty based on the pumping energy consumption characteristic function; determining a strength compliance index based on the strength characteristic function; A multidimensional reward function is constructed based on the concentration deviation penalty, the water-cement ratio deviation penalty, the pumping penalty, the strength compliance index, the concentration weight coefficient, the water-cement ratio weight coefficient, the energy consumption weight coefficient, and the strength weight coefficient.
4. The multivariable hierarchical adaptive filling control method based on reinforcement learning according to claim 3 is characterized in that: Determining whether the estimated state value meets the expected standards includes: determining an estimated comparison value based on a comparison result of the state value estimate and a preset state value estimate; The comparison result between the estimated comparison value and the preset comparison value determines whether it meets the expected standards.
5. The multivariable hierarchical adaptive filling control method based on reinforcement learning according to claim 4 is characterized in that: Determine whether to trigger a policy update, including: Based on the comparison result of the control cycle number and the preset cycle number and the comparison result of the current cumulative reward value and the preset cumulative reward value, it is determined whether to trigger a policy update.
6. The multivariable hierarchical adaptive filling control method based on reinforcement learning according to claim 5, characterized in that: Construct a state vector, including: Periodically collecting initial operating parameters of the filling system, and determining parameter change representation values corresponding to each operating parameter based on the initial operating parameters within a target time period; Determining whether to mark the target time period based on a comparison result of each parameter change representation value with a preset change representation value; If marked, the state input variables corresponding to each operating parameter are generated based on the initial operating parameters in the target time period; A state vector is constructed based on each of the state input variables.
7. The multivariable hierarchical adaptive filling control method based on reinforcement learning according to claim 6, characterized in that: Under the condition that it is determined that the strategy update is not triggered, the number of control cycles in the current control period is updated, and the initial operating parameters of the filling system are re-collected to update the state vector.
8. The multivariable hierarchical adaptive filling control method based on reinforcement learning according to claim 7, characterized in that: If the current cumulative reward value is greater than the preset cumulative reward value, the cumulative reward changes within several consecutive control cycles are calculated. Among them, if the number of control cycles is greater than the preset number of cycles, or the cumulative reward changes meet the preset changes, it is determined that the trigger strategy update is triggered.
9. The multivariable hierarchical adaptive filling control method based on reinforcement learning according to claim 8, characterized in that: Adjusting the filling control parameters of the filling system based on the target action parameters also includes: Based on the comparison results of the concentration control deviation and the concentration safety index, the comparison results of the target water-cement ratio and the water-cement ratio safety index, and the comparison results of the target pump speed and the pump speed safety index, it is determined whether the target freezing logic is triggered. If not triggered, the cement content control amount and the water-cement ratio control amount are determined based on the concentration control deviation, the net water addition amount control amount is determined based on the water-cement ratio control deviation, and the delivery pump frequency control amount is determined based on the pump speed control deviation.
10. The multivariable hierarchical adaptive filling control method based on reinforcement learning according to claim 9, characterized in that: Under the condition that it is determined that the target freezing logic is triggered, the filling control parameters of the filling system are adjusted based on the action parameters in the control cycle before the current control cycle.
Citation Information
Patent Citations
High-performance foamed mortar filling method for mining sites
CN102261262A
Method for optimizing filling material ratio
CN106746946A
Intelligent coal mining equipment cooperative control system and method
CN118226790A
Intelligent slurry discharge pressure control method for underground cemented filling
CN120061913A
Precise lossless identification method for real concentration of filling body
CN120195109A
Cited By
Fly ash composite material goaf closed filling parameter intelligent matching method
CN121350635A
Intelligent control method for tailing filling process based on data driving
CN121352527A
Iron tailing paste filling concentration-flow velocity double-closed-loop intelligent regulation and control method
CN122086184A
Layered adaptive filling control method based on strength evolution and temperature coupling
CN122131615A
A hierarchical adaptive filling control method based on intensity evolution and temperature coupling
CN122131615B