A layered self-adaptive filling control method based on intensity evolution and temperature coupling
By introducing a hierarchical adaptive filling control method with equivalent age and deep Q network, the control lag and imbalance caused by neglecting the nonlinearity of strength evolution and the influence of temperature field in mine filling are solved. This method achieves precise control and adaptive adjustment of the filling structure and improves the overall stability of the filling body.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANDONG GOLD MINING TECHNOLOGY CO LTD
- Filing Date
- 2026-05-06
- Publication Date
- 2026-06-02
AI Technical Summary
Existing mine backfill control methods have failed to effectively address the problems of uneven strength distribution, local segregation, and early collapse under complex geological conditions in deep mines. Traditional control methods neglect the nonlinearity of strength evolution, the influence of temperature field, and interlayer differences, resulting in control lag and imbalance, making it difficult to achieve the overall stability of the backfill structure.
A layered adaptive filling control method based on strength evolution and temperature coupling is adopted. Through discrete-time control and multi-step look-ahead optimization, combined with equivalent age and deep Q network, temperature is collected in real time and a composite state vector is constructed. Strength prediction and temperature correction factors are introduced into the optimization model to achieve precise control and adaptive adjustment of the layered strength of the filling body.
It has achieved an overall improvement in the stability of the filling structure, dynamically balanced the strength development relationship between the poured layer and the layer to be poured, overcome the lag and imbalance problems of traditional control methods, and improved the adaptability and structural stability of the filling body.
Smart Images

Figure CN122131615A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of mine control, specifically to a layered adaptive filling control method based on intensity evolution and temperature coupling. Background Technology
[0002] Mine backfilling, as a key means to achieve green mining and safe exploitation of deep resources, has seen its control methods gradually evolve from traditional manual experience-based batching and intermittent pouring to automated and intelligent control. Early backfilling processes had relatively simple control logic, mainly relying on fixed mix proportions and manual adjustments, with the control objective focused on maintaining a static match between material composition and mechanical properties. While this control method could meet basic backfilling requirements, it lacked the ability to respond to dynamic changes in the process when facing complex deep geological conditions. This often led to problems such as uneven strength distribution of the backfill, local segregation, and even early collapse, exposing the shortcomings of traditional control methods in terms of adaptability and precision.
[0003] With the introduction of sensing and data acquisition technologies, the control architecture of filling systems has been substantially improved. By deploying various types of sensors on-site, such as those for temperature, flow rate, and concentration, the system can acquire real-time data on the slurry state and key physical quantities during transport, feeding these quantities back to the control loop to achieve closed-loop regulation of parameters such as the cement-sand ratio and slurry concentration. Simultaneously, the development of numerical simulation and hydration kinetic models provides theoretical support for establishing mechanism-based predictive control methods. For filling operations in deep, high-temperature, high-pressure, and complex surrounding rock environments, the control objective is gradually shifting from stable control of a single parameter (such as concentration or flow rate) to systematic control involving multiple coupled variables. Existing research shows that the influence of the temperature field on the hydration reaction rate is a crucial factor determining the early strength formation of the filling body, while the interaction between the cement-sand ratio, water-cement ratio, and concentration also significantly affects the stability of the filling structure. Simply relying on feedback adjustments of apparent parameters will be insufficient to adapt to the dynamically changing operating conditions at depth.
[0004] Chinese invention patent application CN109732783A discloses an automated control system for tailings backfilling and its usage method. This system uses a central controller to operate the solenoid valves of each pipeline and dynamically adjusts the lime-sand ratio and slurry concentration using feedback information from flow meters and concentration meters, thereby ensuring the stability of slurry preparation and proportioning. The advantage of this system is improved real-time controllability of proportioning adjustments, reducing human error and material waste. However, its control strategy is still limited to immediate concentration and flow feedback, failing to address the nonlinear evolution of the internal strength of the backfill body over time, nor considering the influence of temperature distribution on the hydration rate. Therefore, in deep, high-temperature or significantly temperature-gradient mining environments, this scheme struggles to guarantee strength matching and structural equilibrium between different layers.
[0005] The paper "Effect of Curing Temperature under Deep Mining Conditions on the Mechanical Properties of Cemented Paste Backfill" by Wang, Y. et al. (Minerals 2023, 13, 383) found through experiments at different curing temperatures that when the curing temperature is below about 10°C, low temperature significantly inhibits the strength growth of the backfill; while in the medium-low temperature range, increasing the temperature helps to improve early strength, but exceeding a certain threshold may lead to strength degradation. This study reveals the significant impact of temperature on early strength, but the discussion mainly stays at the level of the overall effect of temperature and strength, without systematically introducing the temperature effect into the mix proportioning control system, nor has it formed a control path that can be corrected and predicted in real time. The paper "Hydration mechanism and mechanical-thermal correlation in cemented paste backfill" by C Zhang et al. (Journal of Building Engineering, 2024) explores the influence of ambient temperature on the hydration reaction process, pointing out that temperature changes can change the formation of hydration products, crystal structure evolution, and microstructure compactness in the cement-tailings system, thereby affecting the spatiotemporal distribution of mechanical properties. Although this work established the correlation between temperature and hydration-mechanical properties, it did not extend further to the control mechanism of layered strength matching of the filling body.
[0006] It is evident that existing technologies still have significant shortcomings in mine backfill control. Traditional backfill strength prediction methods mostly rely on empirical formulas or single-age measured data to establish regression models. They typically only consider the static relationship between material proportions and curing time, neglecting the dynamic coupling effects of multiple factors such as temperature, concentration, and lime-sand ratio. Model parameters are highly dependent on experimental conditions and lack adaptive updating capabilities. Existing temperature effect analyses are mostly limited to experimental observations under constant temperature curing conditions, failing to establish a mathematical mapping relationship between temperature field distribution and hydration rate, making it difficult to accurately reflect the differences in strength growth among different layers under deep temperature gradients. Furthermore, automatic proportioning control systems generally rely on closed-loop adjustments based on apparent parameters such as concentration and flow rate. The control objective is singular, lacking a feedback mechanism directly related to the final mechanical properties of the backfill body. The regulation is lagging and unable to adaptively respond to fluctuations in material properties and temperature changes. Therefore, in the process of mine backfilling, how to achieve precise and adaptive control of the layered strength of the backfill body to overcome the control lag and imbalance caused by the traditional method due to ignoring the nonlinearity of strength evolution, the influence of temperature field and the differences between layers, and thus improve the overall stability of the backfill structure, has become an urgent technical problem to be solved. Summary of the Invention
[0007] This invention proposes a layered adaptive filling control method based on strength evolution and temperature coupling. Its purpose is to achieve precise control and adaptive adjustment of the layered strength of the filling body, so as to overcome the control lag and imbalance caused by the traditional control method due to ignoring the nonlinear law of strength evolution, the influence of temperature field and the differences between layers, thereby improving the overall stability of the filling structure.
[0008] The technical solution of this invention is as follows:
[0009] A layered adaptive filling control method based on strength evolution and temperature coupling is proposed. The filling operation is carried out in layers. The pouring of the same filling layer requires multiple control cycles to complete. Each control cycle corresponds to a control step and a batching and pouring is performed. The number of control cycles required in each layer is predetermined, and the batching amount is the same in each control cycle.
[0010] The method includes the following steps:
[0011] Step S1: Initialize parameters;
[0012] A discrete-time control method is adopted, and hierarchical adaptive filling control is executed according to a preset fixed control cycle. Each control cycle is recorded as a control step. For any given control step... Each control step, let its corresponding pouring layer index be... Perform steps S2 to S4:
[0013] Step S2: Collect the current temperature of each filling layer, update the equivalent age based on the temperature, and construct a composite state vector containing the equivalent age vector, temperature deviation vector, and control input vector from the previous step.
[0014] Step S3: Construct and solve the optimization model based on the composite state vector to obtain the control input vector for this control step, and then execute the control input vector.
[0015] The objective function of the optimization model adopts a parallel weighted structure, which includes: a short-term tracking term to make the strength of each poured layer in the prediction time domain approach the target strength, a stationarity term to suppress control variable jumps, and a deep Q-network term to evaluate the long-term effect of the current control action. Among them, the short-term tracking term and the stationarity term dominate the optimization decision to ensure control safety when the deep Q-network is not fully trained. The deep Q-network term gradually strengthens its guiding role in the optimization decision as it is trained. The three terms work together to achieve a smooth transition from conservative control to adaptive control.
[0016] Step S4: Store the empirical data and update the deep Q network based on the empirical data.
[0017] As a further improvement to the hierarchical adaptive filling control method based on intensity evolution and temperature coupling, in step S2, the equivalent age is updated in the following way:
[0018] Assume that the temperature is collected by a pre-embedded temperature sensor. The current temperature of the filling layer is , Based on the current temperature Equivalent age in the previous step Update the equivalent age as follows:
[0019]
[0020] In the above formula, Indicates the first The filling layer in the first The equivalent age of each control step, which characterizes the cumulative coupling effect of temperature history on intensity evolution rate; Indicates the first The filling layer in the first The equivalent age of the first control step, if the first... The pouring object of the first control step is not the first Filling layer ; Indicates the control cycle; This represents the temperature correction factor, reflecting the accelerating or decelerating effect of the current temperature relative to the reference temperature on the hydration reaction rate. Indicates the first The initial value of the ratio sensitivity coefficient vector of the filling layer is determined by step S1; Indicates up to the Up to the [number] control steps The average proportion vector of all control input vectors actually executed by the layer, including average cement content, average water-cement ratio, and average paste concentration. If the first layer... The pouring object of the first control step is not the first Filling layer ; The reference ratio vector is determined by step S1;
[0021] for ,make , .
[0022] As a further improvement to the hierarchical adaptive filling control method based on intensity evolution and temperature coupling, a temperature correction factor is introduced. The calculation formula is:
[0023]
[0024] In the above formula, Indicates the activation energy of the hydration reaction; Represents the gas constant; Indicates the reference temperature.
[0025] As a further improvement to the hierarchical adaptive filling control method based on intensity evolution and temperature coupling, in step S2, the composite state vector is represented as:
[0026]
[0027] In the above formula, Indicates the first The composite state vector of each control step; Represents the equivalent age vector. ; Represents the temperature deviation vector. ,in Indicates the first The filling layer in the first Temperature deviation of each control step; This represents the control input vector obtained in the previous step.
[0028] As a further improvement to the hierarchical adaptive filling control method based on intensity evolution and temperature coupling, in step S3, the objective function of the optimization model is... as follows:
[0029]
[0030] In the above formula, Indicates the first The composite objective function value for each control step; This indicates the number of control steps corresponding to the prediction time domain; Indicates the prediction step index; Indicates the filling layer index; Indicates the first Importance weighting of the filling layer; Index of the currently poured layer; For the first The filling layer in the first The predicted strength of the step; Indicates the first The target strength of the filling layer is given by the engineering design; This represents the number of control steps corresponding to the control time domain, taken as... ; Indicates the first The control increment vector for each step is the decision variable. , The term to be determined is the first one. The control input vector of the step, ; The control increment weight matrix is a 3×3 positive definite diagonal matrix, whose diagonal elements correspond to the penalty weights for changes in cement content, water-cement ratio and paste concentration, respectively. These represent the weighting coefficients of the Q-function, used to adjust the Q-function output to a magnitude and scale that matches the short-term costs of the first two terms; Q-function Used to represent the output of a deep Q-network, indicating the state. Take control action Subsequently, the expected level of intensity compliance to be achieved within the future control time domain, This represents the trainable parameter vector of a deep Q-network;
[0031] Then solve the following constrained optimization problem:
[0032]
[0033] In the above formula, Indicates from the first Step to the first The optimal control sequence, composed of control increment vectors from each step, is used to optimize decision variables. Based on the control increment vectors of future control steps, any first-order control can be recursively derived. Step control input vector ;
[0034] After solving, based on the first control increment vector of the optimal control sequence... Get the first Control input vector of control step Then according to the control input vector Perform the pouring.
[0035] As a further improvement to the hierarchical adaptive filling control method based on intensity evolution and temperature coupling, in the prediction time domain, for , No. The filling layer in the first Intensity prediction of step Calculated using the following formula:
[0036]
[0037] In the above formula, Indicates the first The final expected strength of the filling layer; Indicates the first The strength development rate parameter of the filling layer, Indicates the first The strength development shape parameters of the filling layer; The first one obtained by recursion The filling layer in the first The equivalent age of the step.
[0038] As a further improvement to the hierarchical adaptive filling control method based on intensity evolution and temperature coupling, and Calculate using the following formula:
[0039]
[0040]
[0041] In the above formula, For filling layer index, , Indicates the lowest level. Indicates the top level; and These are the intensity development rate parameter and shape parameter at the lowest level, respectively; and This is the attenuation coefficient.
[0042] As a further improvement to the aforementioned hierarchical adaptive filling control method based on intensity evolution and temperature coupling, the first The filling layer in the first The equivalent age of the step The recursive method is as follows: Let the initial value be... ,in Calculated from step S2, for ,like Then the recursive formula is:
[0043]
[0044] In the above formula, Indicates the control cycle; The function for calculating the temperature correction factor; Indicates the first The first step of prediction The filling layer in the first Temperature control step; Indicates the first Vector of proportion sensitivity coefficients of filling layer; The term to be determined is the first one. The control input vector for each control step is determined by the decision variables during the optimization process. Recursive calculation yields that when season = ; The reference ratio vector is initialized in step S1;
[0045] like Then the recursive formula is:
[0046]
[0047] In the above formula, Indicates up to the Up to the [number] control steps The average ratio vector of all control input vectors actually executed by the layer.
[0048] As a further improvement to the hierarchical adaptive filling control method based on intensity evolution and temperature coupling, step S4 specifically includes:
[0049] Step S4.1: Calculate the expected return and store the experience;
[0050] Based on the current composite state vector Given the optimal control sequence, calculate the expected return using the following formula. :
[0051]
[0052] In the above formula, Indicates the first Each control step is based on the expected return predicted by the model; Indicates the first The first control step The filling layer in the first The intensity prediction value of step S3 is obtained recursively based on the optimal control sequence obtained in step S3; This represents the penalty coefficient for changes in the control quantity; Indicates the solved first... The control increment vector of the control step. ; The squared Euclidean norm of the control increment vector is represented by .
[0053] Allow the calculation of the next control step to begin, and wait for the next control step to calculate the composite state vector. Then, the empirical data Save to the experience replay area;
[0054] Step S4.2: Determine whether the time interval since the last update of the deep Q network has reached a preset value. If it has, update the deep Q network.
[0055] As a further improvement to the hierarchical adaptive filling control method based on intensity evolution and temperature coupling, the update method in step S4.2 is as follows:
[0056] Randomly sample batches of data from the experience playback area ,in For batch sample index, , For batch index sets, Indicates batch size; For the first The composite state vector of the current control step in each sample. For the first The composite state vector of the next control step in each sample For the first The control input vector for the current control step of each sample; Indicates the first Expected return for each sample; for each sample in the batch Perform CEM iteration: Initialize the candidate action set within the feasible region of the control variable; combine each candidate action with... They are combined separately and then input into the Q-network at the target depth. The output yields the Q-value used to evaluate each candidate action, and the action with the highest Q-value is selected. Candidate actions are set as an elite set; the sampling distribution is updated using the mean and covariance of the elite set; this process is repeated iteratively. Next, the final mean of the elite set is output as the approximate optimal action vector. ;
[0057] Deep Q network parameters Update by minimizing the following loss function:
[0058]
[0059] In the above formula, Represents the parameters of a deep Q-network The loss function value; Indicates the discount factor; Indicates the state The approximate optimal action vector obtained below; Indicates the target depth Q network in state Next, execute the approximate optimal action vector. Q value, The parameters of the target depth Q-network; Indicates the state of the main deep Q network. Next action Q value;
[0060] Target depth Q network parameters This remains unchanged in this update, and its value is slowly tracked through subsequent soft updates to reflect the parameters of the main depth Q network. :
[0061]
[0062] In the above formula, This represents the soft update coefficient.
[0063] Compared with the prior art, the present invention has the following beneficial effects:
[0064] 1. This invention introduces an equivalent age to characterize the cumulative coupling effect of temperature history on the hydration reaction rate, decoupling the hourly control cycle from the day-level intensity response timescale. This allows the control algorithm to perform long-term optimization on observable equivalent age states. Simultaneously, addressing the engineering constraint that the strength of deep infill bodies is difficult to measure directly in real time, this method utilizes temperature sensors to collect the temperature of each layer in real time. Combined with the recursively calculated equivalent age, a composite state vector is constructed, forming a soft measurement framework based on measurable variables. This solves the control bottleneck of not being able to detect strength online while ensuring engineering feasibility. It overcomes the control lag and imbalance problems caused by traditional control methods that ignore the nonlinear laws of intensity evolution and the influence of the temperature field. This achieves precise control and adaptive adjustment of the layered strength of the infill body, improving the overall stability of the infill structure.
[0065] 2. This invention embeds a strength evolution model and a temperature correction factor into the optimization objective function, and performs equivalent age extrapolation for uncast layers in the prediction time domain, enabling optimization decisions to predict the strength growth differences of different layers under temperature gradients. Based on this, the objective function sets importance weights and target strength constraints for each layer, and the cement content, water-cement ratio, and slurry concentration for the current control step are obtained by solving the optimization problem. This multi-step, forward-looking control strategy can dynamically balance the strength development relationship between cast and uncast layers during the casting process, avoiding the interlayer strength mismatch problem caused by relying solely on apparent parameter feedback in traditional mix proportion control.
[0066] 3. To address the large lag characteristics of the deep filling process, this method designs a deep Q-network learning mechanism based on model prediction of long-term returns. Specifically, within the model predictive control framework, the Q-function in the optimization objective is used to evaluate the impact of the current control action on the degree of strength achievement within multiple future control steps, solving the problem of credit allocation in lag systems. Simultaneously, since the control input is a continuous variable, this method employs the cross-entropy method to solve for the approximately optimal action in the continuous action space of the target deep Q-network during the experience playback update phase, and calculates the loss function accordingly to update the main network parameters. This design enables reinforcement learning to adapt to the slow dynamic characteristics and continuous control requirements of the filling process. As the number of control steps accumulates, the prediction accuracy of the Q-function gradually improves, and the adaptability of the control strategy also increases.
[0067] 4. This invention systematically improves the stability of infill structures through a three-pronged synergistic mechanism of "state reconstruction – short-term optimization – long-term evaluation": the equivalent age transforms the intangible intensity evolution into a calculable state variable, solving the state reconstruction problem of the intensity evolution process; model predictive control (the first two terms of the objective function) solves the ratio sequence that minimizes short-term intensity error and achieves stable control within a finite time domain based on this state, realizing multi-step look-ahead optimization of short-term intensity tracking and operational stability; the deep Q-network (the third term of the objective function) evaluates the long-term consequences of short-term actions based on historical experience and guides the optimization solution to shift towards a direction that takes into account long-term intensity achievement through the objective function, solving the long-term credit allocation problem under the large hysteresis characteristics of hydration reaction. These three elements are organically integrated through a unified objective function. Unlike existing technologies that often connect or replace model predictive control and reinforcement learning in series, this invention employs a parallel weighted structure with three terms in the objective function: the first two terms (sum of squared intensity tracking errors and control increment penalty term) ensure the safety and interpretability of the control strategy within the prediction time domain; the third term (deep Q-network output value) is directly added to the first two terms through weighting coefficients to evaluate the impact of the current control action on long-term intensity achievement beyond the prediction time domain. This structure allows the system to rely on the first two terms to output safe actions during the cold start phase when the deep Q-network estimation is inaccurate. With accumulated experience, the long-term guiding role of the deep Q-network gradually comes into play, achieving a smooth transition from conservative model predictive control to adaptive control that considers long-term rewards. Simultaneously, it avoids the problems of model predictive control errors accumulating in the deep Q-network in series structures and the high exploration risk of reinforcement learning in alternative structures. This invention utilizes the interpretability and safety of physical models, as well as the learning ability of data-driven models to handle long-term large lag characteristics, thereby fundamentally overcoming the control imbalance problem caused by the fragmentation of various links in traditional methods and significantly improving the overall stability of the filling structure.
[0068] 5. This invention further achieves two deep-level integration threads within the aforementioned collaborative framework. First, multi-source information fusion at the state construction level: In the equivalent age recursion, the temperature correction term and the proportioning sensitivity term are added, enabling the fusion of thermodynamic and material chemistry factors within the same state variable; the composite state vector includes the control input from the previous step, explicitly encoding the inertial memory of the control system into the state; the intensity prediction parameter decays exponentially with the stratigraphic level, encoding spatial location differences into the intensity evolution law. Second, long-term and short-term coupling at the optimization decision-making level: In Q-network updates, the cross-entropy method is used to solve for approximate optimal actions instead of actual actions, ensuring the network learns policy value rather than historical behavior value, guaranteeing consistency between long-term evaluation and short-term optimization. These designs enable this invention to form a deep coupling across the entire chain from state construction to decision updates. Attached Figure Description
[0069] Figure 1 This is a flowchart of a hierarchical adaptive filling control method based on the coupling of intensity evolution and temperature. Detailed Implementation
[0070] The technical solution of the present invention will now be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.
[0071] This embodiment provides a layered adaptive filling control method based on the coupling of strength evolution and temperature.
[0072] The filling operation is carried out in layers, with each layer having a different target strength value based on its depth and stress requirements. The pouring of the same filling layer typically requires multiple control cycles (i.e., multiple control steps) to complete, during which the index of that layer... This remains unchanged until the current layer is poured and the next layer begins to be poured. Only then does it increment to the next level of index. In this method, the filling operation is divided into... Layers, target intensity at each layer As specified in the engineering design, the bottom layer typically requires higher strength due to the greater overburden pressure. The number of control cycles required for each layer is predetermined, and the amount of material to be prepared is the same in each control cycle. The control system triggers calculations and decisions according to the control cycle, generating a set of control input vectors (including cement content, water-cement ratio, and grout concentration) to guide the filling material preparation and pouring operations within that time period.
[0073] This method is implemented during mine backfilling, and the controlled object is the proportioning parameters of the backfill grout, including the cement content. (Unit: kg / m³), Water-to-Cement Ratio (Dimensionless) and slurry concentration (Unit: mass percentage), the three constitute the control input vector. Known quantities include: the design target strength of each layer (unit: MPa, given by engineering design, the first...). The design target strength of the layer is denoted as ), basic material parameters (the first The final strength of the layer Activation energy of reaction (Etc., obtained through laboratory calibration). Real-time measurable parameters include: the temperature of each buried layer. (Acquired via pre-embedded thermocouples or fiber optic sensors, with signal lines led out along the filling retaining wall or using a wireless transmission module). Quantities that are difficult to measure directly in real time include: The first control step Actual strength of the layer (Data needs to be obtained through core drilling or ultrasonic testing, typically once a day or per shift, and cannot be obtained in real-time at the hourly level.) This method uses temperature as an intermediate variable to establish an indirect control loop, and transforms the historical cumulative effect of intensity evolution coupled with temperature into an observable state through the concept of equivalent age.
[0074] This method includes the following steps:
[0075] Step S1: Initialize parameters.
[0076] Perform this step only on the first run.
[0077] Specifically, step S1 includes:
[0078] (1) Initialize the equivalent age for each filling layer. , , Set the initial equivalent age to the total number of filling layers. Unit: days.
[0079] (2) Initialize model parameters.
[0080] Specifically, the ratio sensitivity coefficient vector (Dimensionless) is used to characterize the sensitivity of changes in cement content, water-cement ratio, and paste concentration to the equivalent age evolution rate. Among them, This represents the influence coefficient of cement admixture variation on the equivalent age evolution rate, taking a positive value, ranging from 0.1 to 0.3, for example... This indicates that for every 10 kg / m³ increase in cement content, the equivalent age evolution rate increases by 15%. This represents the influence coefficient of the water-cement ratio variation on the equivalent age evolution rate, taking a negative value, ranging from -0.2 to -0.05, for example... This indicates that for every 0.1 increase in the water-cement ratio, the equivalent age evolution rate decreases by 8%; The coefficient representing the influence of changes in slurry concentration on the equivalent age evolution rate is positive and ranges from 0.05 to 0.2. For example... This indicates that for every 2 percentage point increase in slurry concentration, the equivalent age evolution rate increases by 10%.
[0081] The reference matching vector also needs to be initialized. , These are reference values for cement content, water-cement ratio, and slurry concentration, respectively, which are determined by the engineering design.
[0082] (3) Initialize the deep Q-network. The parameters of the main deep Q-network are... Random initialization, target depth Q network parameters The experience replay area has been cleared.
[0083] (4) Initialize the temperature prediction model parameters.
[0084] Specifically, initializing the heat of cement hydration (Unit: J / kg) Hydration rate constant The parameters, such as the material's thermal conductivity and convective heat transfer coefficient, are taken from the material's test calibration values.
[0085] This method employs discrete-time control, executing hierarchical adaptive filling control according to a preset fixed control cycle. Each control cycle is denoted as a control step. For any given control step... Each control step, let its corresponding pouring layer index be... Proceed through steps S2 to S4.
[0086] Step S2: Collect the current temperature of each filling layer, update the equivalent age based on the temperature, and construct a composite state vector containing the equivalent age vector, temperature deviation vector, and control input vector from the previous step.
[0087] The method for updating the equivalent age is as follows:
[0088] Assume that the temperature is collected by a pre-embedded temperature sensor. The current temperature of the filling layer is (Unit: K) Based on the current temperature Equivalent age in the previous step Update the equivalent age as follows:
[0089]
[0090] In the above formula, Indicates the first The filling layer in the first The equivalent age (in days) of each control step, which characterizes the cumulative coupling effect of temperature history on the intensity evolution rate; Indicates the first The filling layer in the first The equivalent age (in days) of each control step is calculated from the previous step. If the first step... The pouring object of the first control step is not the first Filling layer ; Indicates the control period (unit: days, e.g., 2 hours corresponds to 0.0833 days); This represents the temperature correction factor (dimensionless), which reflects the accelerating or decelerating effect of the current temperature relative to the reference temperature on the hydration reaction rate. Indicates the first The initial value of the ratio sensitivity coefficient vector (dimensionless) of the filling layer is determined by step S1; Indicates up to the Up to the [number] control steps The average mix vector of all control input vectors actually executed by the layer, including average cement content (unit: kg / m³), average water-cement ratio (dimensionless), and average paste concentration (unit: mass percentage). The pouring object of the first control step is not the first Filling layer ; The reference ratio vector is determined by step S1.
[0091] for ,make , .
[0092] Among them, temperature correction factor The calculation formula is:
[0093]
[0094] In the above formula, The activation energy of the hydration reaction (unit: J / mol) is determined by laboratory differential scanning calorimetry, with a typical range of 30,000 to 50,000 J / mol. This represents the gas constant, with a value of 8.314 J / (mol·K); This indicates the reference temperature (unit: K), which is usually taken as 293.15 K (i.e. 20℃).
[0095] The composite state vector is represented as:
[0096]
[0097] In the above formula, Indicates the first The composite state vector of each control step; Represents the equivalent age vector. ; Represents the temperature deviation vector. ,in Indicates the first The filling layer in the first Temperature deviation per control step (unit: K); This represents the control input vector obtained in the previous step.
[0098] Step S3: Construct and solve an optimization model based on the composite state vector to obtain the control input vector for this control step, and then execute the control input vector. The optimization model is based on a deep Q-network, which is used to predict the future intensity achievement level.
[0099] Specifically, the objective function of the optimization model as follows:
[0100]
[0101] In the above formula, Indicates the first The composite objective function value for each control step is a scalar. This represents the number of control steps corresponding to the prediction time domain, and its value is such that... Coverage for 3 to 7 days; Indicates the prediction step index; Indicates the filling layer index; Indicates the first The importance weight of the filling layer (dimensionless) is set according to the safety level of the project, and the value ranges from 0.1 to 10; Index of the currently poured layer; For the first The filling layer in the first The predicted strength of the step; Indicates the first The target strength of the filling layer (unit: MPa) is given by the engineering design. This represents the number of control steps corresponding to the control time domain, typically taken as... ; Indicates the first The control increment vector for each step is the decision variable. , The term to be determined is the first one. The control input vector of the step, ; This represents the control increment weight matrix (dimensionless), a 3×3 positive definite diagonal matrix. Its diagonal elements correspond to the penalty weights for changes in cement content, water-cement ratio, and paste concentration, respectively. Larger weight values result in smoother changes in the corresponding control quantities. Typical values range from 0.01 to 1. The Q-function weighting coefficients (dimensionless, typically ranging from 0.1 to 10) are used to adjust the Q-function output to match the magnitude and scale of the first two short-run costs; the Q-function Used to represent the output of a deep Q-network, indicating the state. Take control action Subsequently, the expected level of intensity compliance to be achieved within the future control time domain, This represents the trainable parameter vector of a deep Q-network.
[0102] The deep Q-network in this embodiment adopts a three-layer fully connected neural network structure. The input layer receives a composite state vector consisting of the equivalent age of each layer, the temperature deviation of each layer, and the control parameters of the previous moment. The output layer directly provides the strength achievement level after executing a certain set of control actions in the current state. Before formal deployment, a large number of samples of states, actions, and strength achievement levels are generated using historical filling data or mechanistic models for pre-training the network. This allows the network to initially learn to judge the strength achievement level under different ratios and temperature conditions. After pre-training, the network can provide relatively reasonable expectations in the early stages of online control, reducing blind exploration during the cold start phase and accelerating system convergence.
[0103] Furthermore, within the prediction time domain, for , No. The filling layer in the first Intensity prediction of step (Unit: MPa) is calculated using the following formula:
[0104]
[0105] In the above formula, Indicates the first The final expected strength of the filling layer (in MPa) is given by the engineering design. Indicates the first Strength development rate parameter of the filling layer (dimensionless). Indicates the first Strength development shape parameters of the filling layer (dimensionless). The first one obtained by recursion The filling layer in the first The equivalent age of the step.
[0106] To reflect the changes in strength development patterns of infill layers at different depths due to differences in stress environment, and Calculate using the following formula:
[0107]
[0108]
[0109] In the above formula, For filling layer index, , Indicates the lowest level. Indicates the top level; and The bottom layer ( The strength development rate parameters and shape parameters of the sample were obtained by fitting the strength test data under laboratory standard mix conditions. and The attenuation coefficient (dimensionless) has a value range of 0.05 to 0.3 and is determined based on historical measured data or engineering experience. It is used to control the attenuation rate of the parameter as the layer rises.
[0110] No. The filling layer in the first The equivalent age of the step The recursive method is as follows: Let the initial value be... ,in Calculated from step S2, for ,like Then the recursive formula is:
[0111]
[0112] In the above formula, Indicates the control period (unit: days); This is the temperature correction factor function, and its calculation formula is the same as in step S2; Indicates the first The first step of prediction The filling layer in the first Temperature of each control step (unit: K); Indicates the first Vector of proportion sensitivity coefficients of filling layer; The term to be determined is the first one. The control input vector for each control step is determined by the decision variables during the optimization process. Recursive calculation yields that when season = ; This represents the reference ratio vector.
[0113] like Then the recursive formula is:
[0114]
[0115] In the above formula, Indicates up to the Up to the [number] control steps The average ratio vector of all control input vectors actually executed by the layer.
[0116] To obtain the future temperature in the prediction time domain A polynomial fitting extrapolation method based on historical temperature data is employed. Specifically, the data is collected... The most recent filling layer Historical measured temperature data for each control step ,in The historical data window length is given (dimensionless, typically 10 to 20). The following quadratic polynomial is fitted using the least squares method:
[0117]
[0118] In the above formula, Represents the time variable (unit: days), with the current control step. The corresponding time is Historical data correspondence It is a negative value; , , These are the fitting coefficients. After fitting, the future time... ( Substituting these values into the quadratic polynomial above, we can obtain the predicted temperature values for each future control step. .
[0119] If historical data is insufficient For example, in the initial stage of system startup, the isothermal assumption is adopted, i.e. Once enough historical data has been accumulated, it will automatically switch to the fitting extrapolation method.
[0120] The constraints of the optimization model include:
[0121] (1) Physical constraints of control quantities:
[0122]
[0123] In the above formula, This represents the lower limit vector of the control quantity. ,in This indicates the lower limit of cement content (unit: kg / m³). Indicates the lower limit of the water-cement ratio (dimensionless). This represents the lower limit of slurry concentration (unit: mass percentage), which is determined by equipment capacity and process specifications, and corresponds one-to-one with the elements of the control input vector; This represents the upper limit vector of the control quantity. The meanings of each component are similar to the lower limit, and are determined by equipment capacity and process specifications. The term to be determined is the first one. The control input vector for each step.
[0124] (2) Constraint on the rate of change of the control variable:
[0125]
[0126] In the above formula, This represents the vector of maximum change in a single step. ,in This indicates the maximum change in cement dosage in a single step (unit: kg / m³). This represents the maximum change in the water-cement ratio in a single step (dimensionless). This represents the maximum single-step change in slurry concentration (unit: mass percentage), determined by the equipment's response characteristics. This indicates taking the absolute value of the component.
[0127] Solve the following constrained optimization problem:
[0128]
[0129] In the above formula, Indicates from the first Step to the first The optimal control sequence, composed of control increment vectors from each step, is used to optimize decision variables. Based on the control increment vectors of future control steps, any first-order control can be recursively derived. Step control input vector .
[0130] Because the objective function includes a neural network The nonlinear output of this problem indicates that the optimization problem is nonconvex. In this embodiment, the Sequential Quadratic Programming (SQP) method is used to solve it.
[0131] After solving, based on the first control increment vector of the optimal control sequence... Get the first Control input vector of control step Then according to the control input vector To execute and control equipment such as cement feeders, water pumps, and mixing devices.
[0132] Step S4: Store the empirical data and update the deep Q network based on the empirical data.
[0133] Control quantity After execution, perform the relevant learning operations, specifically including:
[0134] Step S4.1: Calculate the expected return and store the experience. Based on the current composite state vector. Given the optimal control sequence, calculate the expected return using the following formula. :
[0135]
[0136] In the above formula, Indicates the first Each control step is based on the expected return predicted by the model, with negative values representing costs; Indicates the first The first control step The filling layer in the first The intensity prediction value (unit: MPa) of step S3 is obtained recursively based on the optimal control sequence obtained in step S3; This represents the penalty coefficient for changes in the control quantity (dimensionless), with a value range of 0.01 to 1; Indicates the solved first... The control increment vector of the control step. ; This represents the squared Euclidean norm of the control increment vector.
[0137] Allow the calculation of the next control step to begin, and wait for the next control step to calculate the composite state vector. Then, the empirical data Save to the experience replay area.
[0138] Step S4.2: Determine if the time interval since the last update of the deep Q-network has reached a preset value. If it has, update the deep Q-network by randomly sampling batches of data from the experience replay area. ,in For batch sample index, , For batch index sets, This indicates the batch size (dimensionless), typically ranging from 32 to 256. For the first The composite state vector of the current control step in each sample. For the first The composite state vector of the next control step in each sample For the first The control input vector for the current control step of each sample; Indicates the first Expected return for each sample.
[0139] Due to control input For three-dimensional continuous variables, in standard Q-learning The operations cannot be precisely computed in a continuous action space. This embodiment uses the cross-entropy method (CEM) to approximate the maximization problem. For each sample in the batch... Perform CEM iteration: Initialize the candidate action set within the feasible region of the control variable; combine each candidate action with... They are combined separately and then input into the Q-network at the target depth. The output yields the Q-value used to evaluate each candidate action, and the action with the highest Q-value is selected. Candidate actions are set as an elite set; the sampling distribution is updated using the mean and covariance of the elite set; this process is repeated iteratively. After 5 to 10 iterations, the final elite set mean is output as the approximate optimal action vector. In this embodiment, the candidate action set size is 500, and the elite ratio is... The covariance matrix is initially set as a diagonal matrix diag(0.1^2).
[0140] Deep Q network parameters Update by minimizing the following loss function:
[0141]
[0142] In the above formula, Represents the parameters of a deep Q-network The loss function value is a scalar; The discount factor (dimensionless, typically ranging from 0.9 to 0.99) defines the degree to which future rewards are discounted. Indicates the status via CEM The approximate optimal action vector obtained below; Indicates the target depth Q network in state Next, execute the approximate optimal action vector. Q value, The parameters of the target depth Q-network; Indicates the state of the main deep Q network. Next action The Q value.
[0143] Target depth Q network parameters This remains unchanged in this update, and its value is slowly tracked through subsequent soft updates to reflect the parameters of the main depth Q network. :
[0144]
[0145] In the above formula, This represents the soft update coefficient (dimensionless), typically ranging from 0.001 to 0.01.
[0146] This method also includes step S5, which is performed daily.
[0147] Step S5: Model parameter calibration. Obtain the actual strength measurements of each layer that has been poured. (Data obtained daily or per shift via core drilling or ultrasonic testing, unit: MPa) Model parameter calibration is then performed. It is the first The actual strength of the layer at the cumulative age from the time of pouring to the current correction.
[0148] The actual intensity measurement value is recursively calculated based on the optimal control sequence obtained in step S3, and the first... The current intensity prediction values of the layer are compared, and the difference is used as a correction signal, with the first layer being the corrected value. The deviation vector between the historical average allocation vector actually executed by the layer and the reference allocation vector is used as the regression value. The recursive least squares method with a forgetting factor is used to correct the allocation sensitivity coefficient. To compensate for model drift, whenever new measured intensity data is obtained, the algorithm calculates a correction step size based on the current correction signal and regression value, reducing the influence of historical data on the current estimate in an exponentially decaying manner (the forgetting factor typically ranges from 0.95 to 0.99), thereby gradually updating the model. The estimated values allow the model's predicted intensity to gradually approach the measured intensity.
[0149] The corrected parameters take effect immediately in the next control step. To avoid sudden changes in the control quantity due to abrupt parameter changes, a parameter smoothing transition mechanism (such as exponential moving average) can be used. If measured intensity data is missing for a certain period, skip that model calibration and correct it when the next data becomes available.
[0150] Through the above steps, this method solves the problem of time scale mismatch between control period and intensity response: by introducing the concept of equivalent age to represent the coupling relationship between intensity evolution and temperature, the hourly control period and the daily intensity response are decoupled, enabling the control algorithm to perform long-term optimization based on observable equivalent age states, avoiding control failure caused by time scale mismatch in traditional methods. Simultaneously, this method achieves indirect adaptive control based on measurable variables: for engineering constraints where the strength of deep filling bodies is difficult to measure directly in real time, a soft measurement framework based on temperature field (measurable in real time) and equivalent age (recursively calculated) is established. The measured intensity is only used for daily model calibration rather than real-time control, resulting in strong engineering feasibility. Furthermore, this method possesses policy learning capabilities with delay compensation: by designing a Q-function learning mechanism based on model prediction of long-term returns, the credit allocation problem in large-lag systems is solved; the cross-entropy method (CEM) solves the Q-function update problem in continuous action space, enabling reinforcement learning to adapt to the slow dynamic characteristics of the filling process and the requirements of continuous control. Finally, this method adopts a robust control architecture with feedforward-feedback separation: the feedforward part performs long-term trajectory optimization based on the physical model, and the feedback part compensates for uncertainties by correcting model parameters online. Combined with the temperature prediction model, it supports multi-step prediction, which synergistically improves the system's robustness to material fluctuations and environmental changes.
[0151] It should be noted that, in all mathematical formulas and calculations involved in this specific implementation method, although each variable has its own dimension in a physical sense (e.g., equivalent age unit is days, temperature unit is Kelvin, cement admixture unit is kg / m³, water-cement ratio is dimensionless, paste concentration unit is mass percentage, strength unit is MPa, etc.), in actual numerical calculations, standardized numerical forms under the International System of Units (SI) are used for calculations. Specifically, all physical quantities have been converted into pure numerical scalars based on basic units (e.g., seconds, kilograms, Kelvin, Pascals, etc.) or agreed multiples of derived units before being substituted into the formula. Mathematical operations such as addition, subtraction, multiplication, division, exponentiation, logarithm, and trigonometric functions in the formulas are all performed on these dimensionless numerical values and do not involve direct algebraic mixing operations between different units. For example, when calculating the exponential term in the temperature correction factor, the ratio of activation energy to the gas constant automatically offsets the dimensional effect of temperature, resulting in a dimensionless pure numerical value. Similarly, in the equivalent age recursive formula, although the control period Δt is expressed in days, it is uniformly converted to a numerical form corresponding to seconds or days during calculation, and directly multiplied with dimensionless quantities such as the temperature correction factor. Therefore, the calculation process is numerically self-consistent and programmable, without any mathematical or logical obstacles caused by inconsistent units.
[0152] The above description is merely a preferred embodiment of the present invention. The scope of the present invention is defined by the claims rather than the foregoing description.
Claims
1. A hierarchical adaptive filling control method based on intensity evolution and temperature coupling, characterized in that: The filling operation is carried out in layers. The pouring of the same filling layer requires multiple control cycles to complete. Each control cycle corresponds to a control step and a batching and pouring is carried out once. The number of control cycles required for each layer is determined in advance, and the batching amount is the same in each control cycle. The method includes the following steps: Step S1: Initialize parameters; A discrete-time control method is adopted, and hierarchical adaptive filling control is executed according to a preset fixed control cycle. Each control cycle is recorded as a control step. For any given control step... Each control step, let its corresponding pouring layer index be... Perform steps S2 to S4: Step S2: Collect the current temperature of each filling layer, update the equivalent age based on the temperature, and construct a composite state vector containing the equivalent age vector, temperature deviation vector, and control input vector from the previous step. Step S3: Construct and solve the optimization model based on the composite state vector to obtain the control input vector for this control step, and then execute the control input vector. The objective function of the optimization model adopts a parallel weighted structure, including: a short-term tracking term to make the strength of each poured layer in the prediction time domain approach the target strength, a stationarity term to suppress control variable jumps, and a deep Q-network term to evaluate the long-term effect of the current control action; among them, the short-term tracking term and the stationarity term dominate the optimization decision to ensure control safety when the deep Q-network is not fully trained, and the deep Q-network term gradually enhances its guiding role in the optimization decision as it accumulates with training. The three terms work together to achieve a smooth transition from conservative control to adaptive control. Step S4: Store the empirical data and update the deep Q network based on the empirical data.
2. The hierarchical adaptive filling control method based on intensity evolution and temperature coupling as described in claim 1, characterized in that, In step S2, the equivalent age is updated as follows: Assume that the temperature is collected by a pre-embedded temperature sensor. The current temperature of the filling layer is , Based on the current temperature Equivalent age in the previous step Update the equivalent age as follows: ; In the above formula, Indicates the first The filling layer in the first The equivalent age of each control step, which characterizes the cumulative coupling effect of temperature history on intensity evolution rate; Indicates the first The filling layer in the first The equivalent age of the first control step, if the first... The pouring object of the first control step is not the first Filling layer ; Indicates the control cycle; This represents the temperature correction factor, reflecting the accelerating or decelerating effect of the current temperature relative to the reference temperature on the hydration reaction rate. Indicates the first The initial value of the ratio sensitivity coefficient vector of the filling layer is determined by step S1; Indicates up to the Up to the [number] control steps The average proportion vector of all control input vectors actually executed by the layer, including average cement content, average water-cement ratio, and average paste concentration. If the first layer... The pouring object of the first control step is not the first Filling layer ; The reference ratio vector is determined by step S1; for ,make , .
3. The hierarchical adaptive filling control method based on intensity evolution and temperature coupling as described in claim 2, characterized in that, Temperature correction factor The calculation formula is: ; In the above formula, Indicates the activation energy of the hydration reaction; Represents the gas constant; Indicates the reference temperature.
4. The hierarchical adaptive filling control method based on intensity evolution and temperature coupling as described in claim 2, characterized in that, In step S2, the composite state vector is represented as: ; In the above formula, Indicates the first The composite state vector of each control step; Represents the equivalent age vector. ; Represents the temperature deviation vector. ,in Indicates the first The filling layer in the first Temperature deviation of each control step; This represents the control input vector obtained in the previous step.
5. The hierarchical adaptive filling control method based on intensity evolution and temperature coupling as described in claim 4, characterized in that, In step S3, the objective function of the optimization model is... as follows: ; In the above formula, Indicates the first The composite objective function value for each control step; This indicates the number of control steps corresponding to the prediction time domain; Indicates the prediction step index; Indicates the filling layer index; Indicates the first Importance weighting of the filling layer; Index of the currently poured layer; For the first The filling layer in the first The predicted intensity of the step; Indicates the first The target strength of the filling layer is given by the engineering design; This represents the number of control steps corresponding to the control time domain, taken as... ; Indicates the first The control increment vector for each step is the decision variable. , The term to be determined is the first one. The control input vector of the step, ; The control increment weight matrix is a 3×3 positive definite diagonal matrix, whose diagonal elements correspond to the penalty weights for changes in cement content, water-cement ratio and paste concentration, respectively. These represent the weighting coefficients of the Q-function, used to adjust the Q-function output to a magnitude and scale that matches the short-term costs of the first two terms; Q-function Used to represent the output of a deep Q-network, indicating the state. Take control action Subsequently, the expected level of intensity compliance to be achieved within the future control time domain, This represents the trainable parameter vector of a deep Q-network; Then solve the following constrained optimization problem: ; In the above formula, Indicates from the first Step to the first The optimal control sequence, composed of control increment vectors from each step, is used to optimize decision variables. Based on the control increment vectors of future control steps, any first-order control can be recursively derived. Step control input vector ; After solving, based on the first control increment vector of the optimal control sequence... Get the first Control input vector of control step Then according to the control input vector Perform the pouring.
6. The hierarchical adaptive filling control method based on intensity evolution and temperature coupling as described in claim 5, characterized in that, In the prediction time domain, for , No. The filling layer in the first Intensity prediction of step Calculated using the following formula: ; In the above formula, Indicates the first The final expected strength of the filling layer; Indicates the first The strength development rate parameter of the filling layer Indicates the first The strength development shape parameters of the filling layer; The first one obtained by recursion The filling layer in the first The equivalent age of the step.
7. The hierarchical adaptive filling control method based on intensity evolution and temperature coupling as described in claim 6, characterized in that, and Calculate using the following formula: ; ; In the above formula, For filling layer index, , Indicates the lowest level. Indicates the top level; and These are the intensity development rate parameter and shape parameter at the lowest level, respectively; and This is the attenuation coefficient.
8. The hierarchical adaptive filling control method based on intensity evolution and temperature coupling as described in claim 6, characterized in that, No. The filling layer in the first The equivalent age of the step The recursive method is as follows: Let the initial value be... ,in Calculated from step S2, for ,like Then the recursive formula is: ; In the above formula, Indicates the control cycle; This is a function for calculating the temperature correction factor. Indicates the first The first step of prediction The filling layer in the first Temperature control step; Indicates the first Vector of proportion sensitivity coefficients of filling layer; The term to be determined is the first one. The control input vector for each control step is determined by the decision variables during the optimization process. Recursive calculation yields that when season = ; The reference ratio vector is initialized in step S1; like Then the recursive formula is: ; In the above formula, Indicates up to the Up to the [number] control steps The average ratio vector of all control input vectors actually executed by the layer.
9. The hierarchical adaptive filling control method based on intensity evolution and temperature coupling as described in claim 8, characterized in that, Step S4 specifically includes: Step S4.1: Calculate the expected return and store the experience; Based on the current composite state vector Given the optimal control sequence, calculate the expected return using the following formula. : ; In the above formula, Indicates the first Each control step is based on the expected return predicted by the model; Indicates the first The first control step The filling layer in the first The intensity prediction value of step S3 is obtained recursively based on the optimal control sequence obtained in step S3; This represents the penalty coefficient for changes in the control quantity; Indicates the solved first... The control increment vector of the control step. ; The squared Euclidean norm of the control increment vector is represented by . Allow the calculation of the next control step to begin, and wait for the next control step to calculate the composite state vector. Then, the empirical data Save to the experience replay area; Step S4.2: Determine whether the time interval since the last update of the deep Q network has reached a preset value. If it has, update the deep Q network.
10. The hierarchical adaptive filling control method based on intensity evolution and temperature coupling as described in claim 9, characterized in that, The update method in step S4.2 is as follows: Randomly sample batches of data from the experience playback area ,in For batch sample index, , For batch index sets, Indicates batch size; For the first The composite state vector of the current control step in each sample. For the first The composite state vector of the next control step in each sample For the first The control input vector for the current control step of each sample; Indicates the first Expected return for each sample; for each sample in the batch Perform CEM iteration: Initialize the candidate action set within the feasible region of the control variable; combine each candidate action with... They are combined separately and then input into the Q-network at the target depth. The output yields the Q-value used to evaluate each candidate action, and the action with the highest Q-value is selected. Candidate actions are set as an elite set; the sampling distribution is updated using the mean and covariance of the elite set; this process is repeated iteratively. Next, the final mean of the elite set is output as the approximate optimal action vector. ; Deep Q network parameters Update by minimizing the following loss function: ; In the above formula, Represents the parameters of a deep Q-network The loss function value; Indicates the discount factor; Indicates the state The approximate optimal action vector obtained below; Indicates the target depth Q network in state Next, execute the approximate optimal action vector. Q value, The parameters of the target depth Q-network; Indicates the state of the main deep Q network. Next action Q value; Target depth Q network parameters This remains unchanged in this update, and its value is slowly tracked through subsequent soft updates to reflect the parameters of the main depth Q network. : ; In the above formula, This represents the soft update coefficient.