A high-temperature layered rock mass micro-parameter automatic calibration method based on a DDPG algorithm

CN122674491APending Publication Date: 2026-09-01NANHUA UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610773397.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-01
Publication Date
2026-09-01

AI Technical Summary

Technical Problem

[0005]本发明针对现有技术中存在的技术问题,提供一种基于DDPG算法的高温层状岩体微观参数自动标定方法,解决了现有技术中矿山高温层状岩体微观参数标定效率低、人工依赖性强以及难以精准表征复杂热损伤演化机制的问题

Benefits of technology

[0018] This invention provides an automatic calibration method, system, electronic device, and storage medium for microscopic parameters of high-temperature layered rock masses based on the DDPG algorithm. It utilizes macroscopic mechanical experimental data of high-temperature layered rock masses as the target benchmark, generates corresponding simulation data through a thermo-coupled numerical simulation model, and maps the deviation between the two to the state input of the DDPG model based on deep reinforcement learning. This DDPG model continuously outputs and executes microscopic parameter adjustment actions, driving the numerical simulation model to perform simulation calculations and obtain feedback rewards. This allows it to learn autonomously in a continuous action space, enabling the simulation results to dynamically approximate the macroscopic mechanical experimental data, and ultimately iteratively optimizing to obtain the optimal combination of microscopic parameters. This invention achieves full automation and intelligence in the microscopic parameter calibration process, avoiding the subjectivity and inefficiency of traditional manual trial and error. It can quickly adapt to different temperatures and bedding conditions, effectively improving the calibration accuracy and efficiency of rock mass numerical simulation parameters under complex conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122674491A_ABST
    Figure CN122674491A_ABST
Patent Text Reader

Abstract

This invention relates to the field of numerical simulation technology in geotechnical engineering, and provides an automatic calibration method for microscopic parameters of high-temperature layered rock masses based on the DDPG algorithm. This method collects macroscopic mechanical experimental data of high-temperature layered rock masses under different working conditions and generates corresponding simulation feature data using a thermo-coupling numerical simulation model. By performing feature alignment and sensitivity analysis on the experimental and simulation data, a set of key microscopic parameters that significantly affect the macroscopic response is identified. The experimental and simulation data are mapped to a deviation to construct a state space, and the key microscopic parameter set serves as the action space. A reward function is designed to construct a complete interactive environment for the deep deterministic policy gradient DDPG model. Through continuous interaction and iterative optimization between the DDPG model and the numerical simulation model, the optimal combination of microscopic parameters is automatically output to drive the simulation model to complete closed-loop calibration. This invention effectively improves the accuracy and efficiency of numerical simulation of high-temperature complex layered rock masses.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of numerical simulation of geotechnical engineering and artificial intelligence, and more specifically, to an automatic calibration method for microscopic parameters of high-temperature layered rock masses based on the DDPG (Depth Deterministic Strategy Gradient) algorithm. Background Technology

[0002] With the deepening of projects such as deep resource extraction, nuclear waste geological disposal, and geothermal energy development, deep rock masses are in high-temperature and complex stress environments. As a typical anisotropic rock, layered sandstone's mechanical properties are not only controlled by its bedding geometry but also severely affected by high-temperature-induced thermal damage.

[0003] The Linear Parallel Bond Model (LPM) in the Discrete Element Method (DEM) can simulate the fracture evolution mechanism of rocks from a microscopic perspective and has become an important tool for studying high-temperature rock mass damage. However, before applying the LPM model for numerical simulation, microscopic parameters such as particle stiffness, bond strength, and friction coefficient must be obtained by inversion from laboratory macroscopic test data. This process is called parameter calibration, and the accuracy of the calibration process directly determines the reliability of the numerical simulation results.

[0004] However, existing micro-parameter calibration techniques mainly rely on manual trial and error or heuristic algorithms such as genetic algorithms (GA) and particle swarm optimization (PSO) when dealing with high-temperature layered rock masses. Manual calibration is inefficient and subjective; while heuristic algorithms need to perform a complete search calculation again when facing new temperature conditions, and cannot achieve rapid, intelligent, and adaptive parameter output by learning from historical calibration experience. Micro-parameters usually fluctuate in a continuous space, and traditional algorithms are prone to getting trapped in local optima when dealing with multi-parameter, high-dimensional continuous action searches, making it difficult to achieve accurate fitting of rock mass behavior under extreme conditions. Summary of the Invention

[0005] This invention addresses the technical problems existing in the prior art by providing an automatic calibration method for microscopic parameters of high-temperature layered rock masses based on the DDPG algorithm. This method solves the problems of low efficiency, strong reliance on manual methods, and difficulty in accurately characterizing complex thermal damage evolution mechanisms in the calibration of microscopic parameters of high-temperature layered rock masses in mines.

[0006] According to a first aspect of the present invention, an automatic calibration method for microscopic parameters of high-temperature layered rock masses based on the DDPG algorithm is provided, comprising: S1. Collect macroscopic mechanical experimental data of high-temperature layered rock mass under different temperatures and bedding conditions, extract characteristic indicators to generate the first characteristic dataset; based on the initial micro parameters, drive the thermo-mechanical coupling numerical simulation model to perform calculations under the same conditions, extract model evolution characteristics to generate the second characteristic dataset. S2, perform feature alignment processing on the first feature dataset and the second feature dataset, and use sensitivity analysis to identify the set of key micro-parameters that have a significant impact on the macroscopic mechanical response from the micro-parameters of the numerical simulation model; take the first feature dataset as the target and perform deviation mapping with the second feature dataset, construct the state space with the obtained deviation information as the core, take the set of key micro-parameters as the action space, and design a reward function. S3, the constructed state space, action space and reward function are applied to the Deep Deterministic Policy Gradient (DDPG) model; through the interaction between the DDPG model and the numerical simulation model, micro-parameter adjustment actions are output, and the key micro-parameter set is iteratively optimized based on the feedback after the action is executed; S4, based on the optimal combination of output microscopic parameters, drive the numerical simulation model to update parameters to complete the closed-loop calibration.

[0007] Based on the above technical solution, the present invention can also be improved as follows.

[0008] Optionally, the macroscopic mechanical experimental data includes macroscopic stress-strain experimental curves and failure mode data.

[0009] Optionally, in step S2, the identification of a set of key micro-parameters that significantly influence the macroscopic mechanical response from the micro-parameters of the numerical simulation model using sensitivity analysis includes: S201. Based on the physical constitutive model of high-temperature layered rock masses, a set of candidate microscopic parameters to be analyzed is selected, and a physical value range is set for each parameter. The set of candidate microscopic parameters includes at least: grain / contact normal stiffness. Tangential stiffness Parallel bond cohesion C, parallel bond friction angle ; S202, orthogonal experimental design is used to extract multiple sets of micro-parameter combinations within the physical value range of the candidate micro-parameter set; S203, input the multiple sets of microscopic parameter combinations into the numerical simulation model for calculation, and record the macroscopic mechanical indicators of each set of simulation outputs; S204, the range values ​​of each microscopic parameter with respect to the macroscopic mechanical index are calculated using the range analysis method. The calculation formula is: in, For the first j The range of each parameter; the larger the R value, the more significant the influence of the micro parameter on the macro result. j Representing the jThere are three micro-parameters, i∈1,2,3…, representing different levels selected for each micro-parameter. Indicates the first j The average of all experimental results for each parameter at the i-th level; S205, based on the descending order of the range values, select the top n micro parameters whose cumulative contribution rate reaches the preset threshold to form the key micro parameter set.

[0010] Optionally, in step S2, the step of using the first feature dataset as the target and performing a deviation mapping with the second feature dataset to construct a state space based on the obtained deviation information includes: S206, Extract the macroscopic mechanical constants used as the target reference from the first feature dataset, including: target elastic modulus. Target peak intensity and target peak strain ; Simultaneously, from the second feature dataset at the current moment, the corresponding real-time simulation variables are extracted, including: real-time simulation elastic modulus. Real-time simulation of peak intensity Real-time simulation of peak strain and real-time microcrack ratio ; S207, Perform corresponding difference calculations between the extracted real-time simulation variables and the target benchmark to obtain the macroscopic mechanical response deviations: And microscopic damage evolution bias: S208, the macroscopic mechanical response deviation, microscopic damage evolution deviation, and current operating condition variables are concatenated in a preset order to obtain a multidimensional state space vector. : Where T is the target temperature in the current thermo-mechanical coupling simulation, and θ is the layering loading angle.

[0011] Optionally, the reward function The calculation formula is based on the dynamic change rate of macroscopic mechanical characteristic error and microscopic damage evolution deviation between adjacent training rounds. in, This represents the difference in macroscopic mechanical error between the current round and the previous round. This represents the difference between the current round and the previous round's micro-mechanism deviations. and These are the corresponding weighting coefficients.

[0012] Optionally, in step S3, the Deep Deterministic Policy Gradient (DDPG) model is obtained through construction and training: The DDPG model construction process includes: establishing a main Actor network as a policy network, a main Critic network as a value evaluation network, a target Actor network and a target Critic network with the same structure as the main network, and an experience replay pool; The DDPG model training process includes: A. Initialize the parameters of the main Actor network, the main Critic network, the target Actor network, and the target Critic network, and initialize the experience replay pool; B, the main Actor network according to the current state Output Action Add exploration noise The numerical simulation model is then input and executed to obtain rewards from environmental feedback. and the next state , will experience tuples ( , , , Store it in the experience replay pool; C. Sample batches of experience data from the experience replay pool after the data volume reaches a preset scale, and calculate the temporal difference target based on the sampled batches of experience data. The main Critic network is updated by minimizing the error of the value function; The main Actor network is updated using the policy gradient provided by the updated main Critic network through policy gradient ascent. D, with soft update rate The parameters of the main Actor network and the main Critic network are smoothly updated to the corresponding target Actor network and target Critic network; E. Repeat steps B to D until the model converges and the goodness of fit between the simulation results and the first feature dataset reaches the preset threshold, thus obtaining the trained DDPG model.

[0013] Optionally, in step C, the main Critic network is updated using the following loss function L: Where N is the number of samples, and i is the number of training rounds. The error between the estimated state-action value function at the current moment and the expected value of the target obtained by the target network is calculated using the following formula: in, As a discount factor, For immediate rewards, Q represents the main Critic network value function. For the target Critic network value function, The main Critic network evaluates the value of the current state-action pair. This is the current state. For the current action, The main Critic network parameters. For the purpose of the Critic network, the value assessment of the next state-action pair is performed. For the target Actor network, based on the next state The predicted next optimal action, The parameters are those of the target Critic network.

[0014] Optionally, in step C, the main Actor network is updated using the following strategy gradient: In the formula, J represents the expected return of the strategy. Indicates the desired operation. Represents a deterministic strategy. , and These are the network parameters of the objective policy network and the objective value function network, respectively; s represents the sample state. Represents a value function.

[0015] According to a second aspect of the present invention, an automatic calibration system for microscopic parameters of high-temperature layered rock masses based on the DDPG algorithm is provided. The system applies the above-mentioned automatic calibration method for microscopic parameters of high-temperature layered rock masses based on the DDPG algorithm and includes a parameter control terminal and a numerical simulation platform that are interconnected. The parameter control terminal includes: The data acquisition and processing module is used to collect macroscopic mechanical experimental data of high-temperature layered rock mass under different temperatures and bedding conditions to generate a first feature dataset, and drive the numerical simulation platform to perform calculations based on initial micro parameters to generate an initial second feature dataset. The reinforcement learning environment construction module is used to perform feature alignment processing on the first feature dataset and the second feature dataset, identify key micro-parameter sets using sensitivity analysis, and perform bias mapping to construct the state space and reward function. The DDPG agent module has a built-in deep deterministic policy gradient model, which is used to receive the state space vector and adjust the action according to the micro parameters output by the reward function. The parameter-driven and verification module is used to send the micro-parameter adjustment action to the numerical simulation platform and receive the optimization results to output the optimal combination of micro-parameters. The numerical simulation platform has a built-in thermo-coupled numerical simulation model, which is used to adjust its actions to update its own model parameters and perform simulation calculations based on the received micro parameters. The simulation feedback data generated by the simulation calculation is returned to the reinforcement learning environment construction module to update the state space.

[0016] According to a third aspect of the present invention, an electronic device is provided, including a memory and a processor, wherein the processor is configured to execute a computer management program stored in the memory to implement the steps of the above-described automatic calibration method for microscopic parameters of high-temperature layered rock masses based on the DDPG algorithm.

[0017] According to a fourth aspect of the present invention, a computer-readable storage medium is provided, on which a computer management program is stored, wherein when the computer management program is executed by a processor, the steps of the above-described automatic calibration method for microscopic parameters of high-temperature layered rock masses based on the DDPG algorithm are implemented.

[0018] This invention provides an automatic calibration method, system, electronic device, and storage medium for microscopic parameters of high-temperature layered rock masses based on the DDPG algorithm. It utilizes macroscopic mechanical experimental data of high-temperature layered rock masses as the target benchmark, generates corresponding simulation data through a thermo-coupled numerical simulation model, and maps the deviation between the two to the state input of the DDPG model based on deep reinforcement learning. This DDPG model continuously outputs and executes microscopic parameter adjustment actions, driving the numerical simulation model to perform simulation calculations and obtain feedback rewards. This allows it to learn autonomously in a continuous action space, enabling the simulation results to dynamically approximate the macroscopic mechanical experimental data, and ultimately iteratively optimizing to obtain the optimal combination of microscopic parameters. This invention achieves full automation and intelligence in the microscopic parameter calibration process, avoiding the subjectivity and inefficiency of traditional manual trial and error. It can quickly adapt to different temperatures and bedding conditions, effectively improving the calibration accuracy and efficiency of rock mass numerical simulation parameters under complex conditions. Attached Figure Description

[0019] Figure 1 A flowchart of an automatic calibration method for microscopic parameters of high-temperature layered rock masses based on the DDPG algorithm is provided for this invention. Figure 2 A schematic diagram illustrating the principle of an automatic calibration method for microscopic parameters of high-temperature layered rock masses based on the DDPG algorithm, provided for one embodiment. Figure 3A schematic diagram of training a DDPG model based on deep reinforcement learning, provided for one embodiment; Figure 4 A functional module block diagram of an automatic calibration system for microscopic parameters of high-temperature layered rock masses based on the DDPG algorithm provided by this invention; Figure 5 A schematic diagram of the hardware structure of a possible electronic device provided by the present invention; Figure 6 This is a schematic diagram of the hardware structure of a possible computer-readable storage medium provided by the present invention. Detailed Implementation

[0020] The specific embodiments of the present invention will be described in further detail below with reference to the accompanying drawings and examples. The following examples are for illustrative purposes only and are not intended to limit the scope of the invention.

[0021] Example 1: Figure 1 The present invention provides a flowchart of an automatic calibration method for microscopic parameters of high-temperature layered rock masses based on the DDPG algorithm. Figure 2 This is a schematic diagram illustrating the principle of the automatic calibration method for microscopic parameters of high-temperature layered rock masses based on the DDPG algorithm provided in this embodiment.

[0022] Combination Figure 1 and Figure 2 As shown, the automatic calibration method for microscopic parameters of high-temperature layered rock masses based on the DDPG algorithm provided in this embodiment of the invention includes steps S1 to S4.

[0023] S1, Data Acquisition and Feature Extraction Steps This step first collects macroscopic mechanical experimental data of high-temperature layered rock mass under different experimental conditions (such as different temperatures and bedding conditions). The macroscopic mechanical experimental data includes macroscopic stress-strain experimental curves and failure mode data. Then, characteristic indicators such as elastic modulus, peak strength, and peak strain are extracted from the macroscopic mechanical experimental data to generate the first characteristic dataset.

[0024] Meanwhile, based on a set of pre-set initial microscopic parameters, the thermo-mechanical coupling numerical simulation model is driven to perform simulation calculations under the same working conditions as the macroscopic experimental conditions. The corresponding model evolution features are extracted from the simulation results, such as the simulated elastic modulus, simulated peak strength, simulated peak strain and microcrack ratio, to generate a second feature dataset.

[0025] S2, Key Parameter Identification and Interactive Environment Construction Steps This step first performs feature alignment preprocessing on the first feature dataset and the second feature dataset to ensure data comparability.

[0026] Next, sensitivity analysis was used to identify the set of key micro-parameters that have a significant impact on the macroscopic mechanical response from all the micro-parameters of the numerical simulation model. Specifically, the set of parameters was selected by orthogonal experimental design, simulation was run, and the influence weight of each parameter was calculated using range analysis.

[0027] In one possible implementation, sensitivity analysis is used to identify a set of key micro-parameters that have a significant impact on the macroscopic mechanical response from the micro-parameters of the numerical simulation model, specifically including sub-steps S201~S205: S201. Based on the physical constitutive model of high-temperature layered rock masses, a set of candidate microscopic parameters to be analyzed is selected, and a physical value range is set for each candidate microscopic parameter. The set of candidate microscopic parameters includes at least: particle / contact normal stiffness. Tangential stiffness Parallel bond cohesion C, parallel bond friction angle ; S202, orthogonal experimental design is used to extract M representative combinations of micro-parameters within the physical value range of the candidate micro-parameter set; compared with blind sampling, this method can cover the entire parameter space with fewer simulations, thereby improving computational efficiency; S203, input the above M sets of microscopic parameter combinations into the thermo-mechanical coupling numerical simulation model for simulation calculation, and record the macroscopic mechanical indicators of each set of simulation outputs; S204 uses range analysis to calculate the influence weights of each microscopic parameter on the macroscopic response index. Taking range analysis as an example, the influence weights here are the range values ​​of each microscopic parameter on the macroscopic mechanical index. The calculation formula is: in, For the first j The range of a micro-parameter, The larger the value, the more significant the impact of the microscopic parameter on the macroscopic result; j Representing the j There are three micro-parameters, i ∈ 1, 2, 3, ..., where i represents the different levels selected for the micro-parameter. This represents the average of all experimental results for the j-th parameter at the i-th level; S205, based on the calculated range or contribution rate, sort them in descending order, and select the top n micro-parameters whose cumulative contribution rate reaches a preset threshold to form the key micro-parameter set. This key micro-parameter set will serve as the action space of the Actor network in the subsequent DDPG algorithm, thereby achieving dimensionality reduction optimization and significantly improving the convergence speed and calibration accuracy of the algorithm.

[0028] After determining the key set of microscopic parameters, combined with Figure 2 As shown, the steps for constructing an interactive environment for a reinforcement learning agent are as follows. Specifically, the first feature dataset is used as the target constraint value, and it is subjected to nonlinear difference operations and deviation mapping with the second feature dataset output by real-time simulation. The calculated macroscopic mechanical response deviation and microscopic damage evolution deviation are combined with the current temperature and stratification angle to construct a state space vector reflecting the thermal coupling damage deviation, thereby achieving a deep correlation between experimental data and the simulation environment in the spatiotemporal dimensions. On this basis, the action space is defined with a set of key microscopic parameters, and a reward function based on the error change rate between adjacent training rounds is designed, thus completing the construction of the DDPG interactive environment for the reinforcement learning agent.

[0029] In one possible implementation, the first feature dataset is used as the target, and a deviation mapping is performed between it and the second feature dataset. The resulting deviation information is used as the core to construct the state space, specifically including sub-steps S206~S208: S206, Extract the macroscopic mechanical constants used as the target reference from the first feature dataset, including: target elastic modulus. Target peak intensity and target peak strain ; Simultaneously, from the second feature dataset at the current moment, the corresponding real-time simulation variables are extracted, including: real-time simulation elastic modulus. Real-time simulation of peak intensity Real-time simulation of peak strain and real-time microcrack ratio ; S207, Perform corresponding difference calculations between the extracted real-time simulation variables and the target benchmark to obtain the macroscopic mechanical response deviation, including the elastic modulus deviation. Peak intensity deviation Peak strain deviation : And microscopic damage evolution bias: S208, the macroscopic mechanical response deviation, microscopic damage evolution deviation, and current operating condition variables are concatenated in a preset order to obtain a multidimensional state space vector. : Where T is the target temperature in the current thermo-mechanical coupling simulation, and θ is the layering loading angle.

[0030] This embodiment constructs a state-space vector that includes multidimensional physical characteristics and external environmental constraints. This can provide sufficient optimization basis for DDPG agents, avoiding getting trapped in local optimal solutions that violate the physical mechanism of rock fracture by simply fitting curves.

[0031] In one possible implementation, the action space is defined using a set of key microscopic parameters, in order to This represents the microscopic parameter adjustment action at time t, which is a set of initial values ​​of microscopic constitutive parameters within a physically reasonable range (such as particle / contact normal stiffness). Tangential stiffness Parallel bond cohesion C, parallel bond friction angle (and coefficient of thermal expansion, etc.).

[0032] In one possible implementation, a reward function is designed. reward function The calculation formula is based on the dynamic change rate of macroscopic mechanical characteristic error and microscopic damage evolution deviation between adjacent training rounds. in, This represents the difference in macroscopic mechanical error between the current round and the previous round. This represents the difference between the current round and the previous round's micro-mechanism deviations. and These are the corresponding weighting coefficients.

[0033] S3, DDPG model-driven iterative optimization steps This step applies the state space, action space, and reward function constructed in step S2 to a Deep Deterministic Policy Gradient (DDPG) model, which is obtained through construction and training.

[0034] refer to Figure 2As shown, the DDPG model includes a main Actor network as the policy network, a main Critic network as the value evaluation network, and target networks (including target Actor networks and target Critic networks) and an experience replay pool corresponding to the main network. Through the interaction between the DDPG model and the numerical simulation model, the Actor network of the DDPG model adjusts its actions based on the micro-parameters output by the current state. These actions drive the numerical simulation model to update its parameters and perform simulation calculations, while the environment provides immediate rewards and new states. The interaction experience between the DDPG model and the numerical simulation model is stored in the experience replay pool and can be used to periodically update the network parameters of the DDPG model. This interaction and learning process is repeated cyclically, thereby achieving iterative optimization of the key micro-parameter set.

[0035] In one possible implementation, refer to Figure 3 As shown, the training process of the Deep Deterministic Policy Gradient (DDPG) model based on deep reinforcement learning is as follows: Steps A through E: Step A: Initialize the parameters of the main Actor network, the main Critic network, the target Actor network, and the target Critic network, and initialize the experience replay pool.

[0036] That is, initialize the state tuple ( ),in, The initial value represents the deviation of the rock mass mechanical state (also known as the state space vector). Specifically, this state space vector is composed of macroscopic mechanical response deviation, microscopic damage evolution deviation, and current working condition environmental variables, and can be expressed as: Here, The initial simulation shows the deviation in the elastic modulus. For peak intensity deviation, For peak strain deviation, This is due to deviations in micro-evolutionary characteristics. The target temperature for the current thermo-coupling simulation is... The current loading angle is defined as follows. By constructing a state vector that includes multi-dimensional physical characteristics and external environmental constraints, the DDPG agent can be provided with sufficient optimization basis, avoiding getting trapped in a local optimum solution that violates the physical mechanism of rock fracture during the parameter optimization process by simply fitting the curve.

[0037] In the initial state tuple, The initial feature vector represents the micro-parameter adjustment action, specifically referring to a set of preset initial values ​​of micro-constitutive parameters (such as contact stiffness, cohesion, friction angle, and coefficient of thermal expansion) within a physically reasonable range. This initial action vector serves as the benchmark starting point for the iterative optimization of the deep reinforcement learning network, used to activate the numerical simulation model and obtain initial mechanical response feedback. This represents the initial value of the reward function. This represents the initial value of the deviation in the rock mass mechanical state at the next moment. In this embodiment, the maximum number of training rounds is designed to be 3000, the maximum step size per round is 400, the experience pool capacity is set to 40000, and the learning rate is soft-updated. The cumulative discount factor is 0.01. Set to 0.9.

[0038] Next, the second feature dataset (containing the stress-strain response and microcrack data from the initial simulation) obtained after preprocessing by the thermo-coupling numerical simulation model is placed into the experience playback pool for data verification preprocessing to form a memory sequence.

[0039] Then, a preset number of experience data are randomly extracted from the experience replay pool as training samples.

[0040] Step B, the main Actor network according to the current state Output Action Add exploration noise The numerical simulation model is then input and executed to obtain rewards from environmental feedback. and the next state , will experience tuples ( , , , Store it in the experience replay pool.

[0041] In practice, the main Actor network acquires the current rock mass state information within the main network. And based on the current deterministic strategy Select micro parameters to adjust actions Execute current rock mass condition information and add noise. To increase the exploration rate, it is represented as: Execute the current action Then, the numerical simulation model performs thermo-coupling calculations in its simulation environment and outputs the reward for the current round. and the next rock mass condition information Reward function The formula is calculated based on the dynamic change rate of the macroscopic mechanical characteristic error and the microscopic damage evolution deviation in the current round, as follows: in, This represents the difference in macroscopic mechanical error between the current round and the previous round. This represents the difference between the current round and the previous round's micro-mechanism deviations. and These are the corresponding weighting coefficients.

[0042] This process is repeated, and the resulting data ( They are stored as a set of experience tuples in the experience replay pool until the experience replay pool is full.

[0043] Step C: When the experience replay pool is full, sample a batch of experience data from the experience replay pool, for example, randomly sample N samples from the experience replay pool. () is used as the current network training data.

[0044] The main Critic network is updated by minimizing the loss function L of the value function error using the following formula: In the formula: N is the number of samples, and i is the number of training rounds. The error between the estimated state-action value function at the current moment and the expected value of the target obtained by the target network is calculated using the following formula: in, As a discount factor, For immediate rewards, Q represents the main Critic network value function. For the target Critic network value function, The main Critic network evaluates the value of the current state-action pair. This is the current state. For the current action, The main Critic network parameters. For the purpose of the Critic network, the value assessment of the next state-action pair is performed. For the target Actor network, based on the next state The predicted next optimal action, The parameters are those of the target Critic network.

[0045] Next, the current main actor network is updated using the policy gradient provided by the updated main Critic network through policy gradient ascent. The policy gradient is shown in the following formula: In the formula, J represents the expected return of the strategy. Indicates the desired operation. Represents a deterministic strategy. , and These are the network parameters of the objective policy network and the objective value function network, respectively; s represents the sample state. Represents a value function.

[0046] Step D: Assign the main network parameters to the target network. The main network parameters are continuously updated while the target network remains unchanged. After a preset time (e.g., a preset period of time), the main network parameters are assigned to the target network again, and the target network performs the corresponding soft update.

[0047] In practice, the soft update rate is used. The parameters of the main Actor network and the main Critic network are smoothly updated to the corresponding target Actor network and target Critic network, as follows: Step E: Repeat steps B to D above until the loss function of the DDPG model converges and the goodness of fit between the simulation results and the first feature dataset reaches the preset threshold, thus obtaining the trained DDPG model.

[0048] The trained DDPG network can output the optimal combination of micro parameters in real time based on the input rock mass conditions.

[0049] S4, Closed-loop calibration steps This step drives the numerical simulation model to update its parameters based on the optimal combination of microscopic parameters output, thereby completing the closed-loop calibration.

[0050] After training, the DDPG network sends micro-parameter configuration instructions containing the optimal combination of micro-parameters to the numerical simulation model based on the real-time input rock mass conditions. After receiving the instructions, the numerical simulation model automatically updates the micro-coefficients in the thermo-mechanical coupling numerical model according to the optimal combination of micro-parameters, thereby realizing the automatic calibration of the mechanical response of high-temperature layered rock mass.

[0051] Example 2: Figure 4 A structural diagram of an automatic calibration system for microscopic parameters of high-temperature layered rock masses based on the DDPG algorithm, provided in an embodiment of the present invention, is shown below. Figure 4As shown, an automatic calibration system for microscopic parameters of high-temperature layered rock masses based on the DDPG algorithm is described. The system applies the automatic calibration method for microscopic parameters of high-temperature layered rock masses based on the DDPG algorithm provided in Example 1. The system includes a parameter control terminal and a numerical simulation platform that are interconnected.

[0052] The parameter control terminal includes: The data acquisition and processing module is used to collect macroscopic mechanical experimental data of high-temperature layered rock mass under different temperatures and bedding conditions to generate a first feature dataset, and drive the numerical simulation platform to perform calculations based on initial micro parameters to generate an initial second feature dataset. The reinforcement learning environment construction module is used to perform feature alignment processing on the first feature dataset and the second feature dataset, identify key micro-parameter sets using sensitivity analysis, and perform bias mapping to construct the state space and reward function. The DDPG agent module has a built-in deep deterministic policy gradient model, which is used to receive the state space vector and adjust the action according to the micro parameters output by the reward function. The parameter-driven and verification module is used to send the micro-parameter adjustment action to the numerical simulation platform and receive the optimization results to output the optimal combination of micro-parameters. The numerical simulation platform has a built-in thermo-coupled numerical simulation model, which is used to adjust its actions to update its own model parameters and perform simulation calculations based on the received micro parameters. The simulation feedback data generated by the simulation calculation is returned to the reinforcement learning environment construction module to update the state space.

[0053] It is understood that the automatic calibration system for microscopic parameters of high-temperature layered rock masses based on the DDPG algorithm provided by this invention corresponds to the automatic calibration method for microscopic parameters of high-temperature layered rock masses based on the DDPG algorithm provided in the foregoing embodiments. The relevant technical features of the automatic calibration system for microscopic parameters of high-temperature layered rock masses based on the DDPG algorithm can be referred to the relevant technical features of the automatic calibration method for microscopic parameters of high-temperature layered rock masses based on the DDPG algorithm, and will not be repeated here.

[0054] Please see Figure 5 , Figure 5 This is a schematic diagram illustrating an embodiment of the electronic device provided in this invention. For example... Figure 5 As shown, this embodiment of the invention provides an electronic device 500, including a memory 510, a processor 520, and a computer program 511 stored in the memory 510 and executable on the processor 520. When the processor 520 executes the computer program 511, it performs the following steps: S1. Collect macroscopic mechanical experimental data of high-temperature layered rock mass under different temperatures and bedding conditions, extract characteristic indicators to generate the first characteristic dataset; based on the initial micro parameters, drive the thermo-mechanical coupling numerical simulation model to perform calculations under the same conditions, extract model evolution characteristics to generate the second characteristic dataset. S2, perform feature alignment processing on the first feature dataset and the second feature dataset, and use sensitivity analysis to identify the set of key micro-parameters that have a significant impact on the macroscopic mechanical response from the micro-parameters of the numerical simulation model; take the first feature dataset as the target and perform deviation mapping with the second feature dataset, construct the state space with the obtained deviation information as the core, take the set of key micro-parameters as the action space, and design a reward function. S3, the constructed state space, action space and reward function are applied to the Deep Deterministic Policy Gradient (DDPG) model; through the interaction between the DDPG model and the numerical simulation model, micro-parameter adjustment actions are output, and the key micro-parameter set is iteratively optimized based on the feedback after the action is executed; S4, based on the optimal combination of output microscopic parameters, drive the numerical simulation model to update parameters to complete the closed-loop calibration.

[0055] Please see Figure 6 , Figure 6 This is a schematic diagram illustrating an embodiment of a computer-readable storage medium provided by the present invention. (See diagram below.) Figure 6 As shown, this embodiment provides a computer-readable storage medium 600, on which a computer program 511 is stored. When the computer program 511 is executed by a processor, it performs the following steps: S1. Collect macroscopic mechanical experimental data of high-temperature layered rock mass under different temperatures and bedding conditions, extract characteristic indicators to generate the first characteristic dataset; based on the initial micro parameters, drive the thermo-mechanical coupling numerical simulation model to perform calculations under the same conditions, extract model evolution characteristics to generate the second characteristic dataset. S2, perform feature alignment processing on the first feature dataset and the second feature dataset, and use sensitivity analysis to identify the set of key micro-parameters that have a significant impact on the macroscopic mechanical response from the micro-parameters of the numerical simulation model; take the first feature dataset as the target and perform deviation mapping with the second feature dataset, construct the state space with the obtained deviation information as the core, take the set of key micro-parameters as the action space, and design a reward function. S3, the constructed state space, action space and reward function are applied to the Deep Deterministic Policy Gradient (DDPG) model; through the interaction between the DDPG model and the numerical simulation model, micro-parameter adjustment actions are output, and the key micro-parameter set is iteratively optimized based on the feedback after the action is executed; S4, based on the optimal combination of output microscopic parameters, drive the numerical simulation model to update parameters to complete the closed-loop calibration.

[0056] This invention provides an automatic calibration method, system, and storage medium for microscopic parameters of high-temperature layered rock masses based on the DDPG algorithm. By employing the deep reinforcement learning DDPG algorithm to determine the optimal microscopic parameters of a numerical simulation model, it can dynamically fit various mechanical responses (such as stress-strain curves and crack evolution characteristics) of high-temperature layered rock masses under complex thermo-mechanical coupling conditions in real time. This results in more accurate microscopic parameter calibration, improving the quality of numerical simulation and ensuring high-reliability evaluation of the stability of deep engineering projects (such as nuclear waste disposal sites). Furthermore, it can effectively suppress unsteady-state fluctuations during the evolution of the numerical simulation model, improving the stability and reliability of the simulation calculation. This invention integrates deep learning networks and deterministic policy gradient theory. The fusion algorithm combines the advantages of deep feature perception and reinforcement learning decision optimization. Through the synergy of the Actor-Critic network framework, it can more effectively find the global optimal solution in the continuous action space. For high-temperature layered rock masses, the DDPG algorithm exhibits better performance than other traditional heuristic optimization algorithms in solving nonlinear, anisotropic constitutive fitting problems.

[0057] This invention leverages both the robustness of deep networks in representing mechanical characteristics and the efficiency of policy gradient algorithms in optimizing continuous parameters. The DDPG algorithm exhibits good control capabilities for strongly nonlinear systems, low sensitivity to initial parameter settings, stability, reliability, and low computational cost. It can automatically select the optimal combination of micro-parameters for different temperatures and stratification conditions, significantly reducing the time consumed by subjective operational errors and repeated trial and error in manual calibration, thus demonstrating extremely high solution efficiency and practical engineering value.

[0058] It should be noted that the descriptions of each embodiment in the above embodiments have different focuses. For parts that are not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0059] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0060] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0061] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0062] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable apparatus for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0063] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.

[0064] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. An automatic calibration method for microscopic parameters of high-temperature layered rock masses based on the DDPG algorithm, characterized in that, include: S1: Collect macroscopic mechanical experimental data of high-temperature layered rock mass under different temperatures and bedding conditions, extract feature indicators to generate the first feature dataset; Based on the initial microscopic parameters, the thermo-coupling numerical simulation model is driven to perform calculations under the same working conditions, and the model evolution features are extracted to generate a second feature dataset. S2, perform feature alignment processing on the first feature dataset and the second feature dataset, and use sensitivity analysis to identify the set of key micro-parameters that have a significant impact on the macroscopic mechanical response from the micro-parameters of the numerical simulation model; take the first feature dataset as the target and perform deviation mapping with the second feature dataset, construct the state space with the obtained deviation information as the core, take the set of key micro-parameters as the action space, and design a reward function. S3, the constructed state space, action space and reward function are applied to the Deep Deterministic Policy Gradient (DDPG) model; through the interaction between the DDPG model and the numerical simulation model, micro-parameter adjustment actions are output, and the key micro-parameter set is iteratively optimized based on the feedback after the action is executed; S4, based on the optimal combination of output microscopic parameters, drive the numerical simulation model to update parameters to complete the closed-loop calibration.

2. The automatic calibration method for microscopic parameters of high-temperature layered rock masses based on the DDPG algorithm according to claim 1, characterized in that, The macroscopic mechanical experimental data includes macroscopic stress-strain experimental curves and failure mode data.

3. The automatic calibration method for microscopic parameters of high-temperature layered rock masses based on the DDPG algorithm according to claim 1, characterized in that, In step S2, the use of sensitivity analysis to identify a set of key micro-parameters that significantly influence the macroscopic mechanical response from the micro-parameters of the numerical simulation model includes: S201. Based on the physical constitutive model of high-temperature layered rock masses, a set of candidate microscopic parameters to be analyzed is selected, and a physical value range is set for each parameter. The set of candidate microscopic parameters includes at least: grain / contact normal stiffness. Tangential stiffness Parallel bond cohesion C, parallel bond friction angle ; S202, orthogonal experimental design is used to extract multiple sets of micro-parameter combinations within the physical value range of the candidate micro-parameter set; S203, input the multiple sets of microscopic parameter combinations into the numerical simulation model for calculation, and record the macroscopic mechanical indicators of each set of simulation outputs; S204, the range values ​​of each microscopic parameter with respect to the macroscopic mechanical index are calculated using the range analysis method. The calculation formula is: in, For the first j The range of each parameter; the larger the R value, the more significant the influence of the micro parameter on the macro result. j Representing the j There are three micro-parameters, i∈1,2,3…, representing different levels selected for each micro-parameter. This represents the average of all experimental results for the j-th parameter at the i-th level; S205, based on the descending order of the range values, select the top n micro parameters whose cumulative contribution rate reaches the preset threshold to form the key micro parameter set.

4. The automatic calibration method for microscopic parameters of high-temperature layered rock masses based on the DDPG algorithm according to claim 3, characterized in that, In step S2, the step of using the first feature dataset as the target and performing a deviation mapping with the second feature dataset to construct a state space based on the obtained deviation information includes: S206, Extract the macroscopic mechanical constants used as the target reference from the first feature dataset, including: target elastic modulus. Target peak intensity and target peak strain ; Simultaneously, from the second feature dataset at the current moment, the corresponding real-time simulation variables are extracted, including: real-time simulation elastic modulus. Real-time simulation of peak intensity Real-time simulation of peak strain and real-time microcrack ratio ; S207, Perform corresponding difference calculations between the extracted real-time simulation variables and the target benchmark to obtain the macroscopic mechanical response deviations: And microscopic damage evolution bias: S208, the macroscopic mechanical response deviation, microscopic damage evolution deviation, and current operating condition variables are concatenated in a preset order to obtain a multidimensional state space vector. : Where T is the target temperature in the current thermo-mechanical coupling simulation, and θ is the layering loading angle.

5. The automatic calibration method for microscopic parameters of high-temperature layered rock masses based on the DDPG algorithm according to claim 4, characterized in that, The reward function The calculation formula is based on the dynamic change rate of macroscopic mechanical characteristic error and microscopic damage evolution deviation between adjacent training rounds. in, This represents the difference in macroscopic mechanical error between the current round and the previous round. This represents the difference between the current round and the previous round's micro-mechanism deviations. and These are the corresponding weighting coefficients.

6. The automatic calibration method for microscopic parameters of high-temperature layered rock masses based on the DDPG algorithm according to claim 5, characterized in that, In step S3, the Deep Deterministic Policy Gradient (DDPG) model is obtained through construction and training: The DDPG model construction process includes: establishing a main Actor network as a policy network, a main Critic network as a value evaluation network, a target Actor network and a target Critic network with the same structure as the main network, and an experience replay pool; The DDPG model training process includes: A. Initialize the parameters of the main Actor network, the main Critic network, the target Actor network, and the target Critic network, and initialize the experience replay pool; B, the main Actor network according to the current state Output Action Add exploration noise The numerical simulation model is then input and executed to obtain rewards from environmental feedback. and the next state , will experience tuples ( , , , Store it in the experience replay pool; C. Sample batches of experience data from the experience replay pool after the data volume reaches a preset scale, and calculate the temporal difference target based on the sampled batches of experience data. The main Critic network is updated by minimizing the error of the value function; The main Actor network is updated using the policy gradient provided by the updated main Critic network through policy gradient ascent. D, with soft update rate The parameters of the main Actor network and the main Critic network are smoothly updated to the corresponding target Actor network and target Critic network; E. Repeat steps B to D until the model converges and the goodness of fit between the simulation results and the first feature dataset reaches the preset threshold, thus obtaining the trained DDPG model.

7. The automatic calibration method for microscopic parameters of high-temperature layered rock masses based on the DDPG algorithm according to claim 6, characterized in that, In step C, the main Critic network is updated using the following loss function L: Where N is the number of samples, and i is the number of training rounds. The error between the estimated state-action value function at the current moment and the expected value of the target obtained by the target network is calculated using the following formula: in, As a discount factor, For immediate rewards, Q represents the main Critic network value function. For the target Critic network value function, The main Critic network evaluates the value of the current state-action pair. This is the current state. For the current action, The main Critic network parameters. For the purpose of the Critic network, the value assessment of the next state-action pair is performed. For the target Actor network, based on the next state The predicted next optimal action, The parameters are those of the target Critic network.

8. The automatic calibration method for microscopic parameters of high-temperature layered rock masses based on the DDPG algorithm according to claim 7, characterized in that, In step C, the main Actor network is updated using the following gradient strategy: In the formula, J represents the expected return of the strategy. Indicates the desired operation. Represents a deterministic strategy. , and These are the network parameters of the objective policy network and the objective value function network, respectively; s represents the sample state. Represents a value function.

9. An automatic calibration system for microscopic parameters of high-temperature layered rock masses based on the DDPG algorithm, characterized in that, The automatic calibration method for microscopic parameters of high-temperature layered rock mass based on the DDPG algorithm according to any one of claims 1 to 8, the system includes interconnected parameter control terminals and numerical simulation platforms; The parameter control terminal includes: The data acquisition and processing module is used to collect macroscopic mechanical experimental data of high-temperature layered rock mass under different temperatures and bedding conditions to generate a first feature dataset, and drive the numerical simulation platform to perform calculations based on initial micro parameters to generate an initial second feature dataset. The reinforcement learning environment construction module is used to perform feature alignment processing on the first feature dataset and the second feature dataset, identify key micro-parameter sets using sensitivity analysis, and perform bias mapping to construct the state space and reward function. The DDPG agent module has a built-in deep deterministic policy gradient model, which is used to receive the state space vector and adjust the action according to the micro parameters output by the reward function. The parameter-driven and verification module is used to send the micro-parameter adjustment action to the numerical simulation platform and receive the optimization results to output the optimal combination of micro-parameters. The numerical simulation platform has a built-in thermo-coupled numerical simulation model, which is used to adjust its actions to update its own model parameters and perform simulation calculations based on the received micro parameters. The simulation feedback data generated by the simulation calculation is returned to the reinforcement learning environment construction module to update the state space.

10. An electronic device, characterized in that, The system includes a memory and a processor, wherein the processor is used to execute computer management programs stored in the memory to implement the steps of the automatic calibration method for micro-parameters of high-temperature layered rock masses based on the DDPG algorithm as described in any one of claims 1-8.