An engineering geological risk auxiliary identification method, system, device and medium
By establishing a spatiotemporal knowledge base and neural network model for geological engineering, risk probability prediction and adaptive early warning are carried out, solving the dynamic adaptability problem of geological risk assessment in traditional engineering design. This enables real-time dynamic assessment and adaptive control of geological risks, improving the accuracy of risk control in engineering design and construction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 金成文
- Filing Date
- 2026-03-05
- Publication Date
- 2026-06-02
AI Technical Summary
In traditional engineering design, geological risk assessment relies on static rules and expert experience, which cannot dynamically adapt to changes in the construction environment. This leads to delayed risk assessment, rigid strategies, and difficulty in achieving the best balance between safety management and project progress.
By acquiring multi-source geological exploration data, a spatiotemporal knowledge base for geological engineering is established, parametric hydrodynamic coupling modeling is performed, a neural network model is trained, a geological proxy model is generated, risk probability prediction and adaptive early warning are performed, a risk control action instruction set is generated, and decision optimization is carried out by combining reinforcement learning strategy network and hard rules.
It enables real-time dynamic assessment and adaptive control of geological risks, improves the accuracy of risk management in engineering design and construction, generates flexible and optimal response solutions, and solves the problems of static blind spots and rigid strategies.
Smart Images

Figure CN122134961A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of artificial intelligence-assisted decision-making, and in particular relates to a method, system, device and medium for auxiliary identification of geological risks in engineering design. Background Technology
[0002] With the deep integration of artificial intelligence and numerical simulation technologies, intelligent optimization techniques have emerged capable of real-time modeling and autonomous decision-making for complex dynamic systems. These techniques are characterized by data-driven approaches, online learning, and dynamic trade-offs among multiple objectives, leading to the evolution of engineering risk management from static rules to dynamic adaptive strategies. Traditionally, geological risk assessment during the engineering design phase heavily relies on discrete snapshot data provided by interim exploration reports, and construction management primarily depends on fixed thresholds and response procedures preset by engineers' experience. In this traditional approach, risk handling typically follows a linear pattern of exploration-design-threshold early warning. Geological data is input into the design model as static parameters, and during construction, pre-defined contingency plans are triggered by comparing monitoring data with fixed thresholds. The entire management process heavily relies on expert experience to determine risk levels and select countermeasures, making it difficult to accurately quantify the economic efficiency and timeliness of these measures. This traditional approach based on fixed rules and static analysis has inherent limitations. First, it cannot quantitatively extrapolate and predict the dynamic evolution of geological parameters affected by environmental factors and construction disturbances during the exploration interval, resulting in delayed risk assessment. Secondly, the pre-set static thresholds and tiered response mechanisms lack flexibility, making it difficult to achieve an optimal balance between safety management costs and project schedule efficiency in a rapidly changing construction environment. Finally, the decision-making knowledge is rigid, unable to continuously learn and evolve from historical and real-time data, resulting in management strategies that are difficult to adapt to complex and ever-changing engineering realities, creating static blind spots in risk management and a dilemma of rigid strategies. Summary of the Invention
[0003] Therefore, it is necessary to provide an engineering design geological risk auxiliary identification method, system, equipment, and medium that can accurately assess risks and provide management strategies that adapt to complex and ever-changing engineering realities, addressing the aforementioned technical problems.
[0004] Firstly, this application provides a method for auxiliary identification of geological risks in engineering design, including:
[0005] Multi-source geological exploration data is acquired, and knowledge is extracted from the multi-source geological exploration data to obtain a spatiotemporal knowledge base for geological engineering; the spatiotemporal knowledge base for geological engineering includes historical data and real-time data;
[0006] Based on the spatiotemporal knowledge base of geological engineering, a parametric hydrodynamic coupling model is constructed to obtain a high-fidelity geological physical model.
[0007] Based on a high-fidelity geophysical model, a neural network model is trained to obtain a geological proxy model. Real-time monitoring data is then input into the geological proxy model to predict risk probability and obtain a risk probability map.
[0008] Based on the risk probability map, adaptive early warning is performed to obtain a risk control action instruction set; the risk control action instruction set is used to assist in engineering design and construction.
[0009] Furthermore, based on a high-fidelity geophysical model, a neural network model is trained to obtain a geological proxy model, including:
[0010] Based on the input parameter range of the high-fidelity geophysical model, parameter sampling is performed to obtain the input parameter combination, and the input parameter combination is matched with each target variable to obtain the list of tasks to be simulated;
[0011] Based on the list of tasks to be simulated, the input parameter combinations are input into the high-fidelity geophysical model for numerical simulation, the output results are obtained, and the target variables corresponding to the input parameter combinations are extracted from the output results to obtain input-output paired data.
[0012] By using the combination of input parameters as features and the target variable as a label, the input and output paired data are divided to obtain the training dataset;
[0013] Based on the training dataset, the neural network model is trained to obtain the geological proxy model; the number of input layer nodes of the neural network model is the number of parameters in the combination of input parameters, and the number of output layer nodes of the neural network model is the number of target variables.
[0014] Furthermore, based on the risk probability map, adaptive early warning is performed to obtain a set of risk control action instructions, including:
[0015] Quantitative features are extracted from the risk probability map, and the quantitative features and real-time data are encoded to obtain a state vector;
[0016] The state vector is input into the reinforcement learning policy network to make action decisions, thus obtaining an action draft.
[0017] Based on the hard rule filter, the action draft is matched with the hard rules to obtain the early warning and control actions;
[0018] Based on natural language generation, early warning and control actions are mapped to corresponding recommended decisions, and the early warning and control actions are input into the geological proxy model for rapid simulation to obtain short-term simulation comparison charts;
[0019] By integrating the recommended decision-making and short-term simulation comparison charts, a set of risk management action instructions is obtained.
[0020] Furthermore, the reinforcement learning policy network was obtained through the following method:
[0021] Based on the geological agent model, the random geological response under different decision-making interventions is simulated, and the non-geological state dimension update caused by the decision-making action is configured based on the engineering logic rules to obtain the simulation environment; the non-geological state dimension includes the engineering progress dimension, resource constraint dimension, environmental time sequence dimension, and control state dimension.
[0022] Based on engineering quantization rules, the state space, action space, and reward function are generated in a quantized manner. Based on the state space, action space, and reward function, the policy network architecture is designed to obtain the policy network and value network.
[0023] Based on near-end policy optimization, the policy network and value network are trained and updated in a simulation environment to obtain the trained policy network parameters.
[0024] Based on the test environment set, the performance of the trained policy network parameters is tested to obtain a policy performance evaluation report.
[0025] Based on the policy performance evaluation report, the parameters of the trained policy network are set to those of a reinforcement learning policy network.
[0026] Furthermore, the formula for the reward function is:
[0027]
[0028] in, Let t be the reward function, and t be the time step. For safety weighting coefficients, This is the cost weighting coefficient. This is the schedule weighting coefficient. As a safety reward, As a cost penalty, As a progress reward, As a supplementary reward;
[0029] The safety reward is obtained by comparing the indicators of random geological response with a set of safety thresholds.
[0030] Furthermore, after integrating the recommended decision-making and short-term simulation comparison charts to obtain the risk management action instruction set, it also includes:
[0031] Obtain the actual executed actions and real-time status logs, and calculate the actual reward value based on the real-time status logs and the reward function;
[0032] Based on the decision moment, the complete state vector, the actual action executed, the new state, and the real reward value are integrated to obtain the online experience quadruple;
[0033] Based on a priority strategy, the online experience quadruple and the historical experience pool are prioritized to obtain the online experience pool.
[0034] Based on the online experience pool, the parameters of the reinforcement learning policy network are fine-tuned to obtain the updated policy network parameters; the updated policy network parameters are used to update the reinforcement learning policy network.
[0035] Secondly, this application also provides an engineering design geological risk auxiliary identification system, including:
[0036] The knowledge module is used to acquire multi-source geological exploration data and extract knowledge from the multi-source geological exploration data to obtain a spatiotemporal knowledge base for geological engineering; the spatiotemporal knowledge base for geological engineering includes historical data and real-time data;
[0037] The modeling module is used to perform parametric hydrodynamic coupling modeling based on the geological engineering spatiotemporal knowledge base to obtain a high-fidelity geophysical model.
[0038] The risk module is used to train a neural network model based on a high-fidelity geophysical model to obtain a geological proxy model. Real-time monitoring data is then input into the geological proxy model to predict risk probability and obtain a risk probability map.
[0039] The early warning module is used to provide adaptive early warnings based on the risk probability map and obtain a risk control action instruction set; the risk control action instruction set is used to assist in engineering design and construction.
[0040] Thirdly, this application also provides a computer device including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement any step of the method provided in the first aspect of this application.
[0041] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any step of the method provided in the first aspect of this application.
[0042] The aforementioned method, system, equipment, and medium for auxiliary identification of geological risks in engineering design acquire multi-source geological exploration data and extract knowledge from this data to obtain a spatiotemporal knowledge base for geological engineering. This knowledge base includes historical and real-time data. Based on this knowledge base, a parametric hydrodynamic coupling model is created to obtain a high-fidelity geophysical model. A neural network model is trained based on this model to obtain a geological proxy model. Real-time monitoring data is input into the geological proxy model to predict risk probabilities, resulting in a risk probability map. Based on the risk probability map, adaptive early warning is performed, yielding a risk control action instruction set. This set is used to assist in engineering design and construction. By utilizing the geological proxy model, discrete exploration data is transformed into a continuous digital twin capable of simulating the dynamic evolution of geological parameters, solving the static blind spot and enabling quantitative extrapolation and forward-looking prediction of risks under any construction sequence. Risk control is modeled as a sequential decision-making process, capable of generating flexible and optimal response solutions under multiple constraints based on real-time conditions, effectively improving the accuracy of risk control in engineering design and construction. Attached Figure Description
[0043] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0044] Figure 1 A schematic diagram illustrating the process of an auxiliary identification method for geological risks in engineering design, provided in an embodiment of the present invention;
[0045] Figure 2 This is a schematic diagram of the structure of an engineering design geological risk auxiliary identification system provided in an embodiment of the present invention. Detailed Implementation
[0046] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0047] In one embodiment, such as Figure 1As shown, an auxiliary method for identifying geological risks in engineering design is provided. This embodiment illustrates the application of this method to a terminal. It is understood that this method can also be applied to a server, and further to a system including both a terminal and a server, and implemented through interaction between the terminal and the server. In this embodiment, the method includes the following steps:
[0048] Step 101: Obtain multi-source geological exploration data and extract knowledge from the multi-source geological exploration data to obtain a spatiotemporal knowledge base for geological engineering; the spatiotemporal knowledge base for geological engineering includes historical data and real-time data.
[0049] Multi-source geological exploration data refers to a collection of raw data related to geological conditions and engineering exploration obtained from various channels and using different technical means. Sources include, but are not limited to: field geological mapping records, core logging and testing data from drilling, geophysical exploration data, field monitoring data, laboratory geotechnical mechanics test data, and existing geological reports and literature. Knowledge extraction is a data processing and information enhancement process. It refers to using techniques such as data cleaning, format standardization, information fusion, entity recognition, and relationship construction to identify and extract structured entities with clear engineering geological significance from heterogeneous, multi-source raw exploration data. These entities include strata, faults, and groundwater; attributes such as thickness, strength, and permeability coefficient; and their interrelationships such as inclusion, contact, and influence. The geological engineering spatiotemporal knowledge base refers to a structured database established after the knowledge extraction process. It contains historical and real-time data and is organized using spatial location and time series as key indexes, thereby representing the distribution and evolution of geological conditions and engineering status in the spatiotemporal dimension.
[0050] The terminal collects raw exploration data from multiple channels and aggregates the data into a unified data management platform. The aggregated raw data is cleaned, standardized in format, and conflicts and inconsistencies between multiple data sources are resolved. This achieves preliminary alignment and fusion of the data within a spatiotemporal framework. From the fused data, key geological engineering objects are identified, their attribute parameters are extracted, and spatial and temporal relationships between objects are established. The extracted structured knowledge is stored in a form suitable for querying, analysis, and model invocation. For example, relational database tables, graph database nodes, and edges can be used to construct a spatiotemporal knowledge base for geological engineering. The knowledge base can continuously incorporate new real-time monitoring data to achieve dynamic updates of knowledge.
[0051] Step 102: Based on the spatiotemporal knowledge base of geological engineering, perform parametric hydrodynamic coupling modeling to obtain a high-fidelity geophysical model.
[0052] Specifically, parametric hydromechanical coupling modeling refers to constructing a mathematical model that can simulate the interaction between groundwater seepage and the stress-strain process of soil and rock. Parametric characteristics mean that key coefficients representing specific geological bodies or material properties are defined as a series of adjustable input parameters. Changing these parameter values characterizes different geological conditions or allows for model calibration. A high-precision geophysical model, on the other hand, refers to a high-precision numerical model that has been determined and solidified through modeling and calibration processes. Its parameters and boundary conditions can highly reproduce the geological and mechanical properties of the actual site, and the simulation results show good agreement with measured data, thus being considered to have a high degree of fidelity to real physical processes.
[0053] Based on information from the Geological Engineering Spatiotemporal Knowledge Base, the terminal determines the scope, geological structure, hydrogeological zoning, and mechanical boundaries of the study area, forming a qualitative framework for understanding engineering geological problems. It selects governing equations describing the water-mechanical coupling process; for example, equations based on Darcy's law and the effective stress principle are chosen. Numerical methods are used to discretize the continuous physical domain and equations into computer-computable grid cells and a set of algebraic equations. Relevant attribute parameters are extracted from the Geological Engineering Spatiotemporal Knowledge Base as initial parameter values for the model. Based on the spatiotemporal information in the knowledge base, boundary and initial conditions of the model are set. A preliminary model is run, and its output is repeatedly compared with historical data stored in the Geological Engineering Spatiotemporal Knowledge Base. By inverting the parameterized part of the model, the simulation results are made to fit the known observation data to the maximum extent. Using another part of the knowledge base that was not calibrated, the calibrated model is tested to evaluate its predictive ability, resulting in a reliable and usable high-fidelity geophysical model.
[0054] Step 103: Based on the high-fidelity geophysical model, train the neural network model to obtain the geological proxy model, and input the real-time monitoring data into the geological proxy model to perform risk probability prediction and obtain the risk probability map.
[0055] Specifically, a neural network model refers to a machine learning model that employs a neural network architecture, which serves as the object being trained in this embodiment. Training refers to the process of adjusting the internal parameters of a neural network model by providing it with input data and corresponding expected outputs, enabling it to learn the mapping relationship from input to output. A geological proxy model refers to the neural network model obtained after training, acting as a proxy for a faithful geophysical model. It can approximate the input-output relationship of the faithful model at extremely high speeds, but sacrifices some interpretability of physical details. Real-time monitoring data refers to the continuous and instantaneous observation data acquired from sensors and other equipment at the engineering site. Risk probability prediction refers to using a model to calculate the likelihood of a specific geological risk occurring in the future. A risk probability map is used to graphically visualize the risk probability prediction results, intuitively indicating the probability of risk occurring at different spatial locations within a specific time period.
[0056] The terminal is based on a high-fidelity geophysical model. It generates a large amount of input-output paired data by sampling its input parameters and running simulations. The paired data is used as a training dataset to train a neural network model, resulting in a geological proxy model. Real-time monitoring data from the engineering site is directly input into the trained geological proxy model, which quickly calculates the corresponding risk probability value and generates a visualized risk probability map.
[0057] Step 104: Based on the risk probability map, perform adaptive early warning to obtain a risk control action instruction set; the risk control action instruction set is used to assist in engineering design and construction.
[0058] Adaptive early warning refers to an intelligent early warning mechanism that can dynamically adjust early warning strategies and automatically generate response suggestions based on the current risk situation and real-time data. Risk management action instruction set refers to a specific set of action suggestions or instructions derived from the adaptive early warning process, used to guide engineering designers on what measures to take to avoid or reduce identified risks.
[0059] The terminal uses the risk probability map as the core input to initiate an adaptive early warning process. It extracts key features from the risk probability map and encodes them into states by combining them with other real-time data. Through intelligent decision-making modules such as reinforcement learning strategy networks, it proposes preliminary control action suggestions. These suggestions are then filtered and corrected by hard rules, transforming the actions into recommended decisions described in natural language. The terminal also generates a comparison and projection diagram of the effects before and after taking the action, integrating all the information into a complete set of risk control action instructions.
[0060] This embodiment provides a method for auxiliary identification of geological risks in engineering design. It acquires multi-source geological exploration data and extracts knowledge from this data to obtain a spatiotemporal knowledge base for geological engineering. This knowledge base includes historical and real-time data. Based on this knowledge base, a parametric hydrodynamic coupling model is used to obtain a high-fidelity geophysical model. A neural network model is trained based on this model to obtain a geological proxy model. Real-time monitoring data is input into the geological proxy model to predict risk probabilities, resulting in a risk probability map. Based on the risk probability map, adaptive early warning is performed, resulting in a risk control action instruction set. This set is used to assist in engineering design and construction. By utilizing the geological proxy model, discrete exploration data is transformed into a continuous digital twin capable of simulating the dynamic evolution of geological parameters, solving the static blind spot and enabling quantitative extrapolation and forward-looking prediction of risks under any construction sequence. Risk control is modeled as a sequential decision-making process, capable of generating flexible and optimal response solutions under multiple constraints based on real-time conditions, effectively improving the accuracy of risk control in engineering design and construction.
[0061] In one embodiment, a neural network model is trained based on a high-fidelity geophysical model to obtain a geological proxy model, including:
[0062] Step 201: Based on the input parameter range of the high-fidelity geophysical model, perform parameter sampling to obtain the input parameter combination, and match the input parameter combination with each target variable to obtain the list of tasks to be simulated.
[0063] The input parameter range of a high-fidelity geophysical model refers to the range or set of possible values for key parameters that can be adjusted to characterize different geological or working conditions. Parameter sampling refers to a method of selecting multiple sets of specific parameter values within a given input parameter range according to a specific strategy. An input parameter combination refers to a set of specific values obtained through parameter sampling; this set of values constitutes all the input conditions required for a complete simulation run of the high-fidelity geophysical model, and each combination is a specific parameter assignment scheme. The target variable refers to the physical quantity that we are ultimately concerned with in engineering risk analysis and that needs to be calculated and output by the model. The list of tasks to be simulated refers to a list of multiple simulation tasks to be executed, where each task explicitly associates a set of input parameter combinations with one or more target variables that need to be obtained from the simulation results.
[0064] Based on all the variable input parameters of the high-fidelity geophysical model and their respective ranges, the terminal performs parameter sampling operations within a defined parameter space to generate a large number of different combinations of input parameters. Each combination represents a possible geological condition or engineering scenario. For each set of input parameter combinations, the target variable to be extracted from the model output is specified. This matching relationship is organized into a list to obtain a list of tasks to be simulated.
[0065] Step 202: Based on the list of tasks to be simulated, input the combination of input parameters into the high-fidelity geophysical model, perform numerical simulation, obtain the output results, and extract the target variables corresponding to the combination of input parameters from the output results to obtain input-output paired data.
[0066] Specifically, numerical simulation refers to using a computer to solve the mathematical equations corresponding to a faithful geophysical model, and numerically calculating the process of how the physical response changes over time or space under a given combination of input parameters. Output results refer to the raw calculation result files or data generated after a faithful geophysical model completes a numerical simulation, typically containing detailed physical quantity information for all computational nodes / units across the entire model at each time step. Input-output paired data refers to a structured dataset where each data record explicitly associates a set of input parameter combinations with the specific numerical value of its corresponding target variable extracted from the output results, forming an input-output pairing relationship.
[0067] The terminal, based on the list of tasks to be simulated, iteratively submits each set of input parameter combinations in the list to the high-fidelity geophysical model, initiating numerical simulation calculations. After each simulation, the complete output results generated by the model are saved. For each simulation, the specific values of the target variables specified for the current input parameter combinations in the task list are precisely located and extracted from the output file. The input parameter combinations used in each simulation are recorded one-to-one with the target variable values extracted from that simulation, thus forming input-output pairing data. All the pairing data generated by the simulation tasks are collected together to form a large-scale dataset. Extensive and systematic numerical experiments were conducted using the high-fidelity physical model, solidifying the knowledge of the physical model in a structured, machine-readable input-output pairing dataset. The dataset completely records the model's response patterns under different input conditions, serving as a teaching material for training surrogate models that can approximate the behavior of the physical model.
[0068] Step 203: Using the combination of input parameters as features and the target variable as a label, divide the input and output paired data to obtain the training dataset.
[0069] Specifically, in the context of machine learning, a feature refers to an input variable or attribute used to describe a sample, i.e., an input variable or attribute of a simulation task. In this embodiment, it specifically refers to the values of each parameter in the combination of input parameters. In supervised learning, a label refers to the true or target value corresponding to the sample features, which the model is expected to predict. In this embodiment, it specifically refers to the numerical value of the target variable extracted from the simulation results. The training dataset refers to a subset of data specifically used for training the machine learning model, partitioned from the complete input-output paired data, and consists of multiple feature-label samples.
[0070] The terminal transforms each row of data in the input-output pairing data according to the standard format of machine learning, defines the combination of input parameters in each row as the feature vector of the sample, and defines the target variable corresponding to the same row as the label vector of the sample. All the transformed sample data are then divided into several mutually exclusive subsets according to a certain ratio, optionally 70% for training, 15% for validation, and 15% for testing, randomly or according to specific rules. The largest subset is used for model parameter learning and adjustment, which serves as the training dataset. The remaining portions are usually used as validation and test sets to evaluate model performance during training and to finally test the model's generalization ability.
[0071] Step 204: Based on the training dataset, train the neural network model to obtain the geological proxy model; the number of input layer nodes of the neural network model is the number of parameters in the input parameter combination, and the number of output layer nodes of the neural network model is the number of target variables.
[0072] In this context, a neural network model refers to a computational model composed of a large number of interconnected simple computational units arranged in a specific architecture. In this embodiment, it is the object to be trained, aiming to learn the complex mapping relationship between input parameters and target variables. Training refers to an iterative optimization process. In this embodiment, the internal parameters of the neural network model are continuously adjusted to minimize the error between the model's predicted values on the training dataset and their true labels. A geological proxy model refers to a neural network model that, after training, has learned to quickly predict the mapping relationship between input parameters and target variables. It is a fast and approximate alternative to a faithful geophysical model. The number of input layer nodes refers to the number of neurons in the input layer of the neural network model, determining how many input values the model can receive at once. In this embodiment, it is set to be equal to the number of parameters in the input parameter combination. The number of output layer nodes refers to the number of neurons in the output layer of the neural network model, determining how many predicted values the model can output at once. It is set to be equal to the number of target variables to be predicted.
[0073] The terminal determines the structure of the neural network based on the input and output dimensions of the problem. The key is to set the number of nodes in the input layer to be equal to the number of input parameters, and the number of nodes in the output layer to be equal to the number of target variables. The number of hidden layers and nodes are set according to the complexity of the problem. The training dataset is input into the predefined neural network model, and the internal parameters of the model are iteratively adjusted through optimization algorithms such as backpropagation, so that the predicted value output by the model is as close as possible to the true value of the target variable. When the error of the model on the training set and the validation set decreases to an acceptable level and tends to stabilize, the training process ends. The neural network model obtained at this time, which has stable prediction ability, is the required geological proxy model.
[0074] This embodiment uses machine learning methods to enable a lightweight neural network model to learn the core input-output mapping rules from the massive amounts of data generated by expensive but high-precision geophysical models. The resulting geological proxy model can achieve rapid computation while maintaining high prediction accuracy, thus making frequent risk prediction based on real-time data possible. The strict specification of the number of input and output nodes ensures a strict match between the model interface and the definition of the actual problem.
[0075] In one embodiment, an adaptive early warning is performed based on a risk probability map to obtain a risk control action instruction set, including:
[0076] Step 301: Extract quantitative features from the risk probability map and encode the quantitative features and real-time data to obtain a state vector.
[0077] Among them, the risk probability map refers to a graphical representation of the probability distribution of specific geological risks occurring in different spatial locations in the future, generated through geological proxy models. Quantitative features refer to numerical indicators extracted from the risk probability map that represent key risk information, such as the maximum and average risk probabilities for the entire area, the area of high-risk areas, and the geometric center coordinates of high-risk areas. Real-time data refers to observational data reflecting the current state of the project and its environment, acquired in real time from the engineering site monitoring system, such as current displacement values, water levels, and rainfall. Encoding refers to a data processing operation that converts data of different types and dimensions into a unified, structured numerical representation. The state vector, obtained through encoding, is a multi-dimensional numerical array that provides a comprehensive and compact mathematical description of the entire system at the current moment, including the risk situation and the engineering environment; it serves as a standardized input for intelligent decision-making.
[0078] The terminal analyzes the risk probability map, converting the image data into a grid or vector. It applies image processing and statistical analysis algorithms to calculate a series of quantitative features that summarize the core information. Examples include statistical features such as calculating the mean, variance, skewness, and kurtosis of the overall risk probability; identifying and statistically analyzing the percentage of pixels / regions whose probabilities exceed preset warning and danger thresholds; spatial features such as identifying the hotspots with the highest risk probability and calculating their area, perimeter, and centroid coordinates; calculating the maximum spatial gradient and location of the risk probability; and temporal features such as calculating the differences between the current map and the previous time-series map, including the area of risk expansion and the rate of probability growth. All extracted quantitative features are then aggregated with real-time data synchronously acquired from the monitoring system in a predefined order, ensuring that both are aligned in timestamps to represent the system state at the same moment. The aggregated mixed data is then encoded to handle missing and outlier values. All features and real-time data are normalized or standardized to eliminate dimensional differences and bring them into a similar numerical range, which is beneficial for model processing. All normalized values are then concatenated into a one-dimensional array in a fixed order. The array has a fixed dimension and its length is equal to the number of quantized features plus the number of real-time data variables. The resulting one-dimensional array is the state vector.
[0079] Step 302: Input the state vector into the reinforcement learning policy network to make action decisions and obtain the action draft.
[0080] Specifically, a reinforcement learning policy network refers to a trained neural network that acts as the brain of a reinforcement learning agent, learning and outputting the appropriate action based on the current environmental state to maximize long-term cumulative rewards. Action decision-making refers to the process by which the reinforcement learning policy network, based on the input state vector, internally calculates and outputs a specific action selection. An action draft refers to a preliminary suggested action generated by the reinforcement learning policy network through the action decision-making process; essentially, it is the network's original action instruction, which has not yet been verified for engineering feasibility and safety rules.
[0081] The terminal takes the state vector representing the current system state as input and passes it to the already deployed reinforcement learning policy network. The network performs calculations and inferences based on the policy it has learned internally, and outputs a specific action draft, which is a specific operation instruction.
[0082] Step 303: Based on the hard rule filter, match the action draft with the hard rules to obtain the early warning and control actions.
[0083] Specifically, a hard rule filter refers to a logical verification module built upon a set of explicit and inviolable rules. Hard rules refer to mandatory regulations or constraints that must be strictly followed in engineering practice, such as safety procedures, physical limits, resource constraints, and construction specifications. Early warning and control actions refer to action instructions that, after verification and correction by the hard rule filter, conform to intelligent decision-making suggestions and satisfy all hard constraints, and are practically executable.
[0084] The terminal inputs the draft action into the hard rule filter. The filter compares and makes logical judgments against the preset hard rule library one by one. Optionally, it determines whether the draft will cause a certain monitoring indicator to exceed the safety threshold, or whether the required resources exceed the current inventory. If the draft violates any hard rule, the filter will modify it or directly reject it and replace it with a safe action that conforms to the rule. The action output after all rules are verified is the early warning control action.
[0085] Step 304: Based on natural language generation, the early warning and control actions are mapped to corresponding recommended decisions, and the early warning and control actions are input into the geological proxy model for rapid simulation to obtain a short-term simulation comparison chart.
[0086] Natural language generation (NLP) refers to an artificial intelligence technology that automatically converts structured data or instructions into fluent and easily understandable natural language text. Recommendation-based decision-making refers to using NLP technology to convert structured early warning and control action instructions into textual descriptions and suggestions for engineering management or operational personnel. Short-term simulation comparison charts are comparative visualizations that use early warning and control actions as input, quickly inputting them into a geological proxy model for a simulation, and comparing the simulation results with the current state or a predicted state of no action, thus intuitively demonstrating the expected short-term effects of the action.
[0087] Based on natural language generation technology, the terminal maps machine-readable early warning and control actions into easily understandable text descriptions with explanations and suggestions, i.e., recommended decisions. At the same time, the same early warning and control action is used as a new boundary condition or parameter and input into the geological proxy model to start a rapid simulation. It predicts the changes in key risk indicators after the action is executed in a short period of time. The prediction results are compared with the baseline scenario, i.e., not executing the action, to generate a visual short-term extrapolation comparison chart.
[0088] Step 305: Integrate the recommended decision and short-term simulation comparison chart to obtain a risk management action instruction set.
[0089] Among them, the risk management action instruction set refers to a complete and actionable risk response plan document or data package delivered to the user, which integrates decision-making suggestions in text form and effect simulation in graphical form.
[0090] The terminal combines, encapsulates, and formats text-based recommendation decisions and graphical comparison charts of short-term projections. This includes inserting text and images into a standard report template or packaging them into a structured data object, and then generating a complete set of risk management action instructions through integration.
[0091] This embodiment improves the readability and understandability of decision support information, provides text suggestions, and also provides a visual preview of the effects through rapid simulation, helping engineers to more intuitively understand the expected consequences of the suggested actions, thereby making judgments and decisions.
[0092] In one embodiment, the reinforcement learning policy network is obtained through the following method:
[0093] Step 401: Based on the geological agent model, simulate the stochastic geological response under different decision-making interventions, and based on the engineering logic rules, configure the updates of non-geological state dimensions caused by decision-making actions to obtain the simulation environment; the non-geological state dimensions include engineering progress dimension, resource constraint dimension, environmental time sequence dimension, and control state dimension.
[0094] Among them, geological proxy models refer to neural network models that can quickly approximate geological and physical processes. Decision-making interventions refer to control measures or operations that may be taken in engineering, such as strengthening support, adjusting excavation progress, and initiating drainage. Stochastic geological responses refer to a series of different but reasonable geological reactions that may occur under the same decision-making interventions due to the inherent uncertainty of geological conditions, reflecting the uncertainty of the geological system. Engineering logic rules refer to the logic and rules describing non-geological processes such as engineering management, resource allocation, and planning. For example, how much time and resources are needed to complete an excavation step, or how much material inventory will be consumed by the implementation of a certain support measure. Non-geological state dimensions refer to other engineering state dimensions that need to be tracked and simulated in the simulation environment, in addition to geological responses. These include the engineering progress dimension, which describes how much work has been completed and which construction stage the project is in; the resource constraint dimension, which describes the quantity or status of currently available resources; the environmental time series dimension, which describes the time factors of the external environment, such as date, season, and schedule requirements; and the control status dimension, which describes the status of currently implemented control measures, such as which areas have had their support upgraded and whether the drainage system is operational. A simulation environment refers to a simulation system that integrates a geological proxy model and engineering logic rules, capable of accepting decision-making interventions as input and simulating the complete system state at the next moment.
[0095] The terminal uses a geological proxy model as the core of the simulation environment to simulate the stochastic geological response triggered by decision-making interventions. The stochastic aspect means that random variables representing uncertainty are introduced into the input of the proxy model to simulate the probabilistic nature of the geological response. Based on engineering logic rules, other state update logics besides the geological response are configured for the simulation environment. When a decision-making action is input, in addition to calculating the geological state change through the geological proxy model, the non-geological state dimension is also updated synchronously according to predefined rules. The geological response simulation module and the non-geological state update module are integrated to construct a unified and interactive simulation environment. This environment can receive actions, return new complete states, and can be deduced step by step.
[0096] Step 402: Based on engineering quantization rules, quantize and generate the state space, action space, and reward function. Based on the state space, action space, and reward function, design the policy network architecture to obtain the policy network and value network.
[0097] Specifically, engineering quantification rules refer to rules that transform engineering goals, constraints, and logic into specific mathematical forms. For example, quantifying the "safety first" goal as a reward for displacement not exceeding a certain threshold, and cost control as a negative reward for material consumption. State space refers to the mathematical description of the set of all possible system states in simulation, defining the dimension of the state vector and the range of values for each component. Action space refers to the mathematical description of the agent, i.e., the set of all possible decision actions that the policy to be trained can execute at all time steps, defining the format and range of actions. Reward function is a mathematical function that takes the current state, the action to be executed, and the next state as input, and outputs a scalar value used to quantify whether the result of the action in a given state is good or bad; it is the goal that the reinforcement learning agent wants to optimize. Policy network architecture design refers to determining the specific structure of the policy neural network, such as the number of layers, the number of neurons per layer, and the activation function. A policy network is a neural network whose function is the policy; it takes the state vector in the state space as input, calculates it, and outputs a specific action in the action space; it is the decision-making brain of the agent. A value network refers to another neural network whose function is to evaluate the value of a state or state-action pair. It takes a state as input and outputs a scalar value that represents an estimate of the long-term cumulative reward expected from that state, which is used to guide the update direction of the policy network.
[0098] Based on engineering quantification rules, the terminal performs specific quantification operations, clarifies which specific variables are included in the complete state vector output by the simulation environment, and determines the value range of each variable, thereby formally defining the state space. It also clarifies all the specific actions that the agent can take and formalizes them into an action space that can be processed by a computer. According to the multi-objective engineering, a specific mathematical formula is designed to calculate the immediate reward for each step. Based on the defined state space and action space, the policy network architecture and value network architecture are designed, including selecting the network type, determining the number of layers and neurons, and selecting the activation function, to obtain the model structure of the initializable policy network and value network.
[0099] Step 403: Based on near-end policy optimization, the policy network and value network are trained and updated in the simulation environment to obtain the trained policy network parameters.
[0100] Specifically, proximal policy optimization (PFO) refers to a specific reinforcement learning algorithm, belonging to the policy gradient method. It ensures the stability of the training process by limiting the magnitude of each policy update, and is one of the commonly used and efficient reinforcement learning training algorithms. Training update refers to the process in a simulation environment where the agent interacts through continuous trial and error, and, based on feedback from the reward function, uses the proximal policy optimization algorithm to adjust and optimize the internal parameters of the policy network and value network. Post-training policy network parameters refer to the final set of all weights and biases within the policy network model after the training update process is completed. These parameters determine the behavior of the trained policy network.
[0101] The terminal places the policy network and value network in a simulation environment and runs a proximal policy optimization algorithm for iterative training and updates. The agent uses the current policy network to run multiple rounds or time steps in the simulation environment, collecting a large amount of state-action-reward-new state sequence data. Using the collected data and the value network, the advantage value of each action relative to the average level is calculated. Based on the calculated advantage value and the current policy, an objective function is calculated. A pruning term is used to limit the difference between the old and new policies to be too large to ensure stability. The parameters of the policy network are updated through gradient ascent to make it more inclined to choose actions with high advantage. At the same time, the parameters of the value network are updated to make its state estimation more accurate. The steps are repeated until the performance of the policy converges or reaches the preset number of training rounds. All internal weight values of the policy network at this time are saved, which is the post-trained policy network parameters.
[0102] Step 404: Based on the test environment set, perform performance testing on the trained policy network parameters to obtain a policy performance evaluation report.
[0103] The test environment set refers to a collection of simulation environments, similar in type to the training environment but with different scenarios, used to evaluate the performance of the strategy. These environments simulate different initial geological conditions, engineering design schemes, or external disturbances to test the generalization ability of the trained strategy. Performance testing refers to the process of running an agent equipped with the trained policy network parameters within the test environment set, observing and recording a series of performance indicators. The strategy performance evaluation report is a document or data record summarizing the performance test results. It should include quantitative indicators, optionally including average cumulative reward, number of security violations, average engineering cost, completion time, and qualitative analysis, comprehensively evaluating the strategy's strengths, robustness, and potential defects.
[0104] The terminal prepares a set of test environments, ensuring that the test scenarios in it have not appeared during training, in order to realistically test the generalization ability. In each simulation environment of the test environment set, the trained policy network parameters are loaded into the policy network, and the agent runs a complete engineering simulation from scratch to perform performance testing, that is, recording its decisions at each step, the rewards obtained, and various results. The running data in all test environments are summarized, calculated and analyzed according to predefined evaluation criteria, and a comprehensive policy performance evaluation report is generated. The report will indicate in which aspects the policy performs well and in which extreme or unforeseen situations it may fail.
[0105] Step 405: Based on the policy performance evaluation report, set the parameters of the trained policy network to those of the reinforcement learning policy network.
[0106] In this context, the reinforcement learning policy network refers to a deployed policy network used for online decision-making, which in this embodiment is the target to be updated.
[0107] Based on the policy performance evaluation report, if the report shows that the policy performance meets or exceeds the predetermined deployment criteria, the terminal performs a setup operation to overwrite or load the validated trained policy network parameters into the reinforcement learning policy network used by the online decision-making system.
[0108] In this embodiment, new policy parameters that have undergone thorough training and rigorous testing and exhibit excellent performance are set as an online reinforcement learning policy network. Subsequent real-time adaptive early warnings will employ updated and superior policies for decision-making, thereby improving the level of intelligence in risk management and the quality of decision-making.
[0109] In one embodiment, the reward function is formulated as follows:
[0110]
[0111] in, Let t be the reward function, and t be the time step. For safety weighting coefficients, This is the cost weighting coefficient. This is the schedule weighting coefficient. As a safety reward, As a cost penalty, As a progress reward, As a supplementary reward;
[0112] The safety reward is obtained by comparing the indicators of random geological response with a set of safety thresholds.
[0113] Specifically, in the reinforcement learning framework, the reward function is the core mathematical function used to evaluate the quality of the actions performed by an agent at a specific time step. It outputs a scalar value and is the sole objective for the agent to learn and optimize its behavioral policies.
[0114] The safety weight coefficient is an adjustable positive coefficient used to adjust the relative importance and contribution ratio of safety reward items in the total reward. The cost weight coefficient is an adjustable positive coefficient used to adjust the relative importance and contribution ratio of cost penalty items in the total reward. The schedule weight coefficient is an adjustable positive coefficient used to adjust the relative importance and contribution ratio of schedule reward items in the total reward.
[0115] Safety reward is the component in the reward function that is directly related to engineering safety risk. It compares the indicators of random geological response obtained by simulation or monitoring with a preset safety threshold group, and gives corresponding reward or penalty values based on the comparison results.
[0116] Cost penalty is the component of the reward function related to the consumption of engineering resources. Since performing actions usually consumes resources, this component is often negative to encourage cost saving.
[0117] The progress reward is the component in the reward function that is related to the efficiency of project progress. When the agent's decision is conducive to advancing the project according to plan or faster, a positive reward is given.
[0118] Auxiliary rewards are other auxiliary optimization objectives in the reward function besides the three core objectives of safety, cost, and schedule, such as encouraging decision stability and reducing frequent changes.
[0119] The terminal obtains key indicators of random geological responses from the simulation results of the geological proxy model or real-time monitoring data, compares the indicator values with a preset safety threshold group, and calculates rewards according to preset rules. Optionally, indicators within the safe range are given a small positive reward; indicators exceeding the warning value are given a larger negative reward; and indicators exceeding the danger value are given a very large negative reward. Based on the action selected by the agent at the time step, the resource consumption cost corresponding to the action is queried, and the cost value is either negative or converted into a penalty value proportionally. The contribution of the action to the overall progress of the project is evaluated. For example, if the action is normal excavation, a positive reward is given based on the completion of the planned schedule; if it is a work stoppage for investigation, a zero or negative progress reward may be given. At time step t, the calculated reward of each component is multiplied by its corresponding weight coefficient and summed to calculate the total reward value for that step, which is then fed back to the reinforcement learning agent.
[0120] The complex, multidimensional, and sometimes conflicting management objectives in engineering are unified and quantified into a single scalar value, enabling reinforcement learning algorithms to optimize accordingly. The agent's goal is to maximize long-term cumulative rewards. Through this formula, the agent is guided to learn strategies that bring high safety rewards, low-cost penalties, and high progress rewards, thereby automatically finding an approximate optimal balance point under multiple constraints.
[0121] In one embodiment, after integrating the recommended decision and short-term simulation comparison chart to obtain the risk management action instruction set, the method further includes:
[0122] Step 601: Obtain the actual executed actions and real-time status logs, and calculate the actual reward value based on the real-time status logs and the reward function.
[0123] Among them, "actual execution action" refers to the control operation instructions that are actually executed on the real engineering site after the early warning and control actions have been manually confirmed or adjusted. "Real-time status log" refers to the raw system status data sequence continuously recorded from various monitoring sensors and systems on the engineering site, arranged by timestamps, containing the observed values of all relevant geological and engineering status variables for a period of time before and after the actual execution action. "True reward value" refers to the reward value calculated using the same reward function as in simulation training, based on the results of the actual execution action in the real environment, quantifying the immediate positive or negative impact of the actual action in the real world.
[0124] The terminal obtains confirmed execution records of actual actions from the project management execution end, retrieves real-time status logs covering key periods before and after the execution of the action from the monitoring database, extracts the data required for calculation based on the changes in system status after the decision is executed, substitutes the data into the defined reward function formula for calculation, and obtains the real reward value reflecting the actual effect of the decision.
[0125] Step 602: Based on the decision moment, integrate the complete state vector, the actual action executed, the new state, and the real reward value to obtain the online experience quadruple.
[0126] Specifically, the decision moment refers to the specific point in time when the reinforcement learning policy network makes a decision and outputs a draft action. The complete state vector refers to the state vector input into the reinforcement learning policy network at the decision moment, representing the complete state of the system at that time. The new state refers to the new stable state after the actual action is performed; this state is reflected in subsequent real-time state logs, typically the state vector at the next evaluation moment after the action is executed. The online experience quadruple refers to the standard data structure used in the reinforcement learning framework to describe a complete interactive experience. A quadruple contains information about the agent being in a certain state, taking a certain action, transitioning to a new state, and receiving a reward. In this embodiment, it specifically refers to an experience record integrated from the complete state vector, the actual action performed, the new state, and the actual reward value.
[0127] The terminal takes the decision moment as the alignment benchmark and starting point, takes the complete state vector corresponding to the decision moment, takes the actual action executed after the decision at that moment, takes the new state after the action is executed, and takes the real reward value calculated based on the actual result. The four pieces of information are combined and encapsulated in a fixed order to obtain the online experience quadruple.
[0128] Step 603: Based on the priority strategy, prioritize the online experience quadruple and the historical experience pool to obtain the online experience pool.
[0129] Specifically, a priority strategy refers to a rule or algorithm that assigns different levels of learning importance to different experience data. The core idea is to assign higher sampling probabilities to experiences with higher learning value. The historical experience pool refers to the collection of all experience quadruples collected and stored in previous stages. The online experience pool refers to a data storage pool of experience data that has been prioritized and reorganized, containing some experiences from the historical experience pool and the latest generated online experience quadruples, with each experience in the pool assigned a corresponding sampling weight according to the priority strategy.
[0130] The terminal calculates the priority of online experience quadruples according to the priority strategy. Optionally, a common strategy is based on the time difference error, that is, the difference between the agent's predicted reward and the actual reward for the experience. The larger the difference, the higher the priority is usually. For the existing experience in the historical experience pool, the priority is also recalculated or adjusted according to the new information. The online experience quadruples with priority weights are merged with the historical experience pool. In the merged experience set, a weighted sampling distribution is designed according to the priority of each experience to obtain an online experience pool with priority information that can be used for training and fine-tuning.
[0131] Step 604: Based on the online experience pool, fine-tune the parameters of the reinforcement learning policy network to obtain updated policy network parameters; the updated policy network parameters are used to update the reinforcement learning policy network.
[0132] Parameter fine-tuning refers to making minor adjustments and optimizations to the model parameters based on existing model parameters using a new, typically smaller, dataset to adapt to the new data or distribution, without completely altering the knowledge already learned by the original model. Updating policy network parameters refers to the new set of parameter values formed by the weights and biases within the reinforcement learning policy network after the parameter fine-tuning process.
[0133] The terminal samples a batch of experience data from the online experience pool according to priority weights. Based on the sampled batch of experience data, it uses the same reinforcement learning algorithm as the initial training to perform one or more rounds of gradient updates on the reinforcement learning policy network currently in use online. The characteristics are that the learning rate is usually low and the update amplitude is small, which aims to make the policy slowly adapt to the feedback of the real environment without forgetting the good policies that have been mastered. After the fine-tuning process converges or reaches a specified number of rounds, the parameters inside the policy network are updated to obtain a set of updated policy network parameters.
[0134] This embodiment utilizes the online learning and continuous evolution of the intelligent decision-making system. By leveraging the online experience pool collected from real engineering feedback, the parameters of the deployed strategy network are fine-tuned, enabling the strategy to adapt to the differences between the simulation and real environments. Furthermore, it learns from actual operation and maintenance, continuously improving its decision-making performance in the real world. The updated strategy network parameters are then used to refresh the online system, thereby improving the accuracy of geological risk identification in engineering design.
[0135] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0136] Based on the same inventive concept, this application also provides an engineering design geological risk auxiliary identification system for implementing the above-mentioned engineering design geological risk auxiliary identification method. The solution provided by this system is similar to the implementation scheme described in the above method; therefore, the specific limitations of one or more engineering design geological risk auxiliary identification system embodiments provided below can be found in the limitations of the engineering design geological risk auxiliary identification method described above, and will not be repeated here.
[0137] In one exemplary embodiment, such as Figure 2 As shown, an engineering design geological risk auxiliary identification system 700 is provided, comprising:
[0138] Knowledge module 701 is used to acquire multi-source geological exploration data and extract knowledge from the multi-source geological exploration data to obtain a spatiotemporal knowledge base for geological engineering; the spatiotemporal knowledge base for geological engineering includes historical data and real-time data;
[0139] Modeling module 702 is used to perform parametric hydrodynamic coupling modeling based on the geological engineering spatiotemporal knowledge base to obtain a high-fidelity geophysical model;
[0140] Risk module 703 is used to train a neural network model based on a high-fidelity geophysical model to obtain a geological proxy model, and input real-time monitoring data into the geological proxy model to perform risk probability prediction and obtain a risk probability map.
[0141] The early warning module 704 is used to perform adaptive early warning based on the risk probability map and obtain a risk control action instruction set; the risk control action instruction set is used to assist in engineering design and construction.
[0142] Furthermore, risk module 703 is also used for:
[0143] Based on the input parameter range of the high-fidelity geophysical model, parameter sampling is performed to obtain the input parameter combination, and the input parameter combination is matched with each target variable to obtain the list of tasks to be simulated;
[0144] Based on the list of tasks to be simulated, the input parameter combinations are input into the high-fidelity geophysical model for numerical simulation, the output results are obtained, and the target variables corresponding to the input parameter combinations are extracted from the output results to obtain input-output paired data.
[0145] By using the combination of input parameters as features and the target variable as a label, the input and output paired data are divided to obtain the training dataset;
[0146] Based on the training dataset, the neural network model is trained to obtain the geological proxy model; the number of input layer nodes of the neural network model is the number of parameters in the combination of input parameters, and the number of output layer nodes of the neural network model is the number of target variables.
[0147] Furthermore, the early warning module 704 is also used for:
[0148] Quantitative features are extracted from the risk probability map, and the quantitative features and real-time data are encoded to obtain a state vector;
[0149] The state vector is input into the reinforcement learning policy network to make action decisions, thus obtaining an action draft.
[0150] Based on the hard rule filter, the action draft is matched with the hard rules to obtain the early warning and control actions;
[0151] Based on natural language generation, early warning and control actions are mapped to corresponding recommended decisions, and the early warning and control actions are input into the geological proxy model for rapid simulation to obtain short-term simulation comparison charts;
[0152] By integrating the recommended decision-making and short-term simulation comparison charts, a set of risk management action instructions is obtained.
[0153] Furthermore, the early warning module 704 is also used for:
[0154] Based on the geological agent model, the random geological response under different decision-making interventions is simulated, and the non-geological state dimension update caused by the decision-making action is configured based on the engineering logic rules to obtain the simulation environment; the non-geological state dimension includes the engineering progress dimension, resource constraint dimension, environmental time sequence dimension, and control state dimension.
[0155] Based on engineering quantization rules, the state space, action space, and reward function are generated in a quantized manner. Based on the state space, action space, and reward function, the policy network architecture is designed to obtain the policy network and value network.
[0156] Based on near-end policy optimization, the policy network and value network are trained and updated in a simulation environment to obtain the trained policy network parameters.
[0157] Based on the test environment set, the performance of the trained policy network parameters is tested to obtain a policy performance evaluation report.
[0158] Based on the policy performance evaluation report, the parameters of the trained policy network are set to those of a reinforcement learning policy network.
[0159] Furthermore, the formula for the reward function is:
[0160]
[0161] in, Let t be the reward function, and t be the time step. For safety weighting coefficients, This is the cost weighting coefficient. This is the schedule weighting coefficient. As a safety reward, As a cost penalty, As a progress reward, As a supplementary reward;
[0162] The safety reward is obtained by comparing the indicators of random geological response with a set of safety thresholds.
[0163] Furthermore, the system also includes an update module for:
[0164] Obtain the actual executed actions and real-time status logs, and calculate the actual reward value based on the real-time status logs and the reward function;
[0165] Based on the decision moment, the complete state vector, the actual action executed, the new state, and the real reward value are integrated to obtain the online experience quadruple;
[0166] Based on a priority strategy, the online experience quadruple and the historical experience pool are prioritized to obtain the online experience pool.
[0167] Based on the online experience pool, the parameters of the reinforcement learning policy network are fine-tuned to obtain the updated policy network parameters; the updated policy network parameters are used to update the reinforcement learning policy network.
[0168] In one embodiment, a computer device is provided, including a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the steps of the engineering design geological risk auxiliary identification method as described above.
[0169] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in the above method embodiments.
[0170] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The components described as separate parts may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this disclosure according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0171] The above-described embodiments are merely illustrative of several implementation methods of the embodiments of this application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of the patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the embodiments of this application, and these modifications and improvements all fall within the protection scope of the embodiments of this application.
Claims
1. A method for auxiliary identification of geological risks in engineering design, characterized in that, The method includes: Multi-source geological exploration data is acquired, and knowledge is extracted from the multi-source geological exploration data to obtain a spatiotemporal knowledge base for geological engineering; the spatiotemporal knowledge base for geological engineering includes historical data and real-time data; Based on the aforementioned geological engineering spatiotemporal knowledge base, a parametric hydrodynamic coupling model is constructed to obtain a high-fidelity geological physical model. Based on the aforementioned high-fidelity geophysical model, a neural network model is trained to obtain a geological proxy model. Real-time monitoring data is then input into the geological proxy model to perform risk probability prediction and obtain a risk probability map. Based on the risk probability map, adaptive early warning is performed to obtain a risk control action instruction set; the risk control action instruction set is used to assist in engineering design and construction.
2. The method according to claim 1, characterized in that, The process of training a neural network model based on the fidelity geophysical model to obtain a geological proxy model includes: Based on the input parameter range of the high-fidelity geophysical model, parameter sampling is performed to obtain input parameter combinations, and the input parameter combinations are matched with each target variable to obtain a list of tasks to be simulated; Based on the list of tasks to be simulated, the combination of input parameters is input into the high-fidelity geophysical model to perform numerical simulation, and the output results are obtained. The target variable corresponding to the combination of input parameters is extracted from the output results to obtain input-output paired data. Using the input parameter combination as features and the target variable as a label, the input-output paired data is divided to obtain a training dataset; Based on the training dataset, the neural network model is trained to obtain the geological proxy model; the number of input layer nodes of the neural network model is the number of parameters in the input parameter combination, and the number of output layer nodes of the neural network model is the number of target variables.
3. The method according to claim 1, characterized in that, The adaptive early warning based on the risk probability map yields a risk control action instruction set, including: Quantitative features are extracted from the risk probability map, and the quantitative features and the real-time data are encoded to obtain a state vector; The state vector is input into a reinforcement learning policy network to make action decisions, thereby obtaining an action draft. Based on the hard rule filter, the action draft is matched with the hard rules to obtain the early warning and control actions; Based on natural language generation, the early warning and control actions are mapped to corresponding recommended decisions, and the early warning and control actions are input into the geological proxy model for rapid simulation to obtain a short-term simulation comparison chart. The recommended decision and the short-term simulation comparison chart are integrated to obtain the risk management action instruction set.
4. The method according to claim 3, characterized in that, The reinforcement learning policy network was obtained through the following method: Based on the geological proxy model, the random geological response under different decision-making actions is simulated, and based on engineering logic rules, the non-geological state dimension update caused by the decision-making actions is configured to obtain the simulation environment; the non-geological state dimension includes engineering progress dimension, resource constraint dimension, environmental time sequence dimension and control state dimension. Based on engineering quantization rules, a state space, action space, and reward function are generated. Based on the state space, action space, and reward function, a policy network architecture is designed to obtain a policy network and a value network. Based on near-end policy optimization, the policy network and the value network are trained and updated in the simulation environment to obtain the trained policy network parameters. Based on the test environment set, the performance of the trained policy network parameters is tested to obtain a policy performance evaluation report. Based on the policy performance evaluation report, the parameters of the trained policy network are set as those of the reinforcement learning policy network.
5. The method according to claim 4, characterized in that, The formula for the reward function is: in, Let t be the reward function, and t be the time step. For safety weighting coefficients, This is the cost weighting coefficient. This is the schedule weighting coefficient. As a safety reward, As a cost penalty, As a progress reward, As a supplementary reward; The safety reward is obtained by comparing the indicators of the random geological response with a set of safety thresholds.
6. The method according to claim 3, characterized in that, After integrating the recommended decision and the short-term projection comparison chart to obtain the risk management action instruction set, the method further includes: Obtain the actual executed actions and real-time status logs, and calculate the actual reward value based on the real-time status logs and the reward function; Based on the decision moment, the complete state vector, the actual action executed, the new state, and the real reward value are integrated to obtain the online experience quadruple; Based on a priority strategy, the online experience quadruple and the historical experience pool are prioritized to obtain the online experience pool. Based on the online experience pool, the parameters of the reinforcement learning policy network are fine-tuned to obtain updated policy network parameters; the updated policy network parameters are used to update the reinforcement learning policy network.
7. An engineering design geological risk auxiliary identification system, characterized in that, The system includes: The knowledge module is used to acquire multi-source geological exploration data and extract knowledge from the multi-source geological exploration data to obtain a spatiotemporal knowledge base for geological engineering; the spatiotemporal knowledge base for geological engineering includes historical data and real-time data; The modeling module is used to perform parametric hydrodynamic coupling modeling based on the aforementioned geological engineering spatiotemporal knowledge base to obtain a high-fidelity geophysical model. The risk module is used to train the neural network model based on the high-fidelity geophysical model to obtain a geological proxy model, and input real-time monitoring data into the geological proxy model to perform risk probability prediction and obtain a risk probability map. The early warning module is used to perform adaptive early warning based on the risk probability map and obtain a risk control action instruction set; the risk control action instruction set is used to assist in engineering design and construction.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.