Self-adaptive anchor rod group support control system and method based on deep learning
Through the adaptive anchor group support control method based on deep learning, using real-time sensor data acquisition and CNN/LSTM network, adaptive collaborative decision-making of anchor nodes is realized, which solves the problem that traditional anchor support systems cannot adapt to complex geological conditions in real time and improves the intelligence and stability of the support system.
Patent Information
- Application Number
- CN202511053712.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-09-12
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional anchor support systems are difficult to adapt to complex and changeable geological conditions and engineering environment changes in real time. The existing anchor group control method lacks an effective game mechanism, resulting in conflicts between local optimization and global optimization, making it difficult to meet real-time and intelligent requirements.
An adaptive anchor group support control method based on deep learning is adopted. Data is collected in real time through sensors, anchor nodes are defined as game participants, CNN and LSTM networks are used to extract geological features, generate strategy and value networks, realize adaptive collaborative decision-making, and dynamically update network parameters.
It realizes real-time adaptive adjustment of anchor group support, improves the intelligence level and support effect, can quickly respond to environmental changes, and avoids the risk of failure caused by sudden changes in geological conditions.
Smart Images

Figure CN120630725A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of underground cavern support in geotechnical engineering, and in particular to an adaptive anchor group support control system and method based on deep learning. Background Art
[0002] In geotechnical engineering projects such as mining and tunnel construction, anchor support is an important means to ensure project safety. Traditional anchor support systems usually adopt fixed control strategies, which are difficult to adapt to complex and changeable geological conditions and engineering environment changes in real time. With the continuous expansion of project scale and the improvement of safety requirements, the shortcomings of traditional methods in support effect and intelligence are becoming increasingly prominent. Some existing anchor group control methods rely on central servers for global decision-making, which is difficult to meet real-time requirements. Traditional models (such as PID control) are difficult to cope with the dynamic changes and nonlinear characteristics of geological conditions. There is a lack of effective game mechanism between anchor nodes, which easily leads to conflicts between local optimization and global optimality. Summary of the Invention
[0003] The purpose of the present invention is to solve the above problems and to design an adaptive anchor group support control system and method based on deep learning.
[0004] A first aspect of the present invention provides an adaptive anchor group support control method based on deep learning, the method comprising the following steps:
[0005] The original parameter data of the anchor node is collected in real time by the sensor, and the original parameter data is preprocessed to obtain the preprocessed parameter data;
[0006] Define anchor nodes as game participants, and use reward functions to make anchor nodes converge to the global optimum in the game;
[0007] Adopting an independent Q-learning framework, a collaborative decision-making model for anchor groups is deployed for each anchor node. A CNN network is used to extract geological features, and combined with an LSTM network to model temporal dynamics, generating a policy network and a value network.
[0008] Each anchor node generates control actions through the policy network based on real-time status and game rewards, and dynamically updates the policy network parameters and value network parameters.
[0009] Optionally, in a first implementation of the first aspect of the present invention, the real-time acquisition of raw parameter data of the anchor node by a sensor and preprocessing the raw parameter data to obtain preprocessed parameter data include:
[0010] The 3σ rule combined with the isolation forest algorithm was used to remove abnormal data points that deviated from the mean by 3 times the standard deviation from the original parameter data;
[0011] Apply a 5th-order Butterworth low-pass filter to the time series data in the original parameter data to remove vibration noise;
[0012] The original parameter data after outlier detection and spatiotemporal filtering are normalized to obtain preprocessed parameter data.
[0013] Optionally, in a second implementation of the first aspect of the present invention, defining anchor nodes as game participants and using a reward function to make the anchor nodes converge to a global optimum in the game includes:
[0014] Each anchor node is abstracted as an intelligent agent, and an N-node non-cooperative game model is constructed. The strategy space contains 8 control actions, including 4 levels of preload adjustment and 4 levels of angle adjustment.
[0015] A dual-objective optimization function is defined, including minimizing support costs and maximizing system stability. Based on the reward function, the long-term return is calculated by exponential moving average, so that the anchor nodes converge to the global optimum in the game. The reward function is a combination of immediate reward and long-term return.
[0016] Optionally, in a third implementation of the first aspect of the present invention, the CNN network extracts rock mass structural features through two convolutional layers for the static geological parameters in the preprocessed parameter data, and outputs a spatial feature vector;
[0017] The LSTM network captures the time dependency of the support state through a 128-unit LSTM layer for the time series dynamic parameters in the preprocessed parameter data, and outputs a time series feature vector.
[0018] Optionally, in a fourth implementation of the first aspect of the present invention, the policy network integrates the spatial feature vector extracted by the CNN network with the temporal feature vector output by the LSTM network, and uses a three-layer fully connected network to generate a probability distribution of each action;
[0019] The value network uses a two-layer fully connected network to evaluate the potential value of the current node state and output the state value estimate.
[0020] Optionally, in a fifth implementation of the first aspect of the present invention, each anchor node generates a control action through a policy network based on the real-time state and the game reward, including:
[0021] Each anchor node generates control actions through the policy network based on real-time status and game rewards. After executing the action, the anchor node updates the local experience pool according to the newly observed geological conditions and optimizes the policy network and value network parameters through the gradient descent method.
[0022] The model parameters of each anchor node are fused through weighted averaging to generate a global model, which is then sent to each anchor node.
[0023] Optionally, in a sixth implementation of the first aspect of the present invention, optimizing the policy network and value network parameters by a gradient descent method includes:
[0024] After each action is executed, the error signal is calculated based on the actual reward obtained, and the error signal is reversed from the output layer of the anchor group collaborative decision-making model to the input layer. The contribution of each parameter to the error is determined layer by layer, and the parameters are adjusted according to the direction and gradient of the error.
[0025] A second aspect of the present invention provides an adaptive anchor group support control system based on deep learning, the system comprising:
[0026] The acquisition module is used to collect the original parameter data of the anchor node in real time through the sensor, pre-process the original parameter data, and obtain pre-processed parameter data;
[0027] The convergence module is used to define anchor nodes as game participants and use reward functions to make anchor nodes converge to the global optimum in the game;
[0028] The deployment module is used to deploy an anchor group collaborative decision-making model for each anchor node using an independent Q-learning framework. It uses a CNN network to extract geological features and combines it with an LSTM network to model temporal dynamics and generate a policy network and a value network.
[0029] The generation module is used for each anchor node to generate control actions through the policy network based on real-time status and game rewards, and dynamically update the policy network parameters and value network parameters.
[0030] The third aspect of the present invention provides an adaptive anchor group support control device based on deep learning, and the adaptive anchor group support control device based on deep learning includes a memory and at least one processor, and instructions are stored in the memory; the at least one processor calls the instructions in the memory to enable the adaptive anchor group support control device based on deep learning to perform each step of the adaptive anchor group support control method based on deep learning as described in any one of the above items.
[0031] The fourth aspect of the present invention provides a computer-readable storage medium having instructions stored thereon, which, when executed by a processor, implement the various steps of the adaptive anchor group support control method based on deep learning as described in any of the above items.
[0032] In the technical solution provided by the present invention, the original parameter data of the anchor nodes are collected in real time by sensors, and the original parameter data are preprocessed to obtain the preprocessed parameter data; the anchor nodes are defined as game participants, and the reward function is used to make the anchor nodes converge to the global optimum in the game; an independent Q learning framework is adopted to deploy an anchor group collaborative decision-making model for each anchor node, the CNN network is used to extract geological features, and the LSTM network is combined to model the time series dynamics to generate a policy network and a value network; each anchor node generates a control action through the policy network based on the real-time state and game reward, and dynamically updates the policy network parameters and the value network parameters; the present invention realizes the joint modeling of the spatial structure characteristics of the rock mass and the time series dynamics of the anchor support, improves the state prediction accuracy, and enables each anchor node to work together through a collaborative decision-making mechanism based on game theory to form a globally optimal support strategy with real-time adaptive adjustment capability, and can quickly respond to environmental changes, significantly improving the intelligence level and support effect of the anchor group support control, and has broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Various other advantages and benefits will become apparent to those skilled in the art by reading the following detailed description of the preferred embodiment.The accompanying drawings are only for the purpose of illustrating the preferred embodiment and are not to be considered as limiting the present invention.
[0034] Figure 1 A flowchart of an adaptive anchor group support control method based on deep learning provided by an embodiment of the present invention;
[0035] Figure 2 A schematic diagram of the structure of an adaptive anchor group support control system based on deep learning provided by an embodiment of the present invention;
[0036] Figure 3 A schematic structural diagram of a deep learning-based adaptive anchor group support control device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0037] The terms "first," "second," "third," "fourth," and the like (if any) in the description and claims of the present invention and in the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, apparatus, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to these processes, methods, products, or apparatus.
[0038] For ease of understanding, the specific process of the embodiment of the present invention is described below. Figure 1 The present invention provides a flowchart of a method for controlling an adaptive anchor group support based on deep learning, which specifically includes the following steps:
[0039] Step 101: collecting original parameter data of the anchor node in real time through a sensor, preprocessing the original parameter data to obtain preprocessed parameter data;
[0040] In this embodiment, multiple types of sensors are deployed at each anchor node, including a triaxial stress sensor, a fiber optic displacement sensor, a temperature and humidity sensor, and an acoustic rock detector, to collect 18-dimensional original parameters in real time. The geological parameters include rock compressive strength, elastic modulus, Poisson's ratio, ground stress component, water content, and joint density. The anchor status includes axial force, torque, pull-out displacement, and thread stress concentration factor. The environmental parameters include temperature, humidity, blasting vibration acceleration, and surrounding rock acoustic emission signal intensity.
[0041] In this embodiment, the 3σ rule is combined with the isolation forest algorithm to eliminate abnormal data points that deviate from the mean by 3 times the standard deviation in the original parameter data; a 5th-order Butterworth low-pass filter is applied to the time series data in the original parameter data to remove vibration noise; and the original parameter data after outlier detection and spatiotemporal filtering are normalized to obtain preprocessed parameter data.
[0042] In this embodiment, based on the 3σ rule, the mean and standard deviation of the data series are calculated for parameters that conform to normal distribution characteristics, such as anchor axial force and rock moisture content. Data points that deviate from the mean by more than three standard deviations are initially marked as anomalies. Abnormal data generated in non-normal distribution scenarios, such as sudden changes in ground stress and blasting vibration impact, are further detected using the isolation forest algorithm. This algorithm evaluates the degree of isolation of data points by constructing a random binary tree. It can effectively identify sparsely distributed anomalies, such as jump data caused by transient sensor failures. The two-layer detection mechanism complements each other, covering both significant anomalies under normal operating conditions and hidden anomalies under complex disturbances.
[0043] A 5th-order Butterworth low-pass filter is used to suppress spatiotemporal noise for time-series dynamic parameters such as anchor axial force, displacement, and vibration acceleration. This filter has a flat passband response characteristic and can accurately filter out short-term transient disturbances such as blasting vibration and equipment electromagnetic interference, while retaining slowly changing effective signals such as surrounding rock deformation and stress relaxation. Taking anchor displacement monitoring as an example, when underground blasting causes vibration and high-frequency oscillations in the displacement data, the filter recursively calculates the signal weighted values at the current moment and the historical moment, dynamically smoothes the noise waveform, and makes the displacement curve clearly reflect the true deformation trend of the surrounding rock, avoiding misjudgment of the control strategy due to noise interference.
[0044] Step 102: define anchor nodes as game participants, and use a reward function to make the anchor nodes converge to the global optimum in the game;
[0045] In this embodiment, each anchor node is abstracted as an intelligent agent, and an N-node non-cooperative game model is constructed. The strategy space contains 8 control actions, among which the control actions include 4-level preload adjustment and 4-level angle adjustment. A dual-objective optimization function is defined, including minimizing the support cost and maximizing the system stability. Based on the reward function, the long-term return is calculated by exponential moving average, so that the anchor node converges to the global optimum in the game, and the reward function adopts the form of combining immediate reward and long-term return.
[0046] In this embodiment, each anchor node in the support system is defined as an intelligent agent with autonomous decision-making capabilities. Each intelligent agent can perceive local geological parameters, its own stress state, and its own stress state in real time, and construct an N-node non-cooperative game model. Each node takes maximizing its own local interests as the initial goal when making decisions. However, through mechanical coupling effects and communication interactions, its strategy selection will inevitably affect the stability of the global system. The strategy space of each node contains 8 basic control actions, which are specifically divided into 4 levels of preload adjustment, corresponding to different support strengths, such as low, medium, high, and ultra-high preload, and 4 levels of angle adjustment to adapt to the direction of rock joints, such as 0°, 15°, 30°, and 45° support angles, covering common support demand scenarios in engineering. By executing different action combinations, the node needs to implicitly consider the impact on the load distribution of adjacent nodes and the global stress balance while meeting its own stress safety.
[0047] Minimizing support costs includes hardware execution costs such as motor energy consumption for preload adjustment and hydraulic loss for angle adjustment, as well as collaborative redundancy costs such as energy consumption caused by repeated adjustments of adjacent nodes due to policy conflicts and additional equipment wear caused by stress concentration. By quantifying the energy consumption parameters and equipment loss coefficients of each action, the local costs of individual nodes are accumulated into the global total cost to avoid resource waste caused by excessive adjustment. Maximizing system stability is measured by indicators such as the surrounding rock displacement fluctuation rate and stress distribution uniformity. The deformation and stress concentration of the rock mass around each node are monitored in real time through a sensor network. When multiple node actions form a collaborative support effect, the system stability index is significantly improved. The dual-objective coupling design forces the nodes to seek a balance between high-intensity support and economic adjustment, avoiding system imbalance caused by single-objective optimization.
[0048] Immediate rewards are calculated in real time based on the contribution of the current action to the dual objectives. For example, positive rewards are given for improving system stability, while negative penalties are triggered by increasing global costs. If the action causes excessive stress concentration in adjacent nodes or communication congestion, an additional conflict penalty is added. This immediate feedback mechanism encourages nodes to quickly identify invalid actions and inhibit short-sighted behavior. Historical reward data is weighted using the exponential moving average (EMA) algorithm, giving recent rewards a higher weight while retaining long-term trend memory. For example, if a node repeatedly selects a combination of medium preload and reasonable angle, although the single reward is not outstanding, the long-term accumulated global cost reduction and stability improvement effects will be converted into higher long-term returns through EMA, incentivizing nodes to adhere to such strategies with global optimization potential. Through this mechanism, the influence of local noise on the direction of strategy iteration can be avoided, ensuring that node decisions gradually converge to the Nash equilibrium state, that is, no single node can further reduce the global cost or improve stability by changing its strategy alone, thus achieving the overall optimization of the support system.
[0049] Step 103: Using an independent Q-learning framework, deploy an anchor group collaborative decision-making model for each anchor node, use a CNN network to extract geological features, and combine it with an LSTM network to model temporal dynamics to generate a policy network and a value network.
[0050] In this embodiment, the CNN network extracts rock structure characteristics through two convolutional layers for the static geological parameters in the preprocessed parameter data and outputs a spatial feature vector; the LSTM network captures the time dependency of the support status through a 128-unit LSTM layer for the temporal dynamic parameters in the preprocessed parameter data and outputs a temporal feature vector.
[0051] In this embodiment, the policy network integrates the spatial feature vector extracted by the CNN network with the temporal feature vector output by the LSTM network, and uses a three-layer fully connected network to generate the probability distribution of each action; the value network uses a two-layer fully connected network to evaluate the potential value of the current node state and output the state value estimate.
[0052] In this embodiment, the CNN network is used to automatically mine implicit correlation features within the rock mass for static spatial data such as rock mass strength, ground stress distribution, and joint density. For example, the ground stress component matrix is scanned by multi-layer convolution kernels to identify potential risk patterns in high stress concentration areas, or key spatial features of rock mass stability are extracted from joint density and strike data. For dynamic data that changes over time, such as anchor axial force, surrounding rock displacement, and ambient humidity, the LSTM network is used to capture long-term dependencies in the data. When the axial force is monitored to continue to rise for multiple consecutive cycles, the LSTM can predict the possibility of continuous deformation of the surrounding rock and trigger the preload adjustment strategy in advance to avoid support failure caused by delayed response.
[0053] After the output features are fused, they are input into the policy network and the value network respectively. The policy network generates the selection probability of each control action based on the current geological status and neighborhood collaboration information, and prioritizes the action combination that can reduce the global cost and improve the stability of the system; the value network serves as an evaluation system, quantitatively scores the potential value of the current node status, and feeds back to the policy network to optimize the decision-making direction. Each node interacts with the reward summary and status characteristics through local communication, and perceives the changes in neighborhood strategies while learning independently, forming an intelligent ecosystem with independent decision-making and collaborative optimization that does not rely on the center. In complex geological environments, it can not only quickly respond to local mutations, but also converge to the global optimal solution through strategic game.
[0054] Step 104: Each anchor node generates control actions through the policy network based on the real-time status and game rewards, and dynamically updates the policy network parameters and value network parameters.
[0055] In this embodiment, each anchor node generates a control action through a policy network based on the real-time status and game rewards. After executing the action, the anchor node updates the local experience pool according to the newly observed geological status, and optimizes the policy network and value network parameters through the gradient descent method; the model parameters of each anchor node are fused through weighted averaging to generate a global model, which is then distributed to each anchor node.
[0056] In this embodiment, during real-time operation, each anchor node uses sensors to obtain real-time state vectors, such as the current rock mass strength, its own axial force, and the state of neighboring nodes. These are input into the local policy network to generate an action probability distribution, and specific control actions are executed based on the ε-greedy strategy. After the action is executed, the node immediately collects new observations such as new geological parameters and displacement changes, calculates an immediate reward based on the game reward function, and stores the old state-action-reward-new state quaternary into a local experience pool. The experience pool uses a priority queue mechanism to assign higher weights to samples with high rewards / high errors, ensuring that key data is prioritized in model training.
[0057] When the experience pool data accumulates to a threshold, the node initiates a dual-network optimization process: batches of data are randomly extracted from the experience pool, and gradient descent updates are performed on the policy network and value network. The policy network adjusts the action probability distribution by maximizing the expected reward, and the value network improves state evaluation accuracy by minimizing the error between the predicted value and the actual reward. Parameter clipping and learning rate decay are used during the optimization process to ensure training stability.
[0058] After completing local optimization, each node uploads the updated model parameters to the edge server through a low-bandwidth communication link, such as the convolution kernel weights of CNN and the forget gate parameters of LSTM. The server performs weighted average fusion based on the historical reliability scores of the nodes to generate a globally shared consensus model. This global model contains general strategies verified by data from the entire network, such as standard preload adjustment rules for fault zones. It is then broadcast to all nodes. After receiving it, the nodes perform model fusion at a ratio of 90% global parameters and 10% local parameters. This not only retains the global optimal experience, but also allows fine-tuning of strategies based on local special geological conditions, forming a closed loop of distributed autonomous learning and global knowledge sharing.
[0059] In this embodiment, after each action is executed, the error signal is calculated based on the actual reward obtained, and the error signal is reversed from the output layer of the anchor group collaborative decision-making model to the input layer. The contribution of each parameter to the error is determined layer by layer, and the parameters are adjusted according to the direction and gradient of the error.
[0060] In this embodiment, each time the anchor node performs a control action, such as adjusting the preload or support angle, the deviation from the expected target is calculated based on the actual impact of the action on the global support cost and system stability, forming an error signal, which intuitively reflects the gap between the current decision and the ideal effect. After the error signal is generated, it will be reversely transmitted from the output layer of the anchor group collaborative decision-making model to the input layer. The output layer is the action selection result, and the input layer is the pre-processed geological and stress parameters. The responsibility of each network parameter for the final error is traced layer by layer; for example, if a certain angle adjustment action does not accurately capture the direction of the rock joints, resulting in stress concentration, the reverse transmission process will find If there is a deviation in the parameters of a certain set of convolution kernels responsible for joint feature extraction in the CNN network, or if the weight distribution of historical displacement data by the forget gate in the LSTM network is unreasonable, these key parameters will be marked as high-contribution error sources. According to the direction and gradient of the error, the parameters will be fine-tuned in a targeted manner. For parameters that cause positive errors, they will be adjusted in the direction of enhancing their effects, making them more likely to be activated in similar scenarios; for parameters that cause negative errors (rewards are lower than expected), their weights will be reduced to suppress the recurrence of invalid decisions, so that the decision-making ability of each node will continue to approach the global optimum in continuous trial and error, and ultimately achieve dynamic adaptation of support strategies to complex geological environments.
[0061] In this embodiment, the raw parameter data collected by the sensor in real time is first converted into a standardized feature vector in the interval [-1, 1] through a three-stage processing process of outlier detection, spatiotemporal filtering, and dynamic normalization. These preprocessed data serve as the state input of the intelligent agent in the game model, directly determining each anchor node's perception of the current working condition. For example, when the normalized axial force data exceeds 0.8 for three consecutive cycles, the node will perceive a high stress warning state, triggering a tendency to strengthen support in the subsequent game strategy. The preprocessing process not only improves data quality, but also transforms the complex signals of the physical world into a structured state space suitable for game theory and neural network processing through feature engineering.
[0062] The agent's policy space and dual-objective reward function, defined within the game theory framework, directly determine the output dimensions and training objectives of the neural network. For example, the number of neurons in the policy network's output layer matches the dimensions of the action space: eight neurons correspond to the selection probabilities of eight actions. The output of the value network must be mapped to the global optimization objective in the reward function, such as converting stability scores and cost consumption into value scalars in the 0-1 range. Furthermore, the non-cooperative game assumptions in game theory drive the model's design towards a distributed architecture: each node's CNN-LSTM network processes only locally preprocessed data and obtains reward summaries from neighboring nodes through local communication, forming a lightweight collaborative mechanism between the local model and the global objective, avoiding centralized computing bottlenecks.
[0063] The strategy network of each node outputs an action probability distribution based on the current local state. After the action is converted into actual support adjustment by the actuator, it will generate two kinds of feedback:
[0064] Instant physical feedback: Sensors collect data such as the bolt axial force and surrounding rock displacement after the action, which serves as the new state input for the next cycle, forming a closed loop of perception-decision-execution-reperception.
[0065] Reward signal feedback calculates an immediate reward value based on the actual impact of an action on global cost and system stability. For example, improving stability by +0.5 points increases cost by -0.3 points. This reward value is used to update the local Q table and is shared through neighborhood communication to influence the strategy evaluation of neighboring nodes. For example, if a neighboring node detects that a node has received a high reward due to reasonable angle adjustment, it will increase the selection weight of similar actions when generating its own strategy.
[0066] After the node collects the state-action-reward-new state quadruple data, the gradient descent method is used to fine-tune the parameters of the policy network and the value network. If the policy network obtains a high reward for an action, such as reducing the global cost through collaborative adjustment, the weight of the output neuron corresponding to the action will be enhanced, so that the probability of choosing this action in similar states in the future will be increased; if the value network's stability score for a certain state deviates from the actual reward, such as underestimating the risk under low preload, the network parameters will be adjusted to make subsequent evaluations closer to the actual effect.
[0067] In this embodiment, in terms of anchor rod structure design, an integrated molding process of a hollow threaded steel rod body and a built-in optical fiber channel is adopted to ensure the survival of the sensor during grouting construction; a collaborative decision-making model for anchor rod groups based on game theory is constructed, and local node calculations replace central server decisions, reducing the response delay to less than 0.5 seconds; the hydraulic system innovatively adopts a modular oil circuit design, and each hydraulic pump can independently control a 3×3 anchor rod matrix to achieve millisecond-level precise adjustment of MPa-level tensioning force; it breaks through the technical bottleneck of traditional anchor rod support that cannot adapt dynamically, and is particularly suitable for areas with complex geological conditions such as fault fracture zones, providing an intelligent solution for the safety management of underground caverns throughout the life cycle of hydropower stations.
[0068] In this embodiment, traditional anchor support relies on static design parameters, while this solution uses CNN to extract geological features and LSTM to model time series dynamics, so that anchor nodes can perceive environmental changes in real time and dynamically adjust support strategies to avoid failure risks caused by sudden changes in geological conditions; the reward function design based on game theory enables anchor nodes to achieve collaborative optimization while making independent decisions, overcoming the contradiction between local reinforcement and overall stability in traditional support; traditional support requires manual inspection and delayed maintenance, while this solution uses real-time feedback of sensor data and rapid iteration of the Q learning strategy network to automatically trigger reinforcement actions at the early stage of surrounding rock deformation, so that the support system maintains long-term stability in complex stress environments and reduces downtime caused by periodic maintenance.
[0069] See also Figure 2 , a schematic structural diagram of an adaptive anchor group support control system based on deep learning provided by an embodiment of the present invention, the system includes:
[0070] The acquisition module is used to collect the original parameter data of the anchor node in real time through the sensor, pre-process the original parameter data, and obtain pre-processed parameter data;
[0071] The convergence module is used to define anchor nodes as game participants and use reward functions to make anchor nodes converge to the global optimum in the game;
[0072] The deployment module is used to deploy an anchor group collaborative decision-making model for each anchor node using an independent Q-learning framework. It uses a CNN network to extract geological features and combines it with an LSTM network to model temporal dynamics and generate a policy network and a value network.
[0073] The generation module is used for each anchor node to generate control actions through the policy network based on real-time status and game rewards, and dynamically update the policy network parameters and value network parameters.
[0074] In this embodiment, the system comprises a custom prestressed anchor, an edge computing module, and a hydraulic linkage control system. An axial fiber channel runs through the anchor body, housing a fiber grating sensor array for distributed monitoring of the anchor's axial force and surrounding rock deformation. An innovative edge computing unit is integrated into the bolt's tail. This unit uses a distributed decision-making algorithm to analyze sensor data in real time. When the strain difference between adjacent anchors exceeds a set threshold, the hydraulic control system automatically triggers tension compensation for the target anchor.
[0075] Figure 3: This is a structural diagram of an adaptive anchor group support control device based on deep learning provided by an embodiment of the present invention. The adaptive anchor group support control device 300 based on deep learning may have relatively large differences due to different configurations or performances, and may include one or more processors (central processing units, CPU) 310 (for example, one or more processors) and a memory 320, and one or more storage media 330 (for example, one or more mass storage devices) storing application programs 333 or data 332. Among them, the memory 320 and the storage medium 330 can be short-term storage or persistent storage. The program stored in the storage medium 330 may include one or more modules (not shown in the figure), each module may include a series of instruction operations in the adaptive anchor group support control device 300 based on deep learning. Furthermore, the processor 310 can be configured to communicate with the storage medium 330, and execute a series of instruction operations in the storage medium 330 on the adaptive anchor group support control device 300 based on deep learning to implement the method provided by the above embodiment.
[0076] The deep learning-based adaptive anchor group support control device 300 may further include one or more power supplies 340, one or more wired or wireless network interfaces 350, one or more input and output interfaces 360, and / or one or more operating devices 331, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, etc. It will be understood by those skilled in the art that Figure 3 The structure of the deep learning-based adaptive anchor group support control device shown does not constitute a limitation on the computer device provided by the present invention, and may include more or fewer components than shown in the figure, or a combination of certain components, or a different arrangement of components.
[0077] The present invention also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. The computer-readable storage medium stores instructions. When the instructions are executed on a computer, the computer executes the various steps of the adaptive anchor group support control method based on deep learning provided in the above embodiments.
[0078] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described equipment, devices, and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0079] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, etc., various media that can store program code.
[0080] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions are merely preferred examples of the present invention and are not intended to limit the present invention. Various changes and improvements may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and improvements fall within the scope of the present invention. The scope of protection claimed in the present invention is defined by the appended claims and their equivalents.
Claims
1. An adaptive anchor group support control method based on deep learning, characterized in that: The method comprises the following steps: The original parameter data of the anchor node is collected in real time by the sensor, and the original parameter data is preprocessed to obtain the preprocessed parameter data; Define anchor nodes as game participants, and use reward functions to make anchor nodes converge to the global optimum in the game; Adopting an independent Q-learning framework, a collaborative decision-making model for anchor groups is deployed for each anchor node. A CNN network is used to extract geological features, and combined with an LSTM network to model temporal dynamics, generating a policy network and a value network. Each anchor node generates control actions through the policy network based on real-time status and game rewards, and dynamically updates the policy network parameters and value network parameters.
2. The deep learning-based adaptive anchor group support control method according to claim 1, characterized in that: The method of collecting the original parameter data of the anchor node in real time through the sensor and preprocessing the original parameter data to obtain the preprocessed parameter data includes: The 3σ rule combined with the isolation forest algorithm was used to remove abnormal data points that deviated from the mean by 3 times the standard deviation from the original parameter data; Apply a 5th-order Butterworth low-pass filter to the time series data in the original parameter data to remove vibration noise; The original parameter data after outlier detection and spatiotemporal filtering are normalized to obtain preprocessed parameter data.
3. The deep learning-based adaptive anchor group support control method according to claim 1, characterized in that: The anchor nodes are defined as game participants, and a reward function is used to make the anchor nodes converge to the global optimum in the game, including: Each anchor node is abstracted as an intelligent agent, and an N-node non-cooperative game model is constructed. The strategy space contains 8 control actions, including 4 levels of preload adjustment and 4 levels of angle adjustment. A dual-objective optimization function is defined, including minimizing support costs and maximizing system stability. Based on the reward function, the long-term return is calculated by exponential moving average, so that the anchor nodes converge to the global optimum in the game. The reward function is a combination of immediate reward and long-term return.
4. The method for adaptive anchor group support control based on deep learning according to claim 1, characterized in that: The CNN network extracts rock mass structural features through two convolutional layers for the static geological parameters in the preprocessed parameter data and outputs a spatial feature vector; The LSTM network captures the time dependency of the support state through a 128-unit LSTM layer for the time series dynamic parameters in the preprocessed parameter data, and outputs a time series feature vector.
5. The method for adaptive anchor group support control based on deep learning according to claim 4, characterized in that: The policy network integrates the spatial feature vector extracted by the CNN network with the temporal feature vector output by the LSTM network, and uses a three-layer fully connected network to generate the probability distribution of each action; The value network uses a two-layer fully connected network to evaluate the potential value of the current node state and output the state value estimate.
6. The method for adaptive anchor group support control based on deep learning according to claim 1, characterized in that: Each anchor node generates control actions through a policy network based on real-time status and game rewards, including: Each anchor node generates control actions through the policy network based on real-time status and game rewards. After executing the action, the anchor node updates the local experience pool according to the newly observed geological conditions and optimizes the policy network and value network parameters through the gradient descent method. The model parameters of each anchor node are fused through weighted averaging to generate a global model, which is then sent to each anchor node.
7. The deep learning-based adaptive anchor group support control method according to claim 6, characterized in that: The optimization of the policy network and value network parameters by the gradient descent method includes: After each action is executed, the error signal is calculated based on the actual reward obtained, and the error signal is reversed from the output layer of the anchor group collaborative decision-making model to the input layer. The contribution of each parameter to the error is determined layer by layer, and the parameters are adjusted according to the direction and gradient of the error.
8. An adaptive anchor group support control system based on deep learning, characterized in that: The system includes: The acquisition module is used to collect the original parameter data of the anchor node in real time through the sensor, pre-process the original parameter data, and obtain pre-processed parameter data; The convergence module is used to define anchor nodes as game participants and use reward functions to make anchor nodes converge to the global optimum in the game; The deployment module is used to deploy an anchor group collaborative decision-making model for each anchor node using an independent Q-learning framework. It uses a CNN network to extract geological features and combines it with an LSTM network to model temporal dynamics and generate a policy network and a value network. The generation module is used for each anchor node to generate control actions through the policy network based on real-time status and game rewards, and dynamically update the policy network parameters and value network parameters.
9. An adaptive anchor group support control device based on deep learning, characterized in that: The deep learning-based adaptive anchor group support control device includes a memory and at least one processor, wherein instructions are stored in the memory; the at least one processor calls the instructions in the memory so that the deep learning-based adaptive anchor group support control device executes each step of the deep learning-based adaptive anchor group support control method as described in any one of claims 1-7.
10. A computer-readable storage medium having instructions stored thereon, characterized in that: When the instructions are executed by the processor, the various steps of the adaptive anchor group support control method based on deep learning as described in any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Control method and device for advance support of hydraulic support and electronic equipment
CN121432954A