Deep ground network dynamic adaptive operation and maintenance method and system based on Stackelberg game

By constructing a three-level information integration architecture and a Stackelberg game model in deep Earth networks, and combining deep reinforcement learning, the problem of unified modeling and dynamic operation and maintenance of multi-source heterogeneous data in deep Earth networks is solved. This enables efficient and secure adaptive adjustment of operation and maintenance strategies, improving the real-time performance and security of deep Earth networks.

CN121887812APending Publication Date: 2026-04-17CHINA RAILWAY FIRST SURVEY & DESIGN INST GRP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA RAILWAY FIRST SURVEY & DESIGN INST GRP
Filing Date
2025-11-25
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing deep-earth network operation and maintenance technologies are insufficient to achieve unified modeling of multi-source heterogeneous data, real-time transmission under unstable link conditions, and security assurance in extreme environments such as high hydrostatic pressure, strong electromagnetic interference, and high humidity and temperature. Furthermore, they lack dynamic and adaptive operation and maintenance strategies.

Method used

A dynamic adaptive operation and maintenance method for deep ground networks based on Stackelberg game theory is constructed. Data preprocessing and standardized storage are performed through a three-level information integration architecture (edge ​​nodes, partition gateways, ground and cloud). The Stackelberg master-slave game model and deep reinforcement learning (DQN algorithm) are used to solve the strategy, realizing hierarchical decision-making and dynamic adjustment.

Benefits of technology

It enables efficient integration and unified management of multi-source heterogeneous data, improves the real-time and adaptive nature of operation and maintenance decisions, ensures the safe and stable operation of deep earth networks in high-risk scenarios, and reduces the misjudgment rate and operation and maintenance costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121887812A_ABST
    Figure CN121887812A_ABST
Patent Text Reader

Abstract

The invention relates to a deep ground network dynamic self-adaptive operation and maintenance method and system based on a Stackelberg game. The method comprises the steps that edge nodes are deployed on a deep ground working face, partition gateways are deployed in underground partitions, a cloud data warehouse is deployed on the ground, a three-level information integration framework is formed, and pressure, gas, actuators, communication and personnel positioning data are collected and stored in a layered mode; setting a ground control center and a partition gateway as leader nodes, setting deep ground terminal equipment as follower nodes, constructing a Stackelberg master-slave game model, determining a state and operation and maintenance action set, and establishing a revenue function; and mapping the state and the action into a reinforcement learning space, solving and updating an operation and maintenance strategy table by adopting a DQN algorithm under bandwidth and energy consumption constraints, generating a control instruction by an operation and maintenance platform, issuing the control instruction through a partition gateway, calling a local emergency strategy in a high-risk scene, and realizing dynamic self-adaptive operation and maintenance of the deep ground network. According to the invention, the real-time performance and reliability of deep ground network operation and maintenance can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent operation and maintenance technology for deep earth networks, specifically relating to a dynamic adaptive operation and maintenance method and system for deep earth networks based on Stackelberg game theory. Background Technology

[0002] Deep underground environments are characterized by their great depth, enclosed environment, variable geological structure, and weak external support. Integrated monitoring, communication, and control networks built in such scenarios are generally referred to as deep underground networks. Compared to surface-based industrial internet and conventional production control networks, deep underground networks are subjected to the combined effects of high hydrostatic pressure, strong electromagnetic interference, high humidity and temperature, and even blasting vibrations. This leads to easy attenuation of communication links and significantly faster wear and tear on terminal equipment such as sensors and actuators, resulting in higher demands on the real-time performance, reliability, and security of data during network maintenance. Current deep underground maintenance relies heavily on a combination of manual inspections and periodic data collection, which can maintain equipment operation to a certain extent. However, in the wide-area, three-dimensional environment of deep underground environments, manual methods struggle to identify risks occurring simultaneously at multiple points and to promptly handle data anomalies caused by sudden interference.

[0003] Currently, the information technology tools used to support deep-earth development and safe production mainly focus on three aspects. First, specialized sensing and communication equipment for extreme environments is constantly emerging, such as pressure, gas, and displacement sensors resistant to high pressure, and edge acquisition devices capable of buffering short-term link interruptions. These devices improve data acquisition capabilities in harsh environments, but they are costly and have limited applicability, mostly only suitable for one or a few extreme operating conditions. Second, the fusion and cleaning technology of multi-source monitoring data is maturing. Seismic, electrical, magnetic, and personnel positioning data can be aggregated on the same platform. Combined with deep learning noise reduction and feature extraction algorithms, monitoring accuracy has improved. However, many solutions still rely on centralized ground-based data warehouses, resulting in long transmission paths that cannot fully adapt to the unstable links, limited bandwidth, and dispersed node distribution characteristics of deep-earth scenarios. Third, remote sharing based on cloud platforms and secure links is being promoted. Some systems introduce trusted evidence storage methods to ensure data traceability. However, these platforms are mostly general-purpose solutions deployed on the ground or at wellheads, lacking operation and maintenance decision-making models that match the hierarchical structure of deep-earth equipment.

[0004] Several prominent technical bottlenecks remain in the actual operation of deep-earth networks. First, data sources are highly heterogeneous. The data formats output by sensors, actuators, communication modules, and personnel positioning terminals differ significantly, sampling frequencies are inconsistent, and much data is stored in unstructured formats such as logs and text. Existing integration mechanisms lack sufficient standardization support for this diverse and heterogeneous data, easily leading to data silos. Second, transmission and storage conditions are constrained by the deep-earth environment. Deep-earth links frequently experience attenuation, congestion, or short-term interruptions. If a single centralized upload mode is still used, data will accumulate at the edge, potentially leading to misjudgments in scenarios with strong interference, such as blasting. Third, there are clear coupling and hierarchical relationships between devices. A failure in one device can propagate outwards through coordinated control, triggering cascading failures. However, many existing operation and maintenance methods only optimize individual devices or subsystems independently, failing to model the ground control center, zone control nodes, and terminal devices as a hierarchical and constrained whole. Therefore, it is difficult to coordinate limited bandwidth, energy, and control commands in a timely manner, and to achieve a dynamic balance between security risks and operation and maintenance costs.

[0005] Furthermore, some research on intelligent operation and maintenance (O&M) has attempted to introduce game theory or reinforcement learning into industrial network resource allocation and scheduling to address the constraints of interests among multiple stakeholders. However, most of this work is geared towards terrestrial communication networks or conventional industrial IoT, assuming good network connectivity and abundant node resources. It rarely incorporates the unique risk grading, energy shortages, and timeliness of equipment command issuance specific to deep underground networks into the model. This results in existing methods often failing to truly reflect the hierarchical relationships between ground control centers, regional control nodes, and terminal equipment when migrated to deep underground networks, and also failing to quickly switch to higher-priority O&M strategies during high-risk periods. In summary, existing technologies require a unified modeling and scheduling approach for complex deep underground spaces, capable of simultaneously handling multi-source heterogeneous data, unstable links, distinct equipment hierarchies, and dynamically changing security risks, to support subsequent adaptive O&M methods and system design. Summary of the Invention

[0006] This invention provides a dynamic adaptive operation and maintenance method and system for deep underground networks based on Stackelberg game theory. It is used to realize the hierarchical modeling, dynamic solution and rapid deployment of network operation and maintenance strategies in complex deep underground environments, thereby solving the problem of high concentration of operation and maintenance decisions and difficulty in adapting to the status of multiple source devices in the existing technology.

[0007] To achieve the above objectives, the technical solution of the present invention is as follows:

[0008] In the first aspect, this invention provides a dynamic adaptive operation and maintenance method for deep networks based on Stackelberg game theory, comprising:

[0009] Data Acquisition and Preprocessing: Edge nodes are deployed in deep-ground work areas, equipment supports, and personnel activity areas; zone gateways are deployed in underground zones; and a cloud data warehouse is deployed on the ground, forming a three-tiered information integration architecture consisting of edge nodes, zone gateways, and the ground-based cloud. Data is collected from multiple sources, including pressure sensor data, gas sensor data, hydraulic support actuator data, communication module data, and personnel positioning terminal data. Edge nodes locally cache instantaneous data during link interruptions. Zone gateways compress data uploaded by edge nodes and forward it to the ground-based cloud. The ground-based cloud stores raw data, standardized equipment status data, and data marts divided according to operation and maintenance scenarios in a distributed, layered manner. Data exceeding equipment thresholds is marked and temporarily stored locally.

[0010] The Stackelberg master-slave game model is constructed as follows: The ground control center and underground zone gateways are designated as leader nodes for network operation and maintenance decisions, while various deep-ground terminal devices are designated as follower nodes, establishing a hierarchical control relationship between leaders and followers. The leader's state variables are determined based on the regional risk level and remaining regional energy quantity recorded in the data warehouse, while the follower's state variables are determined based on the terminal device's operating status, data transmission rate, and energy consumption. The regional control commands that the leader can output, as well as the terminal device's executable acquisition frequency, transmission power, communication bandwidth, and location reporting cycle operation and maintenance parameters, are set as the leader's action set and the follower's action set. Leader and follower reward functions are established separately.

[0011] Operation and maintenance strategy solution and update: The leader state, follower state and their action set are mapped to the state space and action space of reinforcement learning. The DQN algorithm in deep reinforcement learning is used to solve the Stackelberg master-slave game model. A reward mechanism including basic reward, constraint reward and emergency reward is constructed. Communication bandwidth constraints, equipment energy consumption constraints and response requirements of high-risk scenarios are introduced into the strategy learning. The operation and maintenance action strategy table is obtained through experience replay and target network update, and the strategy table is updated when the external environment or equipment access changes.

[0012] Platform control and service application: The DeepGen network operation and maintenance platform obtains the leader status and follower status from the data warehouse, matches the status with the operation and maintenance action strategy table, generates control instructions for terminal devices, sends them to the corresponding devices through the partition gateway, and receives the execution results returned by the devices; when a high-risk scenario is determined, a preset emergency strategy subset is invoked to control the system locally on the partition gateway.

[0013] Furthermore, the payoff function of the leader in the Stackelberg master-slave game model is determined according to the following formula:

[0014]

[0015] Wherein, α is the first weighting coefficient used to characterize the importance of regional risk control, β is the second weighting coefficient used to characterize the importance of regional energy consumption control, the regional risk rate comes from the quantification results of deep regional risks in the data collection and preprocessing steps, and the energy consumption rate is the ratio of the energy consumption of the terminal equipment under the leadership node to the total available energy in the current operation and maintenance cycle.

[0016] Furthermore, the payoff function for the follower in the Stackelberg master-slave game model is determined according to the following formula:

[0017]

[0018] Wherein, γ is the third weighting coefficient used to characterize the importance of terminal data collection and uploading quality, δ is the fourth weighting coefficient used to characterize the importance of terminal operation and maintenance overhead, data integrity is the ratio of the amount of data successfully uploaded by the terminal device within the operation and maintenance cycle to the amount of data that should be uploaded, and operation cost is the normalized result of the resource consumption generated by the terminal device after executing the operation and maintenance parameters of the collection frequency, transmission power, communication bandwidth and location reporting cycle.

[0019] Furthermore, the loss function used for training the deep reinforcement learning (DQN) algorithm in the operation and maintenance strategy solution and update step is determined according to the following formula:

[0020]

[0021] Where R is the total reward obtained at the current operation and maintenance moment based on the leader's reward function and the follower's reward function. As a discount factor, The current reinforcement learning state consists of a leader state and a follower state. To reinforce the learning state in the next moment after the action is performed. For the currently selected maintenance action, For the optional maintenance actions in the next moment, The action value output by the target Q network. This represents the action value output by the current Q network.

[0022] Furthermore, the leader state consists of the regional risk level and regional energy reserves stored in the ground-cloud data warehouse, while the follower state consists of the working status of the corresponding deep-ground terminal device, the normalized data transmission rate, the normalized energy consumption value, and the location code. The leader state and the follower state are used together as reinforcement learning states input to the DQN algorithm to solve the operation and maintenance strategy.

[0023] Furthermore, when constructing the leader's action set and the follower's action set, the operation and maintenance parameters of the terminal devices are discretized. The resulting discrete actions include at least: four acquisition actions for sensor-type terminals with acquisition frequencies of 0.5 Hz, 1 Hz, 2 Hz, and 3 Hz; four power actions for actuator-type terminals with transmit power of 50%, 80%, 100%, and sleep mode; three communication actions for communication modules with communication bandwidths of 5 Mbps, 8 Mbps, and 10 Mbps; and three reporting actions for personnel positioning terminals with location reporting periods of 30 s, 10 s, and 5 s. These actions, along with the area control commands that the leader can output, constitute the action set of the Stackelberg master-slave game model.

[0024] Furthermore, the DQN algorithm employs a training method that combines experience replay with target network synchronization. An experience pool is set up to store interaction data tuples consisting of leader state, follower state, operational actions, rewards, and the next time-step state. The capacity of the experience pool is [missing information]. Each training session randomly samples 32 data points to update network parameters. When the experience pool reaches its maximum capacity, the earliest data point is deleted in a first-in-first-out manner. At the same time, a current Q-network and a target Q-network are set up. Every 100 training sessions, the parameters of the current Q-network are copied to the target Q-network.

[0025] Furthermore, when new terminal devices are connected to the deep network, existing terminal devices are deactivated, regional risk levels change, or regional energy reserves change, the deep network operation and maintenance platform re-collects the leader status and the follower status, inputs the updated status into the DQN algorithm to re-solve the operation and maintenance action strategy, generates a new optimal operation and maintenance action strategy table, and replaces the original strategy table.

[0026] Furthermore, during the platform control and service application process, when the DeepEarth Network Operation and Maintenance Platform determines a high-risk scenario based on the leader status, it issues a preset subset of emergency strategies to the corresponding underground partition gateway. The partition gateway then issues control commands directly to its edge nodes and terminal devices locally and collects execution feedback. After the local emergency control is completed, the execution results are sent back to the DeepEarth Network Operation and Maintenance Platform.

[0027] Secondly, this invention provides a dynamic adaptive operation and maintenance system for deep networks based on Stackelberg game theory, comprising:

[0028] The data processing module is used to deploy edge nodes in deep-ground work areas, equipment supports, and personnel activity areas; deploy zone gateways in underground zones; and deploy a cloud data warehouse on the ground, forming a three-level information integration architecture consisting of edge nodes, zone gateways, and a ground-based cloud. It collects multi-source equipment data, including pressure sensor data, gas sensor data, hydraulic support actuator data, communication module data, and personnel positioning terminal data. Edge nodes locally cache instantaneous data during link interruptions, while zone gateways compress and forward data uploaded by edge nodes to the ground-based cloud. The ground-based cloud stores raw data, standardized equipment status data, and data marts divided according to operation and maintenance scenarios in a distributed, hierarchical manner, and marks and temporarily stores data exceeding equipment thresholds.

[0029] The model building module is used to designate the ground control center and underground zone gateways as leader nodes for network operation and maintenance decisions, and various deep-ground terminal devices as follower nodes, establishing a hierarchical control relationship between leaders and followers. It determines the leader's state variables based on the regional risk level and remaining regional energy recorded in the data warehouse, and the follower's state variables based on the terminal device's operating status, data transmission rate, and energy consumption. It sets the regional control commands that the leader can output, as well as the terminal device's executable acquisition frequency, transmission power, communication bandwidth, and location reporting cycle operation and maintenance parameters, as the leader's action set and the follower's action set. It also establishes the leader's reward function and the follower's reward function.

[0030] The strategy solving module is used to map the leader state, follower state and their action set into the state space and action space of reinforcement learning. It uses the DQN algorithm in deep reinforcement learning to solve the Stackelberg master-slave game model, and constructs a reward mechanism that includes basic rewards, constraint rewards and emergency rewards. It introduces communication bandwidth constraints, equipment energy consumption constraints and response requirements of high-risk scenarios into policy learning. It obtains the operation and maintenance action policy table through experience replay and target network update, and updates the policy table when the external environment or equipment access changes.

[0031] Model application module: It is used to obtain the leader status and follower status from the data warehouse by the deep network operation and maintenance platform, match the status with the operation and maintenance action strategy table, generate control instructions for terminal devices, send them to the corresponding devices through the partition gateway, and receive the execution results returned by the devices; when a high-risk scenario is determined, it calls the preset emergency strategy subset to control locally on the partition gateway.

[0032] Thirdly, the present invention provides an electronic device, the device comprising: a processor and a memory;

[0033] The memory is used to store one or more program instructions;

[0034] The processor is used to run one or more program instructions to execute the aforementioned deep network dynamic adaptive operation and maintenance method based on Stackelberg game theory.

[0035] Fourthly, the present invention provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the aforementioned deep network dynamic adaptive operation and maintenance method based on Stackelberg game theory.

[0036] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0037] This invention constructs a three-tiered information integration architecture of "edge nodes, partitioned gateways, and ground-cloud," coupled with a data preprocessing mechanism for deep-earth scenarios. This enables unified access, standardized storage, and hierarchical management of multi-source heterogeneous operation and maintenance data, including sensor signals, actuator operation records, communication module status, and personnel positioning data. This solves the problems of easily interrupted data links, limited bandwidth, data fragmentation, and high misjudgment rates in complex deep-earth environments. Furthermore, this invention abstracts the actual hierarchical control relationship of deep-earth networks into a Stackelberg master-slave game model of "global leader, regional leader, and terminal follower." It constructs a leader payoff function balancing operation and maintenance security and energy consumption, and a follower payoff function balancing data integrity and action costs. Bandwidth constraints, energy consumption constraints, and response requirements for high-risk scenarios are then introduced into the action space and reward design. Combined with the DQN algorithm, this enables dynamic solution and online updating of operation and maintenance strategies, allowing these strategies to adaptively adjust to changes in device access, regional risks, and resource status. Meanwhile, this invention establishes an emergency strategy subset and a local execution mechanism for partition gateways on the platform side. This allows for direct control commands to be issued at the partition level when high-risk conditions, communication link disruptions, or local device failures are detected, ensuring the continuity and security of deep-earth network operations. Furthermore, by comprehensively assessing regional energy status and equipment operation costs, priority is given to ensuring communication and power supply to high-risk areas and critical equipment, while also considering energy conservation and long-term maintenance needs in low-risk areas. Overall, this invention forms a closed-loop, hierarchical operation and maintenance technology system from data acquisition, model building, strategy solving to platform application. It can simultaneously improve data reliability, the real-time and adaptive nature of operation and maintenance decisions, and the responsiveness to high-risk scenarios in the complex space of deep earth networks, providing integrated technical support for the safe and stable operation of deep earth networks.

[0038] Of course, implementing the various technical solutions of this invention does not necessarily require achieving all of the advantages described above at the same time. Attached Figure Description

[0039] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other embodiments can be obtained from these drawings without creative effort.

[0040] Figure 1 This is a diagram illustrating the overall architecture of the deep network operation and maintenance platform according to an embodiment of the present invention.

[0041] Figure 2 This is a diagram illustrating the layered structure of the data warehouse according to an embodiment of the present invention.

[0042] Figure 3 This is a flowchart illustrating the process of solving the platform operation and maintenance strategy in an embodiment of the present invention.

[0043] Figure 4 This is a schematic diagram of the physical structure of an electronic device according to an embodiment of the present invention. Detailed Implementation

[0044] The present invention will now be described in further detail with reference to specific embodiments and accompanying drawings. Similar elements in different embodiments are referred to by associated similar element reference numerals. In the following embodiments, many details are described to facilitate a better understanding of this application. However, those skilled in the art will readily recognize that some features may be omitted in different situations, or may be replaced by other elements, materials, or methods. In some cases, certain operations related to this application are not shown or described in the specification. This is to avoid obscuring the core parts of this application with excessive description. For those skilled in the art, detailed description of these related operations is not necessary; they can fully understand the related operations based on the description in the specification and general technical knowledge in the art.

[0045] Furthermore, the features, operations, or characteristics described in the specification can be combined in any suitable manner to form various embodiments. At the same time, the steps or actions in the method description can be rearranged or adjusted in a manner obvious to those skilled in the art. Therefore, the various orders in the specification and drawings are only for the clear description of a particular embodiment and do not imply a necessary order, unless otherwise stated that a particular order must be followed.

[0046] Network operation and maintenance in deep, complex spaces faces a series of challenges, including diverse data sources and types, susceptibility to link interruptions, transmission limitations, multiple equipment layers, and rapidly changing scenario risks. Existing operation and maintenance methods, primarily based on centralized ground data aggregation and manual configuration, are ill-suited to this environment in a timely and accurate manner. This leads to problems such as difficulties in unified modeling of multi-source heterogeneous data, cascading propagation of incident chains within local areas, low storage and transmission efficiency, and a lack of elastic scalability. Therefore, there is an urgent need for an innovative operation and maintenance technology solution for deep, complex spaces that can simultaneously handle data acquisition, hierarchical modeling, dynamic strategy solving, and rapid local execution.

[0047] This invention is proposed against this backdrop, aiming to construct a dynamic adaptive operation and maintenance method and system for deep underground networks based on Stackelberg game theory. The scheme first constructs a three-tiered information integration architecture of "edge nodes, partition gateways, and ground-cloud," and sets differentiated data processing flows to adapt to multi-source heterogeneous operation and maintenance data such as sensor signals, actuator logs, and personnel positioning data. Simultaneously, it abstracts the actual management relationship of "unified ground control, partitioned collaboration, and end-device execution" in deep underground networks into a Stackelberg master-slave game model of "global leader, regional leader, and followers." Furthermore, it combines this with the deep reinforcement learning (DQN) algorithm to dynamically solve the operation and maintenance strategy, enabling the network to automatically update its strategy when devices access or exit, regional risks change, or resource status changes. This achieves efficient integration of multi-source data in deep underground networks and intelligent scheduling of operation and maintenance services, providing efficient, stable, and secure technical support for real-time monitoring and safe operation and maintenance in complex deep underground spaces.

[0048] On the data side, this invention employs a three-tiered information acquisition and processing architecture—"edge nodes, partition gateways, and ground-cloud"—along with a preprocessing mechanism suitable for complex deep-earth scenarios, to achieve standardized access to multi-source heterogeneous data. Edge nodes can perform local caching during signal interruptions or link jitter; partition gateways compress and forward uplink data; and the ground-cloud utilizes distributed hierarchical storage and data marts divided according to operational scenarios for unified management, thereby improving data reliability and transmission efficiency and avoiding the data fragmentation, transmission delays, and misjudgments that commonly occur in traditional centralized ground data warehouses in deep-earth scenarios. Addressing common deep-earth characteristics such as blasting interference and signal attenuation, this invention also proposes preprocessing procedures including high-precision digitization of analog signals, local temporary storage of outliers, and priority labeling of early warning data, making data integration more aligned with actual deep-earth operating conditions.

[0049] On the decision-making modeling side, this invention constructs a Stackelberg master-slave game model of "global leader, regional leader, and follower," directly integrating the hierarchical management logic of deep-ground equipment ("ground-region-terminal") into the game modeling process. The payoff function on the leader side is used to balance regional energy consumption under global security constraints, while the payoff function on the follower side is used to control terminal action costs while ensuring data integrity. Therefore, it can realistically reflect the collaborative relationship and hierarchical control relationship between deep-ground equipment, ensuring that the final generated operation and maintenance strategy conforms to the global security goals of deep-ground operation and maintenance scenarios. At the same time, this invention introduces actual constraint parameters such as equipment power, communication bandwidth, location reporting cycle, and battery life into the action space, and then uses the DQN algorithm adapted to deep-ground operation and maintenance scenarios to automatically generate the optimal strategy table, forming a closed-loop process of "state awareness, strategy matching, action execution, and feedback optimization," thereby effectively improving the dynamic adaptability and executability of deep-ground operation and maintenance decisions.

[0050] On the security and emergency response side, this invention sets up a three-level risk priority division of high, medium and low, and deploys a subset of locally executable emergency strategies on the partition gateway side. When the platform determines that it is currently in a high-risk scenario or communication is unstable, it can directly issue control commands at the partition level to achieve a rapid or even second-level response to high-risk states. At the same time, it reserves a redundancy guarantee mechanism for communication interruption and equipment failure, reducing the risk of disaster cascading propagation. Compared with the traditional centralized operation and maintenance mode, it can significantly improve the security and continuity of deep-ground operation and maintenance.

[0051] On the resource and cost side, this invention incorporates regional energy status and equipment operation costs into the operation and maintenance service scheduling, enabling dynamic allocation of key resources such as power supply and bandwidth: high-risk areas prioritize the operation of key equipment, while low-risk areas achieve energy conservation by reducing the energy consumption of non-critical equipment; at the same time, it supports dynamic access of equipment and flexible adaptation to various deep-ground scenarios, reducing operation and maintenance complexity and expansion costs, and meeting the long-term operation and maintenance needs under different deep-ground working conditions.

[0052] Furthermore, this invention also optimizes the DQN algorithm to address the limited computing power and storage space constraints of deep edge nodes. It adopts a lightweight network structure and a storage-saving experience pool management method, and integrates basic rewards, constraint rewards, and emergency rewards into the training process. This ensures that the generated policy table can reflect both the bandwidth and energy consumption constraints of deep edges and cover the rapid response requirements of high-risk scenarios.

[0053] This invention ultimately establishes an integrated deep-earth operation and maintenance technology system encompassing "data acquisition and standardization, hierarchical game modeling, reinforcement learning solution, platform control and local emergency response." It forms a collaborative operation and maintenance service model integrating "closed-loop control, hierarchical emergency response, fault redundancy" and "monitoring, prediction, scheduling, and security," achieving seamless integration from data to model, from model to strategy, and from strategy to service. This represents a systematic improvement for complex deep-earth spatial scenarios, rather than a localized improvement of a single link.

[0054] Example 1:

[0055] like Figure 1 As shown, Figure 1 As shown in the figure, the proposed method for dynamic adaptive operation and maintenance of deep networks based on Stackelberg game theory includes four stages: data acquisition and preprocessing, construction of Stackelberg master-slave game model, solution and update of operation and maintenance strategy, and platform control and service application, which correspond to the data processing layer, model construction layer, strategy solution layer and model application layer, respectively.

[0056] 1. Data Processing Layer

[0057] Multi-source data is collected based on a three-tiered acquisition and transmission architecture consisting of "edge nodes, partition gateways, and ground-cloud." A preprocessing workflow adapted for deep-earth space is designed to address the issues of high interference and complex formats in deep-earth data. The processed data is then stored in a data warehouse, including:

[0058] 101. Data collection for core devices is based on a three-tiered data acquisition and transmission architecture consisting of "edge nodes, partition gateways, and ground-cloud." The data acquisition steps include:

[0059] Based on the needs of deep-ground operation and maintenance, the data acquisition frequency is determined, such as the sampling frequency of pressure sensors being 1Hz, gas sensors being 0.5Hz, the status update frequency of hydraulic support actuators being 2Hz, the bandwidth utilization rate of real-time acquisition communication modules, and the location update frequency of personnel positioning terminals being 1 time / 30s.

[0060] Deploy edge nodes that can cache instantaneous data at the work surface and next to equipment supports within a 500m underground range. The edge nodes can adapt to network link interruption or signal attenuation. When the link is interrupted, the edge nodes can upload the cached data in batches after the link is restored.

[0061] One zone gateway is deployed every 2km. The gateway supports 5G / fiber dual-mode transmission. When the bandwidth utilization rate is >80%, the zone gateway can adaptively adjust the real-time communication bandwidth utilization rate (switching from fiber transmission to 5G narrowband transmission) and perform LZ4 compression on the data uploaded by the edge nodes to reduce the transmission bandwidth consumption.

[0062] The ground and cloud platforms utilize cloud servers and the AnalyticDB data warehouse to receive data transmitted from the partition gateway, enabling centralized storage and management.

[0063] 102. Establish unified data standards for deep-earth equipment and remove outliers. The preprocessing procedure includes:

[0064] The sensor's analog current signal is converted into a digital signal with an accuracy of 0.01; unstructured data such as actuator switch records and communication module logs are converted into computer-readable JSON format; and the latitude and longitude data of the personnel positioning terminal is converted into three-dimensional coordinate data (x / y / z, unit m).

[0065] Set specific value range thresholds for different devices, and mark data that exceeds the normal threshold as "outlier data"; temporarily store outlier data locally on the edge node (retain for 7 days) and do not upload it to the data warehouse; mark data that is close to the threshold as "warning data" and add a priority identifier when uploading it to facilitate subsequent analysis.

[0066] 103. The data warehouse is designed using a distributed, layered storage architecture. The data warehouse architecture is as follows: Figure 2 As shown, it includes:

[0067] The raw data layer (ODS) uses file storage to store unprocessed raw data uploaded from edge nodes for data traceability and anomaly investigation.

[0068] The Data Integration Layer (DW) stores standardized device status data, employing AnalyticDB columnar storage for fast column-based queries (with a 3-month retention period). A star schema is constructed, where the fact table `device_state` stores real-time device status, with fields including `device_id`, `state_time`, `pressure`, `gas_concentration`, and `energy_consumption`. Dimension tables include `device_info` for storing `device_type`, `rated_parameters`, and `position`, `time_info` for storing time information, and `region_info` for storing region name, risk level, and energy supply status.

[0069] The Data Mart layer (DM) is built according to the equipment operation and maintenance scenario. Each mart only retains scenario-related data. For example, the pressure monitoring mart only stores pressure values, regional risk levels, and sensor location data.

[0070] The data warehouse adopts a dynamic update mechanism. When the device location changes or a new device is connected, the data warehouse topology update command is automatically triggered and synchronized to the position field in the device_info table. Incremental update technology is used to update only the changed data fields and not the historical status data. Version identifiers are added to the device status data in the data warehouse to record the time, reason and operator of each update, which facilitates data backtracking.

[0071] 2. Model Building Layer

[0072] Based on the hierarchical relationship of deep-earth equipment, the game's main players and roles are defined. A two-tiered leader design (global and regional) balances global security with local efficiency. Follower classification ensures that the actions of different equipment align with their functional roles, including:

[0073] 201. Definition of Game Players and Roles

[0074] The leader nodes are divided into two levels: global leaders (ground control center) and regional leaders (underground zone gateways). The global leader is responsible for formulating and outputting global control commands based on the overall risk status and total energy supply of deep underground space. The regional risk level of the leader can adapt to the real-time environmental parameters of the deep underground scenario and has the highest priority. One regional leader is deployed every 2km. It is responsible for converting global goals into local commands and adjusting the goals based on the energy status and equipment status of the local area, and outputting regional control commands. Regional leaders must obey global control commands, and terminal devices must obey regional commands.

[0075] Follower nodes are categorized based on the functions of deep-ground terminal equipment, including sensor types (pressure, gas, and temperature sensors), actuator types (hydraulic supports, ventilation valves, and water pumps), communication modules (fiber optics and wireless relays), and personnel positioning terminals.

[0076] 202. Quantifying the states, actions, and reward functions of both sides in a game.

[0077] The state space (S) includes leader states and follower states. The leader state includes the regional risk level and remaining energy. The regional risk level is based on gas concentration and pressure values ​​in the data warehouse, with high risk (gas or...) (The text abruptly ends here, so the translation stops as well.) Concentration > 1%, or pressure > 50 MPa), medium risk (gas or Concentration between 0.5% and 1%, pressure between 30 and 50 MPa, low risk (gas or Concentration < 0.5%, pressure < 30 MPa), one-hot coding is used to encode high, medium and low risk (high risk [1,0,0], medium risk [0,1,0], low risk [0,0,1]); energy surplus is based on underground power station or zone power supply load division, energy sufficient (load ≥ 80%), energy medium (load between 30% and 80%), energy tight (load < 30%), one-hot coding is used to encode energy level (sufficient [1,0,0], medium [0,1,0], tight [0,0,1]).

[0078] The follower status includes device operating status, data transmission rate, and energy consumption. Device operating status is divided into normal (no fault, parameters meet standards), warning (parameters approaching threshold), and fault (parameters exceeding standards or device offline). One-hot encoding is used to encode device operating status (normal [1,0,0], warning [0,1,0], fault [0,0,1]), and the data is obtained from the work_state field of the device_state table. The data transmission rate is the actual transmission rate of the communication module (unit Mbps), obtained from the data_transfer_rate field of the device_state table, and the actual transmission rate is normalized to the [0,1] range by dividing the actual rate by the rated bandwidth. The energy consumption is the actual energy consumption of the device (unit W / h), obtained from the energy_consumption field of the device_state table, and the actual energy consumption is normalized to the [0,1] range by dividing the actual energy consumption by the rated energy consumption.

[0079] Action space (A) includes leader actions and follower actions, where leader actions include global leader actions and regional leader actions, specifically including:

[0080] Sensor commands: "Low power mode" (sampling frequency 0.5Hz), "Normal mode" (1Hz), "High risk monitoring mode" (2Hz), "Emergency mode" (3Hz);

[0081] Actuator commands: "Energy Saving Mode" (50% power), "Normal Mode" (80%), "Load Support Mode" (100%), "Sleep Mode" (No Load);

[0082] Communication module instructions: "Low priority transmission" (5Mbps), "Normal transmission" (8Mbps), "High priority transmission" (10Mbps);

[0083] Location terminal instructions: "Routine report" (30s), "High-risk report" (10s), "Emergency report" (5s);

[0084] The follower (terminal device) selects specific action parameters according to the leader's instructions. The action set corresponds one-to-one with the leader's instructions, and the follower's action space can adapt to the control instructions output by the leader.

[0085] The payoff function (U) includes the leader's payoff and the follower's payoff, where the leader's payoff... The formula used to measure the balance between safety and efficiency in a leader's decision-making is:

[0086]

[0087] in , indicating risk weight, This indicates energy consumption weight; the regional risk rate is calculated based on the region_info table, and the regional risk rate = area of ​​high-risk areas / total area; the energy consumption rate is calculated from the device_state table, and the energy consumption rate = actual total energy consumption / rated total energy consumption; leader benefits. The range is [0,1], where 0 represents the lowest return (high risk and full energy consumption in the entire area) and 1 represents the highest return (no high risk and lowest energy consumption).

[0088] Follower benefits are used to measure the balance between the data value and cost of follower actions. The formula is:

[0089]

[0090] in (Data weights) (Cost weighting); Data integrity is statistically derived from the data warehouse: Data integrity = Number of successful data uploads within 10 consecutive seconds by the device / Total number of data collections; Action cost = Actual energy consumption of the device when performing the action / Rated energy consumption; Follower benefits The range is [0,1], where 0 represents the lowest benefit for the follower (no data upload and full cost), and 1 represents the highest benefit for the follower (full data upload and zero cost).

[0091] 203. Game Equilibrium Conditions

[0092] Game equilibrium refers to a state in which both the leader's and followers' payoffs reach their minimum acceptable levels, and the system is in a safe and stable state. Specific conditions include:

[0093] Leader Benefits ,in This indicates the leader's minimum acceptable return, ensuring a certain level of overall or regional risk. 20% and energy consumption rate 20%; Follower Benefits ,in This indicates the minimum acceptable return for followers, ensuring data integrity. 90% and action cost 30%;

[0094] When both conditions are met, the game reaches Nash equilibrium. At this point, the leader does not need to adjust its goals, followers do not need to change their actions, and the system's operational efficiency is optimal. If the risk level is less than 0.8, leaders need to strengthen risk control measures, such as expanding the scope of high-risk monitoring; if If the energy consumption is less than 0.7, the follower needs to adjust its actions, such as reducing energy consumption.

[0095] 3. Strategy Solving Layer

[0096] Solving deep operation and maintenance strategies based on the reinforcement learning DQN algorithm includes:

[0097] 301. DQN Algorithm Solution Strategy Model

[0098] The DQN algorithm learns the optimal action policy dynamically and adaptively by constructing an interaction model between the agent and the environment, such as... Figure 3 As shown, it includes:

[0099] Intelligent agent definition: Each follower node is an independent intelligent agent (e.g., each gas sensor is one intelligent agent). Each intelligent agent only focuses on its own state and the state of the corresponding leader, learns action strategies independently, and shares the global state through a data warehouse, such as the risk level of the area where the device is located.

[0100] State space reconstruction: Reconstructing the leader's state space in a game theory model Follower status Spatial fusion into reinforcement learning state space One-hot encoding and numerical normalization are used to ensure consistent input dimensions.

[0101] Taking the gas sensor agent as an example, the leader state (6-dimensional) is composed of 3-class one-hot encoding of risk level (3-dimensional) + 3-class one-hot encoding of remaining energy (3-dimensional); the follower state (6-dimensional) is composed of 3-class one-hot encoding of equipment working status (3-dimensional) + data transmission rate normalization (1-dimensional) + energy consumption normalization (1-dimensional) + location encoding (1-dimensional). The dimension is the result of reconstructing (summing) the leader's state space dimension (6 dimensions) and the follower's state space dimension (6 dimensions);

[0102] Action space reconstruction: Discretize the follower's continuous actions into a finite set of selectable actions, including:

[0103] Sensor-based intelligent agent action set ={0.5Hz acquisition (low power consumption), 1Hz acquisition (normal), 2Hz acquisition (high risk monitoring), 3Hz acquisition (emergency)}, a total of 4 discrete actions;

[0104] Action set of actuator-type intelligent agents ={50% power (energy saving), 80% power (normal), 100% power (load support), sleep (no load)}, a total of 4 discrete actions;

[0105] Communication module action set ={5Mbps bandwidth (low priority), 8Mbps bandwidth (normal), 10Mbps bandwidth (high priority data transmission)}, a total of 3 discrete actions;

[0106] Positioning terminal action set ={30s reporting (routine), 10s reporting (high-risk areas), 5s reporting (emergency)}, a total of 3 discrete actions.

[0107] Reward Function Design: Based on the dual requirements of security and efficiency in deep-ground operations and maintenance, a basic reward is designed. Constraints and Rewards Emergency Rewards Three-tiered reward mechanism:

[0108] Basic rewards for intelligent agents Enhance data value, reduce operational costs, and Indicates basic reward Equivalent to follower payoff in a game theory model ;

[0109] Agent uses constraint rewards To prevent the agent from choosing actions that violate constraints, the constraint rule is: if the constraint simultaneously satisfies the communication bandwidth... 10Mbps and power consumption of all devices For the corresponding rated value, then =0.2; if only one constraint is satisfied, then =0; if none of these conditions are met, then =-0.5;

[0110] Rewards for intelligent agents using deep-earth emergency scenarios To incentivize agents to respond quickly to high-risk states, the rule is: when the leader node triggers a high-risk state, the agent executes an emergency action. =0.3; if no emergency action is taken, then =-1.0 (severe penalty); in non-emergency situations, =0;

[0111] Total reward of the agent The range is [-1.5, 1.5]. The agent learns to maximize the total reward, thereby achieving a balance between security and efficiency in deep-earth operations.

[0112] 302. Implementation of the DQN Algorithm

[0113] The implementation of the DQN algorithm includes four core steps: network structure design, experience replay mechanism, target network and update strategy, training iteration and convergence conditions. To address the limited computing power and storage resources of deep edge nodes, the DQN algorithm implementation needs optimization. Specifically:

[0114] The network structure adopts a single-Q network architecture, consisting of an input layer, two hidden layers, and an output layer. The input layer receives normalized state data, and its dimension is the same as that of the state space. The dimensions are consistent; the first hidden layer has 64 neurons and the second layer has 32 neurons, both using the ReLU activation function; the dimension of the output layer is consistent with the size of the corresponding agent's action set, such as 4 dimensions for sensor classes, and the output result is the Q value (action value) of each discrete action, with the action with the largest Q value being the current optimal action.

[0115] The experience replay mechanism uses a capacity of Experience pool storage Intelligent agent interaction data tuple, in which, Encode the currently selected discrete action. The new state after the action is executed; binary format is used for storage to save space; in order to synchronize with the data warehouse's update frequency of 1 second / time, the experience pool samples 32 data points each time according to a uniform random strategy, and when the experience pool is full, the earliest data is eliminated according to the "first-in, first-out" (FIFO) principle, which effectively addresses the data correlation problem;

[0116] Introducing the target Q network With the current Q network The hard update method directly copies the current network parameters to the target network every 100 iterations. After the update, the Q-value error is calculated using a validation set composed of 20% of historical data. If the error exceeds 5%, the parameters are rolled back to the previous version.

[0117] The training iterations initialize the network parameters using the Xavier initialization method, then clear the experience pool and configure the learning rate. = 5×10 -4Discount Factor =0.9, Exploration Rate The initial value of hyperparameters such as 0.9 is set, and then... - A greedy strategy is used to select actions for the agent, and the network exploration rate is... Decrease by 0.05 every 100 training rounds, until it decreases to 0.1; when the experience pool capacity... At 1000, a batch update of the network is triggered, which implements a parameter update using the mean squared error (MSE) loss function and the Adam optimizer every 32 data samples. The formula for the MSE function is as follows:

[0118]

[0119] in The target network predicts the optimal Q-value for the next state. Finally, the convergence criterion is that the target Q-value fluctuation is ≤3% for 500 consecutive rounds and the optimal action selection probability is ≥85%. After stopping training, the optimal action policy table in JSON format is output for easy platform use.

[0120] 4. Model Application Layer

[0121] To address the operational and maintenance needs of deep-ground scenarios, a network platform control and service mechanism adapted to the complex environment of deep-ground environments is designed. The platform takes the control mechanism of "closed-loop control and emergency response" as its core, and combines it with an integrated service mechanism of status monitoring, fault prediction, resource scheduling and personnel safety to comprehensively ensure the real-time performance, security and scalability of deep-ground operations and maintenance.

[0122] The control mechanism, centered on the policy table generated by DQN, implements a closed-loop process of "data acquisition - policy matching - command execution - result feedback." For high-risk scenarios in complex deep-earth spaces, the platform is designed with an adaptive emergency response mechanism, including:

[0123] 401. A closed-loop process of "data acquisition, strategy matching, instruction execution, and result feedback".

[0124] The data warehouse collects the leader status (regional risk level, remaining energy) and follower status (equipment working status, transmission rate, energy consumption) every 1 second and updates the corresponding data table. Then, the two types of status are merged into a real-time status vector according to a specified format. When data is missing, it is filled by linear interpolation. If the missing data exceeds 30 seconds, an alarm is triggered.

[0125] The strategy matching process quickly completes the matching and converts the optimal action into a device-executable instruction after querying the hash index. If multiple devices need to coordinate actions, a "batch instruction distribution" mechanism is adopted to generate instruction packages in a unified manner.

[0126] The instructions are transmitted to the terminal device via the partition gateway through the TCP / IP protocol. The terminal device completes the parameter adjustment and sends back the results. The partition gateway verifies whether the action performed by the intelligent agent conforms to the deep-earth scenario constraints. If it does, the data warehouse is updated synchronously. If it does not, the suboptimal action is automatically matched and the instructions are reissued.

[0127] 402. Platform Emergency Response Mechanism

[0128] The platform's emergency response mechanism needs to adapt to the risk types in deep-earth scenarios. For high-risk deep-earth scenarios, the platform divides emergency response priorities into high, medium, and low levels according to the severity of the risk and sets corresponding response time limits. It also pre-trains emergency strategy subsets for high-risk scenarios such as gas leaks, rock collapses, and personnel intrusion, and stores them locally on the partition gateway to shorten the call path. High-risk states are triggered in real time by the data warehouse at a frequency of 100ms / time. After triggering, the platform directly skips the regular strategy matching process and issues emergency commands to the partition gateway. The partition gateway calls the locally stored emergency strategy subset to generate batch control commands to ensure rapid response. Once the risk level drops to medium or low, the platform automatically switches back to the regular closed-loop control process.

[0129] To ensure control continuity in extreme situations, when communication between the zone gateway and the ground server is interrupted, the zone gateway will automatically switch to local emergency control mode and make decisions independently based on the locally stored policy table and real-time collected data. If critical equipment, such as the zone gateway, fails, the adjacent zone gateway will automatically take over its control range and relay commands through edge nodes to ensure that the control link is not interrupted.

[0130] The service mechanism revolves around the full-scenario needs of deep-ground operations and maintenance, including four core services: status monitoring, fault prediction, resource scheduling, and personnel safety, as detailed below:

[0131] The status monitoring service enables real-time visualization of device status and key parameters, as well as personnel location display through web and mobile interfaces. Based on preset multi-level alarm thresholds, it automatically generates alarm information including alarm type, location of occurrence, real-time parameter value, and targeted handling suggestions to help maintenance personnel quickly locate problems.

[0132] The fault prediction service uses historical data from the data warehouse for the past three months to train an LSTM model and combines it with the trend of action cost changes in the DQN strategy table to predict fault risks.

[0133] The resource scheduling service dynamically allocates power supply and bandwidth resources based on regional energy status and equipment operating costs, achieving a balance between prioritizing high-risk areas and energy conservation in low-risk areas.

[0134] Personnel safety services include area access control and emergency rescue support. When unauthorized personnel mistakenly enter a high-risk area, an alarm is triggered. When personnel make an emergency call, the platform simultaneously displays the location, adjusts surrounding equipment, and generates a rescue route.

[0135] Example 2:

[0136] This application provides a deep network dynamic adaptive operation and maintenance system based on Stackelberg game theory, which corresponds to the steps of the deep network dynamic adaptive operation and maintenance method based on Stackelberg game theory in Embodiment 1. It realizes the organic collaboration between various modules. The system includes the following modules:

[0137] Data Processing Module: Deploys edge nodes in deep-ground work areas, equipment supports, and personnel activity areas; deploys zone gateways in underground zones; and deploys a cloud data warehouse on the ground, forming a three-tiered information integration architecture consisting of edge nodes, zone gateways, and the ground-cloud. It collects multi-source equipment data, including pressure sensor data, gas sensor data, hydraulic support actuator data, communication module data, and personnel positioning terminal data. Edge nodes locally cache instantaneous data during link interruptions. Zone gateways compress data uploaded by edge nodes and forward it to the ground-cloud. The ground-cloud stores raw data, standardized equipment status data, and data marts categorized by operation and maintenance scenarios in a distributed, hierarchical manner, and marks and temporarily stores data exceeding equipment thresholds. Model Building Module: Designates the ground control center and underground zone gateways as leader nodes for network operation and maintenance decisions, and various deep-ground terminal devices as follower nodes, establishing a hierarchical control relationship between leaders and followers. It determines the leader's state variables based on the regional risk level and remaining regional energy recorded in the data warehouse, and determines the follower's state variables based on the terminal device's operating status, data transmission rate, and energy consumption. It also converts the data output by the leader into a model. Domain control commands and executable collection frequency, transmission power, communication bandwidth, and location reporting cycle maintenance parameters for terminal devices are set as leader action sets and follower action sets; leader reward functions and follower reward functions are established respectively; the strategy solving module is used to map the leader state, follower state, and their action sets into the state space and action space of reinforcement learning, and uses the DQN algorithm in deep reinforcement learning to solve the Stackelberg master-slave game model, constructing a reward mechanism including basic rewards, constraint rewards, and emergency rewards, and introducing communication bandwidth constraints, device energy consumption constraints, and response requirements for high-risk scenarios into policy learning, obtaining the maintenance action strategy table through experience replay and target network updates, and updating the strategy table when the external environment or device access changes; the model application module is used to obtain the leader state and follower state from the data warehouse from the deep network maintenance platform, match the state with the maintenance action strategy table, generate control commands for terminal devices, send them to the corresponding devices through the partition gateway, and receive the execution results returned by the devices; when a high-risk scenario is determined, a preset emergency strategy subset is invoked to control locally on the partition gateway.

[0138] Based on the same inventive concept, embodiments of this disclosure also provide an electronic device for dynamic adaptive operation and maintenance of deep networks based on Stackelberg game theory. Figure 4This is a schematic diagram of the physical structure of an electronic device provided in an embodiment of the present invention. The electronic device may include: a processor 301, a communications interface 302, a memory 303, and a bus 304. The processor 301, communications interface 302, and memory 303 communicate with each other via the bus 304. The processor 301 can call a computer program stored in the memory 303 and executable on the processor 301 to perform the deep-earth network dynamic adaptive operation and maintenance method based on Stackelberg game theory provided in the above embodiment.

[0139] Furthermore, the logical instructions in the aforementioned memory 303 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention embodiment, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0140] Based on the same inventive concept, embodiments of this disclosure also provide a computer program product, including a computer program that, when executed by a processor, can implement the deep network dynamic adaptive operation and maintenance method based on Stackelberg game theory described above.

[0141] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0142] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

Claims

1. A dynamic adaptive operation and maintenance method for deep earth networks based on Stackelberg game theory, characterized in that, include: Data Acquisition and Preprocessing: Edge nodes are deployed in deep-ground work areas, equipment supports, and personnel activity areas; zone gateways are deployed in underground zones; and a cloud data warehouse is deployed on the ground, forming a three-tiered information integration architecture consisting of edge nodes, zone gateways, and a ground-based cloud. Data is collected from pressure sensors, gas sensors, hydraulic support actuators, communication modules, and personnel positioning terminals. Edge nodes locally cache instantaneous data when links are interrupted. Zone gateways compress data uploaded by edge nodes and forward it to the ground-based cloud. The ground-based cloud stores raw data, standardized equipment status data, and data marts divided according to operation and maintenance scenarios in a distributed and hierarchical manner, and marks and temporarily stores data that exceeds equipment thresholds. The Stackelberg master-slave game model is constructed as follows: The ground control center and underground zone gateways are designated as leader nodes for network operation and maintenance decisions, while various deep-ground terminal devices are designated as follower nodes, establishing a hierarchical control relationship between leaders and followers. The leader's state variables are determined based on the regional risk level and remaining regional energy quantity recorded in the data warehouse, while the follower's state variables are determined based on the terminal device's operating status, data transmission rate, and energy consumption. The regional control commands that the leader can output, as well as the terminal device's executable acquisition frequency, transmission power, communication bandwidth, and location reporting cycle operation and maintenance parameters, are set as the leader's action set and the follower's action set. Leader and follower reward functions are established separately. Operation and maintenance strategy solution and update: The leader state, follower state and their action set are mapped to the state space and action space of reinforcement learning. The DQN algorithm in deep reinforcement learning is used to solve the Stackelberg master-slave game model. A reward mechanism including basic reward, constraint reward and emergency reward is constructed. Communication bandwidth constraints, equipment energy consumption constraints and response requirements of high-risk scenarios are introduced into the strategy learning. The operation and maintenance action strategy table is obtained through experience replay and target network update, and the strategy table is updated when the external environment or equipment access changes. Platform control and service application: The DeepGen network operation and maintenance platform obtains the leader status and follower status from the data warehouse, matches the status with the operation and maintenance action strategy table, generates control instructions for terminal devices, sends them to the corresponding devices through the partition gateway, and receives the execution results returned by the devices; when a high-risk scenario is determined, a preset emergency strategy subset is invoked to control the system locally on the partition gateway.

2. The deep network dynamic adaptive operation and maintenance method based on Stackelberg game theory as described in claim 1, characterized in that, In the Stackelberg master-slave game model, the leader's payoff function is determined by the following formula: Wherein, α is the first weighting coefficient used to characterize the importance of regional risk control, β is the second weighting coefficient used to characterize the importance of regional energy consumption control, the regional risk rate comes from the quantification results of deep regional risks in the data collection and preprocessing steps, and the energy consumption rate is the ratio of the energy consumption of the terminal equipment under the leadership node to the total available energy in the current operation and maintenance cycle.

3. The deep network dynamic adaptive operation and maintenance method based on Stackelberg game theory as described in claim 1, characterized in that, In the Stackelberg master-slave game model, the payoff function of the follower is determined by the following formula: Wherein, γ is the third weighting coefficient used to characterize the importance of terminal data collection and uploading quality, δ is the fourth weighting coefficient used to characterize the importance of terminal operation and maintenance overhead, data integrity is the ratio of the amount of data successfully uploaded by the terminal device within the operation and maintenance cycle to the amount of data that should be uploaded, and operation cost is the normalized result of the resource consumption generated by the terminal device after executing the operation and maintenance parameters of the collection frequency, transmission power, communication bandwidth and location reporting cycle.

4. The deep network dynamic adaptive operation and maintenance method based on Stackelberg game theory as described in claim 1, characterized in that, The loss function used to train the deep reinforcement learning (DQN) algorithm in the operation and maintenance strategy solution and update step is determined according to the following formula: Where R is the total reward obtained at the current operation and maintenance moment based on the leader's reward function and the follower's reward function. As a discount factor, The current reinforcement learning state consists of a leader state and a follower state. To reinforce the learning state in the next moment after the action is performed. For the currently selected maintenance action, For the optional maintenance actions in the next moment, The action value output by the target Q network. This represents the action value output by the current Q network.

5. The deep network dynamic adaptive operation and maintenance method based on Stackelberg game theory as described in claim 1, characterized in that, The leader state consists of the regional risk level and regional energy reserves stored in the ground-cloud data warehouse. The follower state consists of the working status of the corresponding deep-ground terminal device, the normalized data transmission rate, the normalized energy consumption value, and the location code. The leader state and the follower state are used together as reinforcement learning states input to the DQN algorithm to solve the operation and maintenance strategy.

6. The deep network dynamic adaptive operation and maintenance method based on Stackelberg game theory as described in claim 1, characterized in that, When constructing the leader's action set and the follower's action set, the operation and maintenance parameters of the terminal devices are discretized. The resulting discrete actions include at least: four acquisition actions for sensor-type terminals with acquisition frequencies of 0.5 Hz, 1 Hz, 2 Hz, and 3 Hz; four power actions for actuator-type terminals with transmit power of 50%, 80%, 100%, and sleep mode; three communication actions for communication modules with communication bandwidths of 5 Mbps, 8 Mbps, and 10 Mbps; and three reporting actions for personnel positioning terminals with location reporting periods of 30 s, 10 s, and 5 s. These actions, along with the area control commands that the leader can output, constitute the action set of the Stackelberg master-slave game model.

7. The deep network dynamic adaptive operation and maintenance method based on Stackelberg game theory as described in claim 4, characterized in that, The DQN algorithm adopts an experience replay and target network synchronization training mode, sets an experience pool for storing interactive data tuples composed of leader state, follower state, operation and maintenance action, reward and next time state, the capacity of the experience pool is 5x10 5 Data, 32 data are randomly sampled for network parameter update each time training is performed, when the capacity of the experience pool reaches the upper limit, the earliest entered data is deleted in a first-in-first-out manner; meanwhile, a current Q network and a target Q network are set, and the parameters of the current Q network are copied to the target Q network in whole after 100 rounds of training are completed.

8. The deep network dynamic adaptive operation and maintenance method based on Stackelberg game theory as described in claim 1, characterized in that, When new terminal devices are connected to the deep network, existing terminal devices are deactivated, regional risk levels change, or regional energy reserves change, the deep network operation and maintenance platform re-collects the leader status and the follower status, inputs the updated status into the DQN algorithm to re-solve the operation and maintenance action strategy, generates a new optimal operation and maintenance action strategy table, and replaces the original strategy table.

9. The deep network dynamic adaptive operation and maintenance method based on Stackelberg game theory as described in claim 1, characterized in that, During platform control and service application, when the DeepEarth Network Operation and Maintenance Platform determines a high-risk scenario based on the leader status, it issues a preset subset of emergency strategies to the corresponding underground partition gateway. The partition gateway then issues control commands to its edge nodes and terminal devices locally and collects execution feedback. After the local emergency control is completed, the execution results are sent back to the DeepEarth Network Operation and Maintenance Platform.

10. A dynamic adaptive operation and maintenance system for deep earth networks based on Stackelberg game theory, characterized in that, include: The data processing module is used to deploy edge nodes in deep-ground work areas, equipment supports, and personnel activity areas; deploy zone gateways in underground zones; and deploy a cloud data warehouse on the ground, forming a three-level information integration architecture consisting of edge nodes, zone gateways, and a ground-based cloud. It collects multi-source equipment data, including pressure sensor data, gas sensor data, hydraulic support actuator data, communication module data, and personnel positioning terminal data. Edge nodes locally cache instantaneous data during link interruptions, while zone gateways compress and forward data uploaded by edge nodes to the ground-based cloud. The ground-based cloud stores raw data, standardized equipment status data, and data marts divided according to operation and maintenance scenarios in a distributed, hierarchical manner, and marks and temporarily stores data exceeding equipment thresholds. The model building module is used to designate the ground control center and underground zone gateways as leader nodes for network operation and maintenance decisions, and various deep-ground terminal devices as follower nodes, establishing a hierarchical control relationship between leaders and followers. It determines the leader's state variables based on the regional risk level and remaining regional energy recorded in the data warehouse, and the follower's state variables based on the terminal device's operating status, data transmission rate, and energy consumption. It sets the regional control commands that the leader can output, as well as the terminal device's executable acquisition frequency, transmission power, communication bandwidth, and location reporting cycle operation and maintenance parameters, as the leader's action set and the follower's action set. It also establishes the leader's reward function and the follower's reward function. The strategy solving module is used to map the leader state, follower state and their action set into the state space and action space of reinforcement learning. It uses the DQN algorithm in deep reinforcement learning to solve the Stackelberg master-slave game model, and constructs a reward mechanism that includes basic rewards, constraint rewards and emergency rewards. It introduces communication bandwidth constraints, equipment energy consumption constraints and response requirements of high-risk scenarios into policy learning. It obtains the operation and maintenance action policy table through experience replay and target network update, and updates the policy table when the external environment or equipment access changes. Model application module: It is used to obtain the leader status and follower status from the data warehouse by the deep network operation and maintenance platform, match the status with the operation and maintenance action strategy table, generate control instructions for terminal devices, send them to the corresponding devices through the partition gateway, and receive the execution results returned by the devices; when a high-risk scenario is determined, it calls the preset emergency strategy subset to control locally on the partition gateway.