An edge cooperation-based intelligent coal mine computing offloading method and system

Through the edge collaboration-based smart coal mine computing offloading method, the deep reinforcement learning model is used to optimize communication resource allocation and load balancing, which solves the delay problem of computing task offloading of terminal equipment in smart mines and realizes low-latency and efficient computing task processing.

CN120614643BActive Publication Date: 2025-10-21SOUTHEAST UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511102383.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-07
Publication Date
2025-10-21
Estimated Expiration
2045-08-07

AI Technical Summary

Technical Problem

In smart mines, existing technologies are unable to effectively solve the latency problem when computing tasks are offloaded from terminal devices, especially in terms of multi-node collaborative communication resource scheduling and load balancing, which leads to uneven distribution of computing resources and excessive communication latency.

Method used

An edge-collaboration-based smart coal mine computing offloading method is adopted. The communication resource allocation and load balancing decision-making are constructed through the deep reinforcement learning model. The Markov decision process is used to optimize the communication resource allocation and load balancing, and task collaboration between terminal devices and edge computing nodes is realized. The multi-agent flexible actor-critic algorithm is used to train the communication resource allocation and the flexible actor-critic algorithm is used to train the load balancing decision model to achieve low-latency computing task processing.

Benefits of technology

It significantly reduces task processing delay, improves system resource utilization, enhances task execution efficiency and system stability, adapts to the complexity and variability of the mine environment, and ensures the continuity and intelligence level of mining operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120614643B_ABST
    Figure CN120614643B_ABST
Patent Text Reader

Abstract

The application discloses an intelligent coal mine computing unloading method and system based on edge cooperation, the method comprising: a terminal device in a mining site generates time-sensitive audio, video and sensor computing tasks during operation, constructs a Markov decision process of communication resource allocation, selects the optimal transmission power and transmission subchannel, and realizes intelligent decision of a local computing unloading strategy; a center edge computing node considers edge cooperation and constructs a load balancing optimization model based on global load information, and formulates an optimal computing task migration scheme through a deep reinforcement learning algorithm, thereby realizing efficient cooperation between edge nodes. The application divides the overall optimization problem into two subproblems of communication resource allocation and load balancing, jointly solves the two problems through a two-stage deep reinforcement learning method, and can quickly generate a low-latency unloading and scheduling strategy after training, thereby significantly improving the utilization efficiency of computing resources and the task processing performance in an intelligent coal mine environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computing offloading, and in particular to a smart coal mine computing offloading method and system based on edge collaboration. Background Art

[0002] In the mining areas of smart mines, a variety of sensors are deployed, including panoramic video capture systems, microphone arrays, and sensors for monitoring location, environment, equipment status, human behavior, and geological changes. These terminal devices, through real-time perception of the surrounding operating environment, collect a wealth of on-site status information and further perform large-scale computational tasks locally. Therefore, faced with the massive amount of data generated in mining scenarios, how to achieve low-latency data processing while ensuring rapid operational response has become a core technical challenge that requires urgent attention in the construction of smart mines.

[0003] To address the challenges of high data processing demands and the timeliness of computational feedback, Mobile Edge Computing (MEC) technology has become a key research focus. By shifting computing power from the central cloud to edge nodes closer to data sources, end devices can offload some compute-intensive tasks to edge servers. This not only alleviates the pressure of insufficient computing power on end devices but also effectively reduces the communication latency associated with transmitting data to the cloud platform. In this architecture, the cloud is responsible for overall planning and macro-data analysis, while the edge focuses on real-time data and local business responses. End devices focus on perception and execution, thus forming an efficient and collaborative computing ecosystem.

[0004] While edge computing provides the technical foundation for achieving three-level collaboration across the cloud, edge, and device, practical deployments still face numerous challenges, such as determining when to offload operations on terminal devices, scheduling communication resources for computing task transmission, and load balancing across edge computing nodes. Currently, research on edge computing specifically for mining face scenarios is relatively scarce, and existing resource scheduling solutions struggle to optimize the allocation of collaborative communication and computing resources across multiple nodes, and they fail to consider task migration between multiple edge nodes. Summary of the Invention

[0005] Purpose of the invention: In response to the above-mentioned problems existing in the prior art, the present invention proposes a smart coal mine computing offloading method and system based on edge collaboration, which realizes the low-latency requirement of computing-intensive task processing by reasonably scheduling communication resources for terminal computing task offloading of the smart mine mining working face and utilizing task collaboration between edge computing nodes.

[0006] Technical solution: The present invention provides a smart coal mine computing offloading method based on edge collaboration, which specifically includes the following steps:

[0007] (1) The multi-type delay-sensitive computing tasks generated by the terminal equipment in the mining area during operation are abstracted into specific task information based on task type, computing requirements, and maximum tolerable delay; the terminal equipment makes computing offloading decisions based on the task information and local computing resources;

[0008] (2) The terminal device regularly estimates the channel state information, and the central edge server regularly collects global load information, establishing a communication model and a computation model that includes the terminal device and the edge computing node. This then constructs a resource allocation optimization problem that minimizes the computation offloading delay, and decomposes the optimization problem into a communication resource allocation sub-problem and a load balancing sub-problem.

[0009] (3) For the communication resource allocation sub-problem, the terminal device establishes a Markov decision process locally to describe the communication resource allocation process and defines its state space, action space, and reward function;

[0010] (4) The local terminal device trains the communication resource allocation decision model based on the multi-agent flexible actor-critic algorithm;

[0011] (5) For the load balancing sub-problem, the central edge node collects the status information of all terminal devices, establishes a Markov decision process to describe the task migration process, and defines its state space, action space, and reward function;

[0012] (6) The central edge node trains the load balancing decision model based on the flexible actor-critic algorithm;

[0013] (7) The terminal device observes the channel status information and task information in real time, obtains the optimal communication resource allocation decision based on the trained communication resource allocation decision model, and sends the computing task to the edge computing node based on this decision; then each edge computing node uploads the local load information to the central edge node, and the central edge node obtains the optimal load balancing decision based on the trained load balancing decision model and delegates it to all edge nodes to complete the inter-node migration of the task.

[0014] Furthermore, the implementation process of establishing the communication model and computing model including the terminal device and the edge computing node in step (2) is as follows:

[0015] Establish a smart coal mine computing and unloading system, Represents the set of edge nodes, a total of R edge nodes, Represents the set of terminal devices within the coverage of edge node r, terminal devices The arrival of computing tasks can be described as the arrival rate The Poisson process generates a computing task in each time slot with available triples express, The amount of data for the calculation task, The amount of computation required to complete this task is The maximum tolerable delay for task execution;

[0016] Establish communication model, terminal equipment The task upload delay between the edge computing node r is expressed as:

[0017]

[0018] in, For terminal devices The task upload rate between the edge computing node r is obtained by the following Shannon formula:

[0019]

[0020] Where B is the selected subchannel bandwidth, For terminal devices Set the transmission power, is the additive white Gaussian noise power, Indicates terminal device The channel gain on subchannel k is, represents the interference between terminal devices and is calculated using the following formulas:

[0021]

[0022]

[0023] in, For terminal devices and the small-scale fading of edge node r on subchannel k, For terminal devices The Euclidean distance to the edge node r, is the path loss, using Represents the subchannel set, there are K subchannels in total, Indicates terminal device Whether to select subchannel k transmission task, Indicates terminal device Do not select subchannel k transmission task, Indicates terminal device Select subchannel k transmission task;

[0024] When the terminal device After the computing task reaches the corresponding edge node r, the task is further migrated to the edge node u with lower load, where u=r means that the computing task does not perform migration. represents the set of load balancing decisions, Indicates terminal device Whether the computing task is migrated from edge node r to edge node u for execution, Indicates that the task is migrated from edge node r to edge node u, It means no migration. The task migration delay is specifically expressed as:

[0025]

[0026] in, represents the task transmission rate between edge node r and edge node u;

[0027] Considering the queuing delay caused by task transfer, the queuing process is modeled as an M / M / 1 queue. The queuing delay caused by task migration is expressed as:

[0028]

[0029] in, is the service rate of the LAN;

[0030] Therefore, the total communication delay of computing task offloading is expressed as:

[0031]

[0032] in, and They are the total task migration delay and the total task migration queuing delay of the terminal device respectively;

[0033] Establish a computing model and consider the queuing process of computing tasks waiting for processing at the edge node. This process can be modeled as an M / M / 1 queue. The computing task is offloaded to the edge node r. When the computing task is further migrated to the edge node u for execution, the computing delay of the task is expressed as:

[0034]

[0035] in, is the computing resource of edge node u, Indicates the amount of computation required by edge node u to complete all tasks to be performed; terminal device The task calculation delay is expressed as:

[0036] .

[0037] Furthermore, the resource allocation optimization problem of minimizing computation offloading latency in step (2) is specifically:

[0038]

[0039] in:

[0040] ;

[0041] in, The total latency of the task offloading, The task upload delay, is the task migration delay, The task queue delay, Calculate the latency for the task, Represents the set of edge nodes, a total of R edge nodes, represents the set of terminal devices within the coverage of edge node r, represents the transmission power set, where For terminal devices The transmission power, represents the channel resource allocation set, where is a subchannel set, with a total of K subchannels. Indicates terminal device Whether to select subchannel k transmission task, Indicates terminal device Do not select subchannel k transmission task, Indicates terminal device Select subchannel k transmission task, represents the set of load balancing decisions, Indicates terminal device Whether the computing task is migrated from edge node r to edge node u for execution, Indicates that the task is migrated from edge node r to edge node u, means no migration. represents the computing power of edge node r, Indicates terminal device The maximum transmission power, Indicates terminal device The task arrival rate is represented by a Poisson process. Indicates the execution terminal device The computing resources required for the computing task, Indicates the service rate of task transmission between edge nodes, Indicates terminal device The maximum tolerable delay of the computing task.

[0042] Furthermore, the implementation process of step (3) is as follows:

[0043] The communication resource allocation subproblem is described as a Markov decision process, whose state space includes the task data volume of the terminal device, the task calculation volume, the maximum tolerable task delay and the observed channel state information. The state space of is represented as:

[0044]

[0045] in, For terminal devices The K-dimensional vector of observed channel small-scale fading;

[0046] Define the action space of the Markov decision process. The action space of the communication resource allocation subprocess includes subchannel selection and transmission power setting. The action space is expressed as:

[0047]

[0048] in, Indicates the sub-channel selection of the terminal device, which is a K-dimensional vector;

[0049] Define the reward function of the Markov decision process to minimize the overall task upload delay, and define the reward function as the negative value of the upload delay of all end-user tasks.

[0050] Furthermore, the implementation process of step (4) is as follows:

[0051] (41) Initialize the policy network, Q-value network, and target Q-value network for each terminal device. Each agent has an independent policy network and an independent Q-value network. The policy network outputs the joint discrete action probability distribution of subchannel selection and discrete power setting. The Q-value network estimates the expected value of each action in a specific state. All network parameters are initialized randomly.

[0052] (42) Initialization temperature coefficient , the number of training rounds and the time slot t in the round;

[0053] (43) Each terminal device observes the state information of the current environment, inputs the state information into the local policy network, and randomly samples actions according to the output probability distribution of the policy network;

[0054] (44) The terminal device applies the action to the environment, obtains the corresponding reward value, and transfers to the state of the next time slot, transferring the state quadruple Stored in the local experience replay pool, where is the observed environmental state, For sampling action, is the corresponding reward value, is the state of the next time slot;

[0055] (45) If the number of steps in the current round reaches the maximum number of steps, a small batch of state transition quads are randomly sampled from the local experience replay pool as the data set for optimizing the policy network and the Q-value network;

[0056] (46) The next state of all quadruplets is input into the target Q-value network to obtain the target Q-value for a specific state-action pair, and then the target value at that moment is calculated as follows:

[0057]

[0058] in, is the discount factor, which indicates the importance of future rewards, Indicates status The output action of the next policy network The probability value of

[0059] (47) Calculate the Q network loss function, which is expressed as:

[0060]

[0061] Update the Q-value network parameters using mini-batch gradient descent method;

[0062] (48) To adaptively adjust the temperature value, the following temperature loss function needs to be minimized:

[0063]

[0064] in, is the pre-set policy entropy threshold;

[0065] (49) Every certain number of steps, the target Q value network parameters are updated using the soft update method:

[0066]

[0067] in, is the Q value network parameter, is the target Q value network parameter;

[0068] (410) Determine whether the number of rounds and time slots t are less than the total number of training rounds and total time slots. If so, proceed to step (43), otherwise terminate the training.

[0069] Furthermore, the implementation process of step (5) is as follows:

[0070] The central edge node is responsible for solving the load balancing subproblem. The load balancing subproblem is described as a Markov decision process. The load intensity indicator is defined to comprehensively reflect the load situation of the edge node:

[0071]

[0072] Define the three elements of the Markov decision process. Its state space includes the computing task information on the edge nodes and the load intensity of all edge nodes. Therefore, the state space of the load balancing sub-process is expressed as:

[0073]

[0074] in, Represent the vectors consisting of task data volume, computational workload, and maximum tolerable delay of all computing tasks, respectively. is the vector of load intensities of all edge nodes;

[0075] Define the action space of the Markov decision process. The action space of the load balancing sub-process includes the migration decisions of computing tasks on all edge computing nodes. Therefore, the action space of this process is expressed as:

[0076]

[0077] Define the reward function of the Markov decision process. To minimize the load balancing delay, define the reward function as the negative value of the sum of all end-user task migration delays, task queuing delays, and task calculation delays.

[0078] Furthermore, the implementation process of step (6) is as follows:

[0079] (61) Randomly initialize the policy network, Q value network and target Q value network for the central edge nodes;

[0080] (62) Initialization temperature coefficient , the number of rounds and the time slot t in the round;

[0081] (63) All edge nodes send the offloaded computing task information to the central edge node. The central edge node sequentially inputs the computing task information into the policy network, randomly samples actions according to the output probability distribution of the policy network, and specifies the migration target node for the task;

[0082] (64) The central edge node sends the migration decision of all computing tasks to the edge nodes in the form of control instructions. Each edge node transfers the task to the target node or local calculation according to the control instruction. When the task migration and task calculation are completed, the corresponding reward value is obtained. The central edge node transfers the state quadruple Deposit into the experience replay pool;

[0083] (65) If the number of steps in the current round reaches the maximum number of steps, a small batch of quads is randomly sampled from the experience replay pool as the data set for optimizing the policy network and the Q-value network;

[0084] (66) The next state of all quadruplets is input into the target Q-value network to obtain the target Q-value for a specific state-action pair, and then the target value at that moment is calculated as follows:

[0085]

[0086] in, is the discount factor, which indicates the importance of future rewards, Indicates status The output action of the next policy network The probability value of

[0087] (67) Calculate the Q network loss function, which is expressed as:

[0088]

[0089] Update the Q-value network parameters using mini-batch gradient descent method;

[0090] (68) To adaptively adjust the temperature value, the following temperature loss function needs to be minimized:

[0091]

[0092] in, is the pre-set policy entropy threshold;

[0093] (69) Every certain number of steps, the target Q value network parameters are updated using the soft update method:

[0094]

[0095] in, is the Q value network parameter, is the target Q value network parameter;

[0096] (610) Determine whether the number of rounds and time slots t are less than the total number of training rounds and total time slots. If so, proceed to step (63), otherwise terminate the training.

[0097] The present invention provides a smart coal mine computing unloading system based on edge collaboration, including edge nodes, central edge nodes and terminal devices deployed on the working face of the mine; the edge nodes are used to receive task unloading requests from terminal devices within their coverage, regularly report unloading task information to the central edge nodes, and receive load balancing decisions from the central edge node side, and then perform task migration and task execution; the central edge node is located at the topological center point of the geographical location of all edge nodes, has more powerful computing resources, allocates load balancing decisions to all edge nodes, and provides edge computing services to terminal devices within its coverage; the terminal devices generate real-time computing tasks, observe channel information and task information locally, and allocate channel resources and set power values ​​for uploading unloading tasks based on a deep reinforcement learning model.

[0098] Beneficial effects: Compared with the prior art, the beneficial effects of the present invention are as follows: the present invention is aimed at the computing architecture of multi-edge node collaboration in coal mines. In the application scenario of smart coal mines, it makes intelligent offloading decisions on the real-time computing tasks generated by terminal devices, and realizes the joint optimization of channel resources and power to meet the computing requirements of high real-time performance and high reliability of underground equipment; the edge nodes in the system are deployed at key positions to form a distributed computing power network, and the central edge node is located at the topological center point for global load balancing, which can effectively solve the problems of unstable channels and uneven distribution of computing resources in the mine environment; the terminal equipment autonomously perceives the channel status through the deep reinforcement learning model and dynamically adjusts the offloading strategy to achieve efficient diversion of computing tasks; the deep reinforcement learning framework adopted by the present invention can significantly reduce task processing delay, improve system resource utilization, and achieve a good balance between algorithm efficiency and computing performance; in summary, in the distributed edge computing scenario of smart coal mines, the present invention has significant advantages in improving task execution efficiency and ensuring system stability. BRIEF DESCRIPTION OF THE DRAWINGS

[0099] Figure 1 Flowchart of the calculation unloading method for smart coal mine based on edge collaboration;

[0100] Figure 2 This is a schematic diagram of the training process of the deep reinforcement learning algorithm proposed in the present invention;

[0101] Figure 3 The graph of training reward value changes of different algorithms in the communication resource allocation stage;

[0102] Figure 4 This is a graph showing the changes in training reward values ​​for different algorithms during the load balancing phase.

[0103] Figure 5 The simulation results of different algorithms under different task data amounts are shown;

[0104] Figure 6The simulation results of different algorithms under different edge node computing resources are shown. DETAILED DESCRIPTION

[0105] The present invention is further described in detail below with reference to the accompanying drawings.

[0106] like Figure 1 As shown, the present invention proposes a smart coal mine computing offloading method based on edge collaboration, which includes the following steps:

[0107] Step 1: The terminal device at the smart mine excavation face generates real-time computing tasks and collects status information; the status information includes the task data volume of the terminal device, the task computing volume, the maximum tolerable delay of the task, and the observed channel status information.

[0108] Step 2: Establish the system communication model and computing model, and then establish the resource allocation optimization problem of minimizing the computing offloading delay. Considering the separability of the optimization problem, the optimization problem is decomposed into the communication resource allocation sub-problem and the load balancing sub-problem.

[0109] (2a) Establishing a smart coal mine computing and unloading system, Represents the set of edge nodes, a total of R edge nodes, Represents the set of terminal devices within the coverage of edge node r, terminal devices The arrival of computing tasks can be described as the arrival rate The Poisson process generates a computing task in each time slot with available triples express, The amount of data for the calculation task, The amount of computation required to complete this task is The maximum tolerable delay for task execution.

[0110] (2b) Establish system communication model, terminal equipment The task upload delay between the edge computing node r can be expressed as:

[0111]

[0112] in, For terminal devices The task upload rate between the edge computing node r can be obtained by the following Shannon formula:

[0113]

[0114] Where B is the selected subchannel bandwidth, For terminal devices Set the transmission power, is the additive white Gaussian noise power, Indicates terminal device The channel gain on subchannel k is, represents the interference between terminal devices and can be calculated using the following formulas:

[0115]

[0116]

[0117] in, For terminal devices and the small-scale fading of edge node r on subchannel k, For terminal devices The Euclidean distance to the edge node r, is the path loss, using Represents the subchannel set, there are K subchannels in total, Indicates terminal device Whether to select subchannel k transmission task, Indicates terminal device Do not select subchannel k transmission task, Indicates terminal device Select subchannel k to transmit the task. Since the amount of data of the calculation result is smaller than the original data, the return time of the calculation result can be ignored.

[0118] When the terminal device After the computing task reaches the corresponding edge node r, the task can be further migrated to the edge node u with lower load, where u=r means that the computing task does not perform migration. represents the set of load balancing decisions, Indicates terminal device Whether the computing task is migrated from edge node r to edge node u for execution, Indicates that the task is migrated from edge node r to edge node u, It means no migration. The task migration delay can be specifically expressed as:

[0119]

[0120] in, It represents the task transmission rate between edge node r and edge node u, which depends on the physical properties of the local area network. Considering the additional control cost generated by multiple rounds of migration, it is limited that a task can only be migrated once in the local area network.

[0121] In addition to the task migration delay, due to the serial nature of task transmission in the LAN, the queuing delay caused by task transmission also needs to be considered. This queuing process can be modeled as an M / M / 1 queue. Therefore, the queuing delay caused by task migration can be expressed as:

[0122]

[0123] in, The service rate of the LAN.

[0124] Therefore, the total communication delay of computing task offloading can be expressed as:

[0125]

[0126] in, and They are the total task migration delay and the total task migration queuing delay of the terminal device respectively.

[0127] (2c) Establish a system computing model, considering the queuing process of computing tasks waiting for processing at the edge nodes. This process can be modeled as an M / M / 1 queue. When the computing task is offloaded to the edge node r and is further migrated to the edge node u for execution, the computing delay of the task can be expressed as:

[0128]

[0129] in, is the computing resource of edge node u, represents the amount of computation required by the edge node u to complete all the tasks that need to be performed. The task computation delay can be expressed as:

[0130] .

[0131] (2d) Establish a communication resource allocation and load balancing optimization model that minimizes the total latency of computation offloading. In summary, the optimization problem can be expressed as:

[0132]

[0133] in:

[0134]

[0135] Among them, among them, The total latency of the task offloading, The task upload delay, is the task migration delay, The task queue delay, Calculate the latency for the task, Represents the set of edge nodes, a total of R edge nodes, represents the set of terminal devices within the coverage of edge node r, represents the transmission power set, where For terminal devices The transmission power, represents the channel resource allocation set, where is a subchannel set, with a total of K subchannels. Indicates terminal device Whether to select subchannel k transmission task, Indicates terminal device Do not select subchannel k transmission task, Indicates terminal device Select subchannel k transmission task, represents the set of load balancing decisions, Indicates terminal device Whether the computing task is migrated from edge node r to edge node u for execution, Indicates that the task is migrated from edge node r to edge node u, means no migration. represents the computing power of edge node r, Indicates terminal device The maximum transmission power, Indicates terminal device The task arrival rate is represented by a Poisson process. Indicates the execution terminal device The computing resources required for the computing task, Indicates the service rate of task transmission between edge nodes, Indicates terminal device The maximum tolerable delay of the computing task.

[0136] The optimization problem can be decomposed into two independent sub-problems: the communication resource allocation sub-problem P2 and the load balancing sub-problem P3, which can be expressed as:

[0137]

[0138] .

[0139] Step 3: For the communication resource allocation sub-problem, the terminal device locally establishes a Markov decision process to describe the communication resource allocation process, and defines its state space, action space, and reward function.

[0140] (3a) The communication resource allocation subproblem is described as a Markov decision process, whose state space includes the task data volume of the terminal device, the task computation volume, the maximum tolerable task delay and the observed channel state information. Therefore, the terminal device The state space of is represented as:

[0141]

[0142] in, For terminal devices The K-dimensional vector consisting of the observed small-scale fading of the channel.

[0143] (3b) Define the action space of the Markov decision process. The action space of the communication resource allocation subprocess includes subchannel selection and transmission power setting. Therefore, the terminal device The action space is expressed as:

[0144]

[0145] in, represents the subchannel selection of the terminal device and is a K-dimensional vector. To meet the actual circuit design requirements, the transmission power value is discretized into four fixed values, namely [0, 5, 10, 23] dBm.

[0146] (3c) Define the reward function of the Markov decision process to minimize the overall task upload delay, and define the reward function as the negative of the upload delay of all end-user tasks.

[0147] Step 4: The local terminal device trains the communication resource allocation decision model based on the Multi-Agent Soft Actor-Critic (MASAC) algorithm in the discrete action space of multiple agents. Figure 2 As shown in Figure 2, the algorithm describes the process of updating the deep reinforcement learning network during the training phase, including the following specific steps:

[0148] (4a) Initialize the policy network, Q-value network, and target Q-value network for each terminal device. Each agent has an independent policy network and an independent Q-value network. The policy network outputs the joint discrete action probability distribution of subchannel selection and discrete power setting. The Q-value network estimates the expected value of each action in a specific state. All network parameters are initialized randomly.

[0149] (4b) Initialize temperature coefficient , the number of training rounds and the time slot t in the round.

[0150] (4c) Each terminal device observes the state information of the current environment, inputs the state information into the local policy network, and randomly samples actions according to the output probability distribution of the policy network.

[0151] (4d) The terminal device applies the action to the environment, obtains the corresponding reward value, and transfers to the state of the next time slot, transferring the state quadruple Stored in the local experience replay pool, where is the observed environmental state, For sampling action, is the corresponding reward value, The state of the next time slot.

[0152] (4e) If the number of steps in the current round reaches the maximum number of steps, a small batch of state transition quads are randomly sampled from the local experience replay pool as the dataset for optimizing the policy network and Q-value network.

[0153] (4f) Input the next state of all quadruplets into the target Q-value network to obtain the target Q-value for a specific state-action pair, and then calculate the target value at that moment as follows:

[0154]

[0155] in, is the discount factor, which indicates the importance of future rewards, Indicates status The output action of the next policy network The probability value of .

[0156] (4g) Calculate the Q network loss function, which is expressed as:

[0157]

[0158] Update the Q-value network parameters using mini-batch gradient descent method;

[0159] (4h) To adaptively adjust the temperature value, the following temperature loss function needs to be minimized:

[0160]

[0161] in, is a pre-set policy entropy threshold.

[0162] (4i) Every certain number of steps, the target Q value network parameters are updated using the soft update method:

[0163]

[0164] in, is the Q value network parameter, is the target Q value network parameter.

[0165] (4j) Determine whether the number of rounds and time slots t are less than the total number of training rounds and total time slots. If so, proceed to step (4c), otherwise end the training.

[0166] Step 5: For the load balancing sub-problem, the central edge node collects the status information of all terminal devices, establishes a Markov decision process to describe the task migration process, and defines its state space, action space, and reward function.

[0167] (5a) When all computing tasks are offloaded to the edge computing nodes, the edge computing nodes transmit the task information of the local computing tasks to the central edge nodes for further processing.

[0168] (5b) The central edge node is responsible for solving the load balancing subproblem, which is described as a Markov decision process. Due to the differences in computing power and task queue lengths among different edge nodes, considering computing power and task queue length alone cannot provide the optimal reference for load balancing decisions. Therefore, the load intensity indicator is defined to comprehensively reflect the load situation of edge nodes:

[0169]

[0170] The three elements of the Markov decision process are defined. Its state space includes the computing task information on the edge nodes and the load intensity of all edge nodes. Therefore, the state space of the load balancing sub-process can be expressed as:

[0171]

[0172] in, Represent the vectors consisting of task data volume, computational workload, and maximum tolerable delay of all computing tasks, respectively. It is a vector of the load strength of all edge nodes.

[0173] (5c) Define the action space of the Markov decision process. The action space of the load balancing subprocess includes the migration decisions of computing tasks on all edge computing nodes. Therefore, the action space of this process is expressed as:

[0174] .

[0175] (5d) Define the reward function of the Markov decision process. To minimize the load balancing delay, define the reward function as the negative of the sum of all end-user task migration delays, task queuing delays, and task computation delays.

[0176] Step 6: The central edge node trains the load balancing decision model based on the Soft Actor-Critic (SAC) algorithm.

[0177] (6a) Randomly initialize the policy network, Q-value network, and target Q-value network for the central edge nodes.

[0178] (6b) Initialize temperature coefficient , the number of rounds and the time slot t in the round.

[0179] (6c) All edge nodes send the offloaded computing task information to the central edge node. The central edge node sequentially inputs the computing task information into the policy network, randomly samples actions according to the output probability distribution of the policy network, and specifies the migration target node for the task.

[0180] (6d) The central edge node sends the migration decision of all computing tasks to the edge nodes in the form of control instructions. Each edge node transfers the task to the target node or local calculation according to the control instructions. When the task migration and task calculation are completed, the corresponding reward value is obtained. The central edge node transfers the state quadruple Stored in the experience replay pool.

[0181] (6e) If the number of steps in the current round reaches the maximum number of steps, a small batch of quads is randomly sampled from the experience replay pool as the dataset for optimizing the policy network and the Q-value network.

[0182] (6f) Input the next state of all quadruplets into the target Q-value network to obtain the target Q-value for a specific state-action pair, and then calculate the target value at that moment as follows:

[0183]

[0184] in, is the discount factor, which indicates the importance of future rewards, Indicates status The output action of the next policy network The probability value of .

[0185] (6g) Calculate the Q network loss function, which is expressed as:

[0186]

[0187] The Q-value network parameters are updated using mini-batch gradient descent.

[0188] (6h) To adaptively adjust the temperature value, the following temperature loss function needs to be minimized:

[0189]

[0190] in, is a pre-set policy entropy threshold.

[0191] (6i) Every certain number of steps, the target Q value network parameters are updated using the soft update method:

[0192]

[0193] in, is the Q value network parameter, is the target Q value network parameter.

[0194] (6j) Determine whether the number of rounds and time slots t are less than the total number of training rounds and total time slots. If so, proceed to step (6c), otherwise end the training.

[0195] Step 7: During the execution phase, the terminal device observes the channel state information and task information in real time, obtains the optimal communication resource allocation decision based on the trained MASAC model, and sends the computing task to the edge computing node based on this decision; then each edge computing node uploads the local load information to the central edge node. The central edge node obtains the optimal load balancing decision based on the trained SAC model and delegates it to all edge nodes to complete the inter-node migration of the task.

[0196] The present invention also provides a smart coal mine computing unloading system based on edge collaboration, including edge nodes, central edge nodes and terminal devices deployed on the working face of the mine. The edge node is used to receive task unloading requests from terminal devices within its coverage area, regularly report unloading task information to the central edge node, and receive load balancing decisions from the central edge node side, and then perform task migration and task execution. The central edge node is located at the topological center point of the geographical location of all edge nodes, has more powerful computing resources, allocates load balancing decisions to all edge nodes, and provides edge computing services to terminal devices within its coverage area. The terminal device generates real-time computing tasks, observes channel information and task information locally, and allocates channel resources and sets power values ​​for uploading unloading tasks based on the deep reinforcement learning model.

[0197] In order to verify the effectiveness of the method disclosed in the present invention, a smart coal mine simulation environment was constructed based on the Python platform, and the computational offloading method disclosed in the present invention was implemented based on the Pytorch deep learning framework.

[0198] During the training phase of the algorithm, the simulation results of the communication resource allocation sub-process and the load balancing sub-process are as follows: Figure 3 and Figure 4 shown; among them, Figure 3is the change in the training reward value of the communication resource allocation sub-process. MADDQN is a communication resource allocation algorithm based on a multi-agent double deep Q-network (Multi-Agent Double Deep Q-Network). Figure 4 is the change in the training reward value of the load balancing sub-process, and DDQN is a load balancing algorithm based on the deep double Q network (Double Deep Q-Network). Figure 3 As shown in the figure, as the number of training rounds increases, the reward values ​​of the MASAC algorithm and the MADDQN algorithm gradually increase, but compared with the MADDQN algorithm, the reward value of the MASAC algorithm increases at a higher rate and can reach a higher reward value after convergence, indicating that the MASAC algorithm disclosed in the present invention can achieve lower end-to-end task upload latency in the communication resource allocation stage. Similarly, in Figure 4 The SAC algorithm disclosed in this invention also achieved higher reward values ​​and faster convergence rates than the DDQN algorithm. Based on the simulation results of the above two stages, during the training phase, the edge collaboration-based smart coal mine computation offloading method disclosed in this invention can search for more optimal actions in the action space, thereby minimizing the total computation offloading latency.

[0199] After the model training is completed, the model is deployed to the terminal device and edge node in the testing phase to test the performance of the model under different environment settings. The test results are as follows: Figure 5 and Figure 6 As shown. Among them, comparison method 1 is the edgeless cooperative communication resource allocation algorithm based on MASAC, comparison method 2 is the computational offloading algorithm based on SAC edge cooperation and random communication resource allocation, and comparison method 3 is the communication resource allocation based on MADDQN and the load balancing algorithm based on DDQN. Figure 5 In this paper, the performance of different algorithms under different task data volumes is verified. As the task data volume increases, the total computational offloading delay of all algorithms shows an upward trend, but the ECCO algorithm proposed in this invention can achieve the lowest total delay. Figure 6 The performance of different algorithms under different edge node computing capabilities is demonstrated. The improvement of edge node computing power will lead to a decrease in the computing latency of the same task, thereby reducing the total latency of computing offloading. The simulation results also show that the ECCO algorithm proposed in this invention can show optimal performance under different environmental settings and has stronger generalization performance.

[0200] In summary, the present invention uses deep reinforcement learning to model and optimize the collaborative relationship between terminal device computing tasks and edge computing resources, which can effectively alleviate the problem of uneven distribution of computing resources, reduce communication and queuing delays, and improve task processing efficiency. In the specific implementation process, the terminal device independently decides the transmission power and transmission sub-channel by constructing a Markov decision process for communication resource allocation, while the central edge server optimizes the load balancing strategy from a global perspective, thereby realizing intelligent computing power scheduling among multiple devices and multiple nodes. The deep reinforcement learning-driven computing power scheduling strategy proposed in the present invention not only takes into account the heterogeneous computing needs and dynamic channel characteristics of the terminal device, but also takes into account the global load status of the edge server. It can maximize the overall resource utilization of the system while ensuring the timeliness of the task. Compared with traditional fixed strategies or heuristic methods, the present invention exhibits better robustness and adaptability in multi-task concurrency and highly dynamic environments. The deep reinforcement learning model can be deployed on the terminal device and the central edge node, and the optimization strategy for load balancing is trained by periodically collecting the status information of the edge node. After the model training is completed, it can quickly infer the calculation offloading path and load balancing decision based on the real-time status of the actual operation site, adapt to the complex and changeable communication and computing environment of the mining area, and effectively ensure the continuity and intelligence level of mining operations.

[0201] The above contents are examples of the methods and structures of the present invention. Any modifications or similar substitutions made by technicians in this technical field to the described specific embodiments shall fall within the scope of protection of the present invention as long as they do not deviate from the methods of the present invention or exceed the scope defined by the claims.

Claims

1. A smart coal mine computing offloading method based on edge collaboration, characterized in that: The following steps are involved: (1) The multi-type delay-sensitive computing tasks generated by the terminal equipment in the mining area during operation are abstracted into specific task information based on task type, computing requirements, and maximum tolerable delay; the terminal equipment makes computing offloading decisions based on the task information and local computing resources; (2) The terminal device regularly estimates the channel state information, and the central edge server regularly collects global load information, establishing a communication model and a computation model that includes the terminal device and the edge computing node. This then constructs a resource allocation optimization problem that minimizes the computation offloading delay, and decomposes the optimization problem into a communication resource allocation sub-problem and a load balancing sub-problem. (3) For the communication resource allocation sub-problem, the terminal device establishes a Markov decision process locally to describe the communication resource allocation process and defines its state space, action space, and reward function; (4) The local terminal device trains the communication resource allocation decision model based on the multi-agent flexible actor-critic algorithm; (5) For the load balancing sub-problem, the central edge node collects the status information of all terminal devices, establishes a Markov decision process to describe the task migration process, and defines its state space, action space, and reward function; (6) The central edge node trains the load balancing decision model based on the flexible actor-critic algorithm; (7) The terminal device observes the channel state information and task information in real time, obtains the optimal communication resource allocation decision based on the trained communication resource allocation decision model, and sends the computing task to the edge computing node based on this decision; then each edge computing node uploads the local load information to the central edge node, and the central edge node obtains the optimal load balancing decision based on the trained load balancing decision model and delegates it to all edge nodes to complete the inter-node migration of the task; The implementation process of establishing the communication model and computing model including the terminal device and the edge computing node in step (2) is as follows: Establish a smart coal mine computing and unloading system, Represents the set of edge nodes, a total of R edge nodes, Represents the set of terminal devices within the coverage of edge node r, terminal devices The arrival of computing tasks can be described as the arrival rate The Poisson process generates a computing task in each time slot with available triples express, The amount of data for the calculation task, The amount of computation required to complete this task is The maximum tolerable delay for task execution; Establish communication model, terminal equipment The task upload delay between the edge computing node r is expressed as: ; in, For terminal devices The task upload rate between the edge computing node r is obtained by the following Shannon formula: ; Where B is the selected subchannel bandwidth, For terminal devices Set the transmission power, is the additive white Gaussian noise power, Indicates terminal device The channel gain on subchannel k is, represents the interference between terminal devices and is calculated using the following formulas: ; ; in, For terminal devices and the small-scale fading of edge node r on subchannel k, For terminal devices The Euclidean distance to the edge node r, is the path loss, using Represents the subchannel set, there are K subchannels in total, Indicates terminal device Whether to select subchannel k transmission task, Indicates terminal device Do not select subchannel k transmission task, Indicates terminal device Select subchannel k transmission task; When the terminal device After the computing task reaches the corresponding edge node r, the task is further migrated to the edge node u with lower load, where u=r means that the computing task does not perform migration. represents the set of load balancing decisions, Indicates terminal device Whether the computing task is migrated from edge node r to edge node u for execution, Indicates that the task is migrated from edge node r to edge node u, It means no migration. The task migration delay is specifically expressed as: ; in, represents the task transmission rate between edge node r and edge node u; Considering the queuing delay caused by task transfer, the queuing process is modeled as an M / M / 1 queue. The queuing delay caused by task migration is expressed as: ; in, is the service rate of the LAN; Therefore, the total communication delay of computing task offloading is expressed as: ; in, and They are the total task migration delay and the total task migration queuing delay of the terminal device respectively; Establish a computing model and consider the queuing process of computing tasks waiting for processing at the edge node. This process can be modeled as an M / M / 1 queue. The computing task is offloaded to the edge node r. When the computing task is further migrated to the edge node u for execution, the computing delay of the task is expressed as: ; in, is the computing resource of edge node u, Indicates the amount of computation required by edge node u to complete all tasks to be performed; terminal device The task calculation delay is expressed as: 。 2. The method for intelligent coal mine computing offloading based on edge collaboration according to claim 1 is characterized in that: The resource allocation optimization problem of minimizing computation offloading latency in step (2) is specifically: ; in: ; in, The total latency of the task offloading, The task upload delay, is the task migration delay, The task queue delay, Calculate the latency for the task, Represents the set of edge nodes, a total of R edge nodes, represents the set of terminal devices within the coverage of edge node r, represents the transmission power set, where For terminal devices The transmission power, represents the channel resource allocation set, where is a subchannel set, with a total of K subchannels. Indicates terminal device Whether to select subchannel k transmission task, Indicates terminal device Do not select subchannel k transmission task, Indicates terminal device Select subchannel k transmission task, represents the set of load balancing decisions, Indicates terminal device Whether the computing task is migrated from edge node r to edge node u for execution, Indicates that the task is migrated from edge node r to edge node u, means no migration. represents the computing power of edge node r, Indicates terminal device The maximum transmission power, Indicates terminal device The task arrival rate is represented by a Poisson process. Indicates the execution terminal device The computing resources required for the computing task, Indicates the service rate of task transmission between edge nodes, Indicates terminal device The maximum tolerable delay of the computing task.

3. The method for intelligent coal mine computing offloading based on edge collaboration according to claim 1 is characterized in that: The implementation process of step (3) is as follows: The communication resource allocation subproblem is described as a Markov decision process, whose state space includes the task data volume of the terminal device, the task calculation volume, the maximum tolerable task delay and the observed channel state information. The state space of is represented as: ; in, For terminal devices The K-dimensional vector of observed channel small-scale fading; Define the action space of the Markov decision process. The action space of the communication resource allocation subprocess includes subchannel selection and transmission power setting. The action space is expressed as: ; in, Indicates the sub-channel selection of the terminal device, which is a K-dimensional vector; Define the reward function of the Markov decision process to minimize the overall task upload delay, and define the reward function as the negative value of the upload delay of all end-user tasks.

4. The method for intelligent coal mine computing offloading based on edge collaboration according to claim 1 is characterized in that: The implementation process of step (4) is as follows: (41) Initialize the policy network, Q-value network, and target Q-value network for each terminal device. Each agent has an independent policy network and an independent Q-value network. The policy network outputs the joint discrete action probability distribution of subchannel selection and discrete power setting. The Q-value network estimates the expected value of each action in a specific state. All network parameters are initialized randomly. (42) Initialization temperature coefficient , the number of training rounds and the time slot t in the round; (43) Each terminal device observes the state information of the current environment, inputs the state information into the local policy network, and randomly samples actions according to the output probability distribution of the policy network; (44) The terminal device applies the action to the environment, obtains the corresponding reward value, and transfers to the state of the next time slot, transferring the state quadruple Stored in the local experience replay pool, where is the observed environmental state, For sampling action, is the corresponding reward value, is the state of the next time slot; (45) If the number of steps in the current round reaches the maximum number of steps, a small batch of state transition quads are randomly sampled from the local experience replay pool as the data set for optimizing the policy network and the Q-value network; (46) The next state of all quadruplets is input into the target Q-value network to obtain the target Q-value for a specific state-action pair, and then the target value at that moment is calculated as follows: ; in, is the discount factor, which indicates the importance of future rewards, Indicates status The output action of the next policy network The probability value of (47) Calculate the Q network loss function, which is expressed as: ; Update the Q-value network parameters using mini-batch gradient descent method; (48) To adaptively adjust the temperature value, the following temperature loss function needs to be minimized: ; in, is the pre-set policy entropy threshold; (49) Every certain number of steps, the target Q value network parameters are updated using the soft update method: ; in, is the Q value network parameter, is the target Q value network parameter; (410) Determine whether the number of rounds and time slots t are less than the total number of training rounds and total time slots. If so, proceed to step (43), otherwise terminate the training.

5. The method for intelligent coal mine computing offloading based on edge collaboration according to claim 1 is characterized in that: The implementation process of step (5) is as follows: The central edge node is responsible for solving the load balancing subproblem. The load balancing subproblem is described as a Markov decision process. The load intensity indicator is defined to comprehensively reflect the load situation of the edge node: ; Define the three elements of the Markov decision process. Its state space includes the computing task information on the edge nodes and the load intensity of all edge nodes. Therefore, the state space of the load balancing sub-process is expressed as: ; in, Represent the vectors consisting of task data volume, computational workload, and maximum tolerable delay of all computing tasks, respectively. is the vector of load intensities of all edge nodes; Define the action space of the Markov decision process. The action space of the load balancing sub-process includes the migration decisions of computing tasks on all edge computing nodes. Therefore, the action space of this process is expressed as: ; Define the reward function of the Markov decision process. To minimize the load balancing delay, define the reward function as the negative value of the sum of all end-user task migration delays, task queuing delays, and task calculation delays.

6. The method for intelligent coal mine computing offloading based on edge collaboration according to claim 1 is characterized in that: The implementation process of step (6) is as follows: (61) Randomly initialize the policy network, Q value network and target Q value network for the central edge nodes; (62) Initialization temperature coefficient , the number of rounds and the time slot t in the round; (63) All edge nodes send the offloaded computing task information to the central edge node. The central edge node sequentially inputs the computing task information into the policy network, randomly samples actions according to the output probability distribution of the policy network, and specifies the migration target node for the task; (64) The central edge node sends the migration decision of all computing tasks to the edge nodes in the form of control instructions. Each edge node transfers the task to the target node or local calculation according to the control instruction. When the task migration and task calculation are completed, the corresponding reward value is obtained. The central edge node transfers the state quadruple Deposit into the experience replay pool; (65) If the number of steps in the current round reaches the maximum number of steps, a small batch of quads is randomly sampled from the experience replay pool as the data set for optimizing the policy network and the Q-value network; (66) The next state of all quadruplets is input into the target Q-value network to obtain the target Q-value for a specific state-action pair, and then the target value at that moment is calculated as follows: ; in, is the discount factor, which indicates the importance of future rewards, Indicates status The output action of the next policy network The probability value of (67) Calculate the Q network loss function, which is expressed as: ; Update the Q-value network parameters using mini-batch gradient descent method; (68) To adaptively adjust the temperature value, the following temperature loss function needs to be minimized: ; in, is the pre-set policy entropy threshold; (69) Every certain number of steps, the target Q value network parameters are updated using the soft update method: ; in, is the Q value network parameter, is the target Q value network parameter; (610) Determine whether the number of rounds and time slots t are less than the total number of training rounds and total time slots. If so, proceed to step (63), otherwise terminate the training.

7. A smart coal mine computing and unloading system based on edge collaboration using the method according to any one of claims 1 to 6, characterized in that: Including edge nodes, central edge nodes and terminal equipment deployed at the mine working face; The edge node is used to receive task offloading requests from terminal devices within its coverage area, regularly report offloading task information to the central edge node, and receive load balancing decisions from the central edge node, thereby performing task migration and task execution; the central edge node is located at the topological center point of the geographical locations of all edge nodes, has more powerful computing resources, allocates load balancing decisions to all edge nodes, and provides edge computing services to terminal devices within its coverage area; the terminal device generates real-time computing tasks, observes channel information and task information locally, and allocates channel resources and sets power values ​​for uploading offloading tasks based on a deep reinforcement learning model.

Citation Information

Patent Citations

  • Mine Internet of Things intelligent calculation unloading method assisted by energy collection

    CN115686669A

  • Calculation unloading and resource allocation method for vehicle-road cooperation system

    CN116896561A