Distributed anti-interference collaborative optimization method under mine wireless network based on multi-agent reinforcement learning

By employing a distributed anti-interference collaborative optimization method based on multi-agent reinforcement learning in a mine wireless network, the electromagnetic interference problem of the underground wireless communication system was solved, achieving efficient anti-interference and dynamic routing optimization, and improving communication reliability and efficiency.

CN121174170APending Publication Date: 2025-12-19NANJING UNIV OF SCI & TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511402057.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-28
Publication Date
2025-12-19

AI Technical Summary

Technical Problem

Underground wireless communication systems are susceptible to complex electromagnetic interference, leading to increased bit error rate, data packet loss, and communication interruptions. Existing anti-interference technologies are insufficient in terms of dynamic adaptability and accuracy.

Method used

A distributed anti-interference collaborative optimization method based on multi-agent reinforcement learning is adopted. Electromagnetic wave propagation is simulated in the NS-3 simulation environment. The interference source is located by combining the improved TDOA algorithm and the generalized cross-correlation algorithm. A multi-agent reinforcement learning framework is constructed to dynamically optimize the routing strategy and channel allocation.

Benefits of technology

It improves the anti-interference capability and communication efficiency of wireless communication in mines, reduces packet loss rate, reduces path optimization steps, and enhances system robustness and communication quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121174170A_ABST
    Figure CN121174170A_ABST
Patent Text Reader

Abstract

The invention provides a distributed anti-interference collaborative optimization method under a mine wireless network based on multi-agent reinforcement learning, and provides a distributed routing and interference suppression collaborative optimization solution for solving the problem that the communication reliability of a wireless sensor network is reduced due to transient electromagnetic interference generated by large-scale electrical equipment under a mine. According to the method, typical mine electromagnetic interference characteristics such as impulse noise and sweep frequency interference are integrated through multi-mode interference, a multi-agent reinforcement learning (MARL) architecture is designed, collaborative optimization of distributed routing and interference suppression is achieved, and interference source positioning, dynamic channel hopping and intelligent rerouting strategies based on spectrum sensing are included. The anti-interference routing protocol is intelligently and dynamically optimized, an interference area is effectively avoided while the shortest path is guaranteed, the anti-interference capability and the communication stability of the mine wireless sensor network are remarkably improved, and an interpretable and extensible intelligent communication solution is provided for a complex industrial environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of underground communication technology, specifically to a distributed anti-interference collaborative optimization method for underground wireless networks based on multi-agent reinforcement learning. Background Technology

[0002] As a core industry in the nation's energy supply, coal mining's safe production heavily relies on the reliability of underground wireless communication systems. With the advancement of intelligent mining, numerous wireless sensor networks and IoT devices have been deployed underground for environmental monitoring, equipment control, and personnel positioning. However, the mining environment suffers from complex electromagnetic interference, especially the transient electromagnetic pulses generated during the startup, operation, or shutdown of large electrical equipment, which severely interfere with wireless communication systems. Studies have shown that such interference can cause the bit error rate of wireless links to surge by 3-4 orders of magnitude, leading to data packet loss, communication interruptions, and even equipment malfunctions, threatening underground operational safety.

[0003] Currently, anti-interference technologies for mine wireless communication mainly rely on frequency domain avoidance, spatial domain avoidance, and traditional filtering techniques. Frequency domain avoidance employs frequency hopping or adaptive channel selection, but its switching efficiency is insufficient under transient broadband interference. Spatial domain avoidance is based on route optimization of fixed topology and cannot dynamically adapt to the mobility of interference sources. Traditional filtering techniques, such as FIR / IIR filters, have limited effectiveness in suppressing non-stationary transient interference. Although existing methods have alleviated the interference problem in mine communication to some extent, they still have key shortcomings such as insufficient interference modeling accuracy and poor dynamic adaptability. Summary of the Invention

[0004] The purpose of this invention is to propose a distributed anti-interference collaborative optimization method for mine wireless networks based on multi-agent reinforcement learning, so as to solve the problem of decreased wireless communication reliability caused by electromagnetic interference in mines and improve the anti-interference capability and communication efficiency of distributed wireless sensor networks.

[0005] To achieve the above objectives, the technical solution adopted by this invention is as follows: A distributed anti-interference cooperative optimization method based on multi-agent reinforcement learning in a mine wireless network, comprising the following steps:

[0006] Step 1: Input the 3D model file of the mine roadway, the list of interference source parameters, and the electromagnetic property parameters of the roadway wall into the Network Simulator 3 platform. Configure the ray tracing propagation model to simulate the multipath reflection, diffraction, and attenuation behavior of electromagnetic waves in the roadway. Instantiate the impulse noise source and the frequency sweep interference source according to the parameters and place them in the 3D spatial model of the roadway. Finally, an NS-3 simulation environment containing real electromagnetic propagation characteristics and dynamic interference sources is constructed.

[0007] Step 2: In the constructed NS-3 simulation environment, deploy distributed wireless sensor network nodes according to the preset number of nodes, initial locations, and communication radius; each node deploys a real-time spectrum feature sensing module to continuously scan the communication frequency band and output the operating status of the distributed wireless sensor network, which includes node deployment information and real-time spectrum interference detection data.

[0008] Step 3: Utilize the spectrum interference detection data collected in real time by the spectrum sensing modules of each node, combined with the known node locations and the current wireless sensor network topology, to perform interference source localization using an improved TDOA algorithm. Introduce a sliding window to process transient signals, use a generalized cross-correlation algorithm to improve the time delay estimation accuracy, and calculate the three-dimensional coordinates of the interference source based on the weighted least squares method. Finally, integrate the obtained interference source localization results, wireless sensor network topology status, and node load information to form the state space of multi-agent reinforcement learning.

[0009] Step 4: Based on the state space and action space of the available next-hop neighbor node set, the action of the multi-agent reinforcement learning framework is constructed using the design method of bi-objective reward function R = α·path quality + β·interference avoidance parameter. Furthermore, the control signaling mechanism for policy update information during the learning process is integrated into the packet routing simulator, allowing agents to interact with neural network weights or target values ​​through model sharing or value sharing, thereby achieving collaborative optimization of routing policies.

[0010] Step 5: Connect the multi-agent reinforcement learning framework to the NS-3 simulation environment using the NS3-gym interface, and set the learning rate and discount factor; train the agents using the Asynchronous Advantage (A3C) algorithm: calculate the advantage value based on the reward function, construct the Actor-Critic network loss function, and iteratively update the neural network parameters through gradient descent and backpropagation until the model converges, thereby obtaining the trained reinforcement learning model;

[0011] Step 6: Load the trained reinforcement learning model into each actual wireless sensor network node. Based on the real-time collected wireless sensor network status and interference source information, dynamically execute anti-interference strategies: apply Q-learning to achieve adaptive channel hopping to avoid frequency domain interference, establish a spatial avoidance radius based on the location of the interference source, trigger local rerouting, and dynamically update the node routing table and channel allocation strategy.

[0012] Further, in step 1, the 3D model file of the mine roadway, the list of interference source parameters, and the electromagnetic property parameters of the roadway wall are input into the Network Simulator 3 platform. A ray tracing propagation model is configured to simulate the multipath reflection, diffraction, and attenuation behavior of electromagnetic waves in the roadway. Based on the parameters, impulse noise sources and frequency sweeping interference sources are instantiated and deployed in the 3D spatial model of the roadway. Finally, an NS-3 simulation environment containing realistic electromagnetic propagation characteristics and dynamic interference sources is constructed. The specific method is as follows:

[0013] Step 11: Based on the collected real mine tunnel scene, generate a 3D model file of the mine tunnel, import it into the Network Simulator 3 platform, and generate a 3D spatial model of the tunnel.

[0014] Step 12: Based on the multipath reflection, diffraction and attenuation behavior of electromagnetic waves in the actual roadway, input the dielectric constant of the coal and rock medium of the roadway wall, the electromagnetic properties of the metal support and the roadway geometric parameters into the three-dimensional spatial model of the roadway. The roadway geometric parameters include cross-sectional shape, dip angle change and branch structure.

[0015] Step 13: In the NS-3 simulation environment that contains real electromagnetic propagation characteristics, input the pulse noise source parameters with amplitude, duration, time domain and frequency domain characteristics, and the sweep frequency interference source parameters with scanning range, rate and waveform to generate a dynamic interference source.

[0016] Further, in step 2, within the constructed NS-3 simulation environment, distributed wireless sensor network nodes are deployed according to the preset number of nodes, initial locations, and communication radius. Each node deploys a real-time spectrum characteristic sensing module to continuously scan the communication frequency band and output the wireless sensor network operating status, including node deployment information and real-time spectrum interference detection data.

[0017] Node deployment information includes the preset number of nodes, initial node locations, and communication radius;

[0018] Real-time spectrum interference detection data includes wireless sensor network operation status information such as interference event identifiers, interference information characteristic parameters, and interference warning signs.

[0019] Further, in step 3, using the real-time spectrum interference detection data collected by the spectrum sensing modules of each node, combined with the known node locations and the current wireless sensor network topology, an improved TDOA algorithm is used to locate the interference source. A sliding window is introduced to process transient signals, a generalized cross-correlation algorithm is used to improve the time delay estimation accuracy, and the three-dimensional coordinates of the interference source are calculated based on the weighted least squares method. Finally, the obtained interference source location results, wireless sensor network topology status, and node load information are integrated to form the state space for multi-agent reinforcement learning. The specific method is as follows:

[0020] Step 31: Based on the real-time spectrum interference detection data collected by the node spectrum sensing module and the node location and the current wireless sensor network topology, the improved TDOA algorithm is used to locate the interference source.

[0021] (1) Sliding window for handling transient signals

[0022] The node performs analog-to-digital conversion on the received spectral interference detection data and buffers it in a fixed-length FIFO window in real time. Then, it calculates the total energy of the signal within the current window, E = ∑|x[k]|. 2 It is compared with a preset threshold in real time; if the energy is lower than the threshold, the window slides one step to continue monitoring; if it exceeds the threshold, the window data is immediately triggered and locked to form a transient signal data slice containing a complete pulse waveform. Finally, the slice is sent to the generalized cross-correlation module for high-precision time delay estimation.

[0023] (2) Improve the accuracy of time delay estimation by using the generalized cross-correlation algorithm.

[0024] Input transient signal data slices, and then apply the Fourier transform formula. Calculate the power spectral density of the two signals, using the calculated signal power spectral density function, and then apply the cross-power spectrum formula P. xy (f)=X(f)·Y * (f) Calculate the cross-power spectrum between the two signals, and adjust the weighting function according to the PHAT phase change. Multiply by the cross-power spectrum, and then apply the multiplied function to the inverse Fourier transform formula. The generalized cross-correlation function R is obtained. xy (T), which outputs the peak position of the generalized cross-correlation function, i.e., the high-precision time delay estimate τ between the two signals;

[0025] (3) Calculate the three-dimensional coordinates of the interference source based on the weighted least squares method

[0026] The i-th high-precision time delay estimate τ between multiple nodes obtained by the generalized cross-correlation algorithm i As input, a hyperbolic equation system is first established based on the time delay difference; due to the different measurement accuracies at each node in the mining environment, a weighted residual objective function is constructed by assigning weights to each measurement value. Where w i Let θ be the weight coefficient between nodes that is positively correlated with the signal-to-noise ratio, and let θ be the coordinates of the interference source to be determined. i Given the known node positions, the gradient descent iterative algorithm is finally used to minimize the objective function, thus solving for the three-dimensional spatial coordinates θ of the interference source that minimizes the weighted error. *=argminJ(θ) outputs the precise location of the interference source, and finally obtains θ, which is the precise coordinate location of the interference source;

[0027] Step 32: The precise location of the interference source, the obtained node load, and the wireless sensor network topology obtained by the TDOA algorithm are integrated and input into the state space of reinforcement learning, i.e., s_t = [precise location of interference source, node load, wireless sensor network topology].

[0028] Furthermore, in step 4, based on the aforementioned state space and the action space of the available next-hop neighbor node set, a design method using the bi-objective reward function R = α·path quality + β·interference avoidance parameter is adopted to construct the action of the multi-agent reinforcement learning framework. The control signaling mechanism used for policy update information during the learning process is integrated into the packet routing simulator, allowing agents to interact with neural network weights or target values ​​through model sharing or value sharing, thereby achieving collaborative optimization of the routing policy. The specific method is as follows:

[0029] Step 41: Based on each agent's periodic maintenance of its neighbor table through beacon exchange, and combined with the interference warning flags output by the real-time spectrum sensing module, dynamically exclude those nodes within the interference avoidance radius or with extremely poor channel quality from all its neighbor nodes within its communication range, forming the set of available next-hop neighbor nodes that can be used for routing decisions at the current moment, i.e., the action space.

[0030] Step 42: Based on the state space and action space, design a bi-objective reward function R = α·path quality + β·interference avoidance parameter, where path quality includes link delay, inverse throughput or hop count, interference avoidance parameter includes distance between node and interference source or interference intensity, and α and β are weighting coefficients.

[0031] Step 43: Build an NS-3 packet routing simulator, add a dedicated control signaling application module to each node, generate and broadcast specific control signaling packets according to a preset signaling cycle; the signaling format is defined and encapsulated using Google Protobuf;

[0032] Step 44: Agents collaborate using either model sharing or value sharing based on the configuration mode. If configured in model sharing mode, after each agent updates the target neural network weights locally, its control signaling application serializes the latest neural network weight parameters, encapsulates them into control signaling packets, and broadcasts them to all neighboring nodes in the next signaling cycle. If configured in value sharing mode, each agent, after calculating the target value used to update the Q-function... Then, the target value Q π (s n ,a n ) and the corresponding state-action pairs (sn ,a n The signaling packet is encapsulated and sent to the relevant neighboring nodes.

[0033] Step 45, each agent maintains an experience replay buffer: during the training phase, the agent not only stores its own state transition experience (s n ,a n ,r n ,s n' It also stores shared information received from neighbors, including model weights or target value experience, into a buffer; when updating the neural network, it samples small batches of data from the hybrid buffer, calculates gradients, and performs backpropagation, thereby achieving collaborative strategy optimization that integrates local experience and neighbor knowledge.

[0034] Further, in step 5, the multi-agent reinforcement learning framework is connected to the NS-3 simulation environment via the NS3-gym interface, and training hyperparameters are set; the agents are trained using the Asynchronous Dominance (A3C) algorithm, the dominance value is calculated based on the reward function, the Actor-Critic network loss function is constructed, and the neural network parameters are iteratively updated through gradient descent and backpropagation until the model converges, thereby obtaining the trained reinforcement learning model. The specific method is as follows:

[0035] Step 51: Instantiate OpenGymInterface in the NS-3 simulation script and bind it to a custom wireless sensor network environment class; in this environment class, override GetObservation(), GetReward(), and ExecuteAction() to implement wireless sensor network state information, including latency, packet loss rate, interference intensity, agent observation vector, reward value calculation, and actions; and in the reinforcement learning training script on the Python side, explicitly define and set a set of training hyperparameters, including learning rate, discount factor, initial exploration rate, entropy regularization coefficient specific to the A3C algorithm, and global optimizer.

[0036] Step 52: Use the PyTorch framework to build a neural network with two output heads, Actor and Critic; in the main training loop, start multiple parallel environment processes worker, each worker is responsible for collecting empirical data in an independent NS-3 simulation environment;

[0037] Step 53: In each worker, based on the state-value function V(s) output by the Actor-Critic network and the actual reward R, calculate the advantage estimate A(s,a) = RV(s); subsequently, construct the total loss function L = L3 according to the A3C algorithm. policy +L value ×0.5-L entropy×β, where the policy loss L policy = -logπ(a|s)×A(s,a), value loss L value =(RV(s)) 2 Entropy loss L entropy =∑π(a|s)×logπ(a|s);

[0038] Step 54: After each worker completes a certain number of steps of experience collection, calculate the gradient of the loss function with respect to the neural network parameters; submit the gradients of all workers asynchronously to the global neural network server, perform weighted averaging and execute one step of gradient descent update; after the update is completed, synchronize the latest global neural network parameters to all workers, clear their local gradients, and start a new round of data collection and learning until the average cumulative reward converges.

[0039] Further, in step 6, the trained model is loaded into each actual wireless sensor network node. Based on the real-time collected wireless sensor network status and interference source information, an anti-interference strategy is dynamically executed. Q-learning is applied to achieve adaptive channel hopping to avoid frequency domain interference. At the same time, a spatial avoidance radius is established based on the location of the interference source, triggering local rerouting and dynamically updating the node routing table and channel allocation strategy. The specific method is as follows:

[0040] Step 61: Load the final reinforcement learning model parameters obtained after training into the actual wireless sensor network node, initialize the neural network inference engine and establish communication links with the local spectrum sensing module and the forwarding plane.

[0041] Step 62: The node enters normal operation state, and its agent program starts a periodic decision loop. In each cycle, the agent generates the latest channel interference intensity vector based on real-time spectrum sensing data, obtains the current neighbor list and link quality index from the routing protocol stack, and combines them with the latest known interference source location information to form the current observation state vector s_t; inputs this vector into the loaded neural network, performs forward inference, and obtains the output action a_t;

[0042] Step 63: Maintain a channel quality Q table based on Q-learning for each node, with its state being the channel ID and its action being "switching" or "holding". Dynamically update the Q value based on the real-time perceived bit error rate and throughput of each channel. Once the interference power of the current channel exceeds the threshold and the Q table indicates the existence of a better channel, immediately trigger the channel switching action and synchronously switch to the target channel after coordinating with neighboring nodes through signaling.

[0043] Step 64: The node continuously monitors its Euclidean distance from known interference sources. If the distance is less than the preset avoidance radius, it immediately marks itself as an "interference zone node" and sends a route update announcement to its neighboring nodes. After receiving the announcement, the neighboring nodes recalculate the routing table based on the updated neural network topology and dynamically reroute all service flows that were originally planned to be forwarded through the "interference zone node" to other optimal paths.

[0044] In step 65, the agent sends the decision action to the node's data packet forwarding engine; the forwarding engine immediately updates its kernel-mode forwarding table and the channel configuration of the wireless network card to ensure that subsequent data packets can be forwarded according to the latest anti-interference strategy, thereby achieving continuous optimization of the communication link.

[0045] A distributed anti-interference cooperative optimization method for a mine wireless network based on multi-agent reinforcement learning includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the distributed anti-interference cooperative optimization method for a mine wireless network based on multi-agent reinforcement learning, thereby realizing distributed anti-interference cooperative optimization for a mine wireless network based on multi-agent reinforcement learning.

[0046] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the aforementioned distributed anti-interference cooperative optimization method for mine wireless networks based on multi-agent reinforcement learning, thereby realizing distributed anti-interference cooperative optimization for mine wireless networks based on multi-agent reinforcement learning.

[0047] A computer-readable storage medium storing a computer program thereon, wherein when the computer program is executed by a processor, it implements the aforementioned distributed anti-interference cooperative optimization method for mine wireless networks based on multi-agent reinforcement learning, thereby realizing distributed anti-interference cooperative optimization for mine wireless networks based on multi-agent reinforcement learning.

[0048] Compared with existing technologies, the significant advantages of this invention are as follows: 1) The high-fidelity mine electromagnetic environment modeling method based on NS-3 can accurately simulate transient pulse interference and frequency sweep interference generated by equipment such as frequency converters and fully mechanized mining units, solving the problem that traditional simulation platforms (such as MATLAB) cannot realistically reproduce the complex electromagnetic environment of mines, and improving the reliability of anti-interference algorithms in actual deployment. 2) The improved TDOA interference source localization algorithm of this invention improves the time delay estimation accuracy of transient signals through generalized cross-correlation (GCC) and sliding window mechanisms, solving the problem of large localization errors in strong noise environments of traditional TDOA, and providing key support for dynamic route optimization. 3) The multi-agent reinforcement learning (MARL) distributed decision framework of this invention, by defining the state space s_t = [interference source accurate localization, node load, wireless sensor network topology] and the bi-objective reward function R = α·path quality + β·interference avoidance parameter, solves the problems of slow response of centralized control and poor adaptability of static strategies, reducing packet loss rate while significantly reducing the number of path optimization convergence steps. 5) The interference-aware shortest path algorithm of this invention incorporates interference avoidance constraints, solving the problem of frequent packet loss in high-interference areas by traditional shortest path algorithms, and achieving coordinated optimization of communication quality and transmission efficiency. 6) The distributed asynchronous collaborative mechanism of this invention, through independent thread operation of nodes and exchange of q-tables with adjacent nodes, solves the control signaling delay problem caused by multipath effects in mine roadways, ensuring system robustness. Attached Figure Description

[0049] Figure 1 Flowchart of a distributed anti-interference optimization method based on multi-agent reinforcement learning

[0050] Figure 2 A schematic diagram of the architecture and communication process of a multi-agent reinforcement learning system.

[0051] Figure 3 To reinforce the node architecture graph in learning

[0052] Figure 4 Communication module protocol diagram

[0053] Figure 5 Changes in mean squared error loss during training

[0054] Figure 6 The image shows a comparison between the results with and without the anti-interference algorithm. Detailed Implementation

[0055] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific examples described herein are merely illustrative and not intended to limit the scope of this application.

[0056] This invention discloses a distributed anti-interference cooperative optimization method for mine wireless networks based on multi-agent reinforcement learning. The overall method flow is as follows: Figure 1 As shown, it specifically includes:

[0057] Step 1: Input the 3D model file of the mine roadway, the list of interference source parameters, and the electromagnetic property parameters of the roadway wall into the Network Simulator 3 platform. Configure the ray tracing propagation model to simulate the multipath reflection, diffraction, and attenuation behavior of electromagnetic waves in the roadway. Instantiate the impulse noise source and the frequency sweep interference source according to the parameters and place them in the 3D spatial model of the roadway. Finally, an NS-3 simulation environment containing real electromagnetic propagation characteristics and dynamic interference sources is constructed.

[0058] Step 11: Based on the collected real mine tunnel scene, generate a 3D model file of the mine tunnel, import it into the Network Simulator 3 platform, and generate a 3D spatial model of the tunnel.

[0059] Step 12: Based on the multipath reflection, diffraction and attenuation behavior of electromagnetic waves in the actual roadway, input the dielectric constant of the coal and rock medium of the roadway wall, the electromagnetic properties of the metal support and the roadway geometric parameters into the three-dimensional spatial model of the roadway. The roadway geometric parameters include cross-sectional shape, dip angle change and branch structure.

[0060] Step 13: In the NS-3 simulation environment that contains real electromagnetic propagation characteristics, input the pulse noise source parameters with amplitude, duration, time domain and frequency domain characteristics, and the sweep frequency interference source parameters with scanning range, rate and waveform to generate a dynamic interference source.

[0061] Step 2: In the constructed NS-3 simulation environment, deploy distributed wireless sensor network nodes according to the preset number of nodes, initial locations, and communication radius. Each node deploys a real-time spectrum characteristic sensing module to continuously scan the communication frequency band and output the operating status of the distributed wireless sensor network, including node deployment information and real-time spectrum interference detection data.

[0062] Node deployment information includes the preset number of nodes, initial node locations, and communication radius;

[0063] Real-time spectrum interference detection data includes wireless sensor network operation status information such as interference event identifiers, interference information characteristic parameters, and interference warning signs.

[0064] Step 3: Utilize the spectrum interference detection data collected in real time by the spectrum sensing modules of each node, combined with the known node locations and the current wireless sensor network topology, to perform interference source localization using an improved TDOA algorithm. Introduce a sliding window to process transient signals, use a generalized cross-correlation algorithm to improve the time delay estimation accuracy, and calculate the three-dimensional coordinates of the interference source based on the weighted least squares method. Finally, integrate the obtained interference source localization results, wireless sensor network topology status, and node load information to form the state space of multi-agent reinforcement learning.

[0065] Step 31: Based on the real-time spectrum interference detection data collected by the node spectrum sensing module and the node location and the current wireless sensor network topology, the improved TDOA algorithm is used to locate the interference source.

[0066] (1) Sliding window for handling transient signals

[0067] The node performs analog-to-digital conversion on the received spectral interference detection data and buffers it in a fixed-length FIFO window in real time. Then, it calculates the total energy of the signal within the current window, E = ∑|x[k]|. 2 It is compared with a preset threshold in real time; if the energy is lower than the threshold, the window slides one step to continue monitoring; if it exceeds the threshold, the window data is immediately triggered and locked to form a transient signal data slice containing a complete pulse waveform. Finally, the slice is sent to the generalized cross-correlation module for high-precision time delay estimation.

[0068] (2) Improve the accuracy of time delay estimation by using the generalized cross-correlation algorithm.

[0069] Input transient signal data slices, and then apply the Fourier transform formula. Calculate the power spectral density of the two signals, using the calculated signal power spectral density function, and then apply the cross-power spectrum formula P. xy (f)=X(f)·Y * (f) Calculate the cross-power spectrum between the two signals, and adjust the weighting function according to the PHAT phase change. Multiply by the cross-power spectrum, and then apply the multiplied function to the inverse Fourier transform formula. The generalized cross-correlation function R is obtained. xy (T), which outputs the peak position of the generalized cross-correlation function, i.e., the high-precision time delay estimate τ between the two signals;

[0070] (3) Calculate the three-dimensional coordinates of the interference source based on the weighted least squares method

[0071] The i-th high-precision time delay estimate τ between multiple nodes obtained by the generalized cross-correlation algorithm iAs input, a hyperbolic equation system is first established based on the time delay difference; due to the different measurement accuracies at each node in the mining environment, a weighted residual objective function is constructed by assigning weights to each measurement value. Where w i Let θ be the weight coefficient between nodes that is positively correlated with the signal-to-noise ratio, and let θ be the coordinates of the interference source to be determined. i Given the known node positions, the gradient descent iterative algorithm is finally used to minimize the objective function, thus solving for the three-dimensional spatial coordinates θ of the interference source that minimizes the weighted error. * =argminJ(θ) outputs the precise location of the interference source, and finally obtains θ, which is the precise coordinate location of the interference source;

[0072] Step 32: The precise location of the interference source, the obtained node load, and the wireless sensor network topology obtained by the TDOA algorithm are integrated and input into the state space of reinforcement learning, i.e., s_t = [precise location of interference source, node load, wireless sensor network topology].

[0073] Step 4: Based on the state space and action space of the available next-hop neighbor node set, the action of the multi-agent reinforcement learning framework is constructed using the design method of bi-objective reward function R = α·path quality + β·interference avoidance parameter. Furthermore, the control signaling mechanism for policy update information during the learning process is integrated into the packet routing simulator, allowing agents to interact with neural network weights or target values ​​through model sharing or value sharing, thereby achieving collaborative optimization of routing policies.

[0074] Step 41: Based on each agent's periodic maintenance of its neighbor table through beacon exchange, and combined with the interference warning flags output by the real-time spectrum sensing module, dynamically exclude those nodes within the interference avoidance radius or with extremely poor channel quality from all its neighbor nodes within its communication range, forming the set of available next-hop neighbor nodes that can be used for routing decisions at the current moment, i.e., the action space.

[0075] Step 42: Based on the state space and action space, design a bi-objective reward function R = α·path quality + β·interference avoidance parameter, where path quality includes link delay, inverse throughput or hop count, interference avoidance parameter includes distance between node and interference source or interference intensity, and α and β are weighting coefficients.

[0076] Step 43: Build an NS-3 packet routing simulator, add a dedicated control signaling application module to each node, generate and broadcast specific control signaling packets according to a preset signaling cycle; the signaling format is defined and encapsulated using Google Protobuf;

[0077] Step 44: Agents collaborate using either model sharing or value sharing based on the configuration mode. If configured in model sharing mode, after each agent updates the target neural network weights locally, its control signaling application serializes the latest neural network weight parameters, encapsulates them into control signaling packets, and broadcasts them to all neighboring nodes in the next signaling cycle. If configured in value sharing mode, each agent, after calculating the target value used to update the Q-function... Then, the target value Q π (s n ,a n ) and the corresponding state-action pairs (s n ,a n The signaling packet is encapsulated and sent to the relevant neighboring nodes.

[0078] Step 45, each agent maintains an experience replay buffer: during the training phase, the agent not only stores its own state transition experience (s n ,a n ,r n ,s n' It also stores shared information received from neighbors, including model weights or target value experience, into a buffer; when updating the neural network, it samples small batches of data from the hybrid buffer, calculates gradients, and performs backpropagation, thereby achieving collaborative strategy optimization that integrates local experience and neighbor knowledge.

[0079] Step 5: Connect the multi-agent reinforcement learning framework to the NS-3 simulation environment using the NS3-gym interface, and set the learning rate and discount factor; train the agents using the Asynchronous Advantage (A3C) algorithm: calculate the advantage value based on the reward function, construct the Actor-Critic network loss function, and iteratively update the neural network parameters through gradient descent and backpropagation until the model converges, thereby obtaining the trained reinforcement learning model;

[0080] Step 51: Instantiate OpenGymInterface in the NS-3 simulation script and bind it to a custom wireless sensor network environment class; in this environment class, override GetObservation(), GetReward(), and ExecuteAction() to implement wireless sensor network state information, including latency, packet loss rate, interference intensity, agent observation vector, reward value calculation, and actions; and in the reinforcement learning training script on the Python side, explicitly define and set a set of training hyperparameters, including learning rate, discount factor, initial exploration rate, entropy regularization coefficient specific to the A3C algorithm, and global optimizer.

[0081] Step 52: Use the PyTorch framework to build a neural network with two output heads, Actor and Critic; in the main training loop, start multiple parallel environment processes worker, each worker is responsible for collecting empirical data in an independent NS-3 simulation environment;

[0082] Step 53: In each worker, based on the state-value function V(s) output by the Actor-Critic network and the actual reward R, calculate the advantage estimate A(s,a) = RV(s); subsequently, construct the total loss function L = L3 according to the A3C algorithm. policy +L value ×0.5-L entropy ×β, where the policy loss L policy = -logπ(a|s)×A(s,a), value loss L value =(RV(s)) 2 Entropy loss L entropy =∑π(a|s)×logπ(a|s);

[0083] Step 54: After each worker completes a certain number of steps of experience collection, calculate the gradient of the loss function with respect to the neural network parameters; submit the gradients of all workers asynchronously to the global neural network server, perform weighted averaging and execute one step of gradient descent update; after the update is completed, synchronize the latest global neural network parameters to all workers, clear their local gradients, and start a new round of data collection and learning until the average cumulative reward converges.

[0084] Step 6: Load the trained reinforcement learning model into each actual wireless sensor network node. Based on the real-time collected wireless sensor network status and interference source information, dynamically execute anti-interference strategies: apply Q-learning to achieve adaptive channel hopping to avoid frequency domain interference, establish a spatial avoidance radius based on the location of the interference source, trigger local rerouting, and dynamically update the node routing table and channel allocation strategy.

[0085] Step 61: Load the final reinforcement learning model parameters obtained after training into the actual wireless sensor network node, initialize the neural network inference engine and establish communication links with the local spectrum sensing module and the forwarding plane.

[0086] Step 62: The node enters normal operation state, and its agent program starts a periodic decision loop. In each cycle, the agent generates the latest channel interference intensity vector based on real-time spectrum sensing data, obtains the current neighbor list and link quality index from the routing protocol stack, and combines them with the latest known interference source location information to form the current observation state vector s_t; inputs this vector into the loaded neural network, performs forward inference, and obtains the output action a_t;

[0087] Step 63: Maintain a channel quality Q table based on Q-learning for each node, with its state being the channel ID and its action being "switching" or "holding". Dynamically update the Q value based on the real-time perceived bit error rate and throughput of each channel. Once the interference power of the current channel exceeds the threshold and the Q table indicates the existence of a better channel, immediately trigger the channel switching action and synchronously switch to the target channel after coordinating with neighboring nodes through signaling.

[0088] Step 64: The node continuously monitors its Euclidean distance from known interference sources. If the distance is less than the preset avoidance radius, it immediately marks itself as an "interference zone node" and sends a route update announcement to its neighboring nodes. After receiving the announcement, the neighboring nodes recalculate the routing table based on the updated neural network topology and dynamically reroute all service flows that were originally planned to be forwarded through the "interference zone node" to other optimal paths.

[0089] In step 65, the agent sends the decision action to the node's data packet forwarding engine; the forwarding engine immediately updates its kernel-mode forwarding table and the channel configuration of the wireless network card to ensure that subsequent data packets can be forwarded according to the latest anti-interference strategy, thereby achieving continuous optimization of the communication link.

[0090] This invention also proposes a distributed anti-interference cooperative optimization method for mine wireless networks based on multi-agent reinforcement learning, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the distributed anti-interference cooperative optimization method for mine wireless networks based on multi-agent reinforcement learning, thereby realizing distributed anti-interference cooperative optimization for mine wireless networks based on multi-agent reinforcement learning.

[0091] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the aforementioned distributed anti-interference cooperative optimization method for mine wireless networks based on multi-agent reinforcement learning, thereby realizing distributed anti-interference cooperative optimization for mine wireless networks based on multi-agent reinforcement learning.

[0092] A computer-readable storage medium storing a computer program thereon, wherein when the computer program is executed by a processor, it implements the aforementioned distributed anti-interference cooperative optimization method for mine wireless networks based on multi-agent reinforcement learning, thereby realizing distributed anti-interference cooperative optimization for mine wireless networks based on multi-agent reinforcement learning.

[0093] Example

[0094] To verify the effectiveness of the present invention, experiments were conducted according to the following specific steps:

[0095] 1) Building a high-fidelity simulation training environment

[0096] Based on the NS-3 network simulation platform, a communication bridge with the Python reinforcement learning framework is built using the NS3-gym interface to achieve real-time data interaction between the agent and the network simulation environment. In this environment, NS3-gym is responsible for converting the agent's action commands (such as selecting the next-hop node and switching channels) into simulation operations in NS-3, and synchronously returning network state observations (including latency, packet loss rate, interference intensity, network topology, etc.) to the agent. This embodiment constructs an Abilene network topology with 11 nodes to simulate the complex multipath and attenuation characteristics of a distributed wireless sensor network in a mine tunnel environment. All link propagation delays are uniformly set to 1ms, and the transmission rate is configured to 500Kbps. The data packet payload is 512B and encapsulated in the UDP over IP protocol for transmission to simulate typical sensor data services.

[0097] After completing the simulation environment setup, immediately configure the node function modules.

[0098] 2) Configure node function modules

[0099] Reference Figure 3 In each distributed wireless sensor network node, the following core functional modules are deployed and configured to specifically implement the functions of steps 2 to 6 in the method described in this invention:

[0100] 2.1 Spectrum Detection Module: In the NS-3 node application, this is implemented by instantiating a real-time spectrum feature sensing module, configuring its scanning frequency band to 2.4 / 5.8GHz, energy detection threshold to 800V / m, and duration threshold to 10μs. After parameter configuration, the module starts continuous spectrum scanning. When interference is detected, it generates an interference warning flag containing a timestamp, interference type, and intensity. The detection result is sent to the improved TDOA interference source localization module via a message queue, triggering the subsequent localization process.

[0101] 2.2 Improved TDOA Interference Source Localization Module: Upon receiving an interference warning flag from the spectrum detection module, this module immediately initiates the localization process. First, a sliding window mechanism is used to buffer transient signal data. Then, a generalized cross-correlation algorithm is called to estimate the time delay. Finally, the coordinates of the interference source are calculated using the weighted least squares method. After localization is complete, the obtained three-dimensional coordinate information is broadcast via UDP to the routing decision module of all nodes for updating the state space.

[0102] 2.3 Packet Generation Module: Deployed during system initialization, this module configures Poisson process parameters according to the experimental load factor (50%-200%) to continuously generate UDP packets with a 512B payload. The generated packets are sent to the network layer via the NS-3 socket interface, providing the routing decision module with a realistic network traffic environment for evaluating path quality and calculating the reward function.

[0103] 2.4 Control Signaling Application Module: After the routing decision module makes its decision, this module is responsible for handling the exchange of policy update information. It triggers periodically at 100ms intervals, using the Google Protobuf protocol, such as... Figure 4 It encapsulates neural network weights or target values ​​and broadcasts them to neighboring nodes via NS-3 UDP sockets. Simultaneously, it receives control signaling from other nodes, deserializes it, and updates the local experience replay buffer.

[0104] 2.5 Routing Decision Module: This module receives interference source coordinates from the improved TDOA module and real-time interference information from the spectrum detection module. Combined with network traffic status generated by the packet generation module, it constructs a complete state space. By loading a pre-trained reinforcement learning model for neural network inference, it outputs channel switching or rerouting decisions. Finally, it calls the NS-3 kernel interface to update the routing table and physical layer channel configuration, completing the execution of the anti-interference strategy.

[0105] After completing the deployment and configuration of all functional modules, the next step is to set up the mechanisms and hyperparameters required for training.

[0106] 3) Setting up the training mechanism and hyperparameters

[0107] The A3C asynchronous advantage algorithm was used for training. Training hyperparameters were set as follows: the Adam optimizer was used, the learning rate was set to 0.001, and the discount factor γ was 0.99. An ε-greedy exploration strategy was adopted, maintaining an exploration rate ε of 0.5 in interference regions and decreasing from 1.0 to 0.01 in non-interference regions. The agents coordinated through the control signaling application module using either model sharing or value sharing.

[0108] Once the hyperparameters are set, the training process will begin immediately.

[0109] 4) Execute the training process

[0110] Multiple parallel environment processes (workers) are launched to collect experience data. Each worker calculates an advantage estimate A(s,a) based on the reward function and constructs a total loss function that includes policy loss, value loss, and entropy loss. Gradients from all workers are asynchronously submitted to the global network for weighted averaging and gradient descent updates. Global model parameters are synchronized to all workers, and the process iterates until the average cumulative reward converges. A total of approximately 120 million to 150 million training steps are completed.

[0111] After the model training is completed, the performance testing and evaluation phase begins.

[0112] 5) Conduct performance testing and evaluation

[0113] During the training phase, see Figure 5 All nodes completed approximately 120 million to 150 million training steps. The mean squared error (MSE) loss decreased with increasing training steps, indicating continuous model optimization. Electromagnetic interference of different intensities (0.5 kV / m, 1.0 kV / m, 1.5 kV / m) and types (pulse, sweep, steady-state) was applied and evaluated under different network loads (50%, 100%, 150%, 200%). Packet loss rate, bit error rate (BER), interference source localization error, and interference avoidance decisions were recorded and compared with traditional shortest path routing algorithms to verify the anti-interference performance and robustness of the proposed solution.

[0114] After completing the performance tests, the experimental results were analyzed in detail.

[0115] 6) Analyze the experimental results

[0116] like Figure 6 As shown, the comparison results indicate that under high service load and strong interference conditions, the packet loss rate of this scheme (anti-interference optimized routing) is significantly lower than that of the traditional shortest path routing scheme, proving its effectiveness and superiority in complex electromagnetic environments.

[0117] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. A distributed anti-interference cooperative optimization method based on multi-agent reinforcement learning in a mine wireless network, characterized in that, Includes the following steps: Step 1: Input the 3D model file of the mine roadway, the list of interference source parameters, and the electromagnetic property parameters of the roadway wall into the Network Simulator 3 platform. Configure the ray tracing propagation model to simulate the multipath reflection, diffraction, and attenuation behavior of electromagnetic waves in the roadway. Instantiate the impulse noise source and the frequency sweep interference source according to the parameters and place them in the 3D spatial model of the roadway. Finally, an NS-3 simulation environment containing real electromagnetic propagation characteristics and dynamic interference sources is constructed. Step 2: In the constructed NS-3 simulation environment, deploy distributed wireless sensor network nodes according to the preset number of nodes, initial locations, and communication radius; each node deploys a real-time spectrum feature sensing module to continuously scan the communication frequency band and output the operating status of the distributed wireless sensor network, which includes node deployment information and real-time spectrum interference detection data. Step 3: Utilize the spectrum interference detection data collected in real time by the spectrum sensing modules of each node, combined with the known node locations and the current wireless sensor network topology, to perform interference source localization using an improved TDOA algorithm. Introduce a sliding window to process transient signals, use a generalized cross-correlation algorithm to improve the time delay estimation accuracy, and calculate the three-dimensional coordinates of the interference source based on the weighted least squares method. Finally, integrate the obtained interference source localization results, wireless sensor network topology status, and node load information to form the state space of multi-agent reinforcement learning. Step 4: Based on the state space and action space of the available next-hop neighbor node set, the action of the multi-agent reinforcement learning framework is constructed using the design method of bi-objective reward function R = α·path quality + β·interference avoidance parameter. Furthermore, the control signaling mechanism for policy update information during the learning process is integrated into the packet routing simulator, allowing agents to interact with neural network weights or target values ​​through model sharing or value sharing, thereby achieving collaborative optimization of routing policies. Step 5: Connect the multi-agent reinforcement learning framework to the NS-3 simulation environment using the NS3-gym interface, and set the learning rate and discount factor; train the agents using the Asynchronous Advantage (A3C) algorithm: calculate the advantage value based on the reward function, construct the Actor-Critic network loss function, and iteratively update the neural network parameters through gradient descent and backpropagation until the model converges, thereby obtaining the trained reinforcement learning model; Step 6: Load the trained reinforcement learning model into each actual wireless sensor network node. Based on the real-time collected wireless sensor network status and interference source information, dynamically execute anti-interference strategies: apply Q-learning to achieve adaptive channel hopping to avoid frequency domain interference, establish a spatial avoidance radius based on the location of the interference source, trigger local rerouting, and dynamically update the node routing table and channel allocation strategy.

2. The distributed anti-interference cooperative optimization method for mine wireless networks based on multi-agent reinforcement learning as described in claim 1, characterized in that, Step 1: Input the 3D model file of the mine roadway, the list of interference source parameters, and the electromagnetic property parameters of the roadway wall into the Network Simulator 3 platform. Configure the ray tracing propagation model to simulate the multipath reflection, diffraction, and attenuation behavior of electromagnetic waves in the roadway. Instantiate the impulse noise source and the frequency sweep interference source according to the parameters and place them in the 3D spatial model of the roadway. Finally, an NS-3 simulation environment containing real electromagnetic propagation characteristics and dynamic interference sources is constructed. The specific method is as follows: Step 11: Based on the collected real mine roadway scenes, generate a 3D model file of the mine roadway and import it into the NetworkSimulator 3 platform to generate a 3D spatial model of the roadway. Step 12: Based on the multipath reflection, diffraction and attenuation behavior of electromagnetic waves in the actual roadway, input the dielectric constant of the coal and rock medium of the roadway wall, the electromagnetic properties of the metal support and the roadway geometric parameters into the three-dimensional spatial model of the roadway. The roadway geometric parameters include cross-sectional shape, dip angle change and branch structure. Step 13: In the NS-3 simulation environment that contains real electromagnetic propagation characteristics, input the pulse noise source parameters with amplitude, duration, time domain and frequency domain characteristics, and the sweep frequency interference source parameters with scanning range, rate and waveform to generate a dynamic interference source.

3. The distributed anti-interference cooperative optimization method for mine wireless networks based on multi-agent reinforcement learning as described in claim 1, characterized in that, Step 2: In the constructed NS-3 simulation environment, deploy distributed wireless sensor network nodes according to the preset number of nodes, initial locations, and communication radius. Each node deploys a real-time spectrum characteristic sensing module to continuously scan the communication frequency band and output the wireless sensor network operating status, including node deployment information and real-time spectrum interference detection data. Node deployment information includes the preset number of nodes, initial node locations, and communication radius; Real-time spectrum interference detection data includes wireless sensor network operation status information such as interference event identifiers, interference information characteristic parameters, and interference warning signs.

4. The distributed anti-interference cooperative optimization method for mine wireless networks based on multi-agent reinforcement learning according to claim 1, characterized in that, Step 3: Utilizing the real-time spectrum interference detection data collected by the spectrum sensing modules of each node, combined with the known node locations and the current wireless sensor network topology, an improved TDOA algorithm is used to locate the interference source. A sliding window is introduced to process transient signals, a generalized cross-correlation algorithm is used to improve the time delay estimation accuracy, and the three-dimensional coordinates of the interference source are calculated based on the weighted least squares method. Finally, the obtained interference source location results, wireless sensor network topology status, and node load information are integrated to form the state space for multi-agent reinforcement learning. The specific method is as follows: Step 31: Based on the real-time spectrum interference detection data collected by the node spectrum sensing module and the node location and the current wireless sensor network topology, the improved TDOA algorithm is used to locate the interference source. (1) Sliding window for handling transient signals The node performs analog-to-digital conversion on the received spectral interference detection data and buffers it in a fixed-length FIFO window in real time. Then, it calculates the total energy of the signal within the current window, E = ∑|x[k]|. 2 It is compared with a preset threshold in real time; if the energy is lower than the threshold, the window slides one step to continue monitoring; if it exceeds the threshold, the window data is immediately triggered and locked to form a transient signal data slice containing a complete pulse waveform. Finally, the slice is sent to the generalized cross-correlation module for high-precision time delay estimation. (2) Improve the accuracy of time delay estimation by using the generalized cross-correlation algorithm. Input transient signal data slices, and then apply the Fourier transform formula. Calculate the power spectral density of the two signals, using the calculated signal power spectral density function, and then apply the cross-power spectrum formula P. xy (f)=X(f)·Y * (f) Calculate the cross-power spectrum between the two signals, and adjust the weighting function φ according to the PHAT phase change. p (f)= Multiply by the cross-power spectrum, and then apply the multiplied function to the inverse Fourier transform formula. The generalized cross-correlation function R is obtained. xy (T), which outputs the peak position of the generalized cross-correlation function, i.e., the high-precision time delay estimate τ between the two signals; (3) Calculate the three-dimensional coordinates of the interference source based on the weighted least squares method The i-th high-precision time delay estimate τ between multiple nodes obtained by the generalized cross-correlation algorithm i As input, a hyperbolic equation system is first established based on the time delay difference; due to the different measurement accuracies at each node in the mining environment, a weighted residual objective function is constructed by assigning weights to each measurement value. Where w i Let θ be the weight coefficient between nodes that is positively correlated with the signal-to-noise ratio, and let θ be the coordinates of the interference source to be determined. i Given the known node positions, the gradient descent iterative algorithm is finally used to minimize the objective function, thus solving for the three-dimensional spatial coordinates θ of the interference source that minimizes the weighted error. * =argminJ(θ) outputs the precise location of the interference source, and finally obtains θ, which is the precise coordinate location of the interference source; Step 32: The precise location of the interference source, the obtained node load, and the wireless sensor network topology obtained by the TDOA algorithm are integrated and input into the state space of reinforcement learning, i.e., s_t = [precise location of interference source, node load, wireless sensor network topology].

5. The distributed anti-interference cooperative optimization method for mine wireless networks based on multi-agent reinforcement learning according to claim 1, characterized in that, Step 4: Based on the state space and action space of the available next-hop neighbor nodes, a bi-objective reward function R = α·path quality + β·interference avoidance parameter is used to construct the action of the multi-agent reinforcement learning framework. Furthermore, the control signaling mechanism for policy update information during the learning process is integrated into the packet routing simulator, allowing agents to interact with neural network weights or target values ​​through model sharing or value sharing, thereby achieving collaborative optimization of the routing policy. The specific method is as follows: Step 41: Based on each agent's periodic maintenance of its neighbor table through beacon exchange, and combined with the interference warning flags output by the real-time spectrum sensing module, dynamically exclude those nodes within the interference avoidance radius or with extremely poor channel quality from all its neighbor nodes within its communication range, forming the set of available next-hop neighbor nodes that can be used for routing decisions at the current moment, i.e., the action space. Step 42: Based on the state space and action space, design a bi-objective reward function R = α·path quality + β·interference avoidance parameter, where path quality includes link delay, inverse throughput or hop count, interference avoidance parameter includes distance between node and interference source or interference intensity, and α and β are weighting coefficients. Step 43: Build an NS-3 packet routing simulator, add a dedicated control signaling application module to each node, generate and broadcast specific control signaling packets according to a preset signaling cycle; the signaling format is defined and encapsulated using Google Protobuf; Step 44: Agents collaborate using either model sharing or value sharing based on the configuration mode. If configured in model sharing mode, after each agent updates the target neural network weights locally, its control signaling application serializes the latest neural network weight parameters, encapsulates them into control signaling packets, and broadcasts them to all neighboring nodes in the next signaling cycle. If configured in value sharing mode, each agent, after calculating the target value used to update the Q-function... Then, the target value Q π (s n ,a n ) and the corresponding state-action pairs (s n ,a n The signaling packet is encapsulated and sent to the relevant neighboring nodes. Step 45, each agent maintains an experience replay buffer: during the training phase, the agent not only stores its own state transition experience (s n ,a n ,r n ,s n' It also stores shared information received from neighbors, including model weights or target value experience, into a buffer; when updating the neural network, it samples small batches of data from the hybrid buffer, calculates gradients, and performs backpropagation, thereby achieving collaborative strategy optimization that integrates local experience and neighbor knowledge.

6. The distributed anti-interference cooperative optimization method for mine wireless networks based on multi-agent reinforcement learning according to claim 1, characterized in that, Step 5: Connect the multi-agent reinforcement learning framework to the NS-3 simulation environment via the NS3-gym interface and set the training hyperparameters; train the agents using the Asynchronous Dominance (A3C) algorithm, calculate the dominance value based on the reward function, construct the Actor-Critic network loss function, and iteratively update the neural network parameters through gradient descent and backpropagation until the model converges, thereby obtaining the trained reinforcement learning model. The specific method is as follows: Step 51: Instantiate OpenGymInterface in the NS-3 simulation script and bind it to a custom wireless sensor network environment class; in this environment class, override GetObservation(), GetReward(), and ExecuteAction() to implement wireless sensor network state information, including latency, packet loss rate, interference intensity, agent observation vector, reward value calculation, and actions; and in the reinforcement learning training script on the Python side, explicitly define and set a set of training hyperparameters, including learning rate, discount factor, initial exploration rate, entropy regularization coefficient specific to the A3C algorithm, and global optimizer. Step 52: Use the PyTorch framework to build a neural network with two output heads, Actor and Critic; in the main training loop, start multiple parallel environment processes worker, each worker is responsible for collecting empirical data in an independent NS-3 simulation environment; Step 53: In each worker, based on the state-value function V(s) output by the Actor-Critic network and the actual reward R, calculate the advantage estimate A(s,a) = RV(s); subsequently, construct the total loss function L = L3 according to the A3C algorithm. policy +L value ×0.5-L entropy ×β, where the policy loss L policy = -logπ(a|s)×A(s,a), value loss L value =(RV(s)) 2 Entropy loss L entropy =∑π(a|s)×logπ(a|s); Step 54: After each worker completes a certain number of steps of experience collection, calculate the gradient of the loss function with respect to the neural network parameters; submit the gradients of all workers asynchronously to the global neural network server, perform weighted averaging and execute one step of gradient descent update; after the update is completed, synchronize the latest global neural network parameters to all workers, clear their local gradients, and start a new round of data collection and learning until the average cumulative reward converges.

7. The distributed anti-interference cooperative optimization method for mine wireless networks based on multi-agent reinforcement learning according to claim 1, characterized in that, Step 6: Load the trained model into each actual wireless sensor network node. Based on the real-time collected wireless sensor network status and interference source information, dynamically execute anti-interference strategies. Apply Q-learning to achieve adaptive channel hopping to avoid frequency domain interference. Simultaneously, establish a spatial avoidance radius based on the interference source location, trigger local rerouting, and dynamically update the node routing table and channel allocation strategy. The specific method is as follows: Step 61: Load the final reinforcement learning model parameters obtained after training into the actual wireless sensor network node, initialize the neural network inference engine and establish communication links with the local spectrum sensing module and the forwarding plane. Step 62: The node enters normal operation state, and its agent program starts a periodic decision loop. In each cycle, the agent generates the latest channel interference intensity vector based on real-time spectrum sensing data, obtains the current neighbor list and link quality index from the routing protocol stack, and combines them with the latest known interference source location information to form the current observation state vector s_t; inputs this vector into the loaded neural network, performs forward inference, and obtains the output action a_t; Step 63: Maintain a channel quality Q table based on Q-learning for each node, with its state being the channel ID and its action being "switching" or "holding". Dynamically update the Q value based on the real-time perceived bit error rate and throughput of each channel. Once the interference power of the current channel exceeds the threshold and the Q table indicates the existence of a better channel, immediately trigger the channel switching action and synchronously switch to the target channel after coordinating with neighboring nodes through signaling. Step 64: The node continuously monitors its Euclidean distance from known interference sources; if the distance is less than the preset avoidance radius, it immediately marks itself as an "interference zone node" and sends a route update announcement to its neighboring nodes. After receiving the notification, the neighboring nodes recalculate the routing table based on the updated neural network topology and dynamically reroute all service flows that were originally planned to be forwarded through the "interference zone node" to other optimal paths; In step 65, the agent sends the decision action to the node's data packet forwarding engine; the forwarding engine immediately updates its kernel-mode forwarding table and the channel configuration of the wireless network card to ensure that subsequent data packets can be forwarded according to the latest anti-interference strategy, thereby achieving continuous optimization of the communication link.

8. A distributed anti-interference cooperative optimization method for a mine wireless network based on multi-agent reinforcement learning, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, it implements the distributed anti-interference cooperative optimization method for a mine wireless network based on multi-agent reinforcement learning as described in any one of claims 1-7, thereby realizing distributed anti-interference cooperative optimization for a mine wireless network based on multi-agent reinforcement learning.

9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, it implements the distributed anti-interference cooperative optimization method for mine wireless networks based on multi-agent reinforcement learning as described in any one of claims 1-7, thereby realizing distributed anti-interference cooperative optimization for mine wireless networks based on multi-agent reinforcement learning.

10. A computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, it implements the distributed anti-interference cooperative optimization method for mine wireless networks based on multi-agent reinforcement learning as described in any one of claims 1-7, thereby realizing distributed anti-interference cooperative optimization for mine wireless networks based on multi-agent reinforcement learning.

Citation Information

Cited By

  • GIS sensor dynamic networking anti-interference diagnosis method based on K nearest neighbor algorithm

    CN122131216A