Multi-channel routing anti-interference decision-making method for limited buffer area

Through the routing anti-interference decision-making method of hierarchical deep reinforcement learning, network congestion problems caused by limited buffer capacity and dynamic interference in wireless multi-hop networks are solved, and efficient packet transmission and anti-interference capabilities are achieved.

CN120282236AActive Publication Date: 2025-07-08ARMY ENG UNIV OF PLA
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510782921.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-07-08
Estimated Expiration
2045-06-12

AI Technical Summary

Technical Problem

In multi-source concurrency scenarios, in wireless multi-hop networks, the finite buffer capacity and dynamic interference of nodes lead to reduced network congestion and transmission reliability. It is difficult for existing routing algorithms to achieve multi-target coordination with minimum number of hops, anti-interference and congestion control.

Method used

The hierarchical deep reinforcement learning method is adopted to build a hierarchical deep Q network (HDQN) by decoupling routing planning and channel selection. The upper-layer routing decision is based on the buffer status of neighbor nodes and the destination node address. The lower-layer channel decision is combined with spectrum perception information to optimize path selection and channel use.

Benefits of technology

It improves the successful transmission rate of data packets, reduces congestion and packet loss, and improves the dynamic anti-interference ability and transmission efficiency of the network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120282236A_ABST
    Figure CN120282236A_ABST
Patent Text Reader

Abstract

The invention provides a multi-channel routing anti-interference decision-making method for a limited buffer area, and the method comprises the steps: 1, analyzing the interference of a jammer, the mutual interference between nodes, the influence of channel noise and path loss on a signal-to-interference ratio in an interference scene of a multi-hop wireless network, and calculating the transmission rate between adjacent nodes; 2, calculating the data volume of a node buffer area according to the data packet arrival rate, the interference conflict, the mutual interference conflict, the buffer area overflow packet loss and the maximum transmission frequency packet loss; 3, modeling an anti-interference routing and channel joint optimization problem of the distributed multi-hop network as a partial observable random game; the nodes realize long-term efficiency maximization based on local observation information; step 4, setting a layered DQN architecture; a routing anti-interference problem in a multi-channel scene is split into upper-layer routing planning and lower-layer channel selection. According to the method, the successful transmission rate of the data packet is improved, and the convergence speed of the algorithm is accelerated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of wireless communication, and particularly relates to a multi-channel routing anti-interference decision method for a finite buffer. Background Art

[0002] In an ad hoc network, when multi-source nodes have concurrent services, some key nodes at the intersection of multiple paths need to process multiple data streams, and their finite buffer capacity faces a severe challenge, which is likely to cause network congestion. Simply increasing the buffer capacity can temporarily relieve instantaneous congestion, but it will increase the packet queuing time and extend the end-to-end transmission delay. In addition, due to the openness of the wireless channel, the nodes in the network are easily interfered by the outside world. Interference will exacerbate the competition for spectrum resources and reduce the transmission reliability. Therefore, the nodes need to optimize path planning, dynamic spectrum allocation and congestion control strategies under multiple constraints such as buffer capacity and interference avoidance to achieve global transmission efficiency and local node coordination.

[0003] Traditional routing algorithms such as Ad hoc On-Demand Distance Vector Routing (AODV), Destination-Sequenced Distance-Vector Routing (DSDV), etc., although having advantages in simplifying the decision complexity, are difficult to cope with high-dynamic and multi-constrained complex scenarios. The AODV protocol selects paths based on the shortest hop count criterion. Its "greedy" path selection mechanism can quickly establish low-hop count routes in low-load scenarios, but in high-load scenarios, the traffic is easily overly concentrated on the topological center nodes, causing network congestion. The DSDV protocol, as a typical table-driven routing algorithm, requires all network nodes to periodically broadcast routing tables to maintain path state information. In a dynamic interference environment, the link connection relationship between nodes changes frequently, and the topological information needs to be continuously updated to maintain the routing effectiveness. The DSDV protocol will significantly increase the control message overhead in the network, and more importantly, the path selection is inaccurate due to the lag of the routing table information, making it difficult to adapt to dynamic interference.

[0004] The single-dimensional optimization mechanisms of these traditional routing protocols are difficult to achieve multi-objective coordination such as minimizing the number of hops, anti-interference, and congestion control, seriously restricting the overall network performance. Many improved protocols based on traditional routing protocols have been proposed. An enhanced Greedy Perimeter Stateless Routing (GPSR) protocol based on dynamic adjustment of node buffers calculates the next-hop node selection probability by combining geographical distance and the remaining capacity of the node buffer. This method effectively solves the problem of congestion at key nodes caused by single-dimensional routing decisions in traditional GPSR protocols, significantly reducing the network transmission delay caused by queue accumulation at key nodes in the scenario of concurrent transmission of multi-source data streams. Aiming at the problem that the traditional shortest path algorithm is effective at low load but prone to congestion at high load, a Regularized Routing Optimization (RRO) algorithm has been proposed. This method significantly improves network throughput and reduces latency by combining a congestion function and a regularization term of path length. However, such methods only calculate statically based on the current state, cannot predict future network states, and are difficult to adapt to network environments with dynamic interference and traffic changes. Summary of the Invention

[0005] This application provides a multi-channel routing anti-interference decision method for limited buffers, which can be used to solve the technical problems of routing anti-interference and congestion control with limited buffer capacity in wireless multi-hop networks in multi-source concurrent scenarios.

[0006] This application provides a multi-channel routing anti-interference decision method for limited buffers, and the method includes:

[0007] Step 1: In the interference scenario of a multi-hop wireless network, analyze the influence of the interference of the interferer, the mutual interference between nodes, channel noise, and path loss on the signal-to-interference ratio, and calculate the transmission rate between adjacent nodes.

[0008] Assume that the data packet needs to pass through hops of transmission from the source node to the destination node, and all nodes use the same transmission power , and the transmission power of the interferer is , then the received power of the th-hop node is expressed as:

[0009] (1);

[0010] where the th-hop node and the previous-hop node The channel gain between them is represented by the distance between the two, that is ; The channel gain between the jammer and the hop node is represented by the distance from the jammer to the , that is , , represents the path fading factor; represents the set of neighbor nodes of the hop node , represents the gain from other neighbor nodes to the hop node ; represents the data transmission channel between the hop node and the previous hop node , represents the channel where the interference is located, represents the transmission channel selected by other neighbors of the hop node ; ; is the set of channels; represents the noise of the transmission channel ; and the indicator function is expressed as follows:

[0011] (2);

[0012] (3);

[0013] When , it means that the channel for the routing node to transmit is the same as the channel where the interference is located, and the data link is interfered; when , it means that the channel for the routing node to transmit is the same as the transmission channels of other neighbors, and there is mutual interference in the data link;

[0014] The hop node has a received signal-to-interference ratio of:

[0015] (4);

[0016] According to the received signal-to-interference ratio, the transmission rate between the hop node and the previous hop node is:

[0017] (5);

[0018] where Indicates the node 's received signal-to-interference-plus-noise ratio (SINR), which represents the threshold for successful demodulation by the receiver; when the received SINR is below the demodulation threshold, the communication rate is 0.

[0019] Step 2: Calculate the amount of data in the node buffer based on the packet arrival rate, interference collisions, mutual interference collisions, buffer overflow packet loss, and maximum transmission times packet loss.

[0020] The arrival of source node packets follows a Poisson process with a mean of , and the probability that number of packets arrive at the source node within a time slot is given by:

[0021] (6);

[0022] Assume that each packet has the same length of , then the time for a packet to be transmitted from node to node is:

[0023] (7);

[0024] where represents the transmission rate between node and the previous hop node ;

[0025] Therefore, define the number of successfully transmitted packets at the th hop node in the th time slot as:

[0026] (8);

[0027] where is the maximum data transmission time within a unit time slot, and at most one packet can be transmitted within this time; the maximum number of packets that can be stored in the node buffer is . Assume that at the beginning of time slot , the amount of data in the buffer of the current node is , and the number of received packets is ; if is the source node, it is determined by the source node packet arrival probability, otherwise it is determined by the number of successfully transmitted packets of the previous hop node of the current node in the

[0028] ​​Current node When receiving a data packet, first detect the number of transmissions; if the maximum number of transmissions is reached , if reached, discard the data packet, if not reached, store it in the buffer and wait for transmission; assume the number of lost packets due to reaching the maximum number of transmissions is ; after the node forwards the data packet, delete the data packet from the buffer and release the storage space; therefore, at the nth time slot, the current node buffer data volume is:

[0029] (9);

[0030] When the remaining storage space in the buffer is not enough to store the newly arrived data packet, it will cause packet loss due to buffer packet overflow;

[0031] The goal of all nodes in the network is to maximize the total number of successfully transmitted data packets of all nodes during long-term operation by selecting appropriate channels and next-hop nodes:

[0032] (10);

[0033] In a dynamic interference environment, the channel and next-hop selection of nodes are subject to multiple constraints: First, the node needs to select from a finite set of channels and a set of neighbor nodes ; Second, the transmission time of each hop link must be less than or equal to the maximum data transmission time within a unit time slot, otherwise packet loss will occur; In addition, the data volume in the node buffer is also always restricted by the capacity , where represents the maximum buffer capacity of the node. These constraints together constitute the boundary conditions for node actions, enabling nodes to weigh and play games when making decisions, ultimately affecting the long-term transmission efficiency of the entire network.

[0034] Step 3: Model the anti-interference routing and channel joint optimization problem of the distributed multi-hop network as a partially observable stochastic game; the node realizes the maximization of long-term efficiency based on local observation information.

[0035] In the dynamic optimization scenario for distributed multi-hop cooperative anti-jamming communication, considering the characteristics that nodes only have local sensing and limited information interaction capabilities: when the network is initialized, nodes learn the whole network's static topology through pre-configured coordinate information; during the actual transmission process, due to the lack of a central control unit and limited communication range, nodes cannot detect the dynamic states of non-neighbor nodes in real time. Under this constraint, nodes can autonomously infer the interference pattern and network congestion distribution based on the local spectrum sensing results, the buffer status feedback from neighbor nodes, and the address of the destination node, and make joint routing and channel decisions; in addition, the routing and channel decisions of each node affect other nodes; therefore, the problem of joint optimization of anti-jamming routing and channels in a distributed multi-hop network is modeled as a Partially Observable Stochastic Game (POSG), through described by a six-tuple, specifically defined as:

[0036] Set of agents ( ): The set of all nodes in the network constitutes the set of agents;

[0037] State space ( ): The state space is used to describe the complete information of the network, including the positions of nodes and jammers, the source node and destination node of the data packet, the buffer status of nodes, the node transmission channels, and the interference channels of jammers, expressed as:

[0038] (11);

[0039] where respectively represent the set of positions of communication nodes and jammers, is the set of destination nodes of data packets; represents the set of buffer statuses of all nodes, respectively represent the set of node transmission channels and the set of interference channels of jammers;

[0040] Observation space ( ): The observation space is jointly determined by the interference optional channels, the possible buffer statuses of all neighbor nodes, and the possible destination node addresses of data packets; the observation of node at time slot is expressed as:

[0041] (12);

[0042] where is the destination node of the received data packet , represents the buffer statuses of all neighbor nodes, briefly recorded as , represents the interference channel of the jammer;

[0043] Action space ( ): The action space of node is jointly determined by the set of available channels and the set of neighbor nodes ; In a time slot the node needs to simultaneously select a transmission channel and the next-hop node, ;

[0044] State transition probability ( ): Since the interference channels of the jammer and the channel strategies of other nodes are unknown to node , the state transition probability of the environment is also unknown, and the node needs to continuously interact with the environment to learn this probability;

[0045] Reward function ( ): The influencing factors of the reward function include the cost of routing hops, congestion avoidance, and interference avoidance; However, in a distributed multi-hop routing anti-jamming network, it is difficult for traditional single-objective reward functions to balance multi-dimensional performance metrics. Therefore, a composite reward function is adopted to achieve multi-objective collaborative optimization through decoupled design. The composite reward includes two parts:

[0046] (13);

[0047] (14);

[0048] where represents the reward obtained from the routing decision, represents the reward obtained from the channel decision; represents the number of hops from the current node to the destination node, represents the number of hops from the selected next-hop node to the destination node; represents the buffer overflow of the selected next-hop node , represents that the selected transmission channel has a low transmission rate due to interference or mutual interference, etc., resulting in the inability to transmit a complete data packet within a unit time slot.

[0049] Step 4: In view of the two problems of large state space and difficult multi-objective coordination faced by the ordinary deep network (Deep Q-Network, DQN) in the distributed multi-hop routing anti-jamming problem, a hierarchical DQN architecture is set;

[0050] The routing anti-jamming problem in a multi-channel scenario is split into upper-layer routing planning and lower-layer channel selection to achieve efficient collaborative optimization.

[0051] The input of the upper-layer network is the buffer state obtained from interacting with neighbors and the address of the destination node carried in the data packet, and the output is the selected next-hop node; the input of the lower-layer network is the next-hop node selected by the upper-layer network and the channel where interference is located obtained through spectrum sensing, and the output of the network is the selected transmission channel; both decision-making networks are composed of two fully connected layers.

[0052] For node Define the state-action value function at time , which represents the maximum long-term cumulative reward value that the node can obtain by executing action under the observation state ; The state-action value function consists of two parts:

[0053] (15);

[0054] Among them represents the action value function of routing, represents the observation space of the routing decision-making network, represents the selected next-hop node; represents the action value function of the channel, represents the observation space of the channel decision-making network, represents the selected transmission channel; The update method of the action value function is:

[0055] (16);

[0056] Among them is the learning rate, is the discount factor;

[0057] Define the loss function of the network of node as:

[0058] (17);

[0059] Among them and represent the weight parameters of the prediction network of the node, and represent the optimization objectives of the network, defined as:

[0060] (18);

[0061] Among them and represent all possible actions, and are the weight parameters of the target network;

[0062] The network is trained using the gradient descent method, and the gradient of the loss function is:

[0063] (19);

[0064] Considering that the traditional ε-greedy strategy may make the network environment unstable in a distributed network, the present invention adopts the Boltzmann update strategy and defines the routing selection strategy and the channel selection strategy with the update formula:

[0065] (20);

[0066] where are the relevant parameters of the Boltzmann model:

[0067] (21);

[0068] where represents the initial temperature of the Boltzmann, represents the minimum temperature, which affects the transition time between the exploration and exploitation phases.

[0069] This application proposes a hierarchical deep reinforcement learning method to achieve dynamic anti-interference and congestion control of multi-hop wireless networks by decoupling routing planning and channel selection. At the routing decision layer, nodes dynamically select the next-hop node based on the neighbor buffer status and the destination address, avoiding the aggregation of data packets to saturated nodes and reducing overflow packet loss; at the channel decision layer, non-interfering transmission channels are dynamically accessed by combining the routing results and real-time spectrum sensing. This method improves the successful transmission rate of data packets and speeds up the convergence rate of the algorithm. BRIEF DESCRIPTION OF THE DRAWINGS

[0070] The drawings described herein are used to provide a further understanding of the present invention, form a part of this application, and the illustrative embodiments and descriptions of the present invention are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0071] Figure 1 is a multi-channel routing system diagram with limited buffer provided by an embodiment of this application;

[0072] Figure 2 is a time slot structure diagram of a node and a jammer provided by an embodiment of this application;

[0073] Figure 3 is a hierarchical DQN network framework diagram provided by an embodiment of this application;

[0074] Figure 4 is a simulation topology diagram provided by an embodiment of this application;

[0075] Figure 5 A comparison chart of the success probabilities of different algorithms provided by the embodiments of the present application;

[0076] Figure 6 A comparison chart of the success probabilities of different network parameters provided by the embodiments of the present application;

[0077] Figure 7 A comparison chart of the successful transmission quantities of different algorithms provided by the embodiments of the present application;

[0078] Figure 8 A comparison chart of the successful transmission quantities of different network parameters provided by the embodiments of the present application;

[0079] Figure 9 A node selection probability distribution chart under low load provided by the embodiments of the present application;

[0080] Figure 10 A node selection probability distribution chart after the load increases provided by the embodiments of the present application;

[0081] Figure 11 A node selection probability distribution chart under high load provided by the embodiments of the present application. Detailed implementation manners

[0082] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe in detail the embodiments of the present application in conjunction with the accompanying drawings.

[0083] The present invention discloses a method for solving the problems of routing anti-interference and congestion control with limited buffer capacity in a wireless multi-hop network in a multi-source concurrent scenario. The present application proposes a distributed cooperative anti-interference algorithm based on hierarchical deep reinforcement learning. By decoupling the routing planning and channel selection decisions, a hierarchical deep Q-network (HDQN) is constructed: the upper-layer routing decision network predicts the network congestion trend based on the real-time status of the neighbor node buffers and the destination node address, and dynamically avoids high-load nodes; the lower-layer channel decision network combines the upper-layer routing decision result and the real-time spectrum sensing information to avoid interference frequency points. In the design of the reward value, the hop count cost at the network layer, congestion penalty, and interference avoidance at the physical layer are respectively considered. Through this method of decoupling routing planning and channel decision-making, the coordinated optimization of congestion control, anti-interference transmission, and routing hop count is achieved.

[0084] The following first introduces the embodiments of the present application in conjunction with the accompanying drawings.

[0085] Figure 1 is the system model of the present invention, which consists of a jammer, a group of source nodes, a destination node, and m intermediate nodes to form a wireless multi-hop network, denoted as , where represents the set of source nodes, represents the set of destination nodes, represents the set of intermediate nodes. The system includes channels with a bandwidth of , denoted as . Due to distance limitations, the source node cannot directly transmit information to the destination node , and multi-hop routing forwarding is required. Each node is equipped with multiple antennas and supports parallel multi-channel transceiver in full-duplex mode. However, when a node receives data packets from multiple neighbor nodes on the same channel, packet collisions and losses will occur. Nodes can search for "spectrum holes" through real-time spectrum sensing and dynamically identify idle channels. The data is temporarily stored in a first-in-first-out buffer before transmission. The buffer has a limited capacity and the maximum storage capacity is . The jammer dynamically selects the interference channel using an unknown strategy. Each interference time slot can only interfere with one channel, aiming to disrupt the data link between nodes.

[0086] Figure 2 In the present invention, the nodes and the jammer adopt an asynchronous time slot structure, and the time slot structures of both sides are different and unknown to each other. The time slots of each node are completely synchronized, but not completely synchronized with the jammer. Each node time slot can be divided into four stages: 1. Spectrum sensing stage: The node senses the current spectrum environment, identifies the interference channel, and uses it as the observation input for reinforcement learning; 2. Decision learning stage: The node selects the channel and the next-hop node based on the sensing result, the buffer state feedback from the neighbor nodes in the previous time slot, and the destination node address; 3. Data transmission stage: Transmit on the data link according to the next-hop node and channel determined by the decision; 4. Feedback stage: The node feeds back the transmission result and the current buffer status to the neighbor nodes through the control channel.

[0087] Figure 3 is the algorithm framework diagram of the present invention. Aiming at the two problems of large state space and difficult coordination of multiple objectives in the traditional single-layer DQN for distributed multi-hop routing anti-jamming problems, a hierarchical DQN architecture is proposed, as shown in Figure 3 . This method splits the routing anti-jamming problem in a multi-channel scenario into high-level routing planning and low-level channel selection to achieve efficient collaborative optimization. The input of the upper-layer network is the buffer state obtained from interacting with neighbors and the destination node address carried in the data packet, and the output is the selected next-hop node. The input of the lower-layer network is the next-hop node selected by the upper-layer network and the channel where the interference is located obtained through spectrum sensing, and the output of the network is the selected transmission channel. Both decision-making networks are composed of two fully connected layers.

[0088] Figure 4is the simulation topology diagram of the present invention. To verify the algorithm performance, simulations were carried out in a multi-source node topology environment as shown in Figure 4 which contains 18 nodes and 1 dynamic jammer. Among them, nodes 0, 1, and 2 are source nodes, and the corresponding destination nodes are nodes 14, 15, and 10 respectively.

[0089] Figure 5 The probability comparison of 100 data packets successfully reaching the destination node by different algorithms was statistically analyzed. In the initial stage of learning, the minimum-hop node selection combined with the DQN channel selection algorithm (minimum-hop + DQN) has the best effect because it quickly locks the transmission path with the fewest hops and only needs to optimize the channel selection. However, as the number of learning times increases, the advantages of the proposed algorithm gradually emerge, while the minimum-hop + DQN algorithm exposes defects due to not considering the impact of the buffer on congestion, resulting in an increase in the packet loss rate. Ordinary DQN has a large state space and action space, and the decoupled routing decision and channel decision lead to a slow learning rate and poor performance.

[0090] Figure 6 The success probability comparison of different algorithms under different network loads and buffer capacities was statistically analyzed. The results show that the proposed algorithm is superior to the other two algorithms under different parameters. Hierarchical DQN can more efficiently combine the neighbor buffer state feedback and spectrum sensing results by decoupling the routing and channel decisions, avoiding packet loss caused by buffer overflow or channel conflict, thus significantly improving the data transmission success rate. In addition, the higher the network load, the more scarce the resources, that is, the larger the smaller, the more obvious the superiority of the proposed algorithm.

[0091] Figure 7 The number of data packets successfully transmitted by the three algorithms within every 100 time slots was statistically analyzed. The results show that the proposed algorithm is superior to the other two algorithms in terms of convergence speed and transmission efficiency.

[0092] Figure 8 The comparison of the number of successfully transmitted packets of different algorithms under different network loads and buffer capacities was statistically analyzed. By comparing and the number of successfully transmitted packets, it can be found that even if the buffer capacity is increased, the number of successfully transmitted packets of the minimum-hop DQN algorithm has not been improved. This is because although the increase in the buffer can relieve congestion to a certain extent, it will also cause packet accumulation and reduce the transmission efficiency.

[0093] Figure 9 、 Figure 10 、 Figure 11 The influence of different network loads on the node's next-hop selection strategy was statistically analyzed under the condition of buffer capacity . Figure 9 Among them, at low load When the network resources are sufficient, according to the design of the reward value formula, all nodes tend to choose the path with the shortest hop count. Node 8, being located at the intersection of multiple shortest paths, becomes a common forwarding node with a high probability. However Figure 10 when the load increases in , the nodes need to make a balanced choice between hop count optimization and congestion control. At this time, the nodes will judge the network situation based on the buffer status of the neighbor nodes. When the buffer of the neighbor nodes on the shortest path is occupied to a large extent, the nodes will choose to sacrifice part of the hop count optimization goal, actively avoid the neighbors with saturated buffers, and reduce congestion packet loss. Figure 11 when the load is very high in , the data packets generated by different source nodes will be transmitted separately in a relatively independent manner, trying to avoid overlapping paths and congestion.

[0094] The embodiments of the present application described above do not constitute a limitation on the protection scope of the present application.

Claims

1. A multi-channel routing anti-interference decision-making method for a finite buffer, characterized in that, The method includes: Step 1: In the interference scenario of a multi-hop wireless network, analyze the impacts of the interference from interferers, mutual interference between nodes, channel noise, and path loss on the signal-to-interference ratio, and calculate the transmission rate between adjacent nodes; Step 2: Calculate the node buffer data volume according to the packet arrival rate, interference conflict, mutual interference conflict, buffer overflow packet loss, and packet loss due to the maximum number of transmissions; Step 3: Model the joint optimization problem of anti-interference routing and channel in a distributed multi-hop network as a partially observable stochastic game; the node maximizes the long-term efficiency based on local observation information; Step 4: Set the hierarchical DQN architecture; split the routing anti-interference problem in a multi-channel scenario into upper-layer routing planning and lower-layer channel selection.

2. The method according to claim 1, wherein Step 1: In the interference scenario of a multi-hop wireless network, analyze the impacts of the interference from interferers, mutual interference between nodes, channel noise, and path loss on the signal-to-interference ratio, and calculate the transmission rate between adjacent nodes, including: Assume that the data packet needs to pass through hops from the source node to the destination node, and all nodes use the same transmission power , and the transmission power of the jammer is , then the received power of the th-hop node is expressed as: (1); Among them, the hop node and the previous hop node channel gain between is represented by the distance between the two, that is ; The channel gain between the jammer and the hop node hop node channel gain between is represented by the distance from the jammer to distance is represented, that is , represents the path fading factor; represents the hop node set of neighbor nodes of represents the gain from other neighbor nodes to the hop node ; represents the hop node and the previous hop node data transmission channel between represents the channel where the interference is located, represents the hop node transmission channels selected by other neighbors of ; is the channel set; represents the noise of the transmission channel ; and The indicator function is represented as follows: (2); (3); When it indicates that the channel through which the routing node transmits is the same as the channel where the interference is located, and the data link is interfered; when it indicates that the channel through which the routing node transmits is the same as the transmission channels of other neighbors, and there is mutual interference in the data link; The hop node has a received signal-to-interference-plus-noise ratio of: (4); Determine the hop node and the previous hop node The transmission rate between them is: (5); Among them represents the received signal-to-interference-plus-noise ratio of the node , and represents the threshold for successful demodulation by the receiver; when the received is lower than the demodulation threshold, the communication rate is zero.

3. The method according to claim 1, characterized in that, Step 2: Calculate the node buffer data volume according to the packet arrival rate, interference conflict, mutual interference conflict, buffer overflow packet loss, and packet loss due to the maximum number of transmissions, including: The arrival of data packets at the source node follows a Poisson process with a mean of , The probability that packets arrive at the source node within a time slot is: (6); Assume that each data packet has the same length of , then the time for the data packet to be transmitted from node to node is: (7); Among them represents the node and the previous-hop node the transmission rate between them; Therefore, define the time slot hop node number of successfully transmitted data packets as: (8); Among them is the maximum data transmission time within a unit time slot, at most one data packet can be transmitted within the time; the maximum number of data packets that can be stored in the node buffer is , assuming the time slot initially, the amount of data in the buffer of the current node is , and the number of received data packets is ; if is the source node, it is determined by the arrival probability of the source node packets, otherwise it is determined by the number of data packets successfully transmitted by the previous hop node of the current node in the time slot ; Current node When receiving a data packet, first detect the number of transmissions; if the maximum number of transmissions is reached , if reached, discard the data packet, if not reached, store it in the buffer and wait for transmission; assume the number of packets lost due to reaching the maximum number of transmissions is ; after the node forwards the data packet, delete the data packet from the buffer and release the storage space; therefore, at the time slot, the current node The data volume in the buffer is: (9); When the remaining storage space in the buffer is not enough to store newly arrived packets, buffer packet overflow loss occurs; The goal of all nodes in the network is to maximize the total number of packets successfully transmitted by all nodes during long-term operation by selecting appropriate channels and next-hop nodes: (10); Under dynamic interference environments, the channel and next-hop selection of nodes are subject to multiple constraints: First, a node needs to select from a finite set of channels and a set of neighbor nodes ; Second, the transmission time of each-hop link must be less than or equal to the maximum data transmission time within a unit time slot, otherwise packet loss will occur; In addition, the amount of data in the node buffer is also always restricted by the capacity , where represents the maximum buffer capacity of the node.

4. The method according to claim 1, wherein Step 3: Model the joint optimization problem of anti-interference routing and channel in a distributed multi-hop network as a partially observable stochastic game; the node maximizes the long-term efficiency based on local observation information, including: Nodes can autonomously infer the interference pattern and network congestion distribution based on local spectrum sensing results, the buffer status feedback from neighbor nodes, and the address of the destination node, and make joint routing and channel decisions; in addition, the routing and channel decisions of each node affect other nodes; therefore, the problem of joint optimization of anti-interference routing and channels in a distributed multi-hop network is modeled as a partially observable stochastic game, which is described by a six-tuple, and the specific definition is as follows: Agent set( ): The set of all nodes in the network forms the agent set; State space( ): The state space is used to describe the complete information of the network, including the positions of nodes and jammers, the source node and destination node of data packets, the node buffer status, the node transmission channels, and the interference channels of jammers, which is expressed as: (11); wherein respectively represent the position sets of communication nodes and jammers, is the set of destination nodes of data packets; represents the set of buffer states of all nodes, respectively represent the node transmission channel and the jammer interference channel sets; Observation Space( ): The observation space is jointly determined by the interfering optional channels, the possible buffer states of all neighbor nodes, and the possible destination node addresses of the data packets; the node in the time slot is observed as follows: (12); Among them The received data packet The destination node of Indicates the buffer status of all neighbor nodes, briefly recorded as , Indicates the interference channel of the jammer; Action space ( ): The action space of a node is jointly determined by the set of available channels and the set of neighbor nodes ; at each time slot a node needs to simultaneously select a transmission channel and a next-hop node, ; State transition probability( ): Since the interference channels of the jammer and the channel strategies of other nodes are unknown to node , the state transition probability of the environment is also unknown, and the node needs to continuously interact with the environment to learn this probability; Reward function( ): The influencing factors of the reward function include the cost of routing hops, congestion avoidance, and interference avoidance; a composite reward function is adopted to achieve multi-objective collaborative optimization through decoupled design. The composite reward includes two parts: (13); (14); wherein represents the reward obtained from the routing decision, represents the reward obtained from the channel decision; represents the current node the number of hops to the destination node, represents the selected next-hop node the number of hops to the destination node; represents the selected next-hop node buffer overflow, represents that the selected transmission channel has a low transmission rate due to interference or mutual interference factors, resulting in the inability to transmit a complete data packet within a unit time slot.

5. The method according to claim 1, characterized in that, Step 4: Set the hierarchical DQN architecture; Split the routing anti-interference problem in a multi-channel scenario into upper-layer routing planning and lower-layer channel selection, including: The input of the upper-layer network is the buffer state obtained by interacting with neighbors and the address of the destination node carried in the packet, and the output is the selected next-hop node; The input of the lower-layer network is the next-hop node selected by the upper-layer network and the channel where the interference is located obtained through spectrum sensing, and the output of the network is the selected transmission channel; both decision-making networks are composed of two fully connected layers; For a node Define The state-action value function at a moment , which represents the maximum long-term cumulative return value that the node can obtain by executing the action under the observed state ; The state-action value function consists of two parts: (15); Among them represents the action value function of routing represents the observation space of the routing decision network represents the selected next-hop node represents the action value function of the channel represents the observation space of the channel decision network represents the selected transmission channel; the action value function update method is as follows (16); wherein is the learning rate, is the discount factor; Define the node The loss function of the network is as follows: (17); wherein and represent the weight parameters of the prediction network of the node, and represent the optimization objective of the network, which is defined as: (18); wherein and represent all possible actions, and are the weight parameters of the target network; Use the gradient descent method to train the network, and the gradient of the loss function is: (19); Adopt the Boltzmann update strategy and define the routing selection strategy and the channel selection strategy The update formula is as follows: (20); Among them are the relevant parameters of the Boltzmann model: (21); Among them represents the Boltzmann initial temperature represents the minimum temperature affects the transition time between the exploration and exploitation phases

Citation Information

Patent Citations

  • System-level transmission delay model building method applied to network on chip

    CN102693213A

  • Multi-target batch scheduling method for solar cell module limited relief area based on second DNSGA

    CN104794322A

  • Anti-interference zero sum Markov game model and maximum and minimum depth Q learning method

    CN116866048A

  • Joint channel and route selection cross-layer decision-making method based on Q learning

    CN119484387A

  • Distributed collaborative evolution method, UAV and intelligent routing method therefor, and apparatus

    WO2024021281A1