A multi-channel routing anti-interference decision method for limited buffer zones

Through the hierarchical deep reinforcement learning method, decoupling routing planning and channel selection, the anti-interference and congestion control problems of limited buffer capacity in wireless multi-hop networks are solved, and the packet transmission success rate and network efficiency are improved.

CN120282236BActive Publication Date: 2025-08-12ARMY ENG UNIV OF PLA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510782921.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-08-12
Estimated Expiration
2045-06-12

AI Technical Summary

Technical Problem

In multi-source concurrency scenarios, in wireless multi-hop networks, the anti-interference and congestion control problems of routing with limited buffer capacity are difficult to effectively solve. Traditional routing algorithms are prone to network congestion and transmission delays when high loads are high, and cannot adapt to dynamic interference environments.

Method used

The hierarchical deep reinforcement learning method is adopted to build a hierarchical deep Q network (HDQN) by decoupling routing planning and channel selection. The upper-layer routing decision is based on the buffer status of neighbor nodes and the destination node address to dynamically avoid high-load nodes; the lower-layer channel decision network combines routing results and real-time spectrum perception information to avoid interference frequencies, and designs a composite reward function to achieve multi-objective collaborative optimization.

Benefits of technology

It improves the successful transmission rate of data packets, reduces congestion and packet loss, improves the convergence speed of the algorithm, and optimizes the long-term transmission efficiency of the network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120282236B_ABST
    Figure CN120282236B_ABST
Patent Text Reader

Abstract

The present application provides a multi-channel routing anti-interference decision-making method for a limited buffer, comprising the following steps: Step 1: In an interference scenario of a multi-hop wireless network, analyzing the impact of interference from the jammer, mutual interference between nodes, channel noise, and path loss on the signal-to-interference ratio, and calculating the transmission rate between adjacent nodes; Step 2: Calculating the node buffer data volume based on packet arrival rate, interference conflict, mutual interference conflict, buffer overflow packet loss, and maximum transmission number packet loss; Step 3: Modeling the anti-interference routing and channel joint optimization problem of a distributed multi-hop network as a partially observable random game; Nodes maximize long-term efficiency based on local observation information; Step 4: Setting a hierarchical DQN architecture; Splitting the routing anti-interference problem in a multi-channel scenario into upper-layer routing planning and lower-layer channel selection. This method improves the successful transmission rate of data packets and accelerates the convergence of the algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of wireless communications, and in particular relates to a multi-channel routing anti-interference decision method oriented to a limited buffer zone. Background Art

[0002] In an ad hoc network, when multiple source nodes are concurrently processing traffic, key nodes located at the intersection of multiple paths must process multiple data streams. This poses a significant challenge to their limited buffer capacity, which can easily lead to network congestion. Simply increasing buffer capacity can temporarily alleviate transient congestion, but this increases packet queuing time and increases end-to-end transmission latency. Furthermore, due to the open nature of wireless channels, nodes in the network are susceptible to external interference. This interference intensifies competition for spectrum resources and reduces transmission reliability. Therefore, nodes must optimize path planning, dynamic spectrum allocation, and congestion control strategies within multiple constraints, such as buffer capacity and interference avoidance, to achieve global transmission efficiency and local node coordination.

[0003] Traditional routing algorithms, such as Ad hoc On-Demand Distance Vector Routing (AODV) and Destination-Sequenced Distance-Vector Routing (DSDV), while advantageous in simplifying decision-making, struggle to cope with highly dynamic, multi-constrained scenarios. The AODV protocol selects paths based on the shortest hop count criterion. Its greedy path selection mechanism, while enabling rapid establishment of low-hop routes in low-load scenarios, can lead to excessive traffic concentration at topology nodes in high-load scenarios, causing network congestion. DSDV, a typical table-driven routing algorithm, requires periodic broadcasting of routing tables by all network nodes to maintain path status information. In dynamic interference environments, inter-node link connectivity frequently changes, necessitating continuous topology updates to maintain routing validity. DSDV significantly increases network control message overhead. Furthermore, due to the lag in routing table information, path selection errors occur, making it difficult to adapt to dynamic interference.

[0004] The single-dimensional optimization mechanisms of these traditional routing protocols make it difficult to achieve multi-objective coordination, such as minimizing hop count, mitigating interference, and controlling congestion, severely hindering overall network performance. Numerous improved protocols based on traditional routing protocols have been proposed. For example, an enhanced Greedy Perimeter Stateless Routing (GPSR) protocol, based on dynamic node buffer adjustment, calculates the probability of selecting the next-hop node by combining geographic distance and remaining node buffer capacity. This approach effectively addresses the congestion problem at key nodes caused by the traditional GPSR protocol's single-dimensional routing decisions, significantly reducing network transmission latency caused by queue accumulation at key nodes in scenarios with concurrent transmission of multiple data streams. To address the problem that traditional shortest path algorithms are effective under low loads but prone to congestion under high loads, a regularized routing optimization (RRO) algorithm has been proposed. By combining a congestion function with a regularization term based on path length, this method significantly improves network throughput and reduces latency. However, these methods rely solely on static calculations based on the current state and cannot predict future network states, making them difficult to adapt to dynamic interference and traffic fluctuations. Summary of the Invention

[0005] The present application provides a multi-channel routing anti-interference decision method for limited buffers, which can be used to solve the technical problems of routing anti-interference and congestion control with limited buffer capacity in wireless multi-hop networks in multi-source concurrent scenarios.

[0006] The present application provides a multi-channel routing anti-interference decision method for a limited buffer zone, the method comprising:

[0007] Step 1: In a multi-hop wireless network interference scenario, analyze the impact of jammer interference, mutual interference between nodes, channel noise, and path loss on the signal-to-interference ratio (SIR), and calculate the transmission rate between adjacent nodes.

[0008] Assume that the data packet needs to pass through from the source node to the destination node Hop transmission, all nodes use the same transmission power , the jammer transmit power is , then Jump Node The received power is expressed as:

[0009] (1);

[0010] Among them Jump Node and the previous hop node Channel gain between Use the distance between the two Indicates that ; Jammer and Jump Node Channel gain between Use jammer to distance Indicates that , represents the path fading factor; Indicates the Jump Node The set of neighbor nodes of Indicates other neighbor nodes to the Jump Node Gain; Indicates the Jump Node and the previous hop node The data transmission channel between Indicates the channel where the interference is located, Indicates the Jump Node The transmission channel selected by other neighbors of ; is the channel set; Indicates the transmission channel noise; and The indicator function is expressed as follows:

[0011] (2);

[0012] (3);

[0013] when When , it means that the channel transmitted by the routing node is the same as the interference channel, and the data link is interfered; when When , it means that the transmission channel of the routing node is the same as that of other neighbors, and there is mutual interference in the data link;

[0014] No. Jump Node The received signal-to-interference ratio is:

[0015] (4);

[0016] According to the received signal-to-interference ratio, determine the Jump Node and the previous hop node The transmission rate between them is:

[0017] (5);

[0018] in Representation node The received signal-to-interference ratio, Indicates the threshold for successful demodulation by the receiver; when receiving When the value is lower than the demodulation threshold, the communication rate is 0.

[0019] Step 2: Calculate the node buffer data volume based on the packet arrival rate, interference conflict, mutual interference conflict, buffer overflow packet loss, and maximum transmission number packet loss.

[0020] The data packets arriving at the source node obey the mean Poisson process, Source node in time slot arrive The probability of a data packet is:

[0021] (6);

[0022] Assume that each packet has the same length as , then the data packet from node Transfer to node The time is:

[0023] (7);

[0024] in Representation node and the previous hop node The transmission rate between

[0025] Therefore, define Time Slot Jump Node Number of packets successfully transmitted for:

[0026] (8);

[0027] in is the maximum data transmission time within a unit time slot, At most one data packet can be transmitted within a certain time period; the maximum number of data packets that can be stored in the node buffer is , assuming time slot Initial current node The amount of data in the buffer is , the number of packets received is ;like is the source node, Determined by the probability of the source node packet arriving, otherwise Current node of time slot It is determined by the number of data packets successfully transmitted by the previous hop node;

[0028] Current node When receiving a data packet, the transmission times are checked first; if the maximum transmission times are reached If the number of packets is reached, the data packet is discarded; if it is not reached, it is stored in the buffer and waits for transmission; Assume that the number of packets lost due to reaching the maximum number of transmissions is ; After the node forwards the data packet, it deletes the data packet from the buffer to release the storage space; therefore, Current node of time slot Buffer data volume for:

[0029] (9);

[0030] When the remaining storage space in the buffer is insufficient to store the newly arrived data packets, the buffer overflows and packet loss occurs.

[0031] The goal of all nodes in the network is to maximize the Total number of packets successfully transmitted during the long run:

[0032] (10);

[0033] In a dynamic interference environment, the channel and next hop selection of a node is subject to multiple constraints: first, the node must be in a limited set of channels. and neighbor node set Secondly, the transmission time of each hop link must be less than or equal to the maximum data transmission time within the unit time slot, otherwise packet loss will occur; in addition, the amount of data in the node buffer is always limited by the capacity. constraints, among which Represents the maximum buffer capacity of the node. These constraints together constitute the boundary conditions of node actions, causing nodes to make trade-offs when making decisions, ultimately affecting the long-term transmission efficiency of the entire network.

[0034] Step 3: Model the joint optimization problem of interference-resistant routing and channels in distributed multi-hop networks as a partially observable stochastic game; nodes maximize long-term efficiency based on local observation information.

[0035] In the dynamic optimization scenario for distributed multi-hop collaborative anti-interference communication, the characteristics of nodes that only have local perception and limited information exchange capabilities are considered: when the network is initialized, the nodes obtain the static topology of the entire network through pre-configured coordinate information; in the actual transmission process, due to the lack of a central control unit and limited communication range, the nodes cannot detect the dynamic status of non-neighbor nodes in real time. Under this constraint, the node can autonomously infer the interference pattern and network congestion distribution based on the local spectrum perception results, the buffer status fed back by the neighboring nodes, and the address of the destination node, and make joint routing and channel decisions; in addition, the routing and channel decisions of each node affect other nodes; therefore, the anti-interference routing and channel joint optimization problem of distributed multi-hop networks is modeled as a partially observable stochastic game (POSG). The six-tuple description is defined as:

[0036] Agent Collection( ): All nodes in the network constitute the set of intelligent agents;

[0037] State Space ( ): The state space is used to describe the complete network information, including the location of nodes and jammers, the source and destination nodes of data packets, the node buffer status, the node transmission channel and the jammer's interference channel, which is expressed as:

[0038] (11);

[0039] in represent the location sets of communication nodes and jammers respectively, is the destination node set of the data packet; represents the set of buffer states of all nodes, They represent the node transmission channel and the interference channel set of the jammer respectively;

[0040] Observation space ( ): The observation space is determined by the interference optional channel, the possible buffer states of all neighboring nodes and the possible destination node addresses of the data packets; the node In the time slot The observation is expressed as:

[0041] (12);

[0042] in Received data packets The destination node, Represents the buffer status of all neighbor nodes, abbreviated as , Indicates the jammer's jamming channel;

[0043] Action Space( ):node The action space is composed of the set of available channels and neighbor node set Joint decision-making; time slot The node needs to select the transmission channel and the next hop node at the same time. ;

[0044] State transition probability ( ): Due to the interference channel of the jammer and the channel strategy of other nodes, is unknown, so the state transition probability of the environment is also unknown, and the node needs to learn this probability by continuously interacting with the environment;

[0045] Reward function ( ): The factors affecting the reward function include routing hop cost, congestion avoidance, and interference avoidance. However, in distributed multi-hop routing and anti-interference networks, traditional single-objective reward functions have difficulty balancing multi-dimensional performance indicators. Therefore, a composite reward function is adopted to achieve multi-objective collaborative optimization through decoupling design. The composite reward consists of two parts:

[0046] (13);

[0047] (14);

[0048] in represents the reward obtained by the routing decision, Represents the reward obtained by channel decision; Indicates the current node The number of hops to the destination node, Indicates the selected next hop node Number of hops to the destination node; Indicates the selected next hop node Buffer overflow, This means that the selected transmission channel has a low transmission rate due to interference or mutual interference, resulting in the inability to transmit a complete data packet within the unit time slot.

[0049] Step 4: To address the distributed multi-hop routing anti-interference problem, a conventional deep Q-Network (DQN) faces two problems: a large state space and difficulty coordinating multiple objectives. A hierarchical DQN architecture is established.

[0050] The routing anti-interference problem in multi-channel scenarios is split into upper-layer routing planning and lower-layer channel selection to achieve efficient collaborative optimization.

[0051] The upper-layer network inputs are the buffer status obtained through interaction with neighbors and the destination node address carried in the data packet, and its output is the selected next-hop node. The lower-layer network inputs are the next-hop node selected by the upper-layer network and the channel where the interference is located, obtained through spectrum sensing. The network output is the selected transmission channel. Both decision networks consist of two fully connected layers.

[0052] For nodes definition The state-action-value function at the moment , indicating that the node is in the observation state Next action The maximum long-term cumulative reward value that can be obtained; the state-action value function consists of two parts:

[0053] (15);

[0054] in represents the action-value function of the routing, represents the observation space of the routing decision network, Indicates the selected next hop node; represents the action-value function of the channel, represents the observation space of the channel decision network, represents the selected transmission channel; the action value function is updated as follows:

[0055] (16);

[0056] in is the learning rate, is the discount factor;

[0057] Defining Nodes The loss function of the network is:

[0058] (17);

[0059] in and represents the weight parameter of the prediction network of the node, and represents the optimization goal of the network, which is defined as:

[0060] (18);

[0061] in and represents all possible actions, and is the weight parameter of the target network;

[0062] The network is trained using the gradient descent method, and the gradient of the loss function is:

[0063] (19);

[0064] Considering that the traditional ε-greedy strategy may make the network environment unstable in a distributed network, the present invention adopts the Boltzmann update strategy to define the routing strategy and channel selection strategy The update formula is:

[0065] (20);

[0066] in are the relevant parameters of the Boltzmann model:

[0067] (twenty one);

[0068] in represents the Boltzmann initial temperature, Indicates the minimum temperature, Affects the transition time between exploration and exploitation phases.

[0069] This application proposes a hierarchical deep reinforcement learning method that achieves dynamic interference mitigation and congestion control in multi-hop wireless networks by decoupling route planning and channel selection. At the routing decision layer, nodes dynamically select the next hop based on the state of neighbor buffers and the destination address, preventing packets from congregating at saturated nodes and reducing overflow and packet loss. At the channel decision layer, routing results are combined with real-time spectrum sensing to dynamically access non-interfering transmission channels. This method improves the success rate of packet transmission and accelerates algorithm convergence. BRIEF DESCRIPTION OF THE DRAWINGS

[0070] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0071] Figure 1 A diagram of a multi-channel routing system with limited buffer provided in an embodiment of the present application;

[0072] Figure 2 A diagram showing the node and jammer time slot structure provided in an embodiment of the present application;

[0073] Figure 3 A diagram of the hierarchical DQN network framework provided in an embodiment of the present application;

[0074] Figure 4 The simulation topology diagram provided in the embodiment of the present application;

[0075] Figure 5 A comparison chart of the success probabilities of different algorithms provided in the embodiments of this application;

[0076] Figure 6 A comparison chart of the success probability of different network parameters provided in the embodiment of this application;

[0077] Figure 7 A comparison chart of the number of successful transmissions of different algorithms provided in the embodiments of this application;

[0078] Figure 8 A comparison chart of the number of successful transmissions with different network parameters provided in the embodiments of this application;

[0079] Figure 9 A probability distribution diagram of node selection under low load provided by an embodiment of the present application;

[0080] Figure 10 A probability distribution diagram of node selection after load increase provided in an embodiment of the present application;

[0081] Figure 11 This is a node selection probability distribution diagram under high load provided by an embodiment of the present application. DETAILED DESCRIPTION

[0082] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.

[0083] The present invention discloses a method for solving routing anti-interference and congestion control problems in wireless multi-hop networks with limited buffer capacity in multi-source concurrent scenarios. This application proposes a distributed collaborative anti-interference algorithm based on hierarchical deep reinforcement learning. By decoupling routing planning and channel selection decisions, a hierarchical deep Q-network (HDQN) is constructed: the upper-layer routing decision network predicts network congestion trends based on the real-time status of neighboring node buffers and the destination node address, and dynamically avoids high-load nodes; the lower-layer channel decision network combines the upper-layer routing decision results with real-time spectrum perception information to avoid interference frequencies. The reward value design considers the hop cost of the network layer, congestion penalty, and interference avoidance of the physical layer. Through this method of decoupling routing planning and channel decision-making, the coordinated optimization of congestion control, anti-interference transmission, and routing hop count is achieved.

[0084] The following first introduces the embodiments of the present application with reference to the accompanying drawings.

[0085] Figure 1 The system model of the present invention consists of a jammer, Group source node, destination node and The wireless multi-hop network consists of intermediate nodes, which is represented as ,in represents the set of source nodes, represents the set of destination nodes, Represents a collection of intermediate nodes. The system contains The bandwidth is The channel is denoted as Due to the distance limitation, the source node Cannot send data directly to the destination node Transmitting information requires multi-hop routing. Each node is equipped with multiple antennas and supports multi-channel parallel transmission and reception in full-duplex mode. However, when a node receives data packets from multiple neighboring nodes on the same channel, conflicts and packet loss will occur. Nodes can use real-time spectrum sensing to find "spectrum holes" and dynamically identify idle channels. Data is temporarily stored in a first-in-first-out buffer before being sent. The buffer capacity is limited and the maximum storage capacity is The jammer uses an unknown strategy to dynamically select jamming channels and can only jam one channel in each jamming time slot, aiming to destroy the data link between nodes.

[0086] Figure 2 The nodes and jammers of the present invention adopt an asynchronous time slot structure, and the time slot structures of the two parties are different and unknown to each other. The time slot of each node is completely synchronized, but not completely synchronized with the jammer. Each node time slot can be divided into four stages: 1. Spectrum perception stage: The node perceives the current spectrum environment and identifies the interference channel as the observation input of reinforcement learning; 2. Decision learning stage: The node selects the channel and the next hop node based on the perception results, the buffer status fed back by the neighboring node in the previous time slot, and the destination node address; 3. Data transmission stage: Transmission is performed on the data link based on the next hop node and channel decided; 4. Feedback stage: The node feeds back the transmission results and the current buffer status to the neighboring node through the control channel.

[0087] Figure 3 This is the algorithm framework diagram of the present invention. In order to solve the problem of distributed multi-hop routing anti-interference, the traditional single-layer DQN faces two problems: large state space and difficulty in coordinating multiple targets. A hierarchical DQN architecture is proposed, such as Figure 3 This method splits the routing anti-interference problem in multi-channel scenarios into high-level route planning and low-level channel selection, achieving efficient collaborative optimization. The upper-layer network inputs the buffer status obtained through interaction with neighbors and the destination node address carried in the data packet, and its output is the selected next-hop node. The lower-layer network inputs the next-hop node selected by the upper-layer network and the channel where the interference occurs, as determined through spectrum sensing. The network output is the selected transmission channel. Both decision networks consist of two fully connected layers.

[0088] Figure 4This is the simulation topology diagram of the present invention. To verify the performance of the algorithm, Figure 4 The simulation experiment is carried out in a multi-source node topology environment including 18 nodes and 1 dynamic jammer, where nodes 0, 1, and 2 are source nodes and the corresponding destination nodes are nodes 14, 15, and 10 respectively.

[0089] Figure 5 The probability of 100 packets successfully reaching the destination node after being sent by different algorithms was compared. In the initial learning phase, the minimum hop node selection algorithm combined with the DQN channel selection algorithm (minimum hop + DQN) achieved the best results, as it quickly locked onto the transmission path with the fewest hops, requiring only optimization of channel selection. However, as the number of learning cycles increased, the advantages of the proposed algorithm gradually became apparent, while the minimum hop + DQN algorithm, due to its failure to account for the impact of the buffer on congestion, became vulnerable, leading to increased packet loss rates. Standard DQN, due to its large state and action spaces and its failure to decouple routing and channel decisions, exhibited a slow learning rate and poor performance.

[0090] Figure 6 The success probability of different algorithms under different network loads and buffer capacities was statistically compared. The results show that the proposed algorithm outperforms the other two algorithms under different parameters. By decoupling routing and channel decisions, the hierarchical DQN can more efficiently combine neighbor buffer state feedback and spectrum sensing results, avoiding packet loss caused by buffer overflow or channel conflict, thereby significantly improving the data transmission success rate. In addition, the higher the network load, the more scarce the resources, i.e. The bigger The smaller the , the more obvious the superiority of the proposed algorithm.

[0091] Figure 7 The number of packets successfully transmitted by the three algorithms in every 100 time slots is counted. The results show that the proposed algorithm outperforms the other two algorithms in terms of convergence speed and transmission efficiency.

[0092] Figure 8 The number of successful transmissions under different algorithms is compared with that under different network loads and buffer capacities. and The number of successful transmissions shows that even with increasing the buffer capacity, the number of successful transmissions of the minimum hop DQN algorithm does not improve. This is because, although increasing the buffer capacity can alleviate congestion to a certain extent, it will also cause data packets to accumulate and reduce transmission efficiency.

[0093] Figure 9 、 Figure 10 、 Figure 11 Statistics on the buffer capacity The impact of different network loads on the node's next hop selection strategy under the condition of . Figure 9 Medium to low load When the network resources are sufficient, according to the design of the reward value formula, all nodes tend to choose the path with the shortest hops. Node 8 becomes a public forwarding node with a high probability because it is located at the intersection of multiple shortest paths. However Figure 10 When the load increases When the node is in a state of high hop count, it needs to balance between hop count optimization and congestion control. At this time, the node will judge the network situation based on the buffer status of neighboring nodes. If the buffer of the neighboring node on the shortest path is heavily occupied, the node will choose to sacrifice some of the hop count optimization goals and actively avoid neighbors with nearly saturated buffers to reduce congestion and packet loss. Figure 11 When the load is high When , data packets generated by different source nodes will be transmitted relatively independently, trying to avoid overlapping of selected paths and avoid congestion.

[0094] The above-described embodiments of the present application do not constitute a limitation on the scope of protection of the present application.

Claims

1. A multi-channel routing anti-interference decision method for a limited buffer zone, characterized in that: The method comprises: Step 1: In a multi-hop wireless network interference scenario, analyze the impact of jammer interference, mutual interference between nodes, channel noise, and path loss on the signal-to-interference ratio (SIR), and calculate the transmission rate between adjacent nodes. Step 2: Calculate the node buffer data volume based on the packet arrival rate, interference conflict, mutual interference conflict, buffer overflow packet loss, and maximum transmission number packet loss; Step 3: Model the joint optimization problem of interference-resistant routing and channel in distributed multi-hop networks as a partially observable stochastic game; nodes maximize long-term efficiency based on local observation information; Step 4: Set up a hierarchical DQN architecture; split the routing anti-interference problem in a multi-channel scenario into upper-layer routing planning and lower-layer channel selection; Step 1: In a multi-hop wireless network interference scenario, analyze the impact of jammer interference, mutual interference between nodes, channel noise, and path loss on the signal-to-interference ratio (SIR). Calculate the transmission rate between adjacent nodes, including: Assume that the data packet needs to pass through from the source node to the destination node Hop transmission, all nodes use the same transmission power , the jammer transmit power is , then Jump Node The received power is expressed as: (1); Among them Jump Node and the previous hop node Channel gain between Use the distance between the two Indicates that ; Jammer and Jump Node Channel gain between Use jammer to distance Indicates that , represents the path fading factor; Indicates the Jump Node The set of neighbor nodes of Indicates other neighbor nodes to the Jump Node Gain; Indicates the Jump Node and the previous hop node The data transmission channel between Indicates the channel where the interference is located, Indicates the Jump Node The transmission channel selected by other neighbors of ; is the channel set; Indicates the transmission channel noise; and The indicator function is expressed as follows: (2); (3); when When , it means that the channel transmitted by the routing node is the same as the interference channel, and the data link is interfered; when When , it means that the transmission channel of the routing node is the same as that of other neighbors, and there is mutual interference in the data link; No. Jump Node The received signal-to-interference ratio is: (4); According to the received signal-to-interference ratio, determine the Jump Node and the previous hop node The transmission rate between them is: (5); in Representation node The received signal-to-interference ratio, Indicates the threshold for successful demodulation by the receiver; when receiving When the signal is below the demodulation threshold, the communication rate is zero.

2. The method according to claim 1, characterized in that Step 2: Calculate the node buffer data volume based on the packet arrival rate, interference collision, mutual interference collision, buffer overflow packet loss, and maximum transmission number packet loss, including: The data packets arriving at the source node obey the mean Poisson process, Source node in time slot arrive The probability of a data packet is: (6); Assume that each packet has the same length as , then the data packet from node Transfer to node The time is: (7); in Representation node and the previous hop node The transmission rate between Therefore, define Time Slot Jump Node Number of packets successfully transmitted for: (8); in is the maximum data transmission time within a unit time slot, At most one data packet can be transmitted within a certain time period; the maximum number of data packets stored in the node buffer is , assuming time slot Initial current node The amount of data in the buffer is , the number of packets received is ;like is the source node, Determined by the probability of the source node packet arriving, otherwise Current node of time slot It is determined by the number of data packets successfully transmitted by the previous hop node; Current node When receiving a data packet, the transmission times are checked first; if the maximum transmission times are reached If the number of packets is reached, the data packet is discarded; if it is not reached, it is stored in the buffer and waits for transmission; Assume that the number of packets lost due to reaching the maximum number of transmissions is ; After the node forwards the data packet, it deletes the data packet from the buffer to release the storage space; therefore, Current node of time slot Buffer data volume for: (9); When the remaining storage space in the buffer is insufficient to store the newly arrived data packets, the buffer overflows and packet loss occurs. The goal of all nodes in the network is to maximize the Total number of packets successfully transmitted during the long run: (10); In a dynamic interference environment, the channel and next hop selection of a node is subject to multiple constraints: first, the node must be in a limited set of channels. and neighbor node set Secondly, the transmission time of each hop link must be less than or equal to the maximum data transmission time within the unit time slot, otherwise packet loss will occur; in addition, the amount of data in the node buffer is always limited by the capacity. constraints, among which Indicates the maximum buffer capacity of the node.

3. The method according to claim 1, characterized in that Step 3: Model the joint optimization problem of interference-resistant routing and channels in distributed multi-hop networks as a partially observable stochastic game. Nodes maximize long-term efficiency based on local observation information, including: Based on the local spectrum sensing results, the buffer status fed back by neighboring nodes and the address of the destination node, the node can autonomously infer the interference pattern and network congestion distribution, and make joint routing and channel decisions. In addition, the routing and channel decisions of each node affect other nodes. Therefore, the interference-resistant routing and channel joint optimization problem of distributed multi-hop networks is modeled as a partially observable random game. The six-tuple description is defined as: Agent Collection( ): All nodes in the network constitute the set of intelligent agents; State Space ( ): The state space is used to describe the complete network information, including the location of nodes and jammers, the source and destination nodes of data packets, the node buffer status, the node transmission channel and the jammer's interference channel, which is expressed as: (11); in represent the location sets of communication nodes and jammers respectively, is the destination node set of the data packet; represents the set of buffer states of all nodes, They represent the node transmission channel and the interference channel set of the jammer respectively; Observation space ( ): The observation space is determined by the interference optional channel, the possible buffer states of all neighboring nodes and the possible destination node addresses of the data packets; the node In the time slot The observation is expressed as: (12); in Received data packets The destination node, Represents the buffer status of all neighbor nodes, abbreviated as , Indicates the jammer's jamming channel; Action Space( ):node The action space is composed of the set of available channels and neighbor node set Joint decision-making; time slot The node needs to select the transmission channel and the next hop node at the same time. ; State transition probability ( ): Due to the interference channel of the jammer and the channel strategy of other nodes, is unknown, so the state transition probability of the environment is also unknown, and the node needs to learn this probability by continuously interacting with the environment; Reward function ( ): The factors affecting the reward function include routing hop cost, congestion avoidance, and interference avoidance. A composite reward function is used to achieve multi-objective collaborative optimization through decoupling design. The composite reward consists of two parts: (13); (14); in represents the reward obtained by the routing decision, Represents the reward obtained by channel decision; Indicates the current node The number of hops to the destination node, Indicates the selected next hop node Number of hops to the destination node; Indicates the selected next hop node Buffer overflow, This means that the selected transmission channel has a low transmission rate due to interference or mutual interference, resulting in the inability to transmit a complete data packet within the unit time slot.

4. The method according to claim 1, wherein Step 4: Set up the hierarchical DQN architecture; The routing anti-interference problem in multi-channel scenarios is divided into upper-layer routing planning and lower-layer channel selection, including: The input of the upper layer network is the buffer status obtained by interacting with neighbors and the address of the destination node carried in the data packet, and the output is the selected next hop node; The input of the lower network is the next hop node selected by the upper network and the channel where the interference is located obtained through spectrum sensing. The output of the network is the selected transmission channel. Both decision networks are composed of two fully connected layers. For nodes definition The state-action-value function at the moment , indicating that the node is in the observation state Next action The maximum long-term cumulative reward value that can be obtained; the state-action value function consists of two parts: (15); in represents the action-value function of the routing, represents the observation space of the routing decision network, Indicates the selected next hop node; represents the action-value function of the channel, represents the observation space of the channel decision network, represents the selected transmission channel; the action value function is updated as follows: (16); in is the learning rate, is the discount factor; Defining Nodes The loss function of the network is: (17); in and represents the weight parameter of the prediction network of the node, and represents the optimization goal of the network, which is defined as: (18); in and represents all possible actions, and is the weight parameter of the target network; The network is trained using the gradient descent method, and the gradient of the loss function is: (19); Use the Boltzmann update strategy to define the routing strategy and channel selection strategy The update formula is: (20); in are the relevant parameters of the Boltzmann model: (21); in represents the Boltzmann initial temperature, Indicates the minimum temperature, Affects the transition time between exploration and exploitation phases.

Citation Information

Patent Citations

  • System-level transmission delay model building method applied to network on chip

    CN102693213A

  • Joint channel and route selection cross-layer decision-making method based on Q learning

    CN119484387A