A relay selection method based on buffer status prediction
By using the relay selection method in LSTM-DQN networks to dynamically adjust the receiving and transmitting strategies of relay nodes, the resource waste caused by fixed allocation of relay user buffers is solved, thereby improving communication quality and efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-08
- Publication Date
- 2026-03-24
AI Technical Summary
In existing relay cooperative communication, the buffer allocation for relay users is fixed and does not take into account the user's own needs. This leads to resource waste when the relay user's own needs are large and the forwarding tasks are small, and the communication quality and efficiency are low.
A relay selection method based on LSTM-DQN network is adopted. By constructing a state space, action space and reward function, and combining deep reinforcement learning and long short-term memory network, the receiving and transmitting strategies of relay nodes are dynamically selected by comprehensively considering the buffer requirements of relay users and channel state.
It improves the efficiency of relay user buffer usage, reduces packet loss rate, increases system capacity and user experience, and adapts to the dynamic changes in relay user buffer requirements.
Smart Images

Figure CN116506918B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of cooperative communication technology, and in particular to a relay selection method based on buffer prediction. Background Technology
[0002] Traditional cellular networks rely on cell-based communication, and the resulting large and small signal fading at cell edges leads to poor signal strength. Furthermore, interference between neighboring cells further exacerbates the signal degradation, significantly increasing power consumption for base stations. Relay technology effectively mitigates these issues. It involves placing one or more relay nodes between the originating and destination nodes. These nodes receive and process signals before transmitting them, shortening the transmission distance and thus effectively reducing fading and path loss. This ensures communication quality, expands the signal range, improves the overall performance of the wireless network, increases throughput, and reduces system energy consumption.
[0003] Cooperative communication improves the throughput of wireless networks and expands the communicable range of signals. However, in the half-duplex mode of traditional cooperative networks, relay nodes cannot simultaneously obtain the optimal receive and transmit channels, resulting in compromised signal quality. Buffered relays have been proposed to effectively address these issues. Compared to traditional relay schemes, buffered relay-assisted communication schemes show significant improvements in system throughput, reduced system outage probability, and signal-to-noise ratio.
[0004] Mobile terminals refer to computer devices that can be used while on the move; in the field of communications, this mostly refers to smart devices. However, terminals acting as relays have limited buffer space, and their users also have their own buffering needs. Currently, most buffer-based cooperative communication relay selection only considers the relay's wholehearted cooperative forwarding, meaning all of the relay's buffers assist in communication. It does not consider the buffering needs of the relay users themselves. Relays allocating fixed buffers for forwarding means that the buffer space available to relay users is also fixed. When a relay user's needs are high while the relay's forwarding tasks are low, the relay user's needs are not met, leaving idle buffers. This leads to a poor user experience and wasted buffer resources. Therefore, considering how to first meet user needs and then improve buffer utilization efficiency is a key issue that needs to be addressed in relay cooperative communication. Summary of the Invention
[0005] To address the shortcomings of fixed allocation of limited relay buffers, the present invention aims to provide a relay selection method based on buffer prediction that comprehensively considers packet loss rate and end-users' buffer requirements, thereby improving buffer utilization efficiency in wireless networks while meeting user needs.
[0006] To achieve the above objectives, the present invention adopts the following technical solution: a relay selection method based on buffer prediction, the method comprising the following sequential steps:
[0007] (1) Set the parameters of the communication environment, i.e., the buffer-assisted relay forwarding system: determine the number of relay nodes, the location coordinates of the relay nodes, the location coordinates of the source node, the location coordinates of the destination node, the size of the total buffer, the channel coefficient, the transmission power, the noise power, and the target data rate;
[0008] (2) Construct an LSTM-DQN network and determine the state space, action space, and reward function;
[0009] (3) The agent selects an action in the action space based on the initial state in the state space, that is, makes a decision on the selection of relay nodes in the communication environment and whether to receive or send data from the relay node, and obtains the next state. The above process is repeated continuously, and finally the maximum reward value is obtained, that is, the link capacity is maximized.
[0010] Step (1) specifically refers to the following: the buffer-assisted relay forwarding system consists of a source node S, a destination node D, and a relay node. The system consists of three nodes, 1 ≤ k ≤ K, where k is the number of relay nodes. The relay nodes, source node, and destination node are located within a 100m × 100m area. The relay nodes are composed of end users. It is assumed that each node has an antenna and operates in half-duplex mode. There is no direct link between the source node and the destination node, and communication is completed through relay forwarding. It is assumed that time is divided into time slots of equal duration. In each time slot, the source node S sends a data packet with a fixed power P. The buffer size of each relay node is limited, and the total buffer size is L+1, including the buffer needs of the relay users themselves and the buffer size for assisting forwarding. L is the size of the buffer area used for assisting forwarding. It is assumed that the buffer needs of the relay users themselves occupy at least the size of one data packet. Therefore, in each time slot, the buffer size used for assisting forwarding is at most L.
[0011] Assuming each user's buffer requirement is Lu, the size of the buffer used for forwarding assistance is L+1-Lu; Represents relay node The number of data packets stored in the buffer, 0≤ ≤L, in each time slot, for different Value, relay node The available number of links is also different:
[0012] (1a) = 0: No data packet is sent, and only the source node-relay node link, i.e., the S-R link, is available;
[0013] (1b) 0 < < L + 1 - Lu: Both the source node-relay node link, i.e., the S-R link, and the relay node-destination node link, i.e., the R-D link, are available;
[0014] (1c) = L + 1 - Lu: Only the relay node-destination node link, i.e., the R-D link, is available, and there is no buffer for storing new data packets;
[0015] First, judge according to the past relay channel state and the historical cache requirements of the terminal relay users. If the buffer of this relay can store data packets, then select this relay to send data packets; when the k-th relay node receives the data packet sent by the source node S, the corresponding buffer is occupied by the size of one data packet. When the k-th relay node successfully sends the data packet to the destination node D, the corresponding buffer size is reduced by one data packet; only after the relay node successfully receives the data packet can this relay send the data packet to the destination node D; assume that the source node S always has the task of sending data packets to the destination node D, the channel coefficient follows the Rayleigh distribution, the channel coefficient remains unchanged within one time slot, and varies independently in different time slots. Assume that the signal finally received by the destination node D is affected by additive white Gaussian noise with a mean of zero and a variance of ;
[0016] In a certain time slot, when the selected link is from the source node S to the relay node R, a single data packet is sent from the source node to the corresponding relay of S and stored in the buffer. The received signal at is:
[0017]
[0018] where is the data signal from S, is the additive white Gaussian noise with a variance of , P is the transmission power, is the channel coefficient from the source node to the relay node, is the distance from the source node to the relay node, and α is the path loss exponent; if the relay-to-destination link is selected, then a data packet is sent from the relay buffer to the destination, and the received signal at the destination is:
[0019]
[0020] in, It comes from Data signals, Indicates the prescription difference of the destination node D. Additive white Gaussian noise, It is the channel coefficient from the relay node to the destination node. It is the distance from the relay node to the destination node; the link capacity between node m and node n. for:
[0021]
[0022] In the formula, Let be the channel coefficients from node m to node n. Let m be the distance from node m to node n. This represents the power of additive white Gaussian noise.
[0023] when When the value is less than or equal to η, the corresponding link is interrupted, where η is the target data rate.
[0024] The specific step (2) refers to: adding an LSTM network to the deep reinforcement learning network DQN to form an LSTM-DQN network, inputting data of L consecutive time steps into the LSTM network, which is composed of multiple LSTM units. The LSTM contains three gates, namely the input gate, the forget gate and the output gate.
[0025] The state space, action space, and reward value of the LSTM-DQN network are as follows:
[0026] State space: At time t, the observed state is = [ , , ],in This indicates the user buffer usage at time t−1. It is the channel coefficient from the source node to the relay node. These are the channel coefficients from the relay node to the destination node, and the state space is defined as S1 = [ ,..., ],in, This represents the number of past observed states to be captured;
[0027] Action space: Based on the current finite and changing buffer-assisted relay forwarding system state This requires decisions regarding the selection of a relay and whether that relay should receive or transmit data. The environment is a buffer-assisted relay selection network, and the action is to select a link for data transmission, which is equivalent to determining... , j∈{0,1}, where k represents the number of relay nodes, 0 represents relay receiving data packets, and 1 represents relay sending data packets; if a relay network has k relay nodes, then there are 2k transmission links. In a time slot, one link is selected for transmission, or no link is selected. Therefore, the total number of actions is 2k + 1.
[0028] Reward function: The reward is related to the optimization objective function, and throughput is used as the reward function.
[0029] Step (3) specifically includes the following steps:
[0030] (3a) In a deep reinforcement learning network (DQN), the learner and decision-maker are called agents, and the part that interacts with the agents is called the environment. Assume that in time slot t, the state of the environment is... Based on the current state, the agent decides its next action: which link to choose or not to choose a link for data transmission, using an ε-greedy strategy to determine the state. The action is defined by ε∈(0,1), where ε is the greedy coefficient and n is the number of training iterations. ε is initially set to 1 to obtain good exploration results and gradually decreases with the number of iterations.
[0031] (3b) Once the agent has chosen an action That is, selecting a relay node and determining whether that relay node receives or sends, thereby obtaining a reward value and the next state. If an S→R or R→D link selection occurs, the corresponding buffer length is increased or decreased by 1, respectively; otherwise, the buffer length remains unchanged. On the other hand, if the channel state changes independently from one time slot to another, the state is then transitioned to a new state based on the new buffer length and channel state. ;
[0032] (3c) The current state, the action performed, the reward value obtained after performing the action, and the state at the next moment are combined into a tuple, which is ( , , , ), stored in the experience pool;
[0033] (3d) Return to step (3a), using the state Repeat this process and generate another set of tuples until the state value reaches the termination state and the reward value reaches its maximum value.
[0034] As can be seen from the above technical solution, the beneficial effects of the present invention are as follows: First, when the relay user's own buffer requirement is small, the relay can allocate more buffers to assist in relay forwarding, which can reduce the packet loss rate. When the relay user's own buffer requirement is large, the buffer that the relay can allocate to assist in relay forwarding is quite limited. Reinforcement learning will comprehensively consider the channel state and the relay's historical buffer requirement to select an appropriate link for data packet transmission. Second, compared with the existing relay selection method based on fixed buffers, the present invention adds an LSTM network to the deep reinforcement learning network DQN, making reinforcement learning more suitable for the scenario where the available buffer size for cooperative communication by end users changes. The state is based on the historical user's buffer requirement, the channel state between the source node and the relay node, and the channel state between the relay node and the destination node. Third, it establishes an application scenario where the available buffer for cooperative communication is limited and changes due to the end user's own buffer requirement. When the relay user's own buffer requirement is small, the relay can allocate more buffers to assist in relay forwarding and realize the selection of relay nodes to send and receive data packets. Compared with the prior art, the average available buffer for users is increased, the packet loss rate is reduced, and the system capacity is improved. Attached Figure Description
[0035] Figure 1 This is a flowchart of the method of the present invention;
[0036] Figure 2 This is a schematic diagram of the buffer-assisted relay forwarding system in this invention;
[0037] Figure 3 This is a schematic diagram of an LSTM network;
[0038] Figure 4 This is a schematic diagram of the LSTM unit structure;
[0039] Figure 5 This is a flowchart of the processing flow of an LSTM-DQN network.
[0040] Figure 6 This is a structural diagram of the master network and the destination network in an LSTM-DQN network. Detailed Implementation
[0041] like Figure 1 As shown, a relay selection method based on buffer prediction includes the following sequential steps:
[0042] (1) Set the parameters of the communication environment, i.e., the buffer-assisted relay forwarding system: determine the number of relay nodes, the location coordinates of the relay nodes, the location coordinates of the source node, the location coordinates of the destination node, the size of the total buffer, the channel coefficient, the transmission power, the noise power, and the target data rate;
[0043] (2)Construct the LSTM-DQN network, and determine the state space, action space, and reward function;
[0044] (3)The agent selects an action in the action space according to the initial state in the state space, that is, makes a decision on the selection of the relay node in the communication environment and the reception or transmission of that relay node, obtains the next state, continuously repeats the above process, and finally obtains the maximum reward value, that is, the maximum link capacity.
[0045] The specific content of step (1) is as follows: The buffer-aided relay forwarding system consists of a source node S, a destination node D, and relay nodes composed of, 1 ≤ k ≤ K, where k is the number of relay nodes. The relay nodes, source node, and destination node are located in a 100m × 100m area. The relay nodes are composed of end users. Assume that each node has one antenna and operates in a half-duplex mode. There is no direct link between the source node and the destination node, and communication needs to be completed through relay forwarding. Assume that time is divided into equal-length time slots. In each time slot, the source node S sends a data packet with a fixed power P. The buffer size of each relay node is limited. The total buffer size is L + 1, including the buffer requirements of the relay user itself and the buffer size for assisting forwarding. L is the buffer size used for assisting forwarding. Assume that the buffer requirements of the relay user itself need to occupy at least the size of one data packet. Therefore, in each time slot, the buffer size used for assisting forwarding is at most L;
[0046] Assume that the buffer requirement of each user is Lu. At this time, the buffer size used for assisting forwarding is L + 1 - Lu; use to represent the number of data packets stored in the buffer of relay node , 0 ≤ ≤ L. In each time slot, for different values, the available link numbers of relay node are also different:
[0047] (1a) = 0: No data packet is sent, and only the source node-relay node link, i.e., the S-R link, is available;
[0048] (1b)0 < < L + 1 - Lu: Both the source node-relay node link, i.e., the S-R link, and the relay node-destination node link, i.e., the R-D link, are available;
[0049] (1c) = L + 1 - Lu: Only the relay node-destination node link, i.e., the R-D link, is available, and there is no buffer for storing new data packets;
[0050] First, based on the past relay channel status and the historical buffer requirements of terminal relay users, if the relay's buffer can store data packets, then that relay is selected to send data packets. When the k-th relay node receives a data packet sent by the source node S, the corresponding buffer is occupied by the size of one data packet. When the k-th relay node successfully sends a data packet to the destination node D, the corresponding buffer is reduced by the size of one data packet. Only after a relay node successfully receives a data packet can it send a data packet to the destination node D. Assume that the source node S always has the task of sending data packets to the destination node D, the channel coefficient follows a Rayleigh distribution, remains constant within a time slot, and varies independently in different time slots, and assume that the signal ultimately received by the destination node D has a mean of zero and a variance of... The effect of additive white Gaussian noise;
[0051] In a certain time slot, when the selected link is from source node S to relay node R, the data flows from the source node to the corresponding relay node S. Send a single data packet and store it in a buffer. Received signal at the location for:
[0052]
[0053] in, It is a data signal from S. The variance is Additive white Gaussian noise, where P is the transmitted power. It is the channel coefficient from the source node to the relay node. α is the distance from the source node to the relay node, and α is the path loss exponent. If a relay link to the destination is selected, a data packet is sent from the relay buffer to the destination, and the received signal is given at the destination. for:
[0054]
[0055] in, It comes from Data signals, Indicates the prescription difference of the destination node D. Additive white Gaussian noise, It is the channel coefficient from the relay node to the destination node. It is the distance from the relay node to the destination node; the link capacity between node m and node n. for:
[0056]
[0057] In the formula, Let be the channel coefficients from node m to node n. Let m be the distance from node m to node n. This represents the power of additive white Gaussian noise.
[0058] when When the value is less than or equal to η, the corresponding link is interrupted, where η is the target data rate.
[0059] The specific step (2) refers to: adding an LSTM network to the deep reinforcement learning network DQN to form an LSTM-DQN network, inputting data of L consecutive time steps into the LSTM network, which is composed of multiple LSTM units. The LSTM contains three gates, namely the input gate, the forget gate and the output gate.
[0060] The state space, action space, and reward value of the LSTM-DQN network are as follows:
[0061] State space: At time t, the observed state is = [ , , ],in This indicates the user buffer usage at time t−1. It is the channel coefficient from the source node to the relay node. These are the channel coefficients from the relay node to the destination node, and the state space is defined as S1 = [ ,..., ],in, This represents the number of past observed states to be captured;
[0062] Action space: Based on the current finite and changing buffer-assisted relay forwarding system state This requires decisions regarding the selection of a relay and whether that relay should receive or transmit data. The environment is a buffer-assisted relay selection network, and the action is to select a link for data transmission, which is equivalent to determining... , j∈{0,1}, where k represents the number of relay nodes, 0 represents relay receiving data packets, and 1 represents relay sending data packets; if a relay network has k relay nodes, then there are 2k transmission links. In a time slot, one link is selected for transmission, or no link is selected. Therefore, the total number of actions is 2k + 1.
[0063] Reward function: The reward is related to the optimization objective function, and throughput is used as the reward function.
[0064] Step (3) specifically includes the following steps:
[0065] (3a) In a deep reinforcement learning network (DQN), the learner and decision-maker are called agents, and the part that interacts with the agents is called the environment. Assume that in time slot t, the state of the environment is... Based on the current state, the agent decides its next action: which link to choose or not to choose a link for data transmission, using an ε-greedy strategy to determine the state. The action is defined by ε∈(0,1), where ε is the greedy coefficient and n is the number of training iterations. ε is initially set to 1 to obtain good exploration results and gradually decreases with the number of iterations.
[0066] (3b) Once the agent has chosen an action That is, selecting a relay node and determining whether that relay node receives or sends, thereby obtaining a reward value and the next state. If an S→R or R→D link selection occurs, the corresponding buffer length is increased or decreased by 1, respectively; otherwise, the buffer length remains unchanged. On the other hand, if the channel state changes independently from one time slot to another, the state is then transitioned to a new state based on the new buffer length and channel state. ;
[0067] (3c) The current state, the action performed, the reward value obtained after performing the action, and the state at the next moment are combined into a tuple, which is ( , , , ), stored in the experience pool;
[0068] (3d) Return to step (3a), using the state Repeat this process and generate another set of tuples until the state value reaches the termination state and the reward value reaches its maximum value.
[0069] The key idea of the LSTM-DQN framework proposed in this invention is to ensure effective relay forwarding while maintaining partial state observations due to the relay user's own buffering needs. To achieve this, an LSTM network is added to DQN, which not only preserves the internal state but also aggregates state observations over time. This gives the relay-assisted communication network the ability to infer future states by processing history. Specifically, data of L consecutive time steps is input into the LSTM network, which consists of multiple LSTM units. Generally, LSTM contains three gates: an input gate, a forget gate, and an output gate. The key to LSTM's superiority over RNNs lies in the line running through the unit in the diagram above—the hidden state of the neuron (unit state). The hidden state of a neuron can be simply understood as the "memory" of the input data by the recurrent neural network. This vector represents the "memory" of a neuron after time t. It encompasses the neural network's "summary" of all input information up to time t+1. The forget gate's task is to determine which long-term memory to retain and which to forget. Which part? The function of the memory gate is to determine what new information is stored in the cell state. Finally, based on the cell state, the output value is determined.
[0070] like Figure 2 As shown, the proposed buffer-assisted relay forwarding system consists of a source node S, a destination node D, and k relay nodes. The composition is 1≤k≤K. The relay nodes considered here are composed of end users, whose caches are limited and they also have their own cache requirements.
[0071] Figure 3 This demonstrates an expanded LSTM network. Specifically, it shows how data of L consecutive time steps are input into an LSTM network, which consists of multiple LSTM units, such as... Figure 4 As shown.
[0072] Figure 5 and Figure 6 An LSTM-DQN framework for relay selection in a finite and variable buffer-assisted forwarding environment is presented. The key idea of the proposed LSTM-DQN framework is to ensure efficient relay forwarding while maintaining partial state observations caused by relay users' own buffer requirements.
[0073] In summary, when relay users have limited buffer requirements, relays can allocate more buffers to assist in relay forwarding, reducing packet loss. Conversely, when relay users have large buffer requirements, the available buffers for relay forwarding are quite limited. Reinforcement learning will comprehensively consider channel conditions and historical relay buffer requirements to select appropriate links for data packet transmission. This invention incorporates an LSTM network into the Deep Reinforcement Learning (DQN) network, making reinforcement learning more suitable for scenarios where the available buffer size for collaborative communication by end users changes. It uses historical user buffer requirements, source node-relay node channel conditions, and relay node-destination node channel conditions as states. It establishes an application scenario where end users' buffer requirements lead to limited and changing available buffers for collaborative communication. When relay users have limited buffer requirements, relays can allocate more buffers to assist in relay forwarding and enable relay nodes to select which data packets to send and receive. Compared to existing technologies, this increases the average available buffer for users, reduces packet loss, and improves system capacity.
Claims
1. A relay selection method based on buffer prediction, characterized in that: The method includes the following steps in sequence: (1) Set the parameters of the communication environment, i.e., the buffer-assisted relay forwarding system: determine the number of relay nodes, the location coordinates of the relay nodes, the location coordinates of the source node, the location coordinates of the destination node, the size of the total buffer, the channel coefficient, the transmission power, the noise power, and the target data rate; (2) Construct an LSTM-DQN network and determine the state space, action space, and reward function; (3) The agent selects an action in the action space based on the initial state in the state space, that is, makes a decision on the selection of relay nodes in the communication environment and whether the relay node receives or sends, and obtains the next state. The above process is repeated continuously, and finally the maximum reward value is obtained, that is, the link capacity is maximized. Each relay node has a limited buffer size, and the total buffer size is L+1, which includes the relay user's own buffer needs and the buffer size for assisting forwarding. The specific step (2) refers to: adding an LSTM network to the deep reinforcement learning network DQN to form an LSTM-DQN network, inputting data of L consecutive time steps into the LSTM network, which is composed of multiple LSTM units. The LSTM contains three gates, namely the input gate, the forget gate and the output gate. The state space, action space, and reward value of the LSTM-DQN network are as follows: State space: At time t, the observed state is = [ , , ],in This indicates the user buffer usage at time t−1. It is the channel coefficient from the source node to the relay node. These are the channel coefficients from the relay node to the destination node, and the state space is defined as S1 = [ ,..., ],in, This represents the number of past observed states to be captured; Action space: Based on the current finite and changing buffer-assisted relay forwarding system state This requires decisions regarding the selection of a relay and whether that relay should receive or transmit data. The environment is a buffer-assisted relay selection network, and the action is to select a link for data transmission, which is equivalent to determining... , j∈{0,1}, where k represents the number of relay nodes, 0 represents relay receiving data packets, and 1 represents relay sending data packets; if a relay network has k relay nodes, then there are 2k transmission links. In a time slot, one link is selected for transmission, or no link is selected. Therefore, the total number of actions is 2k + 1. Reward function: The reward is related to the optimization objective function, and throughput is used as the reward function.
2. The relay selection method based on buffer prediction according to claim 1, characterized in that: Step (1) specifically refers to the following: the buffer-assisted relay forwarding system consists of a source node S, a destination node D, and a relay node. The system consists of nodes, where 1 ≤ k ≤ K, and k is the number of relay nodes. The relay nodes, source node, and destination node are located within a 100m × 100m area. The relay nodes are composed of end users. It is assumed that each node has an antenna and operates in half-duplex mode. There is no direct link between the source node and the destination node, and communication is completed through relay forwarding. It is assumed that time is divided into time slots of equal duration. In each time slot, the source node S sends a data packet with a fixed power P. It is assumed that the relay user's own buffer needs to occupy at least the size of one data packet. Therefore, in each time slot, the buffer size used for assisting forwarding is at most L. Assuming each user's buffer requirement is Lu, the size of the buffer used for forwarding assistance is L+1-Lu; Represents relay node The number of data packets stored in the buffer, 0≤ ≤L, in each time slot, for different Value, relay node The number of available links also differs: (1a) =0: No data packets are being sent; only the source node-relay node link (SR link) is available. (1b) 0 < <Both the source node-relay node link, i.e., the S-R link, and the relay node-destination node link, i.e., the R-D link, can be used; (1c) = L+1-Lu: Only the relay node-destination node link, i.e., the RD link, is available; there is no buffer for storing new data packets. First, based on the past relay channel status and the historical buffer requirements of terminal relay users, if the relay's buffer can store data packets, then that relay is selected to send data packets. When the k-th relay node receives a data packet sent by the source node S, the corresponding buffer is occupied by the size of one data packet. When the k-th relay node successfully sends a data packet to the destination node D, the corresponding buffer is reduced by the size of one data packet. Only after a relay node successfully receives a data packet can it send a data packet to the destination node D. Assume that the source node S always has the task of sending data packets to the destination node D, the channel coefficient follows a Rayleigh distribution, remains constant within a time slot, and varies independently in different time slots, and assume that the signal ultimately received by the destination node D has a mean of zero and a variance of... The effect of additive white Gaussian noise; In a certain time slot, when the selected link is from source node S to relay node R, the data flows from the source node to the corresponding relay node S. Send a single data packet and store it in a buffer. Received signal at the location for: , in, It is a data signal from S. The variance is Additive white Gaussian noise, where P is the transmitted power. It is the channel coefficient from the source node to the relay node. α is the distance from the source node to the relay node, and α is the path loss exponent. If a relay link to the destination is selected, a data packet is sent from the relay buffer to the destination, and the received signal is given at the destination. for: , in, It comes from Data signals, Indicates the prescription difference of the destination node D. Additive white Gaussian noise, It is the channel coefficient from the relay node to the destination node. It is the distance from the relay node to the destination node; the link capacity between node m and node n. for: , In the formula, Let be the channel coefficients from node m to node n. Let m be the distance from node m to node n. This represents the power of additive white Gaussian noise. when When the value is less than or equal to η, the corresponding link is interrupted, where η is the target data rate.
3. The relay selection method based on buffer prediction according to claim 1, characterized in that: Step (3) specifically includes the following steps: (3a) In a deep reinforcement learning network (DQN), the learner and decision-maker are called agents, and the part that interacts with the agents is called the environment. Assume that in time slot t, the state of the environment is... Based on the current state, the agent decides its next action: which link to choose or not to choose a link for data transmission, using an ε-greedy strategy to determine the state. The action is defined by ε∈(0,1), where ε is the greedy coefficient and n is the number of training iterations. ε is initially set to 1 to obtain good exploration results and gradually decreases with the number of iterations. (3b) Once the agent has chosen an action That is, selecting a relay node and determining whether that relay node receives or sends, thereby obtaining a reward value and the next state. If an S→R or R→D link selection occurs, the corresponding buffer length is increased or decreased by 1, respectively; otherwise, the buffer length remains unchanged. On the other hand, if the channel state changes independently from one time slot to another, the state is then transitioned to a new state based on the new buffer length and channel state. S is the source node, D is the destination node, and R is the relay node; (3c) The current state, the action performed, the reward value obtained after performing the action, and the state at the next moment are combined into a tuple, which is ( , , , ), stored in the experience pool; (3d) Return to step (3a), using the state Repeat this process and generate another set of tuples until the state value reaches the termination state and the reward value reaches its maximum value.