An Enhanced ALOHA Access Method Based on Deep Reinforcement Learning Algorithm
Through deep reinforcement learning algorithms, the ALOHA access method is adjusted, and the transmission strategy and frame time slot selection is optimized in real time, which solves the problems of communication blockage and fairness during access of massive devices, and improves communication throughput and fairness.
Patent Information
- Application Number
- CN202210438011.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-20
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-04-20
AI Technical Summary
When facing massive equipment access, existing random access solutions are prone to problems such as access overload, scheduling signaling congestion, communication blockage and user communication fairness, especially when the network load is heavy, the user wait time is too long and the throughput is low.
The enhanced ALOHA access method based on deep reinforcement learning algorithm is adopted, and the transmission strategy is adjusted in real time, the collision data is recovered using the DPSA access method, and the users are classified according to the waiting time, and the frame time slot transmission probability is dynamically selected, combined with the reinforcement learning algorithm to determine the optimal transmission probability, optimize communication throughput and fairness.
It effectively improves communication throughput capabilities, reduces the user's waiting time gap, solves the problems of communication blockage and fairness, and realizes the trade-off optimization of user's waiting time performance and throughput.
Smart Images

Figure CN115087131B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an enhanced ALOHA access method based on a deep reinforcement learning algorithm, belonging to the field of wireless communication technology. Background Art
[0002] With the rapid development of mobile communication networks, especially for the application scenarios facing M2M (Machine-to-Machine), the access requirements of a large number of devices need to be met. When a large number of service terminals send access requests to the base station and complete device access in a scheduling manner, due to the excessive number of devices, it is extremely easy to cause access overload and scheduling signaling congestion. This will lead to excessive service access delay and extremely high scheduling overhead, resulting in low resource utilization. In order to reduce the scheduling overhead, random access has become a research hotspot.
[0003] As a classic random access method, ALOHA allows users to complete access by sending data immediately when it arrives, greatly reducing signaling overhead. However, when the network load is heavy, nodes are extremely likely to generate conflicts, resulting in a network throughput of only 0.18. At the same time, some users may not be able to send successfully for a long time, resulting in fairness issues in the communication process. Slotted ALOHA (SA) is an improvement of ALOHA. By dividing equal time slices, each time slice corresponds to a time slot, and this limit reduces the probability of conflicts within the time slice. The channel utilization rate of slotted ALOHA is at most 0.368. Diversity Slotted ALOHA (DSA) is based on SA and adopts a time diversity method. The data packets of each user are divided into several parts and sent repeatedly at different times to improve throughput. However, the above solutions do not consider the index of user waiting delay. In the sending process, users send randomly, which will cause some users to have too long waiting times, resulting in too high AOI. At the same time, the sending strategy cannot be optimized in real time to keep the sending efficiency optimal all the time, resulting in communication congestion problems caused by data conflicts and user communication fairness problems. Summary of the Invention
[0004] Aiming at the existing random access solutions, there are problems such as the sending strategy cannot adapt to the complex changes in the number of access users, resulting in communication congestion problems caused by data conflicts and user communication fairness problems. The main purpose of the present invention is to provide an enhanced ALOHA access method based on a deep reinforcement learning algorithm, which can improve the communication throughput capacity by intelligently changing the sending strategy in real time and provide a gain for communication fairness.
[0005] The object of the present invention is achieved through the following solutions:
[0006] An enhanced ALOHA access method based on a deep reinforcement learning algorithm disclosed by the present invention adopts reinforcement learning to adapt to complex changes in access situations during exploration learning, modulates random access strategies in real time, and has better adaptability. By adopting the DPSA access method, it recovers conflict data, determines the optimal transmission probability for each frame based on the reinforcement learning algorithm, and improves the communication throughput capacity under the condition of low signaling overhead. Classifies users according to the length of waiting time and adopts different transmission strategies for them, reduces the overall AoI gap of users, and solves the problem of communication fairness. It enables the system to adapt to complex changes in user activation status in real time and intelligently, selects the appropriate frame time slot transmission probability, and improves communication throughput and communication fairness.
[0007] An enhanced ALOHA access method based on a deep reinforcement learning algorithm disclosed by the present invention includes the following steps:
[0008] Step 1: Cell users are randomly activated with a predetermined activation probability, and the users send data in each frame according to the enhanced ALOHA access strategy. Decoding is performed at the receiving end to obtain the communication throughput capacity of the current frame and the AoI of the user with the longest waiting time, and a decoded successful user set and a decoded failed user set are respectively constructed. The specific steps are as follows:
[0009] Step 1.1: Cell users are activated with a certain probability, and the activated users form an activated user set;
[0010] Step 1.2: Arrange the AoI in descending order, mark the users in the front part of the sequence as users with too long waiting time, and the rest are marked as users without too long waiting time;
[0011] For users without too long waiting time, with a certain reference frame time slot transmission probability, randomly decide whether to send a data packet in each of the M time slots of the current frame; for users with too long waiting time, under the above reference frame time slot transmission probability, increase the probability and randomly decide whether to send a data packet in each of the M time slots of the current frame. If there are multiple users sending data packets simultaneously in the same time slot and a conflict occurs, the data packets of each user are superimposed and sent;
[0012] Step 1.4: The receiving end decodes the data according to the interference cancellation technology. The users whose data in the current frame is decoded successfully form a decoded successful user set, and the users whose data in the current frame is decoded failed form a decoded failed user set;
[0013] For the decoded successful users, the AoI in the corresponding state is cleared, and for the decoded failed users, the AoI in the corresponding state is increased;
[0014] Step 1.6: Calculate the current frame communication throughput capacity and the AoI of the user with the longest waiting time. The current frame communication throughput capacity is the number of successfully decoded users divided by the number of active users in the current frame, and the AoI of the user with the longest waiting time is the maximum value of the user AoI in the state.
[0015] Step 2: Build a reinforcement learning framework and construct a deep Q network, i.e., Deep Q Network (DQN). The specific steps are as follows:
[0016] Step 2.1: Construct the state space and action space. The state space consists of the current active states of each user and the AoI of each user, and the action space consists of the base transmission probabilities for different frame time slots. Determine the state s as an element in the state space and the action a as an element in the action space.
[0017] Step 2.2: Given a stepped reward according to the current frame communication throughput capacity and the AoI of the user with the longest waiting time.
[0018] Step 2.3: Construct the target Q network and the actual Q network. They have the same structure and adopt a fully connected neural network structure.
[0019] Step 2.4: Use the mean squared error loss as the loss function.
[0020] Step 3: The agent explores in the reinforcement learning framework for the first X frames. By introducing random actions, the experience replay pool is expanded to provide learning data for the agent. The specific steps are as follows:
[0021] Step 3.1: The agent performs greedy learning with a greedy rate of ε and selects the action a according to the greedy algorithm.
[0022] Step 3.2: Use the action a as the frame time slot transmission probability to perform DPSA random access, obtaining the current frame communication throughput capacity, the AoI of the user with the longest waiting time, and the state s′ of the next frame.
[0023] Step 3.3: Determine the reward r of the current frame according to the current frame communication throughput capacity and the AoI of the user with the longest waiting time.
[0024] Step 3.4: Store the current frame state s, the current frame action a, the current frame reward r, and the next frame state s′ in the experience replay pool.
[0025] Step 4: After X frames, randomly sample samples from the experience replay pool for learning, train the neural network to fit the access strategy, and continue to send data. Set the batch size to B, the learning rate to α, and the decay factor to γ. The specific steps are as follows:
[0026] Step 4.1: Randomly sample B samples, input the next frame state s′ among them into the target Q network, and output q n;
[0027] Step 4.2: Based on the reward r, decay factor γ, and q in the sample n obtain q t = r + γ(max(q n ));
[0028] Step 4.3: Input the current frame state s in the above sample into the actual Q-network, and output q e ;
[0029] Step 4.4: Set the loss function as the mean squared error between q e and q t ;
[0030] Step 4.5: Train the actual Q-network according to the gradient descent method, set the learning rate as α, and the batch size as B;
[0031] Step 4.6: Every Y frames, copy its weights from the actual Q-network to update the weights of the target Q-network;
[0032] Step 4.7: After X frames, the agent inputs the current state s into the actual Q-network, obtains the Q-values for each action, selects the one with the largest Q-value as the action a for the current frame, sends data with the transmission probability of action a for the current frame, and obtains the maximum AoI of the user and the communication throughput performance under this frame transmission probability. And after a certain number of frames, the average maximum AoI of the user decreases, and the communication throughput performance is significantly improved.
[0033] Beneficial Effects
[0034] 1. An enhanced ALOHA access method based on a deep reinforcement learning algorithm disclosed by the present invention adopts reinforcement learning, adapts to complex changes in the access situation during exploration and learning, modulates the random access strategy in real time, dynamically selects the frame transmission probability, effectively improves the average communication throughput, and avoids the communication congestion problem existing in the traditional scheme.
[0035] 2. An enhanced ALOHA access method based on a deep reinforcement learning algorithm disclosed by the present invention adopts the DPSA access method to recover the conflict data, determines the optimal transmission probability for each frame based on the reinforcement learning algorithm, and improves the communication throughput capacity under the condition of low signaling overhead.
[0036] 3. An enhanced ALOHA access method based on a deep reinforcement learning algorithm disclosed by the present invention classifies users according to the length of the waiting time, increases the transmission probability of users with too large AoI of the waiting time, reduces the overall AoI gap of users, and solves the communication fairness problem. Description of the Drawings
[0037] Figure 1Schematic diagram of the DPSA random access frame structure for an enhanced ALOHA access method disclosed in the present invention;
[0038] Figure 2 Schematic diagram of the DQN process for an enhanced ALOHA access method disclosed in the present invention based on a deep reinforcement learning algorithm;
[0039] Figure 3 Schematic diagram of the simulation of the communication throughput capacity varying with the number of frames in an enhanced ALOHA access method disclosed in the present invention based on a deep reinforcement learning algorithm;
[0040] Figure 4 Schematic diagram of the simulation of the communication fairness performance in an enhanced ALOHA access method disclosed in the present invention based on a deep reinforcement learning algorithm. Detailed implementation manners
[0041] The present invention will be described in detail below in conjunction with the drawings and embodiments. At the same time, the technical problems solved by the technical solution of the present invention and the beneficial effects are also described. It should be noted that the described embodiments are only for facilitating the understanding of the present invention and do not limit it in any way.
[0042] Embodiment 1
[0043] The scenario of the embodiment is 200 users in a cellular cell, the user equipment is a mobile phone, the users adopt a random access method, each user is randomly activated with a certain probability, and they send in the frame time slots synchronized in the whole cell. Each frame consists of 100 uniform time slots.
[0044] The cell user access scheme adopts an enhanced ALOHA access method disclosed in the present invention based on a deep reinforcement learning algorithm. The specific operation process is as follows:
[0045] Step 1: The cell users are randomly activated with a certain activation probability and send according to the access strategy of the enhanced ALOHA. The specific steps are as follows:
[0046] Step 1.A: The cell users are activated with a certain probability, and the activated users form an activated user set;
[0047] Step 1.B: Arrange in descending order of AoI. The top 10% of the users are marked as users with too long waiting time, and the rest are marked as users without too long waiting time;
[0048] Step 1.C: For users with non-excessive waiting time, the probability of sending a data packet in a certain reference frame time slot is randomly determined for each of the 100 time slots in the current frame. For users with excessive waiting time, on the basis of the above reference frame time slot sending probability, the probability is increased by 10%, and it is randomly determined whether to send a data packet for each of the 100 time slots in the current frame. If there are multiple users sending data packets simultaneously in the same time slot and a conflict occurs, the data packets of each user are superimposed and sent;
[0049] Step 1.D: The receiving end decodes the data according to the interference cancellation technology. The users whose data in the current frame is decoded successfully form the decoded successful user set, and the users whose data in the current frame is decoded failed form the decoded failed user set;
[0050] Step 1.E: The AoI in the state corresponding to the decoded successful user is cleared to zero, and the AoI in the state corresponding to the decoded failed user is incremented by 1;
[0051] Step 1.F: Calculate the communication throughput capacity of the current frame and the AoI of the user with the longest waiting time. Among them, the communication throughput capacity of the current frame is the number of decoded successful users divided by the number of active users in the current frame, and the AoI of the user with the longest waiting time is the maximum value of the user AoI in the state;
[0052] To sum up, users send data in each frame according to the enhanced ALOHA access strategy, decode at the receiving end, obtain the communication throughput capacity of the current frame and the AoI of the user with the longest waiting time, and respectively construct the decoded successful user set and the decoded failed user set.
[0053] Step 2: Build a reinforcement learning framework and design a Deep Q Network (DQN). The specific steps are as follows:
[0054] Step 2.A: Construct the state space and the action space. The state space consists of the current active states of each user and the AoI of each user. The action space consists of the reference transmission probabilities in different frame time slots. Determine that state a is an element in the state space and action a is an element in the action space. Among them, the action space is {0.03, 0.05, 0.07, 0.09, 0.13, 0.15, 0.17, 0.19, 0.21, 0.23, 0.25, 0.27, 0.29};
[0055] Step 2.B: Give a stepped reward according to the communication throughput capacity of the current frame and the AoI of the user with the longest waiting time;
[0056] Step 2.C: Construct a target Q network and an actual Q network. The two have the same structure and adopt a fully connected neural network structure;
[0057] Step 2.D: Use the mean square error loss as the loss function.
[0058] Step 3: The system conducts exploration under the reinforcement learning framework. The specific steps are as follows:
[0059] Step 3.A: In the first 2000 frames, the system performs greedy learning with a greedy rate of 0.8 and selects action a according to the greedy algorithm.
[0060] Step 3.B: Using action a as the frame time slot transmission probability, perform DPSA random access to obtain the current frame communication throughput capacity, the AoI of the user with the longest waiting time, and the state s′ of the next frame.
[0061] Step 3.C: Determine the reward r of the current frame based on the current frame communication throughput capacity and the AoI of the user with the longest waiting time.
[0062] Step 3.D: Store the current frame state s, the current frame action a, the current frame reward r, and the next frame state s′ in the experience replay pool.
[0063] Step 4: After 2000 frames, randomly sample samples from the experience replay pool for learning and continue to send data. Set the batch size to 64, the learning rate to 0.01, and the decay factor to 0.8. The specific steps are as follows:
[0064] Step 4.A: Randomly sample 64 samples and input the next frame state s′ among them into the target Q-network to output q n ;
[0065] Step 4.B: Based on the reward r, the decay factor γ, and q in the sample n obtain q t = r + 0.8(max(q n ));
[0066] Step 4.C: Input the next frame state s′ in the above sample into the actual Q-network to output q e ;
[0067] Step 4.D: Set the loss function as the mean square error between q e and q t ;
[0068] Step 4.E: Train the actual Q-network according to the gradient descent method.
[0069] Step 4.F: Every 50 frames, copy its weights from the actual Q-network to update the weights of the target Q-network.
[0070] Step 4.G: After 2000 frames, the system inputs the current state s into the actual Q-network to obtain the Q-values for each action, and selects the one with the largest Q-value as the current frame action a, and sends data with action a as the current frame transmission probability.
[0071] From step 1 to step 4, an enhanced ALOHA access method based on the deep reinforcement learning algorithm in this embodiment is completed. In this embodiment, after 2000 frames of greedy learning, the communication throughput capacity is gradually increased to around 0.65, as Figure 3 shown, which is stably better than other schemes with a fixed transmission probability. At the same time, its AOI is stable at around 3.1, as Figure 4 shown, which is stably lower than other schemes with a fixed transmission probability.
[0072] Experiments show that an enhanced ALOHA access method based on the deep reinforcement learning algorithm disclosed in this embodiment can intelligently adjust the real-time random access strategy of users through exploration and learning, dynamically select the frame transmission probability, effectively improve the average communication throughput, and avoid the communication congestion problem existing in traditional schemes; at the same time, increase the transmission probability of users with too large waiting time AoI (Age of Information), thereby effectively reducing the probability that some users cannot send for a long time, providing a gain for communication fairness. It realizes the trade-off optimization between the waiting time performance and the throughput performance of users.
[0073] The above specific description further details the purpose, technical solution, and beneficial effects of the invention. It should be understood that the above is only a specific embodiment of the present invention and is not used to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. An enhanced ALOHA access method based on a deep reinforcement learning algorithm, characterized in that, It includes the following steps: Step 1: Cell users are randomly activated with a predetermined activation probability. Users send data in each frame according to the enhanced ALOHA access strategy. At the receiving end, decoding is performed to obtain the current frame communication throughput capacity and the AoI of the user with the longest waiting time, and a set of successfully decoded users and a set of failed decoded users are respectively constructed; The implementation method of Step 1 is as follows: Step 1.1: Cell users are activated with a certain probability, and the activated users form a set of activated users; Step 1.2: The AoI is sorted in descending order, and the users in the front part of the marked sequence are marked as users with too long waiting time, and the rest are marked as users without too long waiting time; Step 1.3: For users without too long waiting time, with a certain base frame time slot transmission probability, randomly decide whether to send a data packet in each of the M time slots of the current frame; for users with too long waiting time, under the above base frame time slot transmission probability, increase the probability and randomly decide whether to send a data packet in each of the M time slots of the current frame; if there are multiple users sending data packets simultaneously in the same time slot resulting in a conflict, then the data packets of each user are superimposed and sent; Step 1.4: The receiving end decodes the data according to the interference cancellation technology; the users with successfully decoded current frame data form a set of successfully decoded users, and the users with failed decoded current frame data form a set of failed decoded users; Step 1.5: The AoI in the corresponding state of the successfully decoded users is cleared to zero, and the AoI in the corresponding state of the failed decoded users increases; Step 1.6: Calculate the current frame communication throughput capacity and the AoI of the user with the longest waiting time, where the current frame communication throughput capacity is the number of successfully decoded users divided by the number of activated users in the current frame, and the AoI of the user with the longest waiting time is the maximum value of the user AoI in the state; Step 2: Build a reinforcement learning framework and construct a deep Q network, namely Deep Q Network, DQN; The implementation method of Step 2 is as follows: Step 2.1: Construct a state space and an action space, where the state space consists of the current activation state of each user and the AoI of each user, and the action space consists of different frame time slot base transmission probabilities. Determine that the state s is an element in the state space and the action a is an element in the action space; Step 2.2: Given a stepped reward according to the current frame communication throughput capacity and the AoI of the user with the longest waiting time; Step 2.3: Construct a target Q network and an actual Q network, with the same structure, adopting a fully connected neural network structure; Step 2.4: Use the mean square error loss as the loss function; Step 3: The intelligent agent explores in the reinforcement learning framework for the first X frames, and expands the experience replay pool by introducing random actions to provide learning data for the intelligent agent; The implementation method of Step 3 is as follows: Step 3.1: The intelligent agent performs greedy learning with a greedy rate of ε and selects the action a according to the greedy algorithm; Step 3.2: The action a is used as the frame time slot transmission probability to perform DPSA random access, and the current frame communication throughput capacity, the AoI of the user with the longest waiting time, and the state s′ of the next frame are obtained; Step 3.3: Determine the reward r of the current frame according to the current frame communication throughput capacity and the AoI of the user with the longest waiting time; Step 3.4: Store the current frame state s, the current frame action a, the current frame reward r, and the next frame state s′ in the experience replay pool; Step 4: After X frames, randomly sample samples from the experience replay pool for learning, train the neural network to fit the access policy, and continue to send data.
2. The enhanced ALOHA access method based on the deep reinforcement learning algorithm according to claim 1, wherein The implementation method of Step 4 is as follows: Step 4.1: Randomly sample B samples, input the next-frame state s′ among them into the target Q-network, and output q n ; Step 4.2: According to the reward r, decay factor γ, and q in the sample n obtain q t = r + γ(max(q n )); Step 4.3: The current frame state s in the above sample is input into the actual Q-network, and q is output e ; Step 4.4: Set the loss function to q e and q t for the mean squared error; Step 4.5: Train the actual Q network according to the gradient descent method, with the learning rate set to α and the batch size set to B; Step 4.6: Every Y frames, copy its weights from the actual Q network to update the weights of the target Q network; Step 4.7: After X frames, the agent inputs the current state s into the actual Q network to obtain the Q values for each action, selects the action with the largest Q value as the current frame action a, sends data with the action a as the current frame sending probability, and obtains the user's maximum AoI and communication throughput performance under this frame sending probability. And after a certain number of frames, the user's average maximum AoI decreases and the communication throughput performance is significantly improved.