Distributed TDMA time slot allocation method based on multi-agent reinforcement learning

By applying the TDMA time slot allocation method of multi-agent reinforcement learning in mobile ad hoc networks, the problem of uneven resource allocation under dynamic and high load conditions is solved, and higher throughput and system stability is achieved.

CN119997209AActive Publication Date: 2025-05-13XIDIAN UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510064608.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-13
Estimated Expiration
2045-01-15

AI Technical Summary

Technical Problem

Traditional TDMA time slot allocation method cannot effectively consider node dynamic changes, network load fluctuations and link quality in dynamic and high-load mobile ad hoc networks, resulting in uneven resource allocation and affecting communication quality and throughput.

Method used

The distributed TDMA time slot allocation method based on multi-agent reinforcement learning is adopted. By building a mobile self-organizing network simulation platform and reinforcement learning module, the multi-agent deep deterministic strategy gradient MADDPG model is trained to realize collaborative work between agents and dynamic adjustment of time slot allocation strategies.

Benefits of technology

It significantly improves the network throughput and resource utilization, can respond to dynamic changes in real time, solves the problem of nodes having no time slot allocation under high load conditions, and enhances the flexibility, stability and fairness of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119997209A_ABST
    Figure CN119997209A_ABST
Patent Text Reader

Abstract

The invention provides a distributed TDMA (Time Division Multiple Access) time slot allocation method based on multi-agent reinforcement learning, which mainly solves the problem of poor performance of a mobile ad hoc network caused by no time slot allocation of nodes under a high-load condition in the prior art, and the scheme comprises the following steps: 1) constructing a mobile ad hoc network scene; 2) determining a frame structure of the TDMA according to a network scene; 3) establishing a mobile ad hoc network simulation platform by using the network scene and the frame structure, and obtaining a performance index; 4) enabling each node to correspond to one agent, and constructing a reinforcement learning module; 5) training the MADDPG model by using a simulation platform and a reinforcement learning module to obtain a parameter file; and 6) loading the trained parameters, and realizing time slot prediction by using the model. According to the method, multi-agent reinforcement learning is introduced, and a priority experience playback and exploration mechanism is combined, so that the time slot allocation efficiency is effectively improved, the throughput of the system is improved, and the overall performance of the mobile ad hoc network is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence technology, and further relates to a time division multiple access (TDMA) time slot allocation method in wireless communication technology, specifically a distributed TDMA time slot allocation method based on multi-agent reinforcement learning, which can be used for a MAC layer in a mobile self-organizing network. Technical Background

[0002] With the rapid development of wireless communication technology, mobile ad hoc networks are widely used in military, disaster relief, and temporary network scenarios due to their flexibility and adaptability. However, the dynamic and uncertain nature of such networks poses challenges to the management and scheduling of network resources. As an effective wireless resource allocation method, TDMA divides time into multiple time slots to achieve orderly communication of nodes, thereby avoiding signal collisions and improving spectrum utilization. In mobile ad hoc networks (MANETs), TDMA time slot allocation not only affects the overall performance of the network, but is also directly related to the communication efficiency and throughput of nodes. Traditional time slot allocation methods are often based on static models and fail to effectively consider factors such as dynamic changes in nodes, fluctuations in network load, and link quality, resulting in uneven resource allocation and ultimately affecting communication quality.

[0003] C. David Young et al. proposed the USAP protocol, or unified time slot allocation protocol, in the paper “USAP: a unifying dynamically distributed multichannel TDMA slot assignment protocol”. The protocol adds a time slot control phase before allocating time slots. In this control phase, all nodes broadcast their contention time slot information to neighboring nodes. After receiving the message, each node broadcasts the message of its neighboring node to its neighboring node again. This mechanism enables each node in the network to obtain the information of nodes within two hops of itself, so that conflicts can be avoided when allocating node time slots, conflicts caused by node movement can be resolved, and dynamic allocation of time slots can be achieved. However, due to the long data frame of the protocol, the access delay is relatively large. Chenxi Zhu et al. proposed the five-phase reservation protocol FPRP (Five-Phase Reservation Protocol) in the paper “A five-phase reservation protocol (FPRP) for mobile ad hoc networks”. The protocol divides time slots into a contention phase and a data phase. When competing, nodes can obtain the right to occupy data time slots through a five-step handshake to ensure that they obtain conflict-free broadcast time slots. However, the protocol has a relatively large overhead and has high requirements for clock synchronization. On this basis, the original author proposed E-TDMA (Evolutionary-TDMA) in the paper "An evolutionary-TDMA scheduling protocol (E-TDMA) for mobile ad hoc networks". This protocol can realize unicast and broadcast hybrid scheduling, which improves the flexibility of the network, but when the network load is large, the performance of the protocol will be greatly reduced. In response to the shortcomings of the USAP protocol, Peng Gexin et al. proposed the P-TDMA (priority-TDMA) protocol in the paper "A conflict-free dynamic time slot allocation algorithm based on fixed TDMA". Compared with USAP, it adds a priority-based time slot competition method, which can dynamically allocate time slots according to business load, and make the network time slot allocation have a certain fairness due to priority. However, the protocol still cannot reduce the node delay and does not solve the fundamental problem of poor scalability of fixed delay.

[0004] In MANET, the traditional TDMA method divides the communication time into multiple fixed-length time slots and allocates them to nodes in the network to avoid direct conflicts between nodes. Although TDMA has high determinism and low communication conflicts in early network environments, in a highly dynamic MANET environment, due to the need for different nodes to compete for limited time slot resources, some nodes may not be able to obtain enough time slots, especially in high-traffic or hot spots. This situation will make the communication needs of individual nodes unmet, further affecting the fairness and reliability of the overall network. In addition, the throughput of the traditional time slot allocation scheme cannot meet the requirements when dealing with complex network scenarios. It is necessary to further improve the rationality of the time slot allocation of each node to improve the overall throughput of the system. Reinforcement learning is a good solution for resource allocation, but due to the large number of nodes and time slot shards in MANET scenarios, the use of reinforcement learning for TDMA time slot allocation training will face the problem of slow training. Summary of the invention

[0005] The purpose of the present invention is to overcome the shortcomings of the above-mentioned prior art and propose a distributed TDMA time slot allocation method based on multi-agent reinforcement learning. By introducing multi-agent reinforcement learning and combining priority experience replay and sample generation optimization technology, it aims to effectively improve the allocation efficiency of TDMA time slots and improve the throughput of the system. At the same time, it solves the problem that nodes have no time slots to allocate under high load conditions, thereby enhancing the overall performance and stability of the mobile self-organizing network.

[0006] To achieve the above purpose, the technical solution adopted by the present invention is as follows:

[0007] (1) Construct a mobile ad hoc network scenario, including nodes, topology, and traffic model;

[0008] (2) Determine the TDMA frame structure based on the constructed mobile ad hoc network scenario;

[0009] (3) Using the mobile ad hoc network scenario and TDMA frame structure, a mobile ad hoc network simulation platform is constructed to obtain network performance indicators;

[0010] (4) Each node corresponds to an intelligent agent, the state space, action space, reward function and network structure are designed, and a reinforcement learning module is constructed; and a communication module for data transmission between the reinforcement learning module and the mobile self-organizing network simulation platform is established using Socket;

[0011] (5) Use the constructed simulation platform and reinforcement learning module to train the multi-agent deep deterministic policy gradient MADDPG model and obtain the trained MADDPG model parameter file;

[0012] (6) Load the trained MADDPG model parameter file and use the model to implement TDMA time slot prediction for mobile ad hoc networks.

[0013] Compared with the prior art, the present invention has the following beneficial effects:

[0014] First, the present invention introduces a multi-agent reinforcement learning framework and uses a simulation platform for joint training. The agents can work together more effectively to achieve fast and accurate time slot allocation, thereby significantly improving the network's throughput and resource utilization. At the same time, it can respond to dynamic changes in mobile self-organizing networks in real time, automatically adjust time slot allocation strategies to adapt to node load fluctuations, and solve the problem of nodes being unable to allocate time slots under high load conditions, thereby enhancing the flexibility, stability and fairness of the system.

[0015] Second, since the present invention uses a priority experience replay mechanism, the intelligent agent gives priority to using important experience samples, thereby accelerating the learning process, improving efficiency, reducing dependence on sample generation, and avoiding waste of resources; at the same time, during the training process, by marking and fine-tuning the "good samples" quickly generated using existing samples, the sample quality is effectively improved, the need to regenerate samples is reduced, and the overall sample generation process is more efficient and reliable; the present invention effectively reduces the reinforcement learning training time and enables the system to converge faster by jointly designing the exploration mechanism and the experience replay algorithm. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.

[0017] Figure 1 Flow chart for realizing the method of the present invention;

[0018] Figure 2 A schematic diagram of a TDMA frame structure provided in an embodiment of the present invention;

[0019] Figure 3 This is a framework diagram of the reinforcement learning simulation platform introduced in the present invention;

[0020] Figure 4 A training flow chart of the multi-agent reinforcement learning model constructed in the present invention;

[0021] Figure 5 A comparison chart of throughput simulation results under different load conditions when the method of the present invention and the existing method are used for time slot division;

[0022] Figure 6 The figure is a comparison diagram of fairness simulation results under different load conditions when the method of the present invention and the existing method are used to divide time slots. DETAILED DESCRIPTION

[0023] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the accompanying drawings in the embodiments of the present invention.

[0024] Example 1: Reference Figure 1 This example provides a distributed TDMA time slot allocation method based on multi-agent reinforcement learning. The specific implementation steps are as follows:

[0025] Step 1. Construct a mobile ad hoc network scenario, including nodes, topology, and traffic model. The implementation is as follows: Assume that the mobile ad hoc network consists of n nodes, and each node is bidirectionally connected to its neighboring nodes; model the entire mobile ad hoc network as an undirected graph G = <V, E>, where V represents the node set, |V| = n, V = {v1, v2, ...v n}, E is the set of edges, Each edge represents a TDMA communication link between two nodes in the network. Each edge has a bidirectional link and the links meet the condition That is, node v i and v j are within the communication transmission range of each other; let L(v i ,t) represents the i-th node v i The load at the current time t, S(v i ,t) represents the i-th node v i The number of time slots allocated at the current time t, T(v i ,t) represents the i-th node v i The throughput at the current time t.

[0026] Step 2. Determine the TDMA frame structure according to the constructed mobile self-organizing network scenario. The TDMA frame structure in this embodiment includes dividing a superframe into N frames, wherein the first frame is a preemption frame and the following N-1 frames are information frames; then dividing a frame into Q time slots, and each time slot includes P time slot slices, each time slot slice lasts for 1ms, that is, the duration of a superframe is N*Q*P milliseconds; setting the basic time slot slice of each node to N-1, that is, each node is allocated at least N-1 time slot slices.

[0027] Step 3. Use the mobile self-organizing network scenario and TDMA frame structure to build a mobile self-organizing network simulation platform to obtain network performance indicators. In this embodiment, the mobile self-organizing network simulation platform is used as the client of the reinforcement learning framework, and the Exata simulation software is used to model the mobile self-organizing network to obtain network performance indicators; the indicators at least include system throughput, load, fairness parameters, etc.

[0028] Step 4. Assign each node to an intelligent agent, design the state space, action space, reward function and network structure, and build a reinforcement learning module; and use Socket to establish a communication module for data transmission between the reinforcement learning module and the mobile self-organizing network simulation platform.

[0029] The design of state space, action space, reward function and network structure, and construction of reinforcement learning module are implemented as follows:

[0030] (4.1) Design a state space containing four-dimensional data, where the four-dimensional data is the node throughput at the previous moment, the number of time slots at the previous moment, the current load of the node, and the sum of the throughput of each node within two hops of the current node;

[0031] (4.2) According to the frame structure of step 2, the maximum number of time slots that can be allocated to each node is (N-1)×Q×P, which is a discrete value. The action space is designed as a discrete value of (N-1)×Q×P dimensions;

[0032] (4.3) The overall system throughput and fairness parameter are weighted as the reward function of the reinforcement learning system; assuming T i is the throughput of the i-th node, W i is the weight of the i-th node, and the fairness parameter is expressed as follows:

[0033]

[0034] The larger the F value, the better the fairness;

[0035] The global reward function can be expressed as:

[0036]

[0037] Where w is the weight value of the fairness parameter, which represents the ratio of the overall system throughput to the fairness parameter;

[0038] (4.4) Use the structure of the MADDPG model as the main framework of reinforcement learning to build the network structure of the reinforcement learning module, train it, and use the prioritized experience replay method for sample generation and extraction.

[0039] Step 5. Use the constructed simulation platform and reinforcement learning module to train the multi-agent deep deterministic policy gradient MADDPG model and obtain the trained MADDPG model parameter file; specifically, input the current load and current number of time slots of each node into the reinforcement learning framework, output the time slot allocation plan at the next moment through the reinforcement learning framework, and transmit the plan to the simulation platform for simulation to obtain the throughput and load at the next moment. The specific implementation steps include the following:

[0040] (5.1) Assume that the maximum total number of iterations is E1, the maximum number of exploration iterations is E2, and E1>E2; initialize the model parameter storage directory and initialize the current number of iterations to 0;

[0041] (5.2) The simulation platform generates the initial load L(v i ,0) and the initial number of time slots S(v i ,0), where the initial value of the number of time slots of each node is N-1, and the initial throughput of the i-th node is T(v i ,0) is 0;

[0042] (5.3) Use the communication module to convert L(v i ,t)、S(v i ,t) and T(v i ,t) is transmitted to the reinforcement learning module, and the reinforcement learning module converts L(v i ,t)、S(v i ,t) and T(v i ,t) is constructed as the state space s at time t t ;

[0043] (5.4) Input the state space into the reinforcement learning module to determine whether the current number of iterations has reached the set maximum number of exploration iterations. If not, enter the exploration module, and the exploration module obtains the action A(v i ,t), and vice versa, the action A(v i ,t); then construct the action space a at time t according to the actions of all nodes t , that is, the time slot allocation scheme.

[0044] The exploration module in this embodiment is a new exploration mechanism designed by the present invention, specifically: in one exploration, if the throughput of this exploration reaches a set threshold, the group of results is marked as "good results", and the random exploration mechanism is not used in the next exploration, but the present exploration mechanism is used; in this exploration mechanism, based on the "good results", the time slot allocation is fine-tuned and then put into the experience pool to obtain a new group of samples. If the throughput of the new samples is improved, fine-tuning is continued on this basis; otherwise, fine-tuning is performed again based on the last "good results"; at the same time, the maximum number of fine-tuning in a fine-tuning stage is set to prevent entering an infinite loop.

[0045] The fine-tuning is done in A(v i ,t) is higher than the average of all node time slot values, adjust A(v i ,t+1)=A(v i ,t)-x, otherwise adjust A(v i ,t+1)=A(v i ,t)+x, thereby achieving; where x is the fine-tuning coefficient.

[0046] The experience pool adopts a priority experience replay mechanism to select samples. The priority experience replay mechanism here uses the time series error TD-error as a measurement standard to evaluate the priority of the sampled data.

[0047] (5.5) A(v i ,t) is input into the simulation platform for simulation, and the throughput, load and number of node time slots T(v i ,t+1),L(v i ,t+1) and S(v i ,t+1), and the reward r at time t t , and construct the state space s for the next moment t+1 ;

[0048] (5.6) t ,a t ,r t ,s t+1 As a sample t =(s t ,a t ,r t ,s t+1 ) into the experience pool;

[0049] (5.7) Determine whether the current number of iterations is greater than the maximum number of exploration iterations. If so, select multiple samples from the experience pool to train the neural network of the MADDPG model and then continue to execute step (5.8); otherwise, directly execute step (5.8);

[0050] (5.8) After adding one to the current number of iterations E3, determine whether it has reached the maximum total number of iterations E1. If not, return to step (5.3). Otherwise, continue to step (5.9);

[0051] (5.9) End the training process and generate the trained MADDPG model parameter file.

[0052] Step 6. Load the trained MADDPG model parameter file and use the model to implement TDMA time slot prediction for mobile ad hoc networks. In this embodiment, the trained model parameter file is imported into the MADDPG model, and the neural network in the model is used to generate the action A(v i ,t), and derive the time slot allocation plan based on the actions of all nodes.

[0053] Embodiment 2: The overall implementation steps of the time slot allocation method proposed in this embodiment are the same as those of Embodiment 1. Figure 2-4 , specific parameter settings are given to further describe the implementation process of the present invention in detail:

[0054] Step 1. Build a mobile ad hoc network scenario:

[0055] Assume that the mobile ad hoc network consists of 8 nodes, and each node can be bidirectionally connected to neighboring nodes. The entire mobile ad hoc network can be modeled as an undirected graph G = <V, E>, where V is the node set, |V| = 8, V = {v1, v2, ... v8}, E is the edge set, Each edge represents a TDMA communication link between two nodes in the network. Each edge has a bidirectional link, and these links all meet one condition: That is, nodes u and v (both belong to set E) must be within each other’s transmission range. Each node has a set of attributes to describe its state, including: the node’s current load L(v i ,t), the number of time slots currently allocated to the node S(v i ,t), the throughput available at the current moment of the node is T(v i ,t).

[0056] Step 2. Determine the TDMA frame structure:

[0057] Compared with the benchmark scheme, a frame structure in the actual scheme is selected as the frame structure of TDMA according to the constructed mobile self-organizing network scenario. In this embodiment, a superframe is preferably divided into 8 time slots, of which the first frame is a preemption frame and the following 7 are information frames; a frame includes 32 time slots, each time slot includes 10 time slot slices, and the duration of each time slot slice is 1ms, that is, the duration of a superframe is 2.56s, and the default basic time slot slice of each node is 7, that is, at least 7 time slot slices are allocated to determine the frame structure; Figure 2 shown.

[0058] Step 3. Get network performance metrics:

[0059] The mobile self-organizing network model is used as the client of the reinforcement learning framework, and Exata simulation software is used to model the mobile self-organizing network, so that network indicators such as system throughput and latency can be obtained.

[0060] Step 4. Build the reinforcement learning module:

[0061] Each node corresponds to an agent, and the state space, action space, reward function and network structure are designed to build a reinforcement learning module; the state S involved represents the state of the environment (State), the reward R represents the feedback value given by the environment (Reward), and the action A represents the action (Action) decided by the agent (Action). Socket is used to build a communication framework to enable data transmission between the reinforcement learning module and the simulation platform.

[0062] The state space, action space, reward function and network structure designed in this embodiment are as follows:

[0063] (4-1) State space design: The state space is four-dimensional data, which includes the node’s throughput at the previous moment, the number of time slots at the previous moment, the node’s current load, and the sum of the throughput of each node within two hops of the current node.

[0064] (4-2) Action space design: According to the frame structure of step 2, the maximum number of time slots that can be allocated to each node is 2240, which is a discrete number. Therefore, the action space is also designed as a discrete value of 2240 dimensions, ranging from 0 to 2239.

[0065] (4-3) Reward function design: Since it is necessary to maximize the system throughput while maintaining a high level of fairness, the overall system throughput and the fairness parameter are weighted as the reward function of the reinforcement learning system, where the fairness parameter uses Jain's Faireness index to represent the fairness of time slot allocation. Assume that there are N nodes in total, T i is the throughput of the i-th node, W i As the weight of each node, the weight in this paper is related to the load, then Jain's Faireness index can be expressed as:

[0066]

[0067] The larger the F value is, the better the fairness is; n represents the number of nodes.

[0068] The global reward function can be expressed as:

[0069]

[0070] Wherein w is the weight, and in this embodiment, it is preferably 200.

[0071] (4-4) Network structure design: Since the node time slot situation depends not only on the state of the node itself, but also on the state of other nodes, the present invention uses a multi-agent reinforcement learning framework for training, in which the MADDPG structure is used as the main framework of reinforcement learning, and the prioritized experience replay method is used for sample generation and extraction.

[0072] The load platform and reinforcement learning module architecture in steps 3 and 4 above are as follows: Figure 3 shown.

[0073] Step 5. Train the MADDPG model:

[0074] The current load and current number of time slots of each node are input into the reinforcement learning framework, which outputs the time slot allocation plan for the next moment and transmits the plan to the simulation platform for simulation to obtain the throughput and load at the next moment:

[0075] (5-1) The simulation platform generates an initial load L(v i ,0) and the initial number of time slots S(v i ,0), where the initial value of the number of time slots for each node is 7, and the initial throughput of the i-th node is T(v i ,0) is 0;

[0076] (5-2) Use the communication module to convert L(v i ,t)、S(v i ,t) and T(v i ,t) is transmitted to the reinforcement learning module, and the reinforcement learning module converts L(v i ,t)、S(v i ,t) and T(v i ,t) is constructed as the state space s at time t t ;

[0077] (5-3) Input the state space into the reinforcement learning module to obtain the action of each node, and then construct the action space a at time t based on the actions of all nodes. t , that is, the time slot allocation scheme is obtained. In this embodiment, the number of nodes in the first 1000 iteration cycles is set to be randomly generated, and the neural network output results of the MADDPG model are used in the subsequent cycles, and some samples are selected from the experience pool for training to obtain the corresponding model file;

[0078] (5-4) Input the time slot allocation scheme parameters into the simulation platform for simulation and obtain the current throughput T(v i ,t), current load L(v i ,t), the number of time slots of the current node S(v i ,t).

[0079] (5-5) Repeat steps (5-2) to (5-4) and convert e t =(s t ,a t ,r t ,s t+1 ) is taken as a sample and put into the experience pool, where r t is the reward at time t, s t+1 is the state space constructed for the next moment.

[0080] Step 6. Import the model file trained in step 5 and execute according to the process of step 5. The difference is that no training is required in step (5-3), and only the allocation scheme needs to be obtained. The training process in step 5 is improved as follows:

[0081] In step (5-3), the number of time slots for each node in the first 1000 cycles is randomly generated. However, for the MANET environment, since the number of optional time slots for each node in TDMA is large and the number of nodes is large, the reinforcement learning action space is also large. At this time, the use of a random exploration mechanism will result in a long time of exploration without being able to find a better time slot allocation result, making it impossible for the data in the experience pool to contain more valid data, and ultimately leading to falling into a local optimal solution during the training phase.

[0082] The present invention proposes a new exploration mechanism: in an exploration, if the throughput of this exploration reaches a set threshold, then this group of results is marked as "good results", and the next exploration will not go through the random exploration mechanism, but turn to the current exploration mechanism. In this exploration mechanism, based on the "good results", the time slot allocation is fine-tuned, and then put into the environment to obtain a new group of samples. If the throughput of the new samples is improved, then fine-tuning will continue on this basis; if the throughput decreases, it will be fine-tuned again on the previous "good samples". In order to prevent entering an infinite loop, a fine-tuning stage is set to perform a limited number of fine-tuning at most.

[0083] In addition, the experience pool sample selection method in step (5-3) adopts the priority experience replay method. The priority experience replay mechanism uses TD-error (time series error) as a criterion to evaluate the priority of the sampled data. TD-error refers to the difference between the current Q value and its target Q value. The larger the error, the more helpful the data is for updating the network parameters. Greedy (selecting the maximum value) sampling of data with large TD-error for training accelerates the convergence of training. The optimized training process is as follows Figure 4 shown.

[0084] The effect of the present invention is further described below in conjunction with simulation experiments.

[0085] 1. Simulation conditions:

[0086] The simulation experiment of the present invention is carried out in the hardware environment of 32G memory, AMD5800H CPU, 3060GPU and the software environment of Exata7.2.0, C++14, Python3.6 and Pytorch1.10.

[0087] 2. Simulation content:

[0088] Using the Exata network simulation platform, the throughput and fairness of the traditional TDMA method, FPRP method and the method of the present invention are compared by changing the packet sending interval to adjust different load conditions, and the load-throughput curve and load-fairness curve are obtained. Figure 5 and Figure 6 shown.

[0089] 3. Simulation results:

[0090] Depend on Figure 5 It can be seen that the time slot allocation method proposed by the present invention has a higher throughput than the traditional TDMA under different load conditions, and is relatively close to the throughput curve of the FPRP protocol.

[0091] pass Figure 6 It can be seen that the fairness of traditional TDMA is the strongest, the fairness of the present invention can also achieve a good effect, and the fairness of FPRP is relatively poor.

[0092] It can be seen that the present invention has stronger competitiveness in practical applications. While having high throughput, it can also have good fairness, can effectively address the shortcomings of traditional technologies, and improve the overall performance of network communications.

[0093] The above simulation analysis proves the correctness and effectiveness of the method proposed in the present invention.

[0094] The parts not described in detail in the present invention belong to the common knowledge of those skilled in the art. The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Obviously, for professionals in this field, after understanding the content and principle of the present invention, it is possible to make various modifications and changes in form and details without departing from the principle and structure of the present invention. However, these solutions based on the idea of ​​the present invention and using the multi-agent reinforcement learning algorithm and the time slot allocation algorithm for joint design are all within the scope of protection of the present invention.

Claims

1. A distributed TDMA time slot allocation method based on multi-agent reinforcement learning, characterized in that: The implementation steps include the following: (1) Construct a mobile ad hoc network scenario, including nodes, topology, and traffic model; (2) Determine the TDMA frame structure based on the constructed mobile ad hoc network scenario; (3) Using the mobile ad hoc network scenario and TDMA frame structure, a mobile ad hoc network simulation platform is constructed to obtain network performance indicators; (4) Assign each node to an intelligent agent, design the state space, action space, reward function and network structure, and build a reinforcement learning module; And use Socket to establish a communication module for data transmission between the reinforcement learning module and the mobile self-organizing network simulation platform; (5) Use the constructed simulation platform and reinforcement learning module to train the multi-agent deep deterministic policy gradient MADDPG model and obtain the trained MADDPG model parameter file; (6) Load the trained MADDPG model parameter file and use the model to implement TDMA time slot prediction for mobile ad hoc networks.

2. The method according to claim 1, characterized in that: The construction of the mobile self-organizing network scenario in step (1) is implemented as follows: Assume that the mobile self-organizing network consists of n nodes, and each node is bidirectionally connected to its neighboring nodes; model the entire mobile self-organizing network as an undirected graph G = <V, E>, where V represents the node set, |V| = n, V = {v1, v2, ... v n }, E is the set of edges, v i 、v j ∈V,i≠j; each edge represents a TDMA communication link between two nodes in the network, each edge has a bidirectional link, and the links all meet the conditions That is, node v i and v j are within the communication transmission range of each other; let L(v i ,t) represents the i-th node v i The load at the current time t, S(v i ,t) represents the i-th node v i The number of time slots allocated at the current time t, T(v i ,t) represents the i-th node v i The throughput at the current time t.

3. The method according to claim 1, characterized in that: The frame structure of the TDMA in step (2) is determined as follows: a superframe is divided into N frames, wherein the first frame is a preemption frame and the following N-1 frames are information frames; a frame is then divided into Q time slots, and each time slot includes P time slots, and each time slot lasts for 1 ms, i.e., the duration of a superframe is N*Q*P milliseconds; the basic time slot of each node is set to N-1, i.e., each node is allocated at least N-1 time slots.

4. The method according to claim 1, characterized in that: The network performance index described in step (3) is obtained by using the mobile self-organizing network simulation platform as the client of the reinforcement learning framework and using Exata simulation software to model the mobile self-organizing network; the index at least includes system throughput, load, and fairness parameters.

5. The method according to claim 1, characterized in that: Step (4) designs the state space, action space, reward function and network structure, and constructs a reinforcement learning module. The implementation steps are as follows: (4.1) Design a state space containing four-dimensional data, where the four-dimensional data is the node throughput at the previous moment, the number of time slots at the previous moment, the current load of the node, and the sum of the throughput of each node within two hops of the current node; (4.2) According to the frame structure of step (2), the maximum number of time slots that can be allocated to each node is (N-1)×Q×P, which is a discrete value. The action space is designed as a discrete value of (N-1)×Q×P dimensions; (4.3) The overall system throughput and fairness parameter are weighted as the reward function of the reinforcement learning system; assuming T i is the throughput of the i-th node, W i is the weight of the i-th node, and the fairness parameter is expressed as follows: The larger the F value, the better the fairness; The global reward function can be expressed as: Where w is the weight value of the fairness parameter, which represents the ratio of the overall system throughput to the fairness parameter; (4.4) Use the structure of the MADDPG model as the main framework of reinforcement learning to build the network structure of the reinforcement learning module, train it, and use the prioritized experience replay method for sample generation and extraction.

6. The method according to claim 1, characterized in that: Step (5) uses the constructed simulation platform and reinforcement learning module to train the MADDPG model, which is to input the current load and current number of time slots of each node into the reinforcement learning framework, output the time slot allocation plan at the next moment through the reinforcement learning framework, and transmit the plan to the simulation platform for simulation to obtain the throughput and load at the next moment. The specific implementation steps include the following: (5.1) Assume that the maximum total number of iterations is E1, the maximum number of exploration iterations is E2, and E1>E2; initialize the model parameter storage directory and initialize the current number of iterations to 0; (5.2) The simulation platform generates the initial load L(v i ,0) and the initial number of time slots S(v i ,0), where the initial value of the number of time slots of each node is N-1, and the initial throughput of the i-th node is T(v i ,0) is 0; (5.3) Use the communication module to convert L(v i ,t)、S(v i ,t) and T(v i ,t) is transmitted to the reinforcement learning module, and the reinforcement learning module converts L(v i ,t)、S(v i ,t) and T(v i ,t) is constructed as the state space s at time t t ; (5.4) Input the state space into the reinforcement learning module to determine whether the current number of iterations has reached the set maximum number of exploration iterations. If not, enter the exploration module, and the exploration module obtains the action A(v i ,t), and vice versa, the action A(v i ,t); then construct the action space a at time t according to the actions of all nodes t , that is, the time slot allocation scheme; (5.5) A(v i ,t) is input into the simulation platform for simulation, and the throughput, load and number of node time slots T(v i ,t+1),L(v i ,t+1) and S(v i ,t+1), and the reward r at time t t , and construct the state space s for the next moment t+1 ; (5.6) t ,a t ,r t ,s t+1 As a sample t =(s t ,a t ,r t ,s t+1 ) into the experience pool; (5.7) Determine whether the current number of iterations is greater than the maximum number of exploration iterations. If so, select multiple samples from the experience pool to train the neural network of the MADDPG model and then continue to execute step (5.8); otherwise, directly execute step (5.8); (5.8) After adding one to the current number of iterations E3, determine whether it has reached the maximum total number of iterations E1. If not, return to step (5.3). Otherwise, continue to step (5.9); (5.9) End the training process and generate the trained MADDPG model parameter file.

7. The method according to claim 1, characterized in that: Step (6) implements TDMA time slot prediction of the mobile ad hoc network, specifically, importing the trained model parameter file into the MADDPG model, using the neural network in the model to generate the action A(v i ,t), and derive the time slot allocation plan based on the actions of all nodes.

8. The method according to claim 6, characterized in that: The exploration module is a new exploration mechanism, specifically: in one exploration, if the throughput of this exploration reaches a set threshold, the group of results is marked as "good results", and the random exploration mechanism is not used in the next exploration, but the exploration mechanism is used; in this exploration mechanism, based on the "good results", the time slot allocation is fine-tuned and then put into the experience pool to obtain a new group of samples. If the throughput of the new samples is improved, fine-tuning is continued on this basis; Otherwise, fine-tune again based on the last "good result"; at the same time, set the maximum number of fine-tuning times in a fine-tuning stage to prevent entering an infinite loop.

9. The method according to claim 8, characterized in that: The fine-tuning is specifically implemented as follows: If A(v i ,t) is higher than the average of all node time slot values, then adjust A(v i ,t+1)=A(v i ,t)-x, otherwise adjust A(v i ,t+1)=A(v i ,t)+x; where x is the fine-tuning coefficient.

10. The method according to claim 8, characterized in that: The experience pool selects samples using a priority experience replay mechanism, and the priority experience replay mechanism uses the timing error TD-error as a criterion to evaluate the priority of the sampled data.

Citation Information

Patent Citations

  • Media access control method for observing underwater acoustic sensor network in UUV (Unmanned Underwater Vehicle) cluster environment

    CN116963286A

  • Topology prediction TDMA (Time Division Multiple Access) time slot allocation method based on dynamic advance

    CN118433883A

  • Resource allocation techniques to support multiple peer-to-peer (P2P) sessions

    US20240292456A1