An intelligent time slot allocation method and system for the space-earth integrated scenario

By adopting the intelligent TDMA slot allocation method based on reinforcement learning in the integrated air-ground network, dynamically allocating time slot resources is solved, and the problems of poor flexibility and resource waste of time slot allocation methods in the existing technology are solved, and efficient resource utilization and throughput are achieved.

CN115551091BActive Publication Date: 2025-06-27UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211142895.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-20
Publication Date
2025-06-27
Estimated Expiration
2042-09-20

AI Technical Summary

Technical Problem

The slot allocation method in the prior art has large access delay, poor flexibility, high transmission delay and low channel utilization. It is impossible to dynamically adjust the slot allocation method according to different service needs, resulting in wasted time slot resources.

Method used

Using the intelligent TDMA slot allocation method based on reinforcement learning, the Markov decision model and deep reinforcement learning neural network model are used to dynamically allocate time slot resources, and the time slot state and allocation strategy are updated in real time according to business needs.

Benefits of technology

It achieves the maximization of system throughput under the conditions of meeting the flexible needs of different services, improves resource utilization efficiency, and reduces time slot resource waste.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115551091B_ABST
    Figure CN115551091B_ABST
Patent Text Reader

Abstract

The present invention discloses an intelligent time slot allocation method and system for the space-ground integrated scenario. The present invention is oriented to the complex and changeable environment and diverse service requirements in the air-ground integrated scenario, as well as the strict requirements for the transmission delay and access rate of different services. The method includes: different users send time slot request information to the base station in real time according to service requirements, and the service requirements include service load requirements, service type requirements, and service delay requirements; the base station, based on the received time slot request information of all users and the time slot status information of the current network, uses an intelligent time slot allocation method based on reinforcement learning to allocate time slots to the time slot request information of all users, obtaining a time slot allocation strategy for the users; and sending the obtained time slot allocation strategy to the corresponding users, while updating the time slot status information. The present invention maximizes the system throughput under the condition of meeting the flexible requirements of different services, and realizes the efficient utilization of resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of time slot allocation, and in particular to an intelligent time slot allocation method and system for a space-ground integrated scenario. Background Art

[0002] In recent years, with the rapid development of aviation technology and unmanned aerial vehicle technology, the air-ground integrated communication technology using high-altitude platforms for data transmission has attracted wide attention in the academic and industrial fields. The air-ground integrated system uses fixed-wing aircraft or unmanned aerial vehicles as airborne base stations, which are integrated with the ground network to jointly provide emergency communication services for users. Compared with satellite communication systems, it has the advantages of low cost, small delay, fast setup, and large capacity; compared with ground communication systems, it has the advantages of small multipath fading, large coverage area, and strong anti-destruction ability.

[0003] However, the network architecture and wireless environment faced by the air-ground integrated network are more complex, and the dynamicity and difference of the environment and services are also more obvious. On the one hand, different from ground communication technologies based on fixed base station deployment, the base stations in the air-ground integrated network have mobility. On the other hand, ground and airborne base stations jointly provide seamless access services for ground and airborne mobile stations in the hybrid network. These mobile stations share wireless resources, and cells with different coverage ranges form a complex heterogeneous access network. At the same time, there are a wide variety of service loads in the network, including voice, short messages, files, videos, etc. There are also obvious differences in the arrival time and space of each service, and the service quality and transmission delay requirements of different services are also significantly different. How to intelligently and efficiently flexibly allocate limited MAC resources in a highly dynamic and complex heterogeneous air-ground integrated network, while ensuring the service quality requirements and transmission delay limitations of services, and maximizing the network throughput and user access rate will be an urgent problem to be solved in air-ground integrated communication technology.

[0004] Currently, MAC protocols are mainly divided into random competition type, allocation scheduling type, and a hybrid type of the first two. The random competition type is derived from classic access protocols such as ALOHA and CSMA / CA. The principle is to rely on asynchronous competition to obtain the right to occupy the channel and use random backoff to alleviate collision problems. However, as the network load increases and collisions increase, the transmission rate and delay performance will seriously decline. TDMA is a typical algorithm based on allocation scheduling. The TDMA protocol divides the channel into several fixed-length time slots, and nodes perform packet transmission in the corresponding time slots according to certain allocation rules. Compared with the random competition MAC protocol, the TDMA protocol has better guarantees in terms of throughput, transmission delay, etc. Therefore, TDMA technology is mostly used in the MAC layer of air-ground integrated systems.

[0005] The existing TDMA time slot allocation methods can be roughly divided into three categories: fixed time slot allocation method, dynamic time slot allocation method, and hybrid time slot allocation method combining fixed and dynamic. Among them, according to the implementation method of the method, the TDMA protocol based on the dynamic allocation method can be further divided into centralized and distributed; the distributed dynamic TDMA protocol can be further divided into two types: topology-dependent and topology-transparent according to whether topological information is required during time slot allocation.

[0006] The TDMA protocol adopting the fixed time slot allocation strategy can ensure that each node obtains fixed time slot resources, effectively guarantee the fairness among users, and better meet the requirements of time delay. However, as the number of nodes in the network increases, the channel utilization rate will decrease significantly, and the throughput of the network will also be limited. In the air-ground integrated network, due to the uneven distribution of node traffic, there may be a situation where some nodes need to send a large amount of data at a certain moment, while some nodes have no traffic to send. In response to this situation, the TDMA protocol adopting the fixed time slot allocation strategy may lead to low channel utilization rate and high service transmission delay. Therefore, a more dynamic time slot allocation scheme needs to be considered.

[0007] However, the current dynamic time slot allocation methods generally have the deficiencies of large access delay, poor flexibility, high transmission delay, low channel utilization rate, inability to dynamically adjust the time slot allocation method according to different service requirements, resulting in a certain waste of time slot resources. Summary of the Invention

[0008] The technical problem to be solved by the present invention is that the existing time slot allocation methods generally have the defects of large access delay, poor flexibility, high transmission delay, low channel utilization rate, inability to dynamically adjust the time slot allocation method according to different service requirements, resulting in a certain waste of time slot resources.

[0009] Facing the complex and changeable environment and diverse service requirements in the air-ground integrated scenario, as well as the strict requirements for the transmission delay and access rate of different services, the purpose of the present invention is to provide an intelligent time slot allocation method and system for the space-ground integrated scenario. The present invention is an intelligent TDMA time slot allocation method based on reinforcement learning, which maximizes the system throughput under the condition of meeting the flexible requirements of different services and realizes the efficient utilization of resources.

[0010] The present invention is realized through the following technical solutions:

[0011] In the first aspect, the present invention provides an intelligent time slot allocation method for the space-ground integrated scenario, and the method includes:

[0012] Different users send time slot request information to the base station (air base station or ground base station) in real time according to service requirements, and the service requirements include service load requirements, service type requirements, and service time delay requirements;

[0013] The base station (air base station or ground base station) performs time slot allocation on the time slot request information of all users based on the received time slot request information of all users and the time slot status information of the current network, using an intelligent time slot allocation method based on reinforcement learning, to obtain the time slot allocation strategy for the users; and sends the obtained time slot allocation strategy to the corresponding users, while updating the time slot status information.

[0014] Furthermore, the method further includes:

[0015] The base station (air base station or ground base station) periodically updates and maintains the time slot status information and time slot request information of all users within the current scope of the base station.

[0016] Furthermore, the specific steps for the intelligent time slot allocation method based on reinforcement learning to perform time slot allocation on the time slot request information of all users are as follows:

[0017] Build a Markov decision model MDP based on time slot allocation, and define the state, action, reward function, set of transition probabilities, discount factor, state value function, and state-action value function of the Markov decision model MDP; where, in the Markov decision model MDP, the base station is used as an entity to autonomously collect environmental state information and determine the time slot allocation strategy according to the time slot request information of the users.

[0018] Build a time slot allocation algorithm based on reinforcement learning in the Markov decision model MDP, and search for the best action that maximizes the global reward function to obtain the optimal time slot allocation strategy.

[0019] Furthermore, the state of the Markov decision model MDP is the time slot request information of the users and the currently available time slot status information, and the action of the Markov decision model MDP is the number of time slots allocated to each service; the action vector of the Markov decision model MDP constitutes the action space, and the Markov decision model MDP uses the reward function to evaluate the action, and the optimization goal is to maximize the total number of service accesses under the condition of meeting different service requirements, so as to maximize the network throughput.

[0020] Furthermore, building a time slot allocation algorithm based on reinforcement learning in the Markov decision model MDP, and searching for the best action that maximizes the global reward function to obtain the optimal time slot allocation strategy includes:

[0021] Establish an intelligent time slot allocation neural network model based on deep reinforcement learning and initialize the model parameters;

[0022] According to the Markov decision model MDP and the slot status request information of all users, collect the status, actions, and reward information of the slots in the network, and use the status, actions, and reward information of the slots in the network as model training data;

[0023] Use the model training data to train the intelligent slot allocation neural network model based on deep reinforcement learning, and search the global action space based on the AC algorithm, and output the slot allocation strategy that maximizes the return function as the optimal slot allocation strategy;

[0024] According to the optimal slot allocation strategy, extract the total number of slots allocated to each service; obtain the number of allocable users according to the time limit request information of each service; and according to the service arrival time, preferentially allocate slots to users with longer waiting times to obtain the slot allocation result of each user;

[0025] The base station broadcasts the slot allocation result of each user and updates the slot status according to the slot allocation result.

[0026] Further, the calculation formula for the number of allocable users is:

[0027] a. If the service type does not support frequency division multiplexing (such as services like video and files), the number of allocable users for this service is = Z / t, where Z is the number of slots allocated to a certain service type obtained from the slot allocation algorithm based on reinforcement learning, and t is the fixed number of slots required to transmit this service;

[0028] b. If the service type supports frequency division multiplexing (such as services like voice and short messages), the number of allocable users for this service is = (Z / t) * M, where Z is the number of slots allocated to a certain service type obtained from reinforcement learning, t is the fixed number of slots required to transmit this service, and M is the maximum number of users that the same slot supports for transmitting the same service in frequency division multiplexing technology.

[0029] Further, the frame structure in the intelligent slot allocation method based on reinforcement learning includes a random access slot, an uplink control slot, a service slot, and a downlink broadcast slot;

[0030] Among them, the random access slot is used for new nodes to join, the uplink control slot is used for users to send slot application information, and the downlink broadcast slot is used for the base station to send slot allocation strategies, synchronization signals, voice, data, etc.

[0031] In a second aspect, the present invention further provides an intelligent slot allocation system for a space-ground integrated scenario, which supports the intelligent slot allocation method for a space-ground integrated scenario described above; the system includes:

[0032] The user terminal is used for different users to access the system and send slot request information to the base station (aerial base station or ground base station) in real time according to service requirements;

[0033] The base station terminal is used for, based on the received slot request information of all users and the slot status information of the current network, using an intelligent slot allocation method based on reinforcement learning to allocate slots for the slot request information of all users, obtaining the slot allocation strategy for users; and sending the obtained slot allocation strategy to the corresponding users, while updating the slot status information; and the base station (aerial base station or ground base station) periodically updates and maintains the slot status information and slot request information of all users within the current scope of the base station.

[0034] Furthermore, the specific execution process of the user terminal is as follows:

[0035] After the initial network construction is completed, unconnected users send access request signaling in random access slots and apply for network access according to the user access algorithm;

[0036] The successfully connected users need to continuously listen for a period of time until they receive the uplink control slot information allocated by the base station for this user;

[0037] According to the current service load situation, the user sends slot request information to the base station in the corresponding control slot and waits for the base station to allocate a service slot address for this user; the slot request information includes the slot demand, the service type sent, the message urgency, and the service waiting time, etc.;

[0038] If the user receives the service slot address broadcast by the base station for it, it will perform service data transmission in the corresponding slot of the current frame; otherwise, it will continue to wait until the data transmission is completed.

[0039] Furthermore, the base station terminal is also used for, after the initial network construction is completed, broadcasting to the users the current number of idle slots and continuously listening for the slot request information sent by the user terminal.

[0040] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0041] An intelligent slot allocation method and system for an integrated space-ground scenario according to the present invention, facing the complex and changeable environment and diverse service requirements in the integrated air-ground scenario, as well as the strict requirements for the transmission delay and access rate of different services. The present invention is an intelligent TDMA slot allocation method based on reinforcement learning, which maximizes the system throughput under the condition of meeting the flexible requirements of different services and realizes the efficient utilization of resources. Description of the Drawings

[0042] The accompanying drawings described herein are used to provide a further understanding of the embodiments of the present invention, form a part of this application, and do not limit the embodiments of the present invention. In the drawings:

[0043] Figure 1 It is a scenario diagram of the air-ground integrated network of the present invention.

[0044] Figure 2 It is a schematic diagram of the frame structure of the present invention.

[0045] Figure 3 It is a flowchart of the operation of the user terminal of the present invention.

[0046] Figure 4 It is a flowchart of the operation of the base station terminal of the present invention.

[0047] Figure 5 It is a basic schematic diagram of the reinforcement learning of the present invention.

[0048] Figure 6 It is a framework diagram of the reinforcement learning for time slot allocation of the present invention.

[0049] Figure 7 It is a framework diagram of the Actor-Critic algorithm of the present invention.

[0050] Figure 8 It is a flowchart of the intelligent time slot allocation method based on reinforcement learning of the present invention.

[0051] Figure 9 It is a flowchart of an intelligent time slot allocation method for an air-ground integrated scenario of the present invention.

[0052] Figure 10 It is a structural block diagram of an intelligent time slot allocation system for an air-ground integrated scenario of the present invention. Detailed implementation manners

[0053] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in combination with embodiments and drawings. The illustrative embodiments and descriptions of the present invention are only used to explain the present invention and do not limit the present invention.

[0054] Embodiment

[0055] Facing the complex and changeable environment and diverse service requirements in the air-ground integrated scenario, as well as the strict requirements for the transmission delay and access rate of different services, the present invention provides an intelligent time slot allocation method and system for an air-ground integrated scenario. The present invention is an intelligent TDMA time slot allocation method based on reinforcement learning, which maximizes the system throughput under the condition of meeting the flexible requirements of different services and realizes the efficient utilization of resources.

[0056] The process of the solution of the present invention is as follows: First, the user sends slot request information to the base station (aerial base station or ground base station) according to requirements such as service load, service type, and service delay. Secondly, the base station executes an intelligent slot allocation method based on reinforcement learning according to the slot request information received by all users, and outputs a slot allocation strategy with the maximum reward function. Finally, this slot allocation strategy is sent to specific users. In order to make slot allocation more flexible to meet the QoS requirements of differentiated services, the solution of the present invention also considers supporting frequency division multiplexing FDMA on the basis of TDMA. In particular, the solution of the present invention considers that each slot can be divided into M orthogonal sub-channels according to spectrum resources, and each sub-channel can be allocated to different users to transmit services. A user can choose to occupy 1, 2 or all sub-channels of a certain slot. Since the bandwidth resource of the sub-channel becomes 1 / M of the original, the transmission rate of the physical channel is reduced. In order to better transmit large-bandwidth services, the solution of the present invention considers that large-bandwidth services always occupy all sub-channels of a slot, that is, large-bandwidth services do not consider FDMA technology. The system model of the solution of the present invention is as follows:

[0057] I. System Model

[0058] 1. Problem Description

[0059] The present invention considers an integrated air-ground scenario, which mainly consists of a central base station and mobile stations. The central base station is used for comprehensive scheduling and dynamically allocates time slots for users. The base station can be a ground base station, a sea base station, a high / low altitude base station, and the mobile station can be a slow-moving sea / ground mobile station and a fast-moving air mobile station, as Figure 1 shown.

[0060] The present invention considers that there are diverse services in the integrated air-ground network. There are huge differences in the priority, message size, and distribution model of each service. How to provide on-demand and flexible radio resources for different services and maximize the system throughput under the condition of meeting the flexible requirements of different services is the problem to be solved by the present invention. In particular, the characteristics of the service types considered by the present invention include:

[0061] 1) Real-time streaming services such as voice and video streams;

[0062] 2) Low-latency services represented by short messages;

[0063] 3) High-bandwidth demand services represented by files;

[0064] At the same time, each type of service has different QoS requirements. Generally speaking, each type of service should have basic requirements for delay and transmission rate. For real-time services such as voice and video streams, there are not only requirements for transmission delay, but also requirements for jitter delay within a certain range to ensure the user experience. The present invention uses the following table to uniformly describe the possible QoS requirements of different types of services.

[0065] Table 1 Business QoS Requirement Description

[0066]

[0067] Among them, the service considering FDMA refers to allowing multiple different users to multiplex and transmit this service in the same time slot through different frequency bands. For high-bandwidth services such as video streams and files, the solution of the present invention does not consider using FDMA technology to multiplex time slot resources. However, for low-bandwidth services such as voice and short messages, it is considered that the time slot resources can be divided into M orthogonal channels.

[0068] The solution of the present invention hopes to flexibly allocate and dynamically adjust the TDMA time slot resources of the space-ground integrated communication system to meet the delay requirements and access rate requirements of different service types in Table 1; at the same time, considering to maximize the network throughput on the premise of meeting the minimum transmission delay and access rate requirements of various services, so that the system performance and capacity reach the joint best state.

[0069] 2. Frame Structure

[0070] The frame structure in the intelligent time slot allocation method of TDMA in the solution of the present invention is fixed. Each frame includes a random access time slot, an uplink control time slot, a service time slot, and a downlink broadcast time slot. The random access time slot is used for new nodes to join, the uplink control time slot is used for users to send time slot application information, and the downlink broadcast time slot is used for the base station to send time slot allocation strategies, synchronization signals, voice, data and other information. The service time slot needs to be applied by the user. The base station executes the intelligent time slot allocation method based on reinforcement learning of the solution of the present invention according to all request information to output a time slot allocation strategy, and distributes the service time slot allocation strategy to the corresponding users. Considering that each frame includes X time slots (each time slot is Lms), N1 access time slots, N2 uplink control time slots, and N3 downlink broadcast time slots, the number of service time slots is X - N1 - N2 - N3, as Figure 2 shown.

[0071] 3. Operation Processes of User Side and Base Station Side

[0072] In the space-ground integrated scenario, the main operations of the user side are divided into access, TDMA time slot resource request, and service transmission. The specific flow chart is as Figure 3As shown in the figure. After the initial network construction is completed, users who have not yet accessed the network send access request signaling in the random access time slot and apply for network access according to the user access algorithm (such as the CSMA algorithm). Users who have successfully accessed the network need to continuously listen for a period of time until they receive the uplink control time slot information allocated by the base station. According to the current service load situation, the user sends time slot request information (including the time slot demand, the type of service sent, the urgency of the message, the service waiting time, etc.) to the base station in the corresponding control time slot and waits for the base station to allocate a service time slot address for it. If the user receives the service time slot address broadcast by the base station for it, it will perform service data transmission in the corresponding time slot of the current frame; otherwise, it will continue to wait until the data transmission is completed.

[0073] In the TDMA time slot allocation algorithm, the base station needs to allocate corresponding service time slots based on all the time slot application requests received from current users, and announce the allocation results in the downlink broadcast time slot of each frame. The flowchart is as Figure 4 shown. Specifically, after the initial network construction is completed, the base station broadcasts the number of current idle time slots to users. Then the base station continuously listens for the time slot application information sent by the user side, and based on all the request information received within each frame and the request information that has not been processed before, executes the intelligent time slot allocation method based on reinforcement learning of the present invention, and broadcasts the time slot allocation result to the user, while updating the time slot allocation table and the number of idle time slots.

[0074] Based on the above discussion, the solution of the present invention focuses on how to intelligently and dynamically allocate the service time slot resources within each frame according to the user request information, maximize the system throughput under the condition of meeting the flexible requirements of different services, and achieve the efficient utilization of resources. In particular, the solution of the present invention considers an intelligent time slot allocation method based on reinforcement learning to achieve the adaptive matching of time slot resources and service load. The intelligent time slot allocation method based on reinforcement learning of the solution of the present invention is as follows:

[0075] II. Intelligent Time Slot Allocation Method Based on Reinforcement Learning

[0076] The solution of the present invention considers modeling the time slot allocation problem as a Markov decision process, and designs a time slot allocation algorithm based on reinforcement learning to intelligently adapt to the changes in the network environment and service load, and meet the QoS requirements of differentiated services.

[0077] 1. Reinforcement Learning Framework for Time Slot Allocation

[0078] Reinforcement learning enables an agent to continuously learn in a system to obtain the maximum reward. Since there is less information available from outside the system, the agent must accumulate its own experiences to perform self-learning in the system. It acquires knowledge and experience through self-learning and further improves its action plan to adapt to the environment. The three key factors of reinforcement learning are state, action, and environmental reward, with the aim of obtaining as much cumulative reward as possible.

[0079] The basic principle of reinforcement learning is as Figure 5 shown. When an agent completes a certain task, it first needs to interact with the system environment through an action. Under the combined action of the action and the environment, the agent reaches a new state S, and at the same time the environment gives an immediate reward R. Through multiple trials, the agent and the environment continuously interact to generate multiple data. By using these generated data to modify its own action selection, then interacting with the environment to generate new data, and continuously improving its behavior with the new data, and so on in a cycle, the agent can learn the optimal action to complete the corresponding task.

[0080] Based on the basic principle of reinforcement learning, the solution of the present invention establishes a slot allocation reinforcement model framework as Figure 6 shown. Among them, the agent is the airborne base station (or ground base station) in the space-ground integration, the environment is the slot resource situation in the current network, and the action is the slot allocation strategy after receiving the service slot request. The solution of the present invention needs to determine how to dynamically and intelligently allocate idle slots at the base station end according to all the received slot request information, and maximize the network throughput on the basis of satisfying the service requirements as much as possible.

[0081] 2. MDP Model

[0082] The solution of the present invention considers modeling the TDMA slot allocation in the air-ground integration scenario as a Markov decision process (MDP), and its specific modeling process is as follows.

[0083] The Markov decision process is a cyclic process in which an agent takes an action to change its own state to obtain a reward and interact with the environment. The Markov decision process is usually represented by a five-tuple (S, A, P, R, γ), where S represents the state of the agent, A represents the set of actions that the agent may choose, P is the state transition probability, R represents the reward function, and γ is the discount factor. Assume that the agent is in state S t ∈S, and at this time the agent takes action at belongs to A and transfers to the next state S with probability P t+1 and receives an immediate reward R from the environment t The agent repeats this process in each decision cycle, and its goal is to find the optimal policy to obtain the maximum cumulative reward.

[0084] Consider that there are four types of services in the network, namely voice, short message, video stream and file, as shown in Table 1. To meet the delay requirements of the services, each user needs to calculate the number of time slots it needs to apply for its service. In the solution of the present invention, we consider that each user determines the number of frames flamenum it needs to last by calculating the remaining time restdelay according to the delay requirement delay of the service, the traffic load of the service, the transmission rate velocity of the service, the waiting time twait of the service, the frame duration flamelength, and the duration slotlength of each time slot. Finally, the minimum number of service time slots slot that the service needs to transmit in one frame is obtained through the total number of time slots slotsum required by the service. The specific calculation method is as follows.

[0085] restdelay = delay - twait (1)

[0086]

[0087]

[0088]

[0089] Note that when the flamenum calculated by Equation (2) is 0, it means that this service must be transmitted within a certain time in this frame. At this time, consider whether there is enough idle space in this frame for this service to send. If there is, continue to send time slot request information for this service. If not, discard it.

[0090] In the MDP model of the solution of the present invention, it is considered that the number of time slots applied by each user for each service is fixed, that is, the number of time slots requested for voice service, short message service, file service, and video stream service is fixed at t1, t2, t3, t4 each time, where the values of t1, t2, t3, t4 can be calculated with reference to Equations (1)-(4); that is, it can be calculated according to the actual physical transmission rate. Considering the change of traffic volume and the possible waiting time of the service, therefore, consider defining the service with too long waiting time, too large traffic volume, and almost not meeting the delay as an emergency service, and define its corresponding emergency time slot request number. For example, the emergency time slot request numbers for voice service, short message service, file service, and video stream service can be set as n1, n2, n3, n4, where the corresponding n i must be greater than ti It is possible to consider n i = t i + 1 or n i = t i + 2. It is also possible to consider having multiple emergency time slot requests for certain services (such as files, video streams), making the requests for time slots more flexible.

[0091] Based on the above discussion, the MDP model of the solution of the present invention is defined as follows:

[0092] 1) State: S = {X1, Y}, where X1 is the number of idle service time slots, y i1 is the number of users applying for a fixed number of time slots for the i-th service, y i2 is the number of users applying for emergency time slots for the i-th service, i = 1,.., 4.

[0093] 2) Action: where Z i1 is the number of fixed time slots allocated to the i-th service, Z i2 is the number of emergency time slots allocated to the i-th service, and there is mod(Z i1 , t i ) = 0, mod(Z i2 , n i ) = 0, i = 1,.., 4. At the same time, there is 0 ≤ Z i1 ≤ min{X1, t i * y i1}, 0 ≤ Z i2 ≤ min{X1, n i * y i2}, ∑ ij Z ij ≤ X1, i = 1,…, 4.

[0094] 3) Reward function: For voice services, since it supports FDMA, when it is allocated Z 11 time slots, M * Z 11 / t1 is the number of users with successful access to the fixed voice time slots applied for, or the number of successful accesses to the fixed time slot voice services. Therefore, the access success rate of the fixed time slot voice service is: M * Z 11 / (y 11 * t1), and the access success rate of the emergency time slot voice service is: M * Z 12 / (y 12 * n1). Thus, the access success rate of the voice service is obtained as:

[0095]

[0096] where α 11, α 12 It depends on the importance between the voice service applying for the fixed time slot and the voice service applying for the emergency time slot.

[0097] Similarly, the access success rate of other services can be calculated. Note that files and video streams do not support FDMA:

[0098]

[0099]

[0100]

[0101] Thus, our reward function can be defined as:

[0102] r = ∑ i β i ·r i (9)

[0103] where β i depends on the requirements of the access success rate of each service. At the same time, we also need to consider the constraint condition of ∑ ij Z ij ≤ X1. Therefore, the reward function can be improved as the following formula:

[0104]

[0105] 4) Set of transition probabilities: The set of transition probabilities is represented by . It represents the probability that when the Agent is in state s and executes action a, it transfers to state s′.

[0106] 5) Discount factor γ: γ indicates the importance of future rewards relative to long-term rewards.

[0107] 6) State value function: The state value function indicates the impact of the current policy π on future benefits at time t and is defined as follows:

[0108]

[0109] 7) State-action value function: The state-action value function is used to represent the long-term benefits of taking a certain action in a certain state. Its definition is as follows:

[0110]

[0111] The decision-making goal of the MDP process is to find an optimal policy such that

[0112]

[0113] Correspondingly, the optimal action-value function is the action-value function that is the largest among all policies, i.e.,

[0114]

[0115] From the optimal state-value function and the optimal action-value function, the Bellman optimal equation can be obtained:

[0116]

[0117] Q * (s,a) = R(s,a) + γ∑ s′∈S p(s′|s,a)max a‘ Q * (s′,a′) (16)

[0118] 3. Finding the Optimal Action in the MDP Based on the Actor-Critic (AC) Algorithm

[0119] The solution of the present invention considers searching for the optimal action in the above MDP model based on the Actor-Critic (AC) algorithm. Specifically, assume that the MDP process starts from the initial state s t ∈S, and a series of actions are executed according to the policy π to form a set of state-action sequences: κ ~ {s t ,a t ,s t+1 ,a t+1 ,...,s t+T ,a t+T}. The optimization objective of the MDP model is to find the optimal policy π and optimize the cumulative return value from time t to t + T Since the definition of the immediate return is related to service access, the optimization objective of the cumulative return is consistent with the aforementioned requirement of maximizing the service. Therefore, the optimization objective of the present invention is to maximize the expected value of the cumulative return of this process, and the objective function can be written as:

[0120]

[0121] For the above optimization objective, the Actor-Critic (AC) algorithm is used for solution. Actor-Critic is an algorithm that combines the advantages of two reinforcement learning algorithms based on value and policy. Its algorithm framework is as Figure 7As shown in the figure. It is mainly divided into the Actor part and the Critic part. Actor-Critic learning approximates the value function and the policy function. Among them, the policy estimation is realized by the Actor part through gradient descent learning using the policy gradient estimation method; while the value function estimation is realized by the Critic part using the TD learning algorithm. For the state s, the Actor selects the action a according to the current policy. After the state s receives the action a, it transfers to the state s+1, and at the same time generates a reward signal r. The state s and the reward signal r are used as the inputs of the Critic, and its output is the estimation of the value function, and a TD error signal is generated, which is used for the update learning of the Critic and the Actor networks, and evaluates the selected action to correct the action selection strategy of the Actor.

[0122] In the reinforcement learning model of AC, it is necessary to update the policy gradient of the Actor. Its core idea is to consider a class of policies with adjustable parameters, calculate the gradient of the performance index function with respect to the policy parameters, and then adjust the parameters according to the direction of the gradient to improve the policy. The performance index function U(π) is a function of the policy and is a scalar to measure the quality of the policy. The larger the value of this function, the better the policy. Therefore, our goal is to maximize the total expected return:

[0123]

[0124] Since we are pursuing the maximum value, the policy gradient method here is a gradient ascent method. In the policy gradient method, the increment of the policy weight vector is approximately proportional to the direction of the gradient:

[0125]

[0126] where β∈R represents the positive step size parameter, is the gradient of the performance index function with respect to the policy weight vector. In order to approximate the policy parameters, first, the gradient needs to be calculated. However, it is very difficult to obtain the functional relationship U(π), especially when the state transition probability of the system is unknown, the analytical formula of the function U(π) cannot be obtained, so the calculation cannot be carried out. To solve this problem, the proposed solution of the present invention presents an equivalent representation method of the gradient as follows:

[0127]

[0128] where represents the probability of reaching the state s when the policy is π in the case of the initial state S0.

[0129] In the AC reinforcement learning model, the role of the evaluator Critic is to estimate the state value function and make it more and more accurate. Through the estimation of the state value function by Critic, the iterative update of the policy by Actor can be made more effective. In the original reinforcement learning framework, due to the discreteness and small dimension of the state set, the state value can be updated and recorded through a table. However, in the time slot allocation MDP model we are concerned with, the state space is large and it is very difficult to store and update it in tabular form. Therefore, the update of the state value can only be approximated by the state value function. The commonly used approximation methods are linear approximation and non-linear approximation. Compared with non-linear approximation, linear approximation is simple and converges faster. Therefore, the proposed solution of this invention uses a linear function to approximate the state value function, which is expressed as follows:

[0130] V θ (s) = θ T φ(s) (21)

[0131] where φ(s) is the feature vector at state s, and θ is the parameter vector. After linear approximation by the parameter, the update of the value function is mainly through iterative update of the parameter vector.

[0132] To effectively update the parameter θ, the TD (temporal difference) deviation between the estimated value and the true value of the state value is introduced:

[0133] δ t = V π (s t ) - V θ (s t ) (22)

[0134] where V π (s t ) = r t+1 + γV θ (s t+1 ), which is calculated according to the typical bootstrapping method in reinforcement learning theory. The goal of Critic is to make the approximation of the state value function more and more accurate so as to guide the Actor policy optimization, which is equivalent to minimizing the TD deviation between the estimated value and the true value of the state value. This optimization goal can be expressed as:

[0135]

[0136] Using the gradient descent method to update θ in the optimization direction, we have:

[0137]

[0138] where α critic is the learning rate for updating the state value function.

[0139] 4. Flow of the time slot allocation algorithm based on reinforcement learning

[0140] Based on the above discussion, the intelligent time slot allocation method of this solution based on reinforcement learning is as Figure 8 shown, and the specific process is as follows:

[0141] (1) Establish an intelligent time slot allocation neural network model based on deep reinforcement learning and initialize the model parameters;

[0142] (2) According to the Markov decision model MDP and the time slot status request information of all users, collect the status, actions, and reward information of the time slots in the network, and use the status, actions, and reward information of the time slots in the network as model training data;

[0143] (3) Use the model training data to train the intelligent time slot allocation neural network model based on deep reinforcement learning, and search the global action space based on the AC algorithm, and output the time slot allocation strategy that maximizes the return function as the optimal time slot allocation strategy;

[0144] (4) According to the optimal time slot allocation strategy, extract the total number of time slots allocated to each service; obtain the number of users that can be allocated according to the time limit request information of each service; and according to the service arrival time, preferentially allocate time slots to users with longer waiting times to obtain the time slot allocation result of each user;

[0145] (5) The base station broadcasts the time slot allocation result of each user and updates the time slot status according to the time slot allocation result.

[0146] Embodiment 1

[0147] As Figure 9 shown, an intelligent time slot allocation method for the space-ground integrated scenario of the present invention includes:

[0148] Different users send time slot request information to the base station (air base station or ground base station) in real time according to service requirements, and the service requirements include service load requirements, service type requirements, and service delay requirements;

[0149] The base station (air base station or ground base station) based on the received time slot request information of all users and the time slot status information of the current network, uses the intelligent time slot allocation method based on reinforcement learning to allocate time slots to the time slot request information of all users to obtain the time slot allocation strategy of the users; and sends the obtained time slot allocation strategy to the corresponding users, and at the same time updates the time slot status information;

[0150] And the base station (air base station or ground base station) periodically updates and maintains the time slot status information and time slot request information of all users within the current scope of the base station.

[0151] Further, the specific steps of the intelligent time slot allocation method based on reinforcement learning for allocating time slots for the time slot request information of all users are as follows:

[0152] Step A, build a Markov decision model MDP based on time slot allocation to output a time slot allocation strategy, and define the state, action, reward function, set of transition probabilities, discount factor, state value function, and state-action value function of the Markov decision model MDP; wherein, in the Markov decision model MDP, the base station is used as an entity to autonomously collect environmental state information and determine the time slot allocation strategy according to the time slot request information of users.

[0153] The state of the Markov decision model MDP is the time slot request information of users and the current available time slot state information, and the action of the Markov decision model MDP is the number of time slots allocated to each service; the action vector of the Markov decision model MDP constitutes the action space, and the Markov decision model MDP uses the reward function to evaluate actions. The optimization goal is to maximize the total number of service accesses under the condition of meeting different service requirements, so as to maximize the network throughput.

[0154] Step B, build a time slot allocation algorithm based on reinforcement learning in the Markov decision model MDP, search for the best action that maximizes the global reward function, and obtain the optimal time slot allocation strategy. As Figure 8 shown, it specifically includes:

[0155] (1) Establish an intelligent time slot allocation neural network model based on deep reinforcement learning and initialize the model parameters;

[0156] (2) According to the Markov decision model MDP and the time slot state request information of all users, collect the state, action, and reward information of time slots in the network, and use the state, action, and reward information of time slots in the network as model training data;

[0157] (3) Use the model training data to train the intelligent time slot allocation neural network model based on deep reinforcement learning, and search the global action space based on the AC algorithm, and output the time slot allocation strategy that maximizes the reward function as the optimal time slot allocation strategy;

[0158] (4) According to the optimal time slot allocation strategy, extract the total number of time slots allocated to each service; obtain the number of allocable users according to the time limit request information of each service; and according to the service arrival time, preferentially allocate time slots to users with longer waiting times to obtain the time slot allocation result of each user; the calculation formula for the number of allocable users is:

[0159] a. If the service type does not support frequency division multiplexing (such as services like video and files), the number of users that can be allocated for this service is = Z / t, where Z is the number of time slots allocated to a certain service type obtained from the time slot allocation algorithm based on reinforcement learning, and t is the fixed number of time slots required to transmit this service;

[0160] b. If the service type supports frequency division multiplexing (such as services like voice and short messages), the number of users that can be allocated for this service is = (Z / t)*M, where Z is the number of time slots allocated to a certain service type obtained from reinforcement learning, t is the fixed number of time slots required to transmit this service, and M is the maximum number of users that the same time slot supports for the transmission of the same service in the frequency division multiplexing technology.

[0161] (5) The base station broadcasts the time slot allocation results of each user and updates the time slot status according to the time slot allocation results.

[0162] Embodiment 2

[0163] As Figure 10 shown, the difference between this embodiment and Embodiment 1 is that this embodiment further provides an intelligent time slot allocation system for the space-ground integrated scenario, and this system supports an intelligent time slot allocation method for the space-ground integrated scenario described in Embodiment 1; this system includes:

[0164] The user side is used for different users to access the system and send time slot request information to the base station (air base station or ground base station) in real time according to service requirements;

[0165] The base station side is used for, based on the received time slot request information of all users and the time slot status information of the current network, using an intelligent time slot allocation method based on reinforcement learning to perform time slot allocation on the time slot request information of all users to obtain the time slot allocation strategy for the users; and sending the obtained time slot allocation strategy to the corresponding users, and at the same time updating the time slot status information; and the base station (air base station or ground base station) periodically updates and maintains the time slot status information and time slot request information of all users within the current scope of the base station.

[0166] Specifically, the specific execution process of the user side is as follows:

[0167] After the initial network construction is completed, the unconnected users send connection request signaling in the random access time slot and apply for connection according to the user access algorithm;

[0168] The successfully connected users need to continuously listen for a period of time until they receive the uplink control time slot information allocated by the base station for this user;

[0169] According to the current service load conditions, the user sends slot request information to the base station in the corresponding control time slot and waits for the base station to allocate a service time slot address for the user; the slot request information includes slot demand, the type of service sent, message urgency, service waiting time, etc.

[0170] If the user receives the service time slot address allocated by the base station through broadcast, the user will perform service data transmission in the corresponding time slot of the current frame; otherwise, the user will continue to wait until the data transmission is completed.

[0171] Specifically, the base station side is further configured to broadcast the current number of idle time slots to the user after the initial network construction is completed, and continuously listen for the slot request information sent by the user side.

[0172] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0173] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0174] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0175] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are performed on the computer or other programmable apparatus to generate a computer-implemented process, thereby providing instructions for implementing the steps of the process Figure 1 in one process or a plurality of processes and / or boxes Figure 1 or steps for implementing the functions specified in one box or a plurality of boxes.

[0176] The specific embodiments described above have further elaborated on the objectives, technical solutions, and beneficial effects of the present invention. It should be understood that the above description is only for the specific embodiments of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. An intelligent time slot allocation method for the space-ground integrated scenario, characterized in that The method includes: Different users send slot request information to the base station in real time according to service requirements, where the service requirements include service load requirements, service type requirements, and service delay requirements; Based on the received slot request information of all users and the slot status information of the current network, the base station uses an intelligent slot allocation method based on reinforcement learning to allocate slots for the slot request information of all users to obtain the slot allocation strategy for the users; and the obtained slot allocation strategy is sent to the corresponding users, and at the same time, the slot status information is updated; The specific steps for the intelligent slot allocation method based on reinforcement learning to allocate slots for the slot request information of all users are as follows: Build a Markov decision model MDP based on slot allocation, and define the state, action, reward function, transition probability set, discount factor, state value function, and state-action value function of the Markov decision model MDP; among them, in the Markov decision model MDP, the base station is used as an entity to autonomously collect environmental state information and determine the slot allocation strategy according to the slot request information of the users; Build a slot allocation algorithm based on reinforcement learning in the Markov decision model MDP, and search for the best action that maximizes the global reward function to obtain the optimal slot allocation strategy; The state of the Markov decision model MDP is the slot request information of the users and the currently available slot status information, and the action of the Markov decision model MDP is the number of slots allocated to each service; the action vector of the Markov decision model MDP constitutes the action space, and the Markov decision model MDP uses the reward function to evaluate the action, and the optimization goal is to maximize the total number of service accesses under the condition of meeting different service requirements, so as to maximize the network throughput; According to the state-action value function, obtain the optimal action value function Q * (s,a), and the formula is: where Q π (s,a) is the state-action value function; From the optimal state value function and the optimal action value function, the Bellman optimal equation is obtained: where, V * (s) is the optimal state value function, R(s, a) is the reward function, γ is the discount factor, p(s′|s, a) is the probability of transitioning to state s′ after taking action a in state s, and V * (s′) is the optimal state value function corresponding to s′. max a‘ Q * (s′, a′) is the maximum value in the optimal action value function obtained for different actions a′; The objective function MaxU(π) of the optimization goal is: In the formula, is the cumulative return value optimized for a period from t to t + T; k is the decision round / cycle, and π ε (k) is the corresponding ε-Greedy policy space at the k-th round of decision-making; E is the mathematical expectation; Use the Actor-Critic algorithm to solve the objective function; Build a slot allocation algorithm based on reinforcement learning in the Markov decision model MDP, and search for the best action that maximizes the global reward function to obtain the optimal slot allocation strategy, including: Establish an intelligent slot allocation neural network model based on deep reinforcement learning and initialize the model parameters; According to the Markov decision model MDP and the slot status request information of all users, collect the status, action, and reward information of the slots in the network, and use the status, action, and reward information of the slots in the network as model training data; Use the model training data to train the intelligent slot allocation neural network model based on deep reinforcement learning, and search the global action space based on the AC algorithm, and output the slot allocation strategy that maximizes the reward function as the optimal slot allocation strategy; According to the optimal slot allocation strategy, extract the total number of slots allocated to each service; obtain the number of allocable users according to the time limit request information of each service; and according to the service arrival time, give priority to allocating slots to users with a long waiting time to obtain the slot allocation result for each user; The base station broadcasts the time slot allocation results for each user and updates the time slot status according to the time slot allocation results.

2. The intelligent time slot allocation method for the space-ground integrated scenario according to claim 1, wherein This method further includes: The base station periodically updates and maintains the time slot status information and time slot request information of all users within the current scope of the base station.

3. The intelligent time slot allocation method for the space-ground integrated scenario according to claim 1, characterized in that The calculation formula for the number of allocable users is: a. If the service type does not support frequency division multiplexing, the number of allocable users for this service = Z / t, where Z is the number of time slots allocated to a certain service type obtained from the time slot allocation algorithm based on reinforcement learning, and t is the fixed number of time slots required to transmit this service. b. If the service type supports frequency division multiplexing, the number of allocable users for this service = (Z / t)*M, where Z is the number of time slots allocated to a certain service type obtained from reinforcement learning, t is the fixed number of time slots required to transmit this service, and M is the maximum number of users that the same time slot supports for the transmission of the same service in frequency division multiplexing technology.

4. The intelligent time slot allocation method for the space-ground integrated scenario according to claim 1, wherein The frame structure in the intelligent time slot allocation method based on reinforcement learning includes a random access time slot, an uplink control time slot, a service time slot, and a downlink broadcast time slot; Among them, the random access time slot is used for the joining of new nodes, the uplink control time slot is used for users to send time slot application information, and the downlink broadcast time slot is used for the base station to send time slot allocation strategies, synchronization signals, voice, data, and other information.

5. An intelligent time slot allocation system for the space-earth integrated scenario, characterized in that, This system supports an intelligent time slot allocation method for the space-ground integrated scenario as described in any one of claims 1 to 4; This system includes: A user side, which is used for different users to access the system and sends time slot request information to the base station in real time according to service requirements; A base station side, which is used to perform time slot allocation on the time slot request information of all users by using an intelligent time slot allocation method based on reinforcement learning based on the received time slot request information of all users and the time slot status information of the current network, to obtain the time slot allocation strategy for users; and send the obtained time slot allocation strategy to the corresponding users, and at the same time update the time slot status information; and the base station periodically updates and maintains the time slot status information and time slot request information of all users within the current scope of the base station.

6. The intelligent time slot allocation system for the space-ground integrated scenario according to claim 5, wherein The specific execution process of the user side is: After the initial network construction is completed, unconnected users send access request signals in the random access time slot and apply for access according to the user access algorithm; The successfully connected users continuously listen until they receive the uplink control time slot information allocated by the base station for this user; According to the current service load situation, the user sends time slot request information to the base station in the corresponding control time slot and waits for the base station to allocate a service time slot address for this user; the time slot request information includes the time slot demand, the service type sent, the urgency of the message, and the service waiting time; If the user receives the service time slot address broadcast by the base station for it, the user will perform service data transmission in the corresponding time slot of the current frame; otherwise, the user will continue to wait until the data transmission is completed.

7. An intelligent time slot allocation system for the space-ground integrated scenario according to claim 5, characterized in that, The base station side is also used to broadcast the current number of idle time slots to users after the initial network construction is completed and continuously listen to the time slot request information sent by the user side.

Citation Information

Patent Citations

  • Channel access method of multi-priority wireless terminal based on deep reinforcement learning

    CN113613339A

  • FDQL-based multi-dimensional resource collaborative optimization method in mobile edge network

    CN114143891A