Online multi-user scheduling method and apparatus based on reinforcement learning, and device and medium

By building a state node diagram and deep reinforcement learning network of the space-space integrated network, calculating feature vectors and action probability vectors, and generating a target scheduling model, the accuracy and load balancing problems of multi-user online scheduling in the space-space integrated network are solved, and the accuracy of user scheduling and the transmission efficiency of data flow are improved.

WO2025138094A1PCT designated stage expired Publication Date: 2025-07-03THE CHINESE UNIV OF HONG KONG (SHENZHEN)

Patent Information

Application Number
PCT/CN2023/143197
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-29
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

In the integrated air-space and earth network, the accuracy of online scheduling of multiple users is poor, especially because the load imbalance caused by the high-altitude platform to undertake multiple tasks at the same time affects the overall performance.

Method used

By constructing a state node diagram of the integrated air-space and earth network, using the deep reinforcement learning network to calculate feature vectors and action probability vectors, and combining real-time rewards to perform gradient updates, and generating a target scheduling model to optimize user scheduling.

Benefits of technology

It improves the accuracy and load balancing of user scheduling, maximizes the transmission efficiency of data flow, and solves the accuracy of multi-user online scheduling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2023143197_03072025_PF_FP_ABST
    Figure CN2023143197_03072025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of communications. Disclosed are an online multi-user scheduling method and apparatus based on reinforcement learning, and a device and a medium. The method comprises: constructing a state node diagram of a space-air-ground integrated network; on the basis of the state node diagram, calculating an eigenvector; on the basis of the eigenvector, generating an action probability vector; on the basis of the action probability vector, calculating an immediate reward; using the immediate reward to perform gradient updating on a deep reinforcement learning network, so as to obtain a target scheduling model; and using the target scheduling model to schedule users of the space-air-ground integrated network. According to the embodiments of the present invention, a state node diagram of a space-air-ground integrated network is constructed, the state node diagram can be used to calculate an immediate reward, and the immediate reward can be used to perform gradient updating on a deep reinforcement learning network, such that an optimal target scheduling model can be obtained; and users are scheduled by means of the target scheduling model, such that a data flow can be maximized while workloads are balanced, thereby improving the accuracy of user scheduling.
Need to check novelty before this filing date? Find Prior Art

Description

Online multi-user scheduling method, device, equipment and medium based on reinforcement learning Technical Field

[0001] The present invention relates to the field of communication technology, and in particular to an online multi-user scheduling method, device, equipment and medium based on reinforcement learning. Background Art

[0002] 6G wireless networks represent the next frontier in wireless communications, aiming to provide global connectivity while ensuring high-quality service across land, air, and sea. To achieve this ambitious vision, integrated air-space-ground networks have emerged as a key technology, offering unprecedented opportunities across various sectors. Compared to geosynchronous (GEO) and medium-Earth orbit (MEO) satellites, low-Earth orbit (LEO) satellites offer significant advantages, such as larger service areas and lower transmission latency. However, it is worth noting that LEO satellites orbit the Earth at high speeds, often resulting in intermittent connectivity with users on the ground. To address this issue, high-altitude platforms (HAPs) provide connectivity support in the stratosphere. Compared to traditional drone-assisted methods, HAPs can provide greater coverage and larger payloads. Therefore, stratospheric platforms are often used as supporting infrastructure, with users, satellites, and HAPs forming the space-ground-integrated network (SAGIN).

[0003] Despite SAGIN's many advantages, its collaborative mechanisms still face various challenges. For example, previous research explored the store-carry-forward mechanism for data collection in SAGIN and proposed an acceleration algorithm to reduce time complexity. However, they did not consider more practical online cases. In another study, the authors explored maximizing the benefits of the entire network by optimizing the online resource allocation of HAPs to users arriving gradually. Building on the above research, they further considered the online situation where multiple users arrive simultaneously. However, when HAPs take on multiple tasks simultaneously, the total task processing time increases and affects the overall performance of the HAPs. Therefore, load balancing needs to be considered, resulting in poor accuracy in online user scheduling.

[0004] Summary of the Invention

[0005] The present invention provides an online multi-user scheduling method, apparatus, device and medium based on reinforcement learning, the main purpose of which is to solve the problem of poor accuracy of multi-user online scheduling.

[0006] To achieve the above-mentioned objectives, the present invention provides an online multi-user scheduling method based on reinforcement learning, comprising: obtaining the current state set of users of an integrated air-space-ground-integrated network, and constructing a state node graph based on the current state set of users; using a pre-built deep reinforcement learning network to calculate the feature vector of the current state set of users based on the state node graph; generating an action probability vector of the current state set of users based on the feature vector; calculating the instant reward of the current state set of users based on the action probability vector; using the instant reward to perform gradient update on the deep reinforcement learning network to obtain a target scheduling model; and using the target scheduling model to schedule users of the integrated air-space-ground-integrated network.

[0007] The present invention also provides an online multi-user scheduling device based on reinforcement learning, including: a state node graph construction module, used to obtain the current state set of users of the integrated air-space-ground-integrated network, and construct a state node graph according to the current state set of users; a feature vector calculation module, used to use a pre-built deep reinforcement learning network to calculate the feature vector of the user's current state set according to the state node graph; an action probability vector generation module, used to generate the action probability vector of the user's current state set according to the feature vector; an instant reward calculation module, used to calculate the instant reward of the user's current state set according to the action probability vector; a deep reinforcement learning network update module, used to use the instant reward to perform gradient update on the deep reinforcement learning network to obtain a target scheduling model; a user scheduling module, used to use the target scheduling model to schedule users of the integrated air-space-ground-integrated network.

[0008] The present invention also provides an electronic device, which includes: a memory communicatively connected to at least one processor; wherein the processor is used to execute a computer program stored in the memory; the memory stores a computer program executable by at least one processor, and the computer program is executed by at least one processor so that the at least one processor can execute the above-mentioned online multi-user scheduling method based on reinforcement learning.

[0009] The present invention also provides a computer-readable storage medium storing a computer program, characterized in that when the computer program is executed by a processor, it implements any one of the above-mentioned online multi-user scheduling methods based on reinforcement learning.

[0010] The embodiment of the present invention proposes an online multi-user scheduling method based on reinforcement learning. By constructing a state node graph of the current state set of users in the air-space-ground integrated network, the graph structure of the state node graph can be used to more effectively extract state features and obtain more accurate feature vectors; by calculating the instant reward of the user's current state set through the feature vector, the reward value under different action probabilities can be calculated, thereby balancing the channel capacity and energy consumption while maximizing the data flow and improving the accuracy of user scheduling; by using the instant reward to perform gradient update on the deep reinforcement learning network, the optimal target scheduling model can be obtained, and by scheduling users through the target scheduling model, the data flow can be maximized while balancing the workload and improving the accuracy of user scheduling. Therefore, the online multi-user scheduling method, device, equipment and medium based on reinforcement learning proposed by the present invention can solve the problem of poor accuracy of online multi-user scheduling. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] FIG1 is a schematic diagram of a flow chart of an online multi-user scheduling method based on reinforcement learning provided by an embodiment of the present invention;

[0012] FIG2 is a schematic diagram of the structure of an air-ground integrated network provided by an embodiment of the present invention;

[0013] FIG3 is a schematic diagram of the principle of a deep reinforcement learning network provided by an embodiment of the present invention;

[0014] FIG4 is a schematic diagram of a process for calculating instant rewards according to an embodiment of the present invention;

[0015] FIG5 is a schematic diagram of data flow of different strategies at different user arrival rates according to an embodiment of the present invention;

[0016] FIG6 is a schematic diagram of workload imbalance according to different strategies provided by an embodiment of the present invention;

[0017] FIG7 is a schematic diagram of average data traffic under different strategies provided by an embodiment of the present invention;

[0018] FIG8 is a schematic diagram showing a comparison of competition ratios under different strategies provided by an embodiment of the present invention;

[0019] FIG9 is a functional module diagram of an online multi-user scheduling device based on reinforcement learning provided by one embodiment of the present invention;

[0020] FIG10 is a schematic diagram of the structure of an electronic device for implementing an online multi-user scheduling method based on reinforcement learning according to an embodiment of the present invention.

[0021] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION

[0022] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0023] The embodiment of the present application provides an online multi-user scheduling method based on reinforcement learning. The execution subject of the online multi-user scheduling method based on reinforcement learning includes but is not limited to at least one of the electronic devices such as the server, the terminal, etc. that can be configured to execute the method provided by the embodiment of the present application. In other words, the online multi-user scheduling method based on reinforcement learning can be executed by software or hardware installed on the terminal device or the server device. The server includes but is not limited to: a single server, a server cluster, a cloud server or a cloud server cluster, etc. The server can be an independent server, or it can be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.

[0024] 1 , which is a flow chart of an online multi-user scheduling method based on reinforcement learning according to an embodiment of the present invention.

[0025] S1. Obtain the user's current state set of the air-ground integrated network and construct a state node graph based on the user's current state set.

[0026] As shown in Figure 2, a space-air-ground integrated network (SAGIN) is an integrated network formed by users, satellites, and high altitude platforms (HAPs). In a SAGIN transmission scenario, multiple users with access to HAPs (high altitude platforms) first transmit data via user-to-HAP links (U2H). The HAPs then relay the data to LEO (low Earth orbit) satellites via HAP-to-satellite links (H2S). Finally, the satellites forward the information to designated satellite ground stations.

[0027] In one embodiment, the user current state set is the connection state between the user, satellite, and HAP at that time. The user, satellite, and HAP in the user current state set are used as nodes in a state node graph, and the current connection relationship between the user, satellite, and HAP is used as an edge to connect the nodes, thereby obtaining a state node graph.

[0028] In one embodiment, constructing a state node graph according to the user's current state set includes: extracting nodes and connection relationships between nodes from the user's current state set; and connecting the nodes according to the connection relationships to obtain the state node graph.

[0029] In one embodiment, the users, satellites, and HAPs in the user current state set are nodes, which can be extracted from the user current state set by predefined names of each user, satellite, and HAP, for example, node V s ={s t |s∈S,t∈T},V h ={h t |h∈H,t∈T},V u ={u t |u∈U,t∈T} represent the satellite, HAP and user node in the entire time range respectively; T is divided into T time slots, each with a length of t. uh ={(u t ,h t )|(u t ,h t )∈U2H,t∈T},A hs ={(h t ,s t )|(h t ,s t )∈H2S,t∈T} represents the edge links across all time periods, so the state node graph can be represented as G={A,V}, where vertex V=V u ∪V h ∪V s , edge link A=A uh ∪A hs .

[0030] S2. Use the pre-built deep reinforcement learning network to calculate the feature vector of the user's current state set based on the state node graph.

[0031] In one embodiment, a deep learning network (Deep Reinforcement Learning) is based on GNN (Graph Neural Networks) and Seq2Seq networks. The specific structure can be seen in Figure 3. The graph neural network can take the state node graph as input, and the hidden layer and point embedded neurons in the GNN can be used to calculate the embedding vector of each node, and the feature vectors of users, HAPs, and satellites as nodes in the state node graph can be extracted respectively.

[0032] In one embodiment, a pre-built deep reinforcement learning network is used to calculate the feature vector of the user's current state set based on the state node graph, including: using the graph neural network in the deep reinforcement learning network to perform graph encoding on the state node graph to obtain the embedding vector of each node in the state node graph; calculating the average embedding vector of the embedding vector, and vector splicing the embedding vector and the average embedding vector to obtain the feature vector of the user's current state set.

[0033] In one embodiment, a message passing neural network (MPNN) is used as an encoder in a graph neural network to effectively generate embedding vectors for newly arrived users, HAPs, and visible satellites at each time t. Each embedding vector is concatenated with the average embedding vector to obtain the feature vectors of the user, HAP, and satellite.

[0034] In one embodiment, the attributes of users, HAPs, and satellites are used as node features through a state node graph, and the utility of the connection is used as the edge weight. The graph structure of the state node graph is used to more effectively extract state features and obtain more accurate feature vectors.

[0035] S3. Generate an action probability vector of the user's current state set based on the feature vector.

[0036] In one embodiment, the action probability vector is a probability vector of different connection relationships between the user, the HAP, and the visible satellites in the user's current state set, and different action probability vectors correspond to different connection relationships.

[0037] In one embodiment, an action probability vector of the user's current state set is generated based on a feature vector, including: feature encoding the feature vector to obtain an encoded feature; feature decoding the encoded feature to obtain a decoded feature; and performing activation calculation on the decoded feature to obtain an action probability vector of the current state set.

[0038] As shown in Figure 3, feature encoding utilizes a pre-built encoder in the Seq2Seq network, followed by feature decoding using a pre-built decoder in the Seq2Seq network to derive decoded features for all available actions for each user. Finally, the output of each Seq2Seq decoder is activated through a softmax function to generate an action probability vector for all available actions. This action probability vector is used to update the environment in the integrated air-ground network to calculate the immediate reward corresponding to the subsequent action probability vector. This immediate reward is then used to update the parameters in the deep reinforcement learning network, resulting in a more accurate target scheduling model.

[0039] Specifically, the encoder and decoder can be word sequence encoders and decoders based on GRUs. GRUs (Gated Recurrent Units) are a variant of RNNs that use gates to record the current state of the sequence. The hidden layer has two gates: a reset gate and an update gate. These two gates together control how much information is updated about the current state. GRUs can address issues such as long-term memory and gradients in backpropagation, while reducing the need for gating and improving the accuracy and efficiency of action probability vector calculations.

[0040] S4. Calculate the immediate reward of the user's current state set based on the action probability vector.

[0041] In one embodiment, the immediate reward is the optimization target of the integrated space-air network under the current action probability in the current time period. That is, the optimization target is obtained by summing all HAPs taking into account the channel capacity of the U2H and H2S links in the t time slot, the energy capacity limit of the HAP, and the overall workload allocated on the HAP, and then calculating the immediate reward under the action probability.

[0042] In one embodiment, it is assumed that users can access HAPs via an Orthogonal Frequency Division Multiple Access (OFDMA) control system, thereby ignoring the interference between U2H links. is a binary variable, indicating that the h-th HAP accepts user u at the t-th time; similarly, represents the binary decision made by the satellite regarding the connection request of the h-th HAP in time t. In addition, each user can only connect to at most one HAP, so the channel capacity of the U2H and H2S links in time period t can be limited.

[0043] Specifically, the main energy consumption of HAP comes from its data transceiver, including the transmission power and received power The energy consumed by the HAP receiver at time t can be defined in a similar form to the energy consumption of the HAP transmitter.

[0044] Specifically, the total energy cost of the h-th HAP at time slot t is denoted as By summing up all the above costs and running costs In addition, each HAP can charge its battery by harvesting solar energy or radio frequency (RF) energy, so the total energy obtained by the HAP through charging depends on the charging rate.

[0045] In one embodiment, referring to FIG4 , calculating the instantaneous reward of the user's current state set based on the action probability vector includes: S41, determining the instantaneous action state of the user's state set based on the action probability vector; S42, constructing a channel capacity model and an energy consumption model based on the instantaneous action state; S43, constructing a target reward function based on the channel capacity model and the energy consumption model; S44, calculating the instantaneous reward of the user's current state set using the target reward function.

[0046] In one embodiment, an activation calculation is performed on the action probability vector to obtain the possible probability of each action. The action corresponding to the maximum probability is selected as the instantaneous action state of the user state set. For example, the connection relationship between the satellite and the HAP at time t, and the connection relationship between the user and the HAP are calculated based on the instantaneous action state to obtain the channel capacity and energy consumption. The channel capacity and energy consumption are constrained by the optimization goal of the ground-to-air-space-ground integrated network to obtain the target reward function.

[0047] In one embodiment, the channel capacity model can be expressed as:

[0048] in, represents the channel capacity between the i-th connection relationship and the j-th connection relationship in the instant action state within time t, represents the set of connection relations in the instant action state at time t, represents the uplink channel bandwidth between the i-th connection relationship and the j-th connection relationship within time t, represents the transmission power of the i-th connection relationship and the j-th connection relationship, They represent the transmission antenna gain during sending and receiving, tr represents sending, re represents receiving, and L represents the preset bus loss. represents the preset free path loss between the i-th connection relationship and the j-th connection relationship within time t, k B represents the Boltzmann constant, T noi represents the preset system noise temperature, and T represents the total time.

[0049] It should be noted that the connection relationship includes the H2S connection relationship between the satellite and the HAP in the immediate action state and the U2H connection relationship between the user and the HAP.

[0050] In one embodiment, transmit power, transmission antenna gain, bus loss, free path loss, system noise temperature, etc. are all preset constant values, and the channel capacity of the U2H and H2S links at time t is limited by the channel capacity function.

[0051] In one embodiment, the energy consumption model can be expressed as:

[0052] in, Expressed as energy consumption, α h Indicates the maximum discharge depth of the h-th high-altitude platform battery, represents the initial battery energy of the h-th high-altitude platform, represents the total energy obtained by charging the h-th high-altitude platform in the total time T, represents the set of h-th high-altitude platforms in time t.

[0053] In one embodiment, the workload is characterized by data size and computation size. The computational load of user u is measured as the number of CPU cycles required to process one bit of data. The data flow size of a task is measured in bits. It is assumed that each task transmission follows a time slot-based approach and that large data flows are broken down into small packets. Therefore, the total workload D distributed on HAP h can be defined as:

[0054] in, represents the workload of the h-th aerial platform in time t, represents the set of users u at time t, represents the data processing density of user u in time t, A vector representing the size of user u’s data at time t, It represents the binary variable that the h-th high-altitude platform accepts user u within time t, represents the set of h-th high-altitude platforms within time t, and T represents the total time.

[0055] Specifically, user scheduling aims to maximize the data flow of the integrated air-ground-space network. However, different user scheduling results may lead to unfair solutions. For example, one HAP is loaded with many tasks while others are idle. Workloads are characterized by data size and computation size. Therefore, it is necessary to consider workload balancing among multiple HAPs in the network to achieve stable operation of multiple HAPs and maximize data flow.

[0056] Specifically, the present invention can consider a logarithmic function, i.e., U(x) = log(x). The logarithmic function is a concave function, meaning that the slope decreases. This characteristic naturally promotes a certain degree of workload balancing, because the utility gain of concentrating all tasks on a single HAP is less than that of distributing tasks across multiple HAPs. Therefore, the logarithmic utility function can be used to sum over all HAPs within time t to obtain the optimization objective, i.e., the target reward function.

[0057] In one embodiment, the objective reward function can be expressed as:

[0058] in, represents the target reward function, P represents the optimization target of the target reward function, represents the set of users u at time t, represents the set of satellites s at time t, A vector representing the size of user u’s data at time t, It represents the binary variable that the h-th high-altitude platform accepts user u within time t, represents the binary decision made by the satellite at time t regarding the connection request from the h-th high-altitude platform, represents the set of h-th high-altitude platforms in time t, The vector representing the transmission speed of the h-th high-altitude platform receiving user u in time t, represents the channel capacity between the i-th connection relationship and the j-th connection relationship in time t, represents energy consumption, α h Indicates the maximum discharge depth of the h-th high-altitude platform battery, represents the initial battery energy of the h-th high-altitude platform, It represents the total energy obtained by charging the h-th high-altitude platform within the total time T, where T represents the total time.

[0059] In one embodiment, the objective reward function constrains the decision variables by connecting capacity. Reliable transmission is guaranteed during the U2H uplink process, and constraints are placed on data traffic and the energy budget of HAPs. Finally, the maximum value P of the objective reward function is defined as a mixed integer programming (MIP) problem, which is classified as an NP-hard problem. As the number of users increases, the computational complexity of the algorithm grows exponentially, making exhaustive enumeration and extended decision time extremely difficult. Since the task connection decision is based only on the current system state, the problem can be converted into a Markov decision process (MDP). The environment state, action, reward, and state transition probability are all defined in this section.

[0060] In one embodiment, under different actions, the air-ground integrated network has different environmental states. Specifically, the environmental state O can be Connect and get the state o of the space-ground integrated network environment at time t t , o t ∈O, where is a vector representing the size of the data, A vector representing the transmission speed, is the vector representing the residual energy of HAP, A vector representing satellite visibility.

[0061] In detail, the comprehensive action a in the air-ground integrated network is obtained through different actions A in the instant action state. t , a t ∈A, which can be expressed as a t =(ξ t ,ψ t ), where ξ t is a row vector indicating whether each user is connected to HAPs, ψ t is a row vector indicating whether each HAP is connected to the satellite; a value of 1 indicates connection and 0 indicates disconnection. The connection state is used to balance the workload between HAPs while maximizing the data flow size using the objective reward function; the instantaneous reward for the total time T is obtained by summing the rewards of all HAPs.

[0062] In one embodiment, by calculating the immediate reward, the reward value under different action probabilities can be calculated, so that the channel capacity and energy consumption can be balanced while maximizing the data flow, thereby improving the accuracy of user scheduling.

[0063] S5. Use the immediate reward to perform gradient update on the deep reinforcement learning network to obtain the target scheduling model.

[0064] In one embodiment, the network parameters in the deep reinforcement learning network are updated by the immediate reward until the value of the immediate reward reaches a maximum, that is, the target reward function converges, and a target scheduling model that can obtain the maximum immediate reward is generated.

[0065] In one embodiment, the immediate reward is used to perform a gradient update on the deep reinforcement learning network to obtain a target scheduling model, including: performing a gradient descent update on the parameters in the deep reinforcement learning network to obtain an updated deep reinforcement learning network; and iteratively updating the immediate reward according to the updated deep reinforcement learning network until the immediate reward converges to obtain the target scheduling model.

[0066] In one embodiment, gradient descent updates the derivative of the parameter at its current position, eventually converging to a local extreme point. With each parameter update, the corresponding curve becomes flatter, the corresponding derivative value decreases, and the update amplitude decreases, eventually settling at the extreme point and obtaining the parameters corresponding to the target scheduling model.

[0067] In one embodiment, after obtaining the updated deep reinforcement learning network, the above steps of calculating the immediate reward are repeated, and the immediate reward is iteratively updated until the value of the immediate reward no longer changes or changes within a certain range, that is, the immediate reward converges and reaches a maximum, thereby obtaining the target scheduling model.

[0068] In one embodiment, by performing gradient updates on the deep reinforcement learning network, an optimal target scheduling model can be obtained. By scheduling users through the target scheduling model, it is possible to maximize data flow while balancing workload and channel capacity, thereby improving the accuracy of user scheduling.

[0069] S6. Use the target scheduling model to schedule users of the integrated air-space-ground network.

[0070] In one embodiment, the target scheduling model is used to perform the above-mentioned action probability vector calculation module on the state node graph corresponding to the current state set of users in the air-space-ground integrated network to obtain a target probability vector. By selecting the action state corresponding to the maximum probability in the target probability vector as the connection state of the air-space-ground integrated network, the connection state between the user and the HAP is obtained. The connection state is used to determine whether the user is connected to the HAP, so as to schedule the users in the air-space-ground integrated network.

[0071] In one embodiment, a greedy algorithm and a gradient policy (GP) can be used as the baseline of the target scheduling model. Referring to FIG5 , the data flow (in Gbit) of the three strategies under different user arrival rates is shown. As can be seen from FIG5 , the performance of the GNN+Seq2Seq strategy of the present invention is significantly better than that of the greedy strategy.

[0072] Figure 6 compares the workload imbalance of the three strategies, with workload imbalance as the vertical axis. It can be seen that with the greedy strategy, the load imbalance between HAPs becomes more pronounced as the user arrival rate increases. The other two deep learning-based strategies effectively balance the workload between HAPs. The GNN+Seq2Seq strategy achieves the highest data transfer volume and the lowest workload imbalance.

[0073] This result can be further illustrated in Figure 7. When the cumulative probability is the same, the greedy strategy lacks the ability to select users with relatively high data traffic from a global perspective, resulting in a large change in the average data traffic.

[0074] In another embodiment, by comparing the three strategies, it can be found that the GNN+Seq2Seq strategy can better achieve user scheduling. For example, the performance of the NN+Seq2Seq strategy is represented by the competition ratio, where CR represents the competition ratio, which can be expressed as:

[0075] where R online and R * These are the target values ​​obtained by the online algorithm and the Gurobi algorithm, respectively. These values ​​are obtained using the same data and optimization objectives.

[0076] As shown in Figure 8, using the ratio of the average contention ratio to the minimum contention ratio as a metric, the performance of the traditional greedy strategy and the FFN+Seq2Seq strategy relative to the offline Gurobi solution degrades as the user arrival rate increases from 2 to 10, corresponding to an increase in the total number of users from 600 to 3000. However, the GNN+Seq2Seq strategy achieves an average CR of 0.91 and a minimum CR of 0.84 when λ = 10. This CR, close to 1, indicates that the performance of the online solution is comparable to that of Gurobi's offline solution.

[0077] In one embodiment, the present invention first develops a comprehensive model for integrated air-ground-space networks by considering various practical constraints, such as transmission and energy consumption, as well as workload models. An architecture comprising GNN and Seq2Seq modules is proposed to optimize the design goal of maximizing user data flow while ensuring workload balancing across HAPs. This architecture can efficiently handle multi-user requests and achieve comparable data flow performance to offline methods, enabling more accurate user scheduling in integrated air-ground-space networks.

[0078] FIG9 is a functional module diagram of an online multi-user scheduling device based on reinforcement learning according to an embodiment of the present invention.

[0079] The reinforcement learning-based online multi-user scheduling device 900 of the present invention can be installed in an electronic device. Depending on the functions implemented, the reinforcement learning-based online multi-user scheduling device 900 may include a state node graph construction module 901, a feature vector calculation module 902, an action probability vector generation module 903, an immediate reward calculation module 904, a deep reinforcement learning network update module 905, and a user scheduling module 906. A module of the present invention, also referred to as a unit, refers to a series of computer program segments that can be executed by an electronic device processor and can perform a fixed function, and is stored in the memory of the electronic device.

[0080] In this embodiment, the functions of each module / unit are as follows: a state node graph construction module 901 is used to obtain the current state set of the user of the integrated air-space-ground network, and construct a state node graph based on the user's current state set; a feature vector calculation module 902 is used to use a pre-built deep reinforcement learning network to calculate the feature vector of the user's current state set based on the state node graph; an action probability vector generation module 903 is used to generate the action probability vector of the user's current state set based on the feature vector; an immediate reward calculation module 904 is used to calculate the immediate reward of the user's current state set based on the action probability vector; a deep reinforcement learning network update module 905 is used to use the immediate reward to perform gradient update on the deep reinforcement learning network to obtain a target scheduling model; a user scheduling module 906 is used to use the target scheduling model to schedule users of the integrated air-space-ground network.

[0081] In detail, in one embodiment, each module in the online multi-user scheduling device 900 based on reinforcement learning adopts the same technical means as the online multi-user scheduling method based on reinforcement learning in the accompanying drawings when in use, and can produce the same technical effects, which will not be repeated here.

[0082] FIG10 is a schematic diagram showing the structure of an electronic device for implementing an online multi-user scheduling method based on reinforcement learning according to an embodiment of the present invention.

[0083] The electronic device 100 may include a processor 101 , a memory 102 , a communication bus 103 , and a communication interface 104 . It may also include a computer program stored in the memory 102 and executable on the processor 101 , such as an online multi-user scheduling program based on reinforcement learning.

[0084] In some embodiments, the processor 101 may be composed of an integrated circuit, for example, a single packaged integrated circuit, or a plurality of packaged integrated circuits with the same or different functions, including a combination of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 101 is the control core (Control Unit) of the electronic device, connecting the various components of the entire electronic device using various interfaces and lines, and executing programs or modules stored in the memory 102 (for example, executing an online multi-user scheduling program based on reinforcement learning, etc.), as well as calling data stored in the memory 102, to perform various functions of the electronic device and process data.

[0085] The memory 102 includes at least one type of readable storage medium, and the readable storage medium includes a flash memory, a mobile hard disk, a multimedia card, a card-type memory (for example, an SD or DX memory, etc.), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the memory 102 may be an internal storage unit of an electronic device, such as a mobile hard disk of the electronic device. In other embodiments, the memory 102 may also be an external storage device of the electronic device, such as a plug-in mobile hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card, etc. equipped on the electronic device. Furthermore, the memory 102 may also include both an internal storage unit of the electronic device and an external storage device. The memory 102 can not only be used to store application software and various types of data installed in the electronic device, such as the code of an online multi-user scheduling program based on reinforcement learning, but can also be used to temporarily store data that has been output or is to be output.

[0086] The communication bus 103 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. The bus may be divided into an address bus, a data bus, a control bus, etc. The bus is configured to enable communication between the memory 102 and at least one processor 101.

[0087] The communication interface 104 is used for communication between the above-mentioned electronic device and other devices, including a network interface and a user interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), which is generally used to establish a communication connection between the electronic device and other electronic devices. The user interface may be a display (Display), an input unit (such as a keyboard (Keyboard)), optionally, the user interface may also be a standard wired interface, a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, and an OLED (Organic Light-Emitting Diode, organic light-emitting diode) touch device, etc. Among them, the display may also be appropriately referred to as a display screen or a display unit, for displaying information processed in the electronic device and for displaying a visual user interface.

[0088] Figure 10 only shows an electronic device with components. Those skilled in the art will understand that the structure shown in Figure 10 does not constitute a limitation on the electronic device 100, and may include fewer or more components than shown in the figure, or combine certain components, or arrange the components differently.

[0089] For example, although not shown, the electronic device may further include a power source (such as a battery) for powering various components. Preferably, the power source may be logically connected to at least one processor 101 via a power management device, thereby implementing functions such as charge management, discharge management, and power consumption management via the power management device. The power source may further include any components such as one or more DC or AC power sources, a recharging device, a power failure detection circuit, a power converter or inverter, and a power status indicator. The electronic device may also include various sensors, Bluetooth modules, Wi-Fi modules, etc., which will not be described in detail here.

[0090] It should be understood that the embodiment is for illustration only and the scope of the patent application is not limited to this structure.

[0091] The present invention also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it can implement the online multi-user scheduling method based on reinforcement learning of any of the above embodiments. It should be noted that the computer-readable storage medium can be volatile or non-volatile. For example, the computer-readable medium may include: any entity or device that can carry computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disk, a computer memory, and a read-only memory (ROM). In the several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of modules is only a logical function division, and there may be other division methods in actual implementation.

[0092] Modules described as separate components may or may not be physically separate, and components shown as modules may or may not be physical units, that is, they may be located in one place or distributed across multiple network elements. Some or all of these modules may be selected to achieve the purpose of this embodiment based on actual needs.

[0093] In addition, the functional modules in various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or hardware plus software functional modules.

[0094] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.

[0095] Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the invention is defined by the appended claims rather than the foregoing description, and all changes that come within the meaning and range of equivalents of the claims are intended to be embraced therein. Any reference to a figure in a claim should not be construed as limiting the claim to which it relates.

[0096] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to achieve optimal results.

[0097] Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices recited in a system claim may also be implemented by a single unit or device through software or hardware. Terms such as "first" and "second" are used to indicate names and do not imply any particular order.

[0098] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. An online multi-user scheduling method based on reinforcement learning, characterized in that, The method includes: Obtaining the current state set of users in the space-air-ground integrated network, and constructing a state node graph according to the current state set of users; Using a pre-constructed deep reinforcement learning network to calculate the feature vector of the current state set of users according to the state node graph; Generating an action probability vector of the current state set of users according to the feature vector; Calculating the immediate reward of the current state set of users according to the action probability vector; Using the immediate reward to perform gradient update on the deep reinforcement learning network to obtain a target scheduling model; Using the target scheduling model to schedule users in the space-air-ground integrated network.

2. The online multi-user scheduling method based on reinforcement learning according to claim 1, wherein, The constructing the state node graph according to the current state set of users includes: Extracting nodes and the connection relationships between the nodes from the current state set of users; Connecting the nodes according to the connection relationships to obtain the state node graph.

3. The online multi-user scheduling method based on reinforcement learning according to claim 1, wherein The using a pre-constructed deep reinforcement learning network to calculate the feature vector of the current state set of users according to the state node graph includes: Using the graph neural network in the deep reinforcement learning network to perform graph encoding on the state node graph to obtain the embedding vector of each node in the state node graph; Calculating the average embedding vector of the embedding vectors, and performing vector splicing on the embedding vectors and the average embedding vector to obtain the feature vector of the current state set of users.

4. The online multi-user scheduling method based on reinforcement learning according to claim 1, characterized in that The generating an action probability vector of the current state set of users according to the feature vector includes: Performing feature encoding on the feature vector to obtain encoded features; Performing feature decoding on the encoded features to obtain decoded features; Performing activation calculation on the decoded features to obtain the action probability vector of the current state set.

5. The online multi-user scheduling method based on reinforcement learning according to claim 1, characterized in that, The calculating the immediate reward of the current state set of users according to the action probability vector includes: Determining the immediate action state of the user state set according to the action probability vector; Constructing a channel capacity model and an energy consumption model according to the immediate action state; The described channel capacity model can be expressed as: Among them, Denote the channel capacity of the i-th connection relationship and the j-th connection relationship in the instant action state within time t. Denote the set of connection relationships in the instant action state at time t, Indicates the uplink channel bandwidth between the i-th connection relationship and the j-th connection relationship within time t. Indicates the transmission power of the i-th connection relationship and the j-th connection relationship. respectively represent the transmission antenna gains during transmission and reception, tr represents transmission, re represents reception, and L represents the preset bus loss Denote the preset free space loss of the i-th connection relationship and the j-th connection relationship within time t, k B Denote the Boltzmann constant, T noi Denote the preset system noise temperature, and T denotes the total time; The energy consumption model can be expressed as: Among them, Denoted as energy consumption, α h Denotes the maximum depth of discharge of the battery of the h-th high-altitude platform, Represents the initial battery energy of the h-th high-altitude platform, denotes the total energy obtained by the h-th high-altitude platform through charging within the total time T, Denote the set of the h-th high-altitude platform within time t; Constructing a target reward function according to the channel capacity model and the energy consumption model; Using the target reward function to calculate the immediate reward of the current state set of users.

6. The online multi-user scheduling method based on reinforcement learning according to claim 5, characterized in that The target reward function can be expressed as: Among them, denote the target reward function, and P denotes the optimization objective of the target reward function, Denotes the set of users u within time t, Denote the set of satellites s within time t, A vector representing the data size of user u within time t A binary variable indicating that the h-th high-altitude platform receives user u within time t, Denotes the binary decision made by the satellite regarding the connection request to the h-th high-altitude platform at time t, Denote the set of the h-th high-altitude platform within time t, A vector representing the transmission speed of the h-th high-altitude platform received from user u within time t, Denote the channel capacity between the i-th connection relationship and the j-th connection relationship within time t. Denote the energy consumption, α h Denote the maximum discharge depth of the h-th high-altitude platform battery, Indicates the initial battery energy of the h-th high-altitude platform, Denote the total energy obtained by the h-th high-altitude platform through charging within the total time T, where T represents the total time.

7. The online multi-user scheduling method based on reinforcement learning according to claim 1, characterized in that, The using the immediate reward to perform gradient update on the deep reinforcement learning network to obtain a target scheduling model includes: Performing gradient descent update on the parameters in the deep reinforcement learning network to obtain an updated deep reinforcement learning network; Performing iterative update on the immediate reward according to the updated deep reinforcement learning network until the immediate reward converges to obtain a target scheduling model.

8. An online multi-user scheduling device based on reinforcement learning, characterized in that, The device includes: A state node graph construction module, configured to obtain the current state set of users in the space-air-ground integrated network, and construct a state node graph according to the current state set of users; A feature vector calculation module, configured to use a pre-constructed deep reinforcement learning network to calculate the feature vector of the current state set of users according to the state node graph; An action probability vector generation module, configured to generate an action probability vector of the current state set of users according to the feature vector; An immediate reward calculation module, configured to calculate the immediate reward of the current state set of the user according to the action probability vector; A deep reinforcement learning network update module, configured to perform gradient update on the deep reinforcement learning network by using the immediate reward to obtain a target scheduling model; A user scheduling module, configured to use the target scheduling model for the users of the space-air-ground integrated network.

9. An electronic device, characterized in that, The electronic device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The processor is configured to execute a computer program stored on the memory; The memory stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, the at least one processor is enabled to execute the online multi-user scheduling method based on reinforcement learning according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the online multi-user scheduling method based on reinforcement learning according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Software-defined space-air-ground integrated network routing optimization method based on deep reinforcement learning

    CN114221691A

  • Comprehensive benefit-oriented resource intelligent collaborative scheduling method in space-ground-air integrated network

    CN114698118A

  • Multi-user scheduling method and system based on reinforcement learning for 5g IoT system

    US20230345451A1

Cited By

  • Multi-dimensional resource management joint optimization method based on wireless edge network

    CN120835006A