Access control and resource allocation method for space-air-ground integrated network
By adopting greedy algorithms and multi-objective deep reinforcement learning methods in the integrated air-space-ground network, user access control and channel allocation are optimized, which solves the shortcomings of access control and resource allocation in existing technologies and improves system performance and user experience.
Patent Information
- Application Number
- CN202411034154.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-30
- Publication Date
- 2025-10-14
- Estimated Expiration
- 2044-07-30
AI Technical Summary
Existing research on integrated air-ground-space networks lacks methods that simultaneously consider access control and resource allocation. Especially in multi-service scenarios, it is unable to flexibly allocate channels and ignores the differences in demand for different services, resulting in insufficient system performance and user experience.
A method based on greedy algorithm and multi-objective deep reinforcement learning is proposed. By optimizing user access control and channel allocation, combining service rate quality, service delay quality and switching overhead, frequency division multiplexing is used for communication, and low-orbit satellite and UAV resources are utilized to achieve multi-objective optimization.
It improves the communication resource utilization efficiency of the integrated air-space-ground network, enhances the user service experience quality, reduces system overhead, and adapts to the needs of different businesses.
Smart Images

Figure CN119031385B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of communication technology, and particularly relates to a space-air-ground integrated network access control and resource allocation method based on multi-target deep reinforcement learning. BACKGROUND
[0002] As an integration of satellite systems, aerial networks and terrestrial communications, space-air-ground integrated networks have become a new network architecture in the past few years and have attracted extensive research interest in academia. Compared with traditional terrestrial networks or satellite networks, the dynamic change of the topology structure of space-air-ground integrated networks will lead to frequent triggering of switching, greatly increasing the signaling overhead and making it difficult to obtain the best performance. Therefore, effective access control in space-air-ground integrated networks is of great significance.
[0003] With the increasing development and popularity of mobile devices, many emerging multimedia services have been born, such as voice recognition, automatic driving, etc. Therefore, space-air-ground integrated networks not only need to provide global seamless communication coverage for users, but also need to provide various service support for users. In real-world scenarios, various actual services and applications have different requirements in terms of throughput, delay, etc., and space-air-ground integrated networks need to efficiently allocate resources according to the needs of users, so as to improve the service quality of experience of users.
[0004] Under the architecture of space-air-ground integrated network, the literature [1] introduces high-altitude platform to cover hotspots. In this context, in order to maximize the throughput, a dynamic multi-beam joint resource optimization method based on low earth orbit satellite motion prediction is proposed, which uses Lagrange dual method and Karush-Kuhn-Tucker (KKT) condition to find the optimal solution. Literature [2] proposes a dynamic channel reservation strategy for multi-service low earth orbit satellite communication system based on deep Q network, and at the same time, the dynamic channel reservation problem of multiple services is represented as a reinforcement learning task. However, the channel allocation scheme of this scheme only includes two kinds of minimum bandwidth and maximum bandwidth, and cannot flexibly allocate channels. Literature [3] proposes a switching method using adaptive learning rate DQN framework (DQN-ALrM), which uses DQN algorithm to solve the optimization target, and uses ALrM method to intelligently adjust the learning rate, optimizes the convergence speed, switching rate, call failure rate and multi-index quality of service (QoS) performance. Literature [4] points out that dynamic resource allocation is the key technology to improve the network performance in resource-limited satellite systems, and proposes a framework based on deep reinforcement learning, which uses a kind of image tensor reconstruction method based on system environment to extract the space-time characteristics of traffic, and verifies the effectiveness of the algorithm in time-varying scene. Literature [5] proposes a user-driven deep reinforcement learning scheme, in which the centralized agent deployed at the backhaul end is responsible for training the parameters of the deep Q network, so that each user can make its own access decision independently according to the trained parameters, in order to improve the long-term throughput of the system and avoid frequent switching between non-ground nodes. However, the above literatures only consider the switching problem or channel allocation problem under the architecture of space-air-ground integrated network, lack of research on access control and resource allocation at the same time, and the existing research ignores the different needs under different services, and lacks research on switching and channel allocation under multi-service classification.
[0005] Different from existing works, (1) the present invention considers a scenario under a space-air-ground integrated network, various services generated by users can be transmitted and processed under the satellite beam and unmanned aerial vehicle covering the user. The downlink of the satellite beam and the unmanned aerial vehicle is equally divided, and different sizes of bandwidth can be allocated according to the service demand of the user in the transmission process. (2) The present invention considers the demand of different services at the same time, considers the service rate quality, service delay quality and switching overhead of the system, optimizes the three targets at the same time, improves the service experience of the user and reduces the system overhead. (3) The present invention proposes two algorithms to solve the proposed system and optimization target. The first is a solution based on the greedy algorithm, the total optimization target function is obtained by weighted sum of the optimization target, then the local optimal strategy under the current state is adopted when each user performs access control and channel allocation decision, so as to approximately obtain the global optimal solution; the second is a solution based on multi-objective deep reinforcement learning, first the access control and channel allocation problem is converted into a multi-objective Markov decision process, then a multi-objective deep Q learning algorithm is proposed through deep reinforcement learning algorithm to solve the optimization problem.
[0006] [1]. Y. Li, N. Deng and W. Zhou, "A Hierarchical Approach to Resource Allocation in Extensible Multi-Layer LEO-MSS," in IEEE Access, vol. 8, pp. 18522-18537, 2020.
[0007] [2]. Z. Li, Z. Xie and X. Liang, "Dynamic Channel Reservation Strategy Based on DQN Algorithm for Multi-Service LEO Satellite Communication System," in IEEE Wireless Communications Letters, vol. 10, no. 4, pp. 770-774, April 2021.
[0008] [3] J. Yang, Z. Xiao, H. Cui, J. Zhao, G. Jiang, and Z. Han, "DQN-ALrM-Based Intelligent Handover Method for Satellite-Ground Integrated Network," in IEEE Transactions on Cognitive Communications and Networking, vol. 9, no. 4, pp. 977-990, Aug. 2023.
[0009] [4] X. Hu, S. Liu, R. Chen, W. Wang, and C. Wang, "A Deep Reinforcement Learning-Based Framework for Dynamic Resource Allocation in Multibeam Satellite Systems," in IEEE Communications Letters, vol. 22, no. 8, pp. 1612-1615, Aug. 2018.
[0010] [5] Y. Cao, S.-Y. Lien, and Y.-C. Liang, "Deep Reinforcement Learning For Multi-User Access Control in Non-Terrestrial Networks," in IEEE Transactions on Communications, vol. 69, no. 3, pp. 1605-1619, March 2021. SUMMARY
[0011] The problem to be solved by the present application is an access control and resource allocation method in a space-air-ground integrated network. In the method proposed in the present application, the total switching overhead of the system, the total service rate quality of the user and the total service delay quality of the user are optimized according to the received signal strength of the user and the service state information of the user, and the optimization strategy of access control and channel allocation is given, so as to maximize the communication resource utilization efficiency of the space-air-ground integrated network.
[0012] To solve the above problems, the technical scheme adopted by the present application is as follows:
[0013] Step 1, a space-air-ground integrated network system model based on low-orbit satellites and unmanned aerial vehicles;
[0014] A space-ground integrated network system consisting of low-orbit satellites, hovering drones, and ground-based mobile users communicates via frequency division multiplexing. The system models handover overhead, quality of service rate, and quality of service delay, introducing three objective functions. Finally, the system models the user's access status and the channel resources allocated to the user as the optimization targets, formulating a multi-objective optimization problem.
[0015] Step 2: Framework of the decision-making algorithm;
[0016] First, before each time slot begins, users can independently obtain the channel quality between themselves and the satellite and drone. The access status and channel allocation status of the satellite and drone are then uniformly distributed by the dispatch center. The user then makes access and channel allocation decisions based on this information. Finally, the user transmits this decision to the satellite or drone, enabling the reallocation of communication resources.
[0017] Step 3: Multi-objective optimization strategy for access control and channel allocation of the air-ground integrated network system;
[0018] For the proposed multi-objective optimization problem, two algorithms are proposed to solve it.
[0019] The first algorithm is a greedy solution. It takes a weighted sum of optimization objectives to derive the overall optimization objective function. Each access control and channel allocation decision is made based on the local optimal solution in the current state. This algorithm has low complexity and can still achieve good performance even when system computing power is insufficient.
[0020] The second algorithm is a solution based on a multi-objective deep reinforcement learning algorithm. This algorithm requires a large amount of computing resources for training. When the system has sufficient computing power, this algorithm can be used to obtain the global optimal solution. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 The system model includes a schematic diagram of the structure of multi-beam satellites, UAVs, and users;
[0022] Figure 2 Schematic diagram of the principle of multi-objective deep reinforcement learning. DETAILED DESCRIPTION
[0023] The present invention will be further described below.
[0024] Step 1: Air-space-ground integrated network system model based on low-orbit satellites and drones
[0025] Figure 1The system model of the application is shown, considering an integrated space-air-ground network system composed of one multi-beam LEO satellite, multiple hovering UAVs and a large number of mobile user terminals, the network architecture is composed of three layers of space layer, air layer and ground layer.
[0026] Space layer: The space layer is in the form of a LEO satellite o, providing extensive network coverage for mobile users on the ground through S2G links. The beams of satellite o are represented by . The height of satellite o is denoted as a constant h o , then the position of satellite o can be represented as (x o (t), y o (t), h o ).
[0027] Air layer: The air layer contains a group of UAVs, denoted by . Each UAV can be flexibly deployed by network operators to hotspots with more service demand, providing wireless transmission services for mobile users through A2G links. Since the UAVs are closer to the mobile users, the service quality of the mobile users is significantly improved. The position of UAV i is represented as (x i (t), y i (t), h i (t)), and its service range is denoted as R i .
[0028] Ground layer: The ground layer is composed of a large number of mobile users, denoted by , and the position of the mobile user is denoted as (x m (t), y m (t), 0). Then the distances between user m and satellite o and UAV i can be represented as
[0029]
[0030] Under this network scenario, the problem to be solved is that when there are multiple available networks during the switching process of the user, a correct switching decision needs to be made to access the most suitable network.
[0031] Step 1.1 transmission model
[0032] The integrated space-air-ground network system completes the coverage of the ground area by satellites and UAVs through frequency reuse, considering that a transmission channel can be reused by different downlinks at different time slots, and multiple users must use different channels to access the same satellite or the same UAV in the same time slot.
[0033] In the transmission process, the total bandwidth of beam k and the downlink of UAV i is denoted as B k and Bi Meanwhile, the bandwidth of the downlink is equally divided into 10 sub-bandwidths, and the downlink sub-bandwidth resource allocated by the beam k to the user m at time t is
[0034] Since the LEO satellite moves very fast relative to the ground, the channel condition between the user and the satellite is rapidly changing. Therefore, considering that the satellite-to-user downlink is a time-varying dynamic channel, without considering rain attenuation and other transmission losses, the channel gain between the user m and the satellite o at time t can be expressed as
[0035]
[0036] where G S is the antenna gain of the satellite transmission, G M is the antenna gain of the user reception, denotes the free space loss between the user and the satellite, and can be further expressed as
[0037]
[0038] According to the Shannon formula, the downlink data rate of the user m accessing the beam k at time t can be expressed as
[0039]
[0040] where P 2 denotes the power allocated by the satellite o to the user m at time t, and σ is the Gaussian white noise.
[0041] For the UAV, the channel gain between the user m and the UAV i at time t can be expressed as
[0042]
[0043] where G U is the antenna gain of the UAV transmission, denotes the free space loss between the user and the UAV, and can be further expressed as
[0044]
[0045] According to the Shannon formula, the downlink data rate of the user m accessing the UAV i at time t can be expressed as
[0046]
[0047] where, denotes the power allocated to user m by UAV i at time t.
[0048] Step 1.2 Handover overhead
[0049] Obviously, a user can be covered by both satellite and UAV at time t, but can only access one network at most, while the same satellite or UAV can connect to multiple users simultaneously.
[0050] Introduce to denote the access state of user m at time t: when , it means that user m accesses beam k; when , it means that user m accesses UAV i. Meanwhile, use to denote that user m accesses satellite, and use to denote that user m accesses UAV. Finally, use to denote which beam of satellite user m accesses, and use to denote which UAV user m accesses.
[0051] Use to denote the switching overhead between beams in LEO satellite, use to denote the switching overhead between LEO satellite and UAV, and use to denote the switching overhead between UAVs. Then the switching overhead of user m at time t can be expressed as
[0052]
[0053] Step 1.3 Service Quality of Experience (QoE)
[0054] Consider a set of service requests , the service request of each user can be expressed as a set s = {r min , r max , d min , d max , p}, where r min is the minimum data rate requirement to meet service s, r max is the data rate when service s reaches the best user experience, d max is the maximum tolerable delay of service s, d min is the delay when service s reaches the best user experience, and p is the data packet size of service content.
[0055] According to the downlink data rate of user m at time t , the downlink transmission time of user m when it accepts service s can be expressed as
[0056]
[0057] Thus, the service rate quality of the user m at the time t can be obtained And the service delay quality Respectively
[0058]
[0059] Step 1.4 problem modeling
[0060] In order to make full use of the channel resources of the satellite and the unmanned aerial vehicle to provide more diversified services for more users, through the proposed dynamic switching algorithm, the maximum service rate quality and service delay quality are taken as the target, the total switching overhead is minimized, the channel resource allocation is dynamically completed by meeting the channel resource constraints, and the correct switching decision is made. Therefore, the corresponding dynamic switching and channel resource optimization problem model is established as follows:
[0061]
[0062] Constraints 1, 2: Each user can only select one satellite beam or unmanned aerial vehicle access;
[0063] Constraint 3: The sum of the channel resources allocated to each user cannot exceed the total bandwidth of the beam;
[0064] Constraint 4: The sum of the channel resources allocated to each user cannot exceed the total bandwidth of the unmanned aerial vehicle.
[0065] Step 3, access control and channel allocation multi-objective optimization strategy of space-air-ground integrated network system;
[0066] Since the optimization target is an optimization considering multiple objectives and a mixed integer programming problem, the problem is np-hard. Due to the dynamic change of existing users, the random arrival of services and the randomness of the wireless communication environment, the problem brings uncertainty, and from the perspective of long-term benefits of the system, an effective optimization strategy needs to be found. Therefore, the present application proposes two decision algorithms, which can be selected according to the computing power of the decision end.
[0067] Step 3.1, optimization strategy based on greedy algorithm
[0068] In order to balance multiple conflicting optimization objectives and obtain the access control and channel allocation optimization strategy of the space-air-ground integrated network system, an optimization strategy based on the greedy algorithm is proposed. The fundamental idea of the greedy algorithm is to obtain a local optimal solution step by step, and to approximate the global optimal solution through the local optimal solution, which is a simple strategy for solving optimization problems.
[0069] To balance all the optimization objectives, let Ω = [ω1, ω2, ω3] be the weight vector of all optimization objectives, and the total optimization objective function of user m at time t is denoted as
[0070]
[0071] Therefore, only the weight vector Ω needs to be changed to achieve different preferences of the system for each optimization objective.
[0072] In this scheme, for all users in the user set , according to the order of the user queue, one user is selected each time to make a decision. All possible nodes that the user can access are denoted as ID ∈ {SAT1, …, SAT k , UAV1, …, UAV i , and all possible channel allocations of the user are denoted as ω ∈ {0, 1, 2, …, 10}, and then the total objective function value of user m at time t is calculated . Finally, the access node and channel allocation fraction with the maximum total optimization objective function value are selected according to the greedy strategy, i.e.
[0073]
[0074] Next, the channel resources allocated to user m are locked, so that subsequent users cannot select the channels allocated to user m. The above operation is repeated until all user decisions are completed, and the decision set of all users is obtained .
[0075] The greedy algorithm has low computational complexity and requires less computing resources, and can quickly obtain a suboptimal decision under low computing resources of the network.
[0076] Step 3.2, solution based on multi-objective deep reinforcement learning algorithm
[0077] To solve the problem that the optimization strategy based on the greedy algorithm cannot effectively obtain the global optimal solution, in this scheme, the optimization problem is converted into a multi-objective Markov decision process (MOMDP), and a multi-objective deep Q-learning algorithm is proposed to solve it, as shown in Figure 2 .
[0078] Step 3.2.1 Multi-objective Markov Decision Process
[0079] MOMDP is represented by a triple , where S, A, are the state space, action space, and reward vector, respectively.
[0080] (1) State space S: Let s m(t) e S represents the state of user m at time t, which includes the following contents:
[0081] 1) The channel quality of each beam and UAV corresponding to user m at the current time t, i.e. . Wherein, is the signal-to-noise ratio between the satellite or UAV and user m at time t.
[0082] 2) The channel resource allocation state of each beam and UAV at the last time t-1, i.e.
[0083]
[0084] Therefore, in s m (t) there are (k+i) x 2 elements, i.e.
[0085]
[0086] (2) Action space A: user m dynamically adjusts the satellite beam or UAV to access at each time, i.e. selects one from the k+i satellite beams and UAVs as the target, and then allocates all beam and UAV channel resources. Therefore, let a m (t) e A represents the action of user m at time t, defined as:
[0087]
[0088] (3) Reward : Let represent the reward obtained by user m at time t, and the three dimensions of the reward vector correspond to the three objectives of the optimization problem, i.e. Wherein respectively represent the reward functions of service rate quality, service delay quality and switching overhead;
[0089] 1) Service rate quality: in order to improve the transmission rate of all users as much as possible, the service rate quality reward of user m at time t is defined as:
[0090]
[0091] 2) Service delay quality: in order to reduce the communication delay of all users as much as possible, the service delay quality reward of user m at time t is defined as:
[0092]
[0093] 3) Switching penalty: in order to reduce the frequent switching caused by improving user experience as much as possible, the switching overhead penalty of user m at time t is defined as:
[0094]
[0095] If the user accesses the same beam or drone as the previous time, the switching cost of this time is 0; if the user performs switching between different beams in the satellite, the switching cost of this time is ; if the user performs switching between the satellite and the drone, the switching cost of this time is ; and if the user performs switching between different drones, the switching cost of this time is .
[0096] Step 3.2.2 Multi-objective Deep Q Learning Algorithm
[0097] The multi-objective deep Q learning algorithm is an evolution of the single-objective deep Q learning algorithm, and is used to optimize multiple conflicting objectives. In this algorithm, MODQN constructs a DQN network for each optimization objective, and when the agent interacts with the environment, the environment gives the agent a reward vector corresponding to each objective, i.e. Since the objectives in the optimization problem are conflicting, MODQN needs to find a strategy that maximizes the total return of these optimization objectives. In this invention, for each optimization objective, the expected discounted reward of action a t in state s t can be represented by an action value function Q(s t ,a t ). Therefore, the action value for all optimization objectives can be uniformly represented by an action value vector , and the weighted sum of the action values for all optimization objectives according to the weight vector Ω of all optimization objectives is
[0098]
[0099] Subsequently, the agent greedily selects the action that maximizes the action value function Q(s t ,a t ) to obtain the optimal access control and channel allocation decision in the current state s t .
[0100]
[0101] In this algorithm, the Q value of each optimization objective is estimated by the evaluation network and the target network, and the parameters of the network are represented as θ and θ - respectively. The evaluation network is used to select the action with the maximum action value function, and the target network is used to calculate the value function of the selected action after the action is selected. Among them, the parameter θ of the evaluation network is updated after each decision, and the parameter θ -The parameters θ of the network are updated periodically at specific times. The iterative formula for evaluating the parameters θ of the network is as follows:
[0102]
[0103] The loss function of the network is calculated in the form of mean square error, defined as follows:
[0104]
[0105] Where y t is the time difference target, calculated as follows:
[0106]
Claims
1. A method for access control and resource allocation in an air-ground integrated network, characterized in that: Based on the user's received signal strength and service status information, the system optimizes the three objectives of total switching overhead, total service rate quality, and total service delay quality. Optimization strategies for access control and channel allocation are provided to maximize the communication resource utilization efficiency of the integrated air-ground-space network. The specific steps are as follows: Step 1: Air-space-ground integrated network system model based on low-orbit satellites and drones; A space-ground integrated network system consisting of low-orbit satellites, hovering drones, and ground mobile users communicates via frequency division multiplexing. The system models handover overhead, quality of service rate, and quality of service delay, respectively, introducing three objective functions. Finally, the system models the user's access status and the channel resources allocated to the user as the optimization targets, formulating a multi-objective optimization problem. Step 2: Framework of the decision-making algorithm; First, before the start of each time slot, users independently obtain the channel quality between the satellite and the drone. The access status and channel allocation status of the satellite and drone need to be uniformly issued by the dispatch center. Then, the user makes access decisions and channel allocation decisions based on the information obtained. Finally, the user sends the decision content to the satellite or drone to achieve the reallocation of communication resources. Step 3: Multi-objective optimization strategy for access control and channel allocation of the air-ground integrated network system; For the proposed multi-objective optimization problem, two algorithms are proposed to solve it; The first algorithm is a greedy algorithm solution that takes a weighted sum of optimization objectives to derive the overall optimization objective function. It adopts the local optimal choice in the current state for each access control and channel allocation decision. This algorithm has low complexity and can still achieve good performance even when the system's computing power is insufficient. The second algorithm is a solution based on a multi-objective deep reinforcement learning algorithm. This algorithm requires a large amount of computing resources for training. When the system has sufficient computing power, this algorithm can obtain the global optimal solution.
2. The method for access control and resource allocation of an air-ground integrated network according to claim 1, characterized in that: In step 1, consider an air-ground integrated network system consisting of one multi-beam LEO satellite, multiple hovering drones, and a large number of mobile user terminals, which consists of three layers: space layer, air layer, and ground layer; Space layer: The space layer provides extensive network coverage for mobile users on the ground through S2G links in the form of LEO satellites. The height of satellite ο is recorded as a constant h o , then the position of satellite ο is expressed as (x o (t),y o (t),h o ); Aerial layer: The aerial layer contains a group of drones, Each drone The network operator flexibly deploys the drone to the hotspot area of service demand and provides wireless transmission services to mobile users through A2G links. The position of drone i is represented by (x i (t),y i (t),h i (t)), whose service range is denoted as R i ; Ground layer: The ground layer consists of a large number of mobile users. The location of the mobile user is represented by (x m (t),y m (t),0); the distances between user m and satellite ο and drone i are expressed as 3. The method for access control and resource allocation of an air-ground integrated network according to claim 2, characterized in that: Step 1 includes: step 1.1 transmission model; The integrated air-ground network system achieves satellite and drone coverage of the ground area through frequency reuse. It is believed that a transmission channel can be reused by different downlinks in different time slots, while multiple users must use different channels to access the same satellite or the same drone in the same time slot; During the transmission process, the total bandwidth of the downlink of beam k and UAV i is denoted as B k and B i , and the downlink bandwidth is divided into 10 sub-bandwidths. The downlink sub-bandwidth resources allocated by beam k to user m at time t are The downlink sub-bandwidth resource allocated by drone i to user m at time t is share, and Since LEO satellites move very fast relative to the ground, the channel conditions between users and satellites change rapidly. Considering that the downlink from satellite to user is a time-varying dynamic channel, without considering the rain attenuation transmission loss, the channel condition between users and satellites changes rapidly. To express the channel gain between user m and satellite o at time t: Among them, G S is the satellite transmitting antenna gain, G M is the antenna gain received by the user, represents the free space loss between the user and the satellite, which is expressed as: According to Shannon's formula, the downlink data rate of user m accessing beam k at time t is expressed as: in, represents the power allocated to user m by satellite ο at time t, σ 2 is Gaussian white noise; For drones, use To express the channel gain between user m and drone i at time t: Among them, G U is the antenna gain of the UAV transmitter, represents the free space loss between the user and the UAV, which is further expressed as: According to Shannon's formula, the downlink data rate of user m accessing drone i at time t is expressed as: in, represents the power allocated to user m by drone i at time t; Step 1.2 switching overhead; Obviously, a user may be covered by both satellites and drones at time t, but can only access one network at most, while the same satellite or drone can be connected to multiple users at the same time; To express the access status of user m at time t: When , it means that user m has accessed beam k; when When , it means that user m has accessed drone i; at the same time, Indicates that user m has accessed the satellite, using Indicates that user m has accessed the drone; finally, Indicates which satellite beam user m accesses, using Indicates which drone user m has connected to; use Denotes the switching overhead between beams within a LEO satellite, using represents the switching overhead between LEO satellite and UAV, and represents the switching cost between drones; then the switching cost of user m at time t is expressed as Step 1.3: Quality of Service Experience (QoE); Consider a set of service requests Service requests per user It can be expressed as a set s = {r min ,r max ,d min ,d max ,p}, where rmin is the minimum data rate requirement for service s, r max is the data rate when service s achieves the best user experience, d max is the maximum tolerable delay of service s, d min is the delay when service s achieves the best user experience, and p is the data packet size of the service content; According to the downlink data rate of user m at time t When user m receives service s, the downlink transmission time of user m is expressed as Thus, the service rate quality of user m at time t is obtained and quality of service delay They are Step 1.4: Problem modeling; The corresponding dynamic switching and channel resource optimization problem model is established as follows: Constraints 1 and 2: Each user can only select one satellite beam or drone to access; Constraint 3: The sum of the channel resources allocated to each user cannot exceed the total bandwidth of the beam; Constraint 4: The sum of the channel resources allocated to each user cannot exceed the total bandwidth of the drone.
4. The method for access control and resource allocation of an air-ground integrated network according to claim 1, characterized in that: The optimization strategy based on the greedy algorithm in step 3 is described as follows: In order to balance all optimization objectives, assuming that Ω = [ω1, ω2, ω3] is the weight vector of all optimization objectives, the total optimization objective function of user m at time t is expressed as By simply changing the weight vector Ω, the system can achieve different preferences for various optimization objectives; For user collections For all users in , one user is selected at a time for decision making according to the order of the user queue; all possible nodes that the user may access are recorded as ID∈{SAT1,…,SAT k ,UAV1,…,UAV i }, record the number of channels that all users may be assigned as ω∈{0,1,2,…,10}, and then calculate the total objective function value of all possible users m at time t Finally, according to the greedy strategy, the access nodes and channel allocation number when the total optimization objective function value is the largest are selected, that is, Lock the channel resources allocated to user m so that subsequent users cannot choose the channel allocated to user m; repeat the above operation until all users have made decisions, and the decision set of all users can be obtained. The greedy algorithm has low computational complexity, requires few computing resources, and can quickly obtain suboptimal decisions when the network has low computing resources.
5. The method for access control and resource allocation of an air-ground integrated network according to claim 1, characterized in that: The multi-objective deep reinforcement learning algorithm in step 3 is described as follows: To solve the problem that the optimization strategy based on the greedy algorithm cannot effectively obtain the global optimal solution, the optimization problem is transformed into a multi-objective Markov decision process MOMDP, and a multi-objective deep Q-learning algorithm is proposed to solve it; Step 3.2.1 Multi-objective Markov decision process; MOMDP using triples To express, where S, A, are state space, action space, and reward vector respectively; (1) State space S: Let s m (t)∈S represents the state of user m at time t, where Includes the following: 1) User m corresponds to the channel quality of each beam and drone at the current time t, that is, in, is the signal-to-noise ratio between the satellite or UAV and user m at time t; 2) The channel resource allocation status of each beam and UAV at the previous moment t-1, that is, In s m There are (k+i)×2 elements in (t), that is, (2) Action space A: User m dynamically adjusts the satellite beam or drone to be accessed at each moment, that is, selects one from k+i satellite beams and drones as a target, and then allocates channel resources to all beams and drones; let a m (t)∈A represents the action of user m at time t, which is defined as: (3) Rewards set up represents the reward obtained by user m at time t. The three dimensions of the reward vector correspond to the three objectives of the optimization problem, namely in The reward functions representing service rate quality, service delay quality and switching overhead respectively; 1) Service rate quality: To improve the transmission rate of all users, the service rate quality reward of user m at time t is defined as: 2) Service delay quality: To reduce the communication delay of all users, the service delay quality reward of user m at time t is defined as: 3) Handover penalty: To reduce frequent handovers due to improved user experience, the handover penalty for user m at time t is defined as: If the user accesses the same beam or drone as the previous moment, the switching cost at that moment is 0; if the user switches between different beams within the satellite, the switching cost at that moment is If the user switches between the satellite and the drone, the switching cost at that moment is If the user switches between drones, the switching cost at that moment is 6. The method for access control and resource allocation of an air-ground integrated network according to claim 1, characterized in that: The multi-objective deep Q learning algorithm in step 3 is described as follows: MODQN constructs a DQN network for each optimization goal. When the agent interacts with the environment, the environment gives the agent a reward vector corresponding to each goal, that is, For each optimization objective, an action-value function Q(s t ,a t ) represents the expected discounted reward of action at in state st; For all optimization objectives, the action value vector To unify the representation, so as to perform weighted summation of the action values under all optimization objectives according to the weight vector Ω of all optimization objectives, that is, Then, the agent greedily chooses the action value function Q(s t ,a t ) The action that reaches the maximum value is used to obtain the current state s t Optimal access control and channel allocation decisions under The Q value of each optimization target is estimated by the evaluation network and the target network, and the network parameters are denoted as θ and θ respectively. - The evaluation network is used to select the action with the largest action value function, and the target network is used to calculate the value function of the selected action after the action is selected; the parameters θ of the evaluation network are updated after each decision step, and the parameters θ of the target network are updated after each decision step. - Updated regularly at specific times; the iterative formula for evaluating the parameters θ of the network is as follows: The loss function of the network is calculated as mean square error and is defined as follows: L(θ t )=E[(y t -Q(s t ,a t ;θ t )) 2 ] Among them, y t is the time difference target, calculated as follows:
Citation Information
Patent Citations
Network resource optimization method and device for space-air-ground integrated network
CN117674958A
Method and device for accessing Internet of Things terminal of space-ground convergence network
CN117750402A