Task scheduling method for unified system sky-ground network
By adopting multi-agent near-end strategy optimization algorithm and reinforcement learning framework in a unified system sky and earth network, a strategy-value neural network architecture is built, which solves the problem that access points are difficult to obtain global information and spectrum allocation is affected by space-time interference, and efficient task scheduling and optimization are achieved.
Patent Information
- Application Number
- CN202510414746.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-03
AI Technical Summary
In a unified system of sky-ground integrated network, access points find it difficult to obtain global information, making it difficult to achieve global optimization of task scheduling, and spectrum allocation is affected by spatiotemporal interference, with high computational complexity, and the existing technology is difficult to adapt to dynamically changing network environments.
The multi-agent proximal strategy optimization algorithm is adopted, combined with the reinforcement learning framework, and the strategy-value neural network architecture is built, and information age is dynamically updated through training and interaction to optimize task allocation.
It effectively improves scheduling efficiency, optimizes task allocation, reduces algorithm complexity, adapts to the dynamically changing network environment, and realizes global coordination needs.
Smart Images

Figure CN119946721A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the field of communications, and in particular to a task scheduling method for a unified system sky-ground network. Background Art
[0002] In a unified sky-ground integrated network, multiple layers of access points share spectrum resources, and the task scheduling problem faces many severe challenges. Since the access points can only obtain limited information within the coverage area, it is difficult to grasp the global state, which makes global optimization extremely difficult. At the same time, spectrum allocation is affected by time and space interference and needs to be completed under dynamically changing constraints, making the resource allocation problem more complicated. More importantly, with the increase in the number of users and channels, the action space of access points grows exponentially, making the design and optimization of scheduling strategies face extremely high computational complexity. These challenges make it difficult for traditional optimization algorithms to adapt to the dynamically changing network environment and the global coordination requirements of multiple access points. Therefore, the existing technology has significant limitations in solving the task scheduling problem of the sky-ground integrated network, and an innovative solution that can efficiently cope with the above complexities and challenges is urgently needed. Summary of the invention
[0003] The purpose of the present invention is to overcome the shortcomings of the prior art and provide a task scheduling method for a unified sky-ground network, which can effectively improve the scheduling efficiency and optimize the task allocation by combining a reinforcement learning framework with network architecture characteristics and time-varying constraints.
[0004] The object of the present invention is achieved through the following technical scheme: a task scheduling method for a unified system sky-ground network, comprising the following steps: Constructing a mission-oriented communication scenario in a unified sky-ground network; Determine the target problem of task scheduling and model the target problem as a Markov decision process; Construct a policy-value neural network architecture based on multi-agent proximal policy optimization; Train the constructed strategy-value neural network architecture to obtain a trained strategy-value neural network, and dynamically update the information age based on the interaction between the trained strategy-value neural network and the real environment. , and apply strategies based on real-time feedback.
[0005] The beneficial effects of the present invention are as follows: the present invention is based on a multi-agent proximal strategy optimization algorithm, can target network architecture characteristics and time-varying constraints, combined with a reinforcement learning framework, effectively improve scheduling efficiency and optimize task allocation, reduce algorithm complexity, and provide a new theoretical basis and technical method for this research field. BRIEF DESCRIPTION OF THE DRAWINGS
[0006] Figure 1 is a flow chart of the method of the present invention; Figure 2 This is a schematic diagram of the unified system sky-ground network scenario; Figure 3 Schematic diagram of the task scheduling architecture based on multi-agent proximal strategy optimization. DETAILED DESCRIPTION
[0007] The technical solution of the present invention is further described in detail below in conjunction with the accompanying drawings, but the protection scope of the present invention is not limited to the following.
[0008] like Figure 1 As shown in the figure, a task scheduling method for a unified sky-ground network is proposed. The method is based on multi-agent proximal strategy optimization. The task scheduling architecture is as follows: Figure 3 As shown, the following steps are included: Constructing a mission-oriented communication scenario in a unified sky-ground network; like Figure 2 As shown, in this scenario, considering U ground users, a communication satellite, M A drone, N Ground base stations, a total of K Access points provide services to users. For the sake of index clarity, the communication satellite index is defined as k = 0, the drone index is , the base station index is . Determine the target problem of task scheduling and model the target problem as a Markov decision process; The scheduling goal is to minimize the user's Age of Information (AoI) and the energy consumption of the access point.
[0009] First, the problem is modeled as a Markov decision process (MDP), which consists of the following four parts: Action: In time t Access Point k The action vector is defined as ,in express The first ( u,p ) elements. Specifically: Indicates access point k Resource Block p In time t Assigned to user u ,and Indicates an unallocated channel. Each resource block is a channel, and the total number of channels is defined as P .
[0010] Status: Access Point k The state space Include information age and transmission record vector Two key components: Information age: Information age measures the freshness of information received by users. For a given user u , at time t When the information age Defined as: in: t is the current time, Is the timestamp of the last time the information was updated or received, indicating when the information was last updated.
[0011] Transmission record vector: Due to the propagation delay between the satellite and the drone, the mission data packet cannot reach the user end immediately. In order to monitor the transmission process of the mission data packet, an access point is defined k With users u The transmission record vector between is: ,in: represents the set of natural numbers, Indicates access point k With users u When the access point k Send the task data package to the user U hour, The corresponding element of will be set to 1 and incremented by 1 at each time step until it reaches Once you arrive , the element will be reset to 0. The d Elements The update steps are summarized as follows: State transfer: In the unified sky-ground network, whether the state transfers depends on whether the user successfully receives the data packet. Based on the transmission record vector, the user's information age transfer can be expressed as: Where: I is an indicator function that satisfies The condition is 1, otherwise it is 0. In the first case, when any base station is a user u Provide services, will be reset to 1; in the second case, if the user u is not served by any base station and does not receive service packets from drones or satellites, then Increment by 1, that is ; In the third case, if no base station is serving the user u service, but successfully received data packets from drones or satellites, The smallest one between the current information age of the data packet and the information age of the received data packet is selected as the new information age.
[0012] Reward: The reward function combines the current information age and energy consumption, expressed as: in, U , P , K Respectively represent the total number of users, the total number of channels and the total number of access points, Indicates access point k Serve u The energy required to consume.
[0013] Based on the above Markov decision process modeling, users u The average information age is defined as: in, represents mathematical expectation. User u The average energy expenditure is defined as: Combining the two, the overall objective function can be written as: in, and are the weights of information age and energy consumption respectively. The joint optimization problem is established as: in, k and For two different access points, and Access Point k and and users u Formula (9) limits the resource block to be used by at most one user, and formula (10) takes into account the time and space interference in the network. Together, they ensure the rationality of task scheduling.
[0014] To solve this problem, an architecture based on Multi-Agent Proximal Policy Optimization (MAPPO) is proposed.
[0015] Policy network (Actor, also called action network): The policy network defines the policy , the parameters are , combined and recorded as . This network is used to optimize access points k The distribution of actions In order to realize resource allocation under mutual blind conditions, the Softmax function is used to reconstruct the actions of the access points into ,in, Indicates access point k Channel p Assign to user u , 0 means no channel is allocated. To further mitigate the impact of spatial and temporal constraints, the input state dimension is filtered according to the access point coverage: specifically, satellites only consider the transmission record vectors of their covered users, while drones and base stations include the information age and transmission record vectors of their covered users. The loss function is defined as: in, and Represent the current strategy and the original strategy respectively, and the advantage function is expressed as , is the set threshold, Indicates that for the input value x , and the given upper and lower limits min and max , the clip function will x Restricted to [ min, max ]; if x If it exceeds the range, it will be truncated to the boundary value; min Represents the minimum value function; Value Network (Critic): The value network has parameters And by using the value function ( is the total state space, is a real number domain) to estimate the expected cumulative reward to evaluate the strategy . When performing an action Afterwards, the access point received an immediate reward and transfer to the next state The input of the value network includes the information age of all users, and the goal is to minimize The mean square error between the cumulative reward and the target. The loss function is expressed as: in, Indicates access point k The state-value function of express In the value function The value below.
[0016] Train the constructed strategy-value neural network architecture to obtain a trained strategy-value neural network, and dynamically update the information age based on the interaction between the trained strategy-value neural network and the real environment. , and apply strategies based on real-time feedback.
[0017] The algorithm implementation mainly consists of two parts: 1) Offline training: It can be divided into two parts: environment simulation and offline training: Environmental simulation: Based on the current state and actions , transfer to the next state , and calculate the average information age and average energy consumption according to formula (5) and formula (6). ,action and global state The immediate reward is calculated by formula (4).
[0018] Offline training: At each time step, store the experience tuple according to the simulated environment ,in: represents the reward vector, Represents the probability distribution of actions. By extracting data from the experience pool, the policy network and value network are trained using formulas (11) and (12).
[0019] 2) Online application: Use the offline trained policy network to interact with the real environment and dynamically update the information age , and apply existing strategies based on real-time feedback: Each access point obtains real-time environmental status based on local observations, including real-time information age and transmission record vector , input the trained action network to generate its own action decision, and dynamically allocate resource blocks for the covered users (that is, each access point chooses whether to use the resource block and which resource block to use to provide services to the user based on the input information), thereby realizing task scheduling.
[0020] Specifically: the trained policy network interacts with the real unified sky-ground network environment, and the policy obtained by offline training is brought into the real business scheduling scenario for real-time testing. Each access point (satellite, drone, base station) obtains the environmental state based on its local observation, and generates its own action decision through the trained Actor network to dynamically allocate resource blocks for the covered users. Information Age According to formula (3), real-time update: when the base station is the user u It is reset to 1 when providing service; it is incremented by 1 when there is no base station service and no data packets from satellites or drones; when a data packet is received from a drone or satellite, the smaller value between the current information age and the information age of the received data packet is selected. The real-time feedback mechanism evaluates the effectiveness of the strategy by monitoring the information age of all users and the energy consumption of all access points. The smaller the weighted sum of the two, the more effective the strategy. The system dynamically adjusts resource allocation based on these indicators to ensure that the overall performance is optimized while meeting coverage constraints and avoiding spatiotemporal interference.
[0021] The trained policy network is directly applied to real-time scheduling decisions, determining which resource block to use for each access point to serve which user at each time step. By optimizing the objective function (Formula 7), the system achieves a balance between the freshness of user information and the energy consumption of the access point.
[0022] By considering the spatiotemporal interference constraints caused by the ripple effect (Formula 10), this scheduling method can effectively avoid channel conflicts and adjust the scheduling according to different task requirements. (Information Age) and The weight of (system energy consumption) is used to achieve flexible task priority management. In addition, through dynamic resource allocation, this method adapts to network changes and ensures the effective execution of time-sensitive tasks in a complex and changeable unified system sky-ground network environment.
[0023] The above is a preferred embodiment of the present invention. It should be understood that the present invention is not limited to the form disclosed herein, and should not be regarded as excluding other embodiments, but can be used in other combinations, modifications and environments, and can be modified within the scope of the concept described herein through the above teachings or the technology or knowledge of the relevant field. The changes and modifications made by those skilled in the art do not depart from the spirit and scope of the present invention, and should be within the scope of protection of the claims attached to the present invention.
Claims
1. A task scheduling method for a unified sky-ground network, characterized in that: The following steps are involved: Constructing a mission-oriented communication scenario in a unified sky-ground network; Determine the target problem of task scheduling and model the target problem as a Markov decision process; Construct a policy-value neural network architecture based on multi-agent proximal policy optimization; Train the constructed strategy-value neural network architecture to obtain a trained strategy-value neural network, and dynamically update the information age based on the interaction between the trained strategy-value neural network and the real environment. , and apply strategies based on real-time feedback.
2. The task scheduling method for a unified sky-ground network according to claim 1 is characterized by: The mission-oriented communication scenario includes a communication satellite, M A drone, N Ground base stations, a total of K An access point provides services to users, wherein the access point is a communication satellite, a drone or a ground base station. K = ; For clear indexing, the communication satellite index is defined as k = 0, the drone index is , the base station index is .
3. The task scheduling method for a unified sky-ground network according to claim 1 is characterized by: The objective problem of the task scheduling is to minimize the information age of users and the energy consumption of access points.
4. The task scheduling method for unified sky-ground network according to claim 3 is characterized by: The modeling of the target problem as a Markov decision process includes the following four parts: (1) Action: in time t Access Point k The action vector is defined as ,in express The first ( u,p ) elements, Indicates access point k Resource Block p In time t Assigned to user u ,and Indicates access point k Resource Block p In time t Not assigned to user u ; Each resource block is a channel, and the total number of channels is defined as P ; (2) Status: Access point k The state space Include information age and transmission record vector Two key components: Information age: Information age measures the freshness of information received by users. For a given user u , at time t When the information age Defined as: in, t is the current time, is the timestamp of the last time the information was updated or received, indicating when the information was last updated; Transmission record vector: Due to the propagation delay between the satellite and the drone, the mission data packet cannot reach the user end immediately. In order to monitor the transmission process of the mission data packet, an access point is defined k With users u The transmission record vector between is: ,in: represents the set of natural numbers, Indicates access point k With users u The propagation delay between When the access point k Send the task data package to the user U hour, The corresponding element of will be set to 1 and incremented by 1 at each time step until it reaches Once you arrive , the element will be reset to 0, The d Elements The update steps are summarized as follows: (3) State transfer: In the unified sky-ground network, whether the state transfers depends on whether the user successfully receives the data packet. Based on the transmission record vector, the user's information age transfer is expressed as: Among them, I is an indicator function that satisfies 1 if condition is met, 0 otherwise; In the first case, when any base station is a user u Provide services, will be reset to 1; In the second case, if the user u is not served by any base station and does not receive service packets from drones or satellites, then Increment by 1, that is ; In the third case, if no base station is serving the user u service, but successfully received data packets from drones or satellites, The smallest one between the current information age of the data packet and the information age of the received data packet is selected as the new information age; (4) Reward: The reward function combines the current information age and energy consumption, expressed as: in, U , P , K Respectively represent the total number of users, the total number of channels and the total number of access points, Indicates access point k Serve u The energy required to be consumed; Based on the above Markov decision process modeling, users u The average information age is defined as: in, represents the mathematical expectation, user u The average energy expenditure is defined as: The overall objective function is: in, and are the weights of information age and energy consumption respectively, and the joint optimization problem is established as: in, k and For two different access points, and Access Point k and and users u Formula (9) limits the resource block to be used by at most one user, and formula (10) takes into account the time and space interference in the network. Together, they ensure the rationality of task scheduling.
5. The task scheduling method for a unified sky-ground network according to claim 1 is characterized by: The strategy-value neural network architecture based on multi-agent proximal strategy optimization includes: Policy Network: The policy network defines the policy , the parameters are , combined and recorded as ; This network is used to optimize access points k The distribution of actions To achieve resource allocation under mutual blind conditions, the Softmax function is used to reconstruct the actions of the access points into ,in, Indicates access point k Channel p Assign to user u , 0 means no channel is allocated; To mitigate the impact of time-space related constraints, the input state dimension is filtered according to the access point coverage: satellites only consider the transmission record vectors of their covered users, while drones and base stations include the information age and transmission record vectors of their covered users. The loss function is defined as: in, and Represent the current strategy and the original strategy respectively, and the advantage function is expressed as , is the set threshold, Indicates that for the input value x , and the given upper and lower limits min and max , the clip function will x Restricted to [ min, max ]; if x If it exceeds the range, it will be truncated to the boundary value; min Represents the minimum value function; Value Network: The value network has parameters And by using the value function Estimate the expected cumulative reward to evaluate the strategy , is the total state space, is the field of real numbers; In action Afterwards, the access point received an immediate reward and transfer to the next state , the input of the value network includes the information age of all users, and the goal is to minimize The mean square error between the cumulative reward and the target, the loss function is expressed as: in, Indicates access point k The state-value function of express In the value function The value below.
6. The task scheduling method for unified sky-ground network according to claim 5 is characterized by: The constructed strategy-value neural network architecture is trained to obtain a trained strategy-value neural network architecture, including: Environmental simulation: Based on the current state and actions , transfer to the next state , and calculate the average information age and average energy consumption according to formula (5) and formula (6), using the current state ,action and global state The immediate reward is calculated by formula (4); Offline training: At each time step, store the experience tuple according to the simulated environment ,in: represents the reward vector, Represents the logarithmic probability distribution of actions. By extracting data from the experience pool, the policy network and value network are trained using formulas (11) and (12).
7. The task scheduling method for unified sky-ground network according to claim 6 is characterized by: The trained policy-value neural network interacts with the real environment, dynamically updates the information age, and applies the policy based on real-time feedback, using only the trained policy network, including: Each access point obtains real-time environmental status based on local observations, including real-time information age and transmission record vector , input the trained policy network to generate its own action decisions, dynamically allocate resource blocks for the covered users, and thus realize task scheduling.
Citation Information
Patent Citations
Air-space-ground integrated unmanned aerial vehicle internet-of-things data acquisition method based on information age
CN114690799A
Unmanned aerial vehicle track adaptive optimization method based on information age
CN115696211A
Optimization method based on information age optimization and considering user transmission energy consumption
CN117726023A
Satellite-ground fusion network information age optimization method based on deep reinforcement learning
CN118157745A
Air-based network deployment method based on two-stage Markov process
CN118843086A
Cited By
Distributed hybrid expert network allocation method based on reinforcement learning
CN121001122A
A distributed hybrid expert network allocation method based on reinforcement learning
CN121001122B