Satellite-ground fusion network information age optimization method based on multi-agent layered deep reinforcement learning
By adopting a multi-agent layered deep reinforcement learning architecture in the satellite communication system and optimizing the joint allocation of spectrum and power, the problem of difficulty in optimizing the information age of the satellite communication system in the existing technology is solved, and a lower information age and more efficient resource management are achieved.
Patent Information
- Application Number
- CN202510394551.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-03-31
AI Technical Summary
The prior art is difficult to effectively optimize the information age of satellite communication systems, especially in the complex and changeable channel environment, the traditional single-layer deep reinforcement learning network architecture cannot handle the hybrid action space challenges brought about by joint optimization of multiple resources.
The satellite-ground fusion network information age optimization method based on multi-agent layered deep reinforcement learning is adopted. Discrete resource block allocation is optimized through the upper DQN network, and the lower PPO network optimizes continuous power allocation to realize the joint allocation scheme of spectrum and power.
It significantly reduces the information age of the satellite communication system, provides an efficient resource management solution, adapts to changes in different beam environments, and achieves stable performance in large-scale user access scenarios.
Smart Images

Figure CN120150802A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of communication technologies, and in particular to an information age optimization method for a satellite-ground integrated network based on multi-agent hierarchical deep reinforcement learning. Background Art
[0002] With the continuous growth of global communication demands, satellite communication will play a key role in achieving the goal of global coverage of 6G networks due to its advantages such as being unrestricted by geographical locations and having a wide coverage range. However, satellite communication also faces many technical challenges, including complex and variable channel environments, difficult modeling, and relatively high communication delays. The Age of Information (AoI) is a key metric for measuring the freshness of information, defined as the time interval (ms) from information generation to receiver reception. Optimizing resource allocation schemes targeting AoI is of great significance for ensuring data freshness and improving system performance.
[0003] In the prior art, research on optimizing the communication quality of satellite-ground networks mostly focuses on optimizing traditional communication metrics such as rate and energy efficiency, and there are few patents studying the information age optimization problem of satellite systems. These research directions only meet the reliability metrics of communication systems and cannot guarantee the immediacy of communication. For example, the patent document with the publication number CN119364423A discloses a congestion control method for an air-space-ground network based on deep reinforcement learning. By obtaining data information of multiple target air-space-ground networks and setting a state space, an action space, and a reward function, as well as implementing state prediction, reward redistribution, and optimization strategies, congestion control of the air-space-ground network based on deep reinforcement learning is achieved. The above solution adopts a single-layer deep reinforcement learning network architecture to solve non-convex optimization problems, but such single-layer network architectures can only handle single-type resource optimization and cannot cope with the challenges of mixed action spaces brought about by jointly optimizing multiple resources. Summary of the Invention
[0004] Aiming at the problems existing in the prior art, the present invention provides an information age optimization method for a satellite-ground integrated network based on multi-agent hierarchical deep reinforcement learning, which overcomes the challenges of complex, dynamic, and variable satellite channels and finds a joint spectrum and power allocation scheme that can minimize the information age of the system.
[0005] The technical solution of the present invention is realized as follows:
[0006] An information age optimization method for a satellite-ground integrated network based on multi-agent hierarchical deep reinforcement learning, the satellite-ground integrated network includes a multi-beam low-earth orbit (LEO) satellite, the LEO includes M orthogonal frequency band resource blocks (RBs), and one RB is occupied by a ground cell base station; the method includes the following steps:
[0007] Step S1: Initialize the network parameters θ, θ′, and correspond to the main network, target network, new policy network, old policy network, and evaluation network respectively; one RB corresponds to one agent;
[0008] In the field of wireless communication, in a communication system, an RB is the smallest unit of frequency-domain and time-domain resource allocation. An RB usually consists of 12 subcarriers in the frequency domain, each subcarrier having a bandwidth of 15 kHz and a total bandwidth of 180 kHz; in the time domain, it corresponds to one time slot (0.5 ms), containing 7 OFDM symbols (normal cyclic prefix mode) or 6 symbols (extended cyclic prefix mode).
[0009] Step S2: Initialize the experience pool and Set the total number of training rounds T, the allocation frequency Ns, and the update frequency F of the old policy network; initialize t = 1;
[0010] Step S3: If t % Ns = 0, then execute Step S4; otherwise, execute Step S5;
[0011] Step S4: For the current state of each agent Use the greedy policy (ε - greedy) to select the action α m (t); m represents the m-th agent; obtain the data representing the state transition process, and store it in ; randomly select multiple samples from to update θ and θ′; α m (t) represents the spectrum allocation strategy of the m-th agent; execute Step S6;
[0012] Step S5: For each agent, input the lower-layer observation state into the PPO new policy network, and output the action probability distribution; obtain the data and store it in ; randomly select multiple samples from to update ζ and p m (t) represents the power allocation strategy of the m-th agent;
[0013] If t % F = 0, then update the old policy network, that is execute Step S6;
[0014] Step S6: If t < T, then t = t + 1, and execute Step S3; otherwise, end the iteration.
[0015] In order to adapt to the changes in different beam environments, this scheme proposes a multi-agent hierarchical DQN-PPO (MHDP) algorithm. The main network θ and the target network θ′ (also called Q network and target Q network, respectively) have the same structure but different update frequencies, and together they constitute the DQN network, also known as the deep Q network, which belongs to the upper network; the new strategy network Old Policy Network and evaluation network It constitutes the proximal policy optimization PPO, which belongs to the lower layer network. Corresponding to the upper network, status Corresponding to the lower network. Correspondingly, α m (t) is the state Selected action; p m (t) is the state Select the action.
[0016] Since it is necessary to simultaneously optimize the discrete spectrum allocation strategy α and the continuous power allocation strategy p, it is difficult to solve and obtain the global optimal solution using traditional optimization algorithms. Therefore, a solution based on hierarchical multi-agent DRL is proposed. The link from each satellite to the user (ground satellite communication equipment) k in the beam m is regarded as an agent, and a hierarchical solution framework is adopted. The upper layer uses DQN to optimize the resource block allocation first, and the lower layer uses the proximal policy optimization (PPO) to optimize the power allocation in each resource block based on the spectrum allocation results. Each agent learns the strategy independently and optimizes the global AoI through local observation and collaboration.
[0017] As a further optimization of the above scheme, there are N satellite communication devices on the ground. In order to reduce interference, one RB allows a maximum of K ground satellite communication devices to access. <N;卫星的下行链路共享地面蜂窝小区的频谱资源为地面上的卫星通讯设备提供服务。地面用户通过距离等因素竞争共享RB资源,卫星为每个RB内的用户独立分配功率。
[0018] Each RB corresponds to an agent, which only observes the access and power allocation status of the RB to which it belongs, avoiding the explosion of the global state-action space.
[0019] Satellite communication equipment sharing the same RB is denoted as The information sent is recorded as {W 1,m ,…,W k,m , …, W K,m}; express The information sent by the next k-th satellite communication device; W k,m The public sub-signal of is the private sub-signal of W k,m ;
[0020] All the public sub-signals of are encoded to form a public stream, denoted as
[0021] All the K sub-signals of are respectively encoded to form K private streams, and the k-th private stream is denoted as
[0022] For the information sent is encoded, denoted as
[0023] wherein, is the transmission power of is the transmission power of
[0024] The signal received by the k-th ground satellite communication device under is z m is Gaussian white noise obeying a normal distribution;
[0025] represents the channel between the k-th ground satellite communication device in and the LEO, wherein, G s is the transmitting antenna gain of the LEO, G k is the transmitting antenna gain of the k-th ground satellite communication device; c represents the speed of light, f c represents the carrier frequency, and d represents the height of the LEO from the ground; represents the path loss in free space; δ k,m represents the Rayleigh small-scale fading distribution.
[0026] As a further optimization of the above solution, the signal-to-interference-plus-noise ratio of the public stream is calculated as:
[0027]
[0028] The signal-to-interference-plus-noise ratios of the private stream are respectively calculated as:
[0029]
[0030] p G represents the transmitting power of the m-th ground cell base station during communication; Denote the interference channel gain of the \(m\)-th ground cell base station to the \(k\)-th satellite communication device, i.e., β k,m and respectively represent the Ricean large-scale fading and Ricean small-scale fading; σ 2 represents the Gaussian white noise power.
[0031] According to the successive interference cancellation principle of RSMA signals, strengthen interference management for each resource block to improve spectral efficiency and communication quality.
[0032] As a further optimization of the above solution, in step S4, the state is expressed as:
[0033]
[0034] In step S5, the state is expressed as:
[0035]
[0036] Among them, is the spectrum allocation factor, represents that in time slot \(t\), the \(n\)-th satellite communication device is allocated to the \(m\)-th RB for receiving satellite signals, represents not being allocated and unable to communicate; respectively multiply with to confirm whether communication is possible;
[0037] represents the power allocation of the \(m\)-th RB to the \(n\)-th satellite communication device in time slot \(t\).
[0038] The present invention mainly minimizes the age of information by optimizing the spectrum allocation strategy α and the transmit power allocation strategy p, so the action space simultaneously includes discrete actions and continuous actions, expressed as a t =
[0039] [[α 1 (t), p 1 (t)], …, [α m (t), p m (t)], …, [α M (t), p M (t)]]。
[0040] As a further optimization of the above solution, the AoI of the \(k\)-th ground satellite communication device under the \(m\)-th RB in time slot \(t\), denoted as The calculation process is:
[0041] After the receiving end successfully receives and correctly demodulates the signal, remain unchanged as Δt, conversely, increase Δt to obtain
[0042] The constraint conditions include:
[0043]
[0044] Among them, represents the actual transmission rate of the common stream under the m-th RB; represents the maximum transmission rate allocated to the common stream; represents the actual transmission rate of the k-th private stream under the m-th RB; represents the maximum transmission rate allocated to the private stream; σ N is the interference power; Δσ is the power threshold.
[0045] That is, the interference power σ N is composed of the interference power of ground base station communication and Gaussian white noise. Δσ is the power threshold at which successive interference cancellation (SIC) can succeed, usually set to
[0046] Only when the rates allocated to the private stream and the common stream are not less than their actual rates can the signals of the private stream and the common stream be successfully demodulated, that is, the age of information does not increase.
[0047] The objective of the present invention is to find a suitable spectrum allocation strategy α and transmit power allocation strategy p under the constraint of the maximum transmission power to optimize the age of information of the satellite communication system. That is, to minimize the total age of information, that is:
[0048]
[0049] Among them, P max is the limit of the maximum transmission power of the satellite.
[0050] As a further optimization of the above scheme, the calculation of the reward value r m (t) is:
[0051]
[0052] The reward function needs to have the effect of guiding the agent to act in the direction of minimizing the age of information. The larger the value of the reward function, the less the AoI increases. The design of a good state space should not contain too much redundant information, otherwise it is difficult for the algorithm to find truly useful features, resulting in difficulty in convergence.
[0053] As a further optimization of the above solution, in step S4, the gradient descent method is used to minimize the first loss function to update θ and θ′;
[0054] The first loss function is expressed as:
[0055]
[0056] where Q predict represents the predicted Q value, and Q target represents the target Q value. represents selecting an action α in the target network θ′ for the state m (t + 1) to maximize ; γ is a hyperparameter representing the discount rate; r m (t + 1) represents the reward value obtained after selecting an action α for the state m (t + 1).
[0057] The upper-level DQN is responsible for dynamically allocating RBs and passing the allocation results to the lower-level PPO network. Specifically, it is a DRL algorithm based on the value function, which is used to handle the discrete RB allocation binary variable α. The core is to approximate the action value function through the Q neural network (QNN) and use the Bellman equation as the training objective, reflecting the expected cumulative discounted reward for executing the action α m (t) in the current state
[0058] The first loss function represents the difference between the predicted Q value and the target Q value.
[0059] The predicted Q value represents the expected cumulative reward for selecting the action a t in the current state s t , that is, the estimated value of the current Q table. The target Q value includes the immediate reward and the maximum Q value after future discounting, that is where γ is the discount factor.
[0060] As a further optimization of the above solution, in step S5, importance sampling is performed on randomly selected samples, which is expressed as:
[0061]
[0062] where x represents any one of the collected samples.
[0063] In the lower-layer network, since the power value is a continuous variable, DQN is no longer applicable, so PPO is adopted. The new policy network (ζ) is responsible for generating actions, and the old policy network (ζ') records historical policies for reference. However, at this time, the data distributions of the two policy networks are different, and it is necessary to correct the sampled samples and weight the data generated by different policies to reflect the probability deviation of their occurrences. This process is importance sampling.
[0064] Sample x from the distribution ζ and calculate the mean of the function f(x) with respect to x:
[0065]
[0066] Since it is impossible to directly sample from the distribution Sampling can only be done from another distribution Sampling data x, so a correction term needs to be added From The data sampled from is multiplied by this importance weight to correct the difference between the two distributions.
[0067] As a further optimization of the above solution, in step S5, the update amplitude of the new policy network is also controlled, that is, the probability ratio of the new policy network relative to the old policy network is calculated And clipped through a clipping function; the calculation formula is:
[0068] If the probability ratio Is within the interval [1 - ε, 1 + ε], the new policy is updated; otherwise, the update amplitude of the new policy network is readjusted until the probability ratio Is within the interval [1 - ε, 1 + ε]; where ε is a hyperparameter representing the clipping amplitude.
[0069] Although importance sampling solves the problem of a large amount of sampled data, in order to prevent the policy update from being too drastic, the gap between the policy network responsible for sampling and the policy network responsible for decision-making cannot be too large. Therefore, the PPO algorithm uses the method of "proximal policy optimization" to control the update amplitude of the policy. The probability ratio Reflects the difference degree before and after the update of the new policy network. Since the new policy network before the update is synchronously copied to the old policy network, the probability ratio Also represents the difference degree between the updated new policy network and the old policy network.
[0070] As a further optimization of the above solution, in step S5, the objective function is also optimized using the stochastic gradient ascent method; the objective function is expressed as:
[0071]
[0072] Among them, the corresponding function is λ is the truncation parameter, γ is the decay factor, δ represents the temporal difference error function, min() represents taking the minimum value, It means to be restricted within [1 - ε, 1 + ε].
[0073] The goal of PPO is to find a new policy such that the actions sampled from the policy are better than those sampled from the old policy. To achieve this goal, PPO defines an objective function and then uses stochastic gradient ascent to optimize this objective function. The objective function is expressed as:
[0074]
[0075] Among them, represents the advantage function, that is, A(s, a) = Q(s, a) - V(s), Q(s, a) refers to the action value function, and V(s) refers to the state value function; the advantage function is the advantage value of taking a certain action in a certain state relative to the average expected reward brought by the agent taking the average action in that state. If the advantage function value of an action is positive, then PPO tends to increase the probability of selecting this action, otherwise, it decreases the probability of selecting this action.
[0076] However, this advantage function is calculated only from the values within a single time step t, and there is a large deviation; therefore, by using Generalized Advantage Estimation (GAE) to balance the rewards in the near future and the far future, the variance of the advantage function estimation is reduced. GAE is calculated from the data of multiple time slots, that is The final objective function is then rewritten as:
[0077]
[0078] The update of the gradient policy is expressed as:
[0079]
[0080] Among them, α is the step size parameter, that is, the gradient of the objective function; finally, the local maximum value corresponding to the objective function is found by the method of gradient ascent.
[0081] In step S4, the loss function corresponding to the evaluation network is the second loss function, that is Among them, refers to the state value of the evaluation network in the state. At this time, the evaluation network is updated by the gradient descent method, that is Among them, that is, the gradient of the loss function.
[0082] Compared with the prior art, the present invention has the following beneficial effects:
[0083] (1) The present invention proposes a hierarchical deep reinforcement learning (DRL) architecture to solve the mixed-integer non-linear programming problem in the satellite communication system through a hierarchical decision-making mechanism. Specifically, the upper layer uses the DQN algorithm to handle the discrete RB (resource block) allocation problem, and the lower layer uses the PPO algorithm to implement continuous power control. The two cooperate to optimize the joint strategy of spectrum resource and power allocation. This architecture significantly reduces the age of information (AoI) by decoupling the discrete and continuous decision spaces hierarchically, providing an efficient resource management solution for the RSMA satellite communication system.
[0084] (2) The present invention abandons the strong dependence on complete CSI (channel state information) in traditional methods and proposes a dynamic power allocation decision method based on intelligent learning. Through the end-to-end environment interaction mechanism of deep reinforcement learning, the agent can directly learn the optimal strategy from partial observation information and adapt to the uncertainty of the channel state. In addition, for the large-scale user access scenario, this solution breaks through the bottleneck of traditional optimization algorithms in high-dimensional data processing through high-dimensional state space modeling and distributed computing, and realizes stable performance in complex scenarios.
[0085] (3) The present invention constructs a multi-agent deep reinforcement learning framework, adopting the centralized learning-distributed execution (CTDE) paradigm. In the offline training stage, the satellite centrally coordinates the global information interaction and policy learning of multiple agents; in the online deployment stage, the trained neural network is independently embedded in each agent, and a power allocation scheme is autonomously generated based on local observations. This framework takes into account both training efficiency and execution real-time performance, minimizes the system AoI through distributed autonomous decision-making, and supports flexible expansion under dynamic network topologies. Description of the Drawings
[0086] Figure 1 is a schematic diagram of the network communication architecture of the satellite-ground integrated network provided by an embodiment of the present invention;
[0087] Figure 2 is a schematic diagram of the deep learning network framework of the satellite-ground integrated network information age optimization method based on multi-agent hierarchical deep reinforcement learning provided by an embodiment of the present invention. Detailed Embodiments
[0088] To make the objectives, technical solutions, and advantages of the present invention more clear and understandable, the following will describe the technical solutions in the embodiments of the present invention clearly and completely in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0089] As Figure 1 , Figure 2 shown, this embodiment provides an information age optimization method for a satellite-ground integrated network based on multi-agent hierarchical deep reinforcement learning. The satellite-ground integrated network includes a multi-beam low-earth orbit satellite LEO. The LEO includes M orthogonal frequency band resource blocks RB, and one RB is occupied by a ground cell base station; one RB corresponds to one agent.
[0090] In this embodiment, there are N satellite communication devices on the ground; to reduce interference, at most K ground satellite communication devices are allowed to access one RB, and the K ground satellite communication devices form a ground user group, K < N; the downlink of the satellite shares the spectrum resources of the ground cellular network to provide services for the satellite communication devices on the ground. The ground users compete for sharing the RB resources based on factors such as distance, and the satellite independently allocates power to the users within each RB.
[0091] Each RB corresponds to one agent, which only observes the access and power allocation status of its affiliated RB to avoid the explosion of the global state-action space.
[0092] The satellite communication devices sharing the same RB (i.e., a ground user group) are denoted as The information sent is denoted as {W 1,m , …, W k,m , …, W K,m}; denotes the information sent by the k-th satellite communication device at is the common sub-signal of W k,m , is the private sub-signal of W k,m ;
[0093] All the common sub-signals of
[0094] are encoded to form a common stream, denoted as
[0095] All the K sub-signals of Encode the transmitted information, denoted as
[0096] wherein is the transmission power of is the transmission power of
[0097] The signal received by the k-th ground satellite communication device at the following is z m is Gaussian white noise obeying a normal distribution;
[0098] denotes the channel between the k-th ground satellite communication device in wherein, G s is the transmitting antenna gain of the LEO, G k is the transmitting antenna gain of the k-th ground satellite communication device; c represents the speed of light, f c denotes the carrier frequency, and d denotes the height of the LEO from the ground; denotes the path loss in free space; δ k,m denotes the Rayleigh small-scale fading distribution.
[0099] In this embodiment, the signal-to-interference-plus-noise ratio of the public stream is calculated as:
[0100]
[0101] The signal-to-interference-plus-noise ratios of the private stream are respectively calculated as:
[0102]
[0103] p G denotes the transmission power of the m-th ground cell base station during communication; denotes the interference channel gain of the m-th ground cell base station to the k-th satellite communication device, i.e., and respectively denote Ricean large-scale fading and Ricean small-scale fading; σ 2 denotes the Gaussian white noise power.
[0104] According to the successive interference cancellation principle of RSMA signals, strengthen interference management for each resource block to improve spectral efficiency and communication quality.
[0105] In this embodiment, the AoI of the k-th terrestrial satellite communication device under the m-th RB at time slot t is denoted as The calculation process is as follows:
[0106] When the receiving end successfully receives and correctly demodulates the signal, it remains unchanged as Δt. Conversely, Δt is increased to obtain
[0107] The constraint conditions include:
[0108]
[0109] Among them, represents the actual transmission rate of the common stream under the m-th RB; represents the maximum transmission rate allocated to the common stream; represents the actual transmission rate of the k-th private stream under the m-th RB; represents the maximum transmission rate allocated to the private stream; σ N is the interference power; Δσ is the power threshold.
[0110] That is, the interference power σ N is composed of the interference power of terrestrial base station communication and Gaussian white noise. Δσ is the power threshold at which successive interference cancellation (SIC) can succeed, usually set to
[0111] Only when the rates allocated to the private stream and the common stream are not less than their actual rates can the signals of the private stream and the common stream be successfully demodulated, that is, the age of information does not increase.
[0112] The objective of the present invention is to find a suitable spectrum allocation strategy α and transmit power allocation strategy p under the constraint of the maximum transmission power to optimize the age of information of the satellite communication system. That is, to minimize the total age of information, that is:
[0113]
[0114] Among them, P max is the limit of the maximum transmission power of the satellite.
[0115] The above method includes the following steps:
[0116] Step S1, initialize the network parameters θ, θ′, and corresponding to the main network, target network, new policy network, old policy network and evaluation network respectively, that is, corresponding to Figure 2 the Q-network, target Q-network, new policy, old policy and critic network in
[0117] Step S2: Initialize the experience pool and Set the total number of training rounds T, the allocation frequency Ns, and the update frequency F of the old policy network; initialize t = 1;
[0118] Step S3: If t % Ns = 0, then execute Step S4; otherwise, execute Step S5;
[0119] Step S4: For the current state of each agent Use the greedy policy (ε - greedy) to select an action α m (t); m represents the m - th agent; obtain the data representing the state transition process, and store it in ;
[0120] In this embodiment, the state is represented as:
[0121] is the spectrum allocation factor, indicating that at time slot t, the n - th satellite communication device is allocated to the m - th RB for receiving satellite signals, indicating not allocated and unable to communicate; are multiplied by respectively to confirm whether communication is possible.
[0122] Randomly select multiple samples from to update θ and θ′; α m (t) represents the spectrum allocation strategy of the m - th agent;
[0123] In this embodiment, the gradient descent method is used to minimize the first loss function to update θ and θ′;
[0124] The first loss function is represented as:
[0125]
[0126] where Q predict represents the predicted Q - value, Q target represents the target Q - value, represents that in the target network θ′, for the state select an action α m (t + 1) to make maximum; γ is a hyper - parameter representing the discount rate; r m (t + 1) represents the reward value obtained after the state selects an action α m (t + 1).
[0127] The upper - layer DQN is responsible for dynamically allocating RBs and passing the allocation results to the lower - layer PPO network. Specifically, it is a DRL algorithm based on the value function, which is used to handle the discrete RB allocation binary variable α. The core is to approximate the action - value function through a Q - neural network (QNN). And use the Bellman equation as the training objective, which reflects the expected cumulative discounted reward for executing the action α at the state under the current policy. m (t).
[0128] The first loss function represents the difference between the predicted Q - value and the target Q - value.
[0129] The predicted Q - value represents the expected cumulative reward for the network θ to select the action a at the current state s t , that is, the estimated value of the current Q - table. The target Q - value includes the immediate reward and the maximum Q - value after future discounting, that is t where γ is the discount factor. Execute step S6.
[0130] For each agent, input the lower - layer observation state
[0131] into the PPO new policy network, and output the action probability distribution; obtain the data and store it in ; In this embodiment, the state
[0132] is represented as: It represents the power allocation of the m - th RB to the n - th satellite communication device at time slot t. p (t) represents the power allocation strategy of the m - th agent. m Randomly select multiple samples from
[0133] and update ζ and In this embodiment, importance sampling is performed on the randomly selected samples, which is represented as:
[0134]
[0135] where x represents any one of the collected samples.
[0136] In the lower-level network, since the power value is a continuous variable, DQN is no longer applicable, so PPO is adopted. The new policy network (ζ) is responsible for generating actions, and the old policy network (ζ') records the historical policies for reference. However, at this time, the data distributions of the two policy networks are different, and it is necessary to correct the sampled samples and weight the data generated by different policies to reflect the probability deviation of their occurrences. This process is importance sampling.
[0137] Sample x from the distribution and calculate the mean of the function f(x) with respect to x:
[0138]
[0139] Since it is impossible to directly sample from the distribution only sample data x from another distribution therefore, a correction term is needed. For the data sampled from multiplying by this importance weight can correct the difference between the two distributions.
[0140] In this embodiment, the stochastic gradient ascent method is also used to optimize the objective function; the objective function is expressed as:
[0141]
[0142] where the corresponding function is λ is the truncation parameter, γ is the attenuation factor, δ represents the temporal difference error function, min() represents taking the minimum value, means restricting within [1 - ε, 1 + ε].
[0143] The goal of PPO is to find a new policy such that the actions sampled from the policy are better than those sampled from the old policy. To achieve this goal, PPO defines an objective function and then uses stochastic gradient ascent to optimize this objective function. The objective function is expressed as:
[0144]
[0145] where represents the advantage function, that is, A(s, a) = Q(s, a) - V(s), Q(s, a) refers to the action value function, and V(s) refers to the state value function; the advantage function is the advantage value of the average expected reward obtained by taking a certain action in a certain state compared to the average action taken by the agent in that state. If the advantage function value of an action is positive, then PPO will tend to increase the probability of selecting this action, and vice versa, it will decrease the probability of selecting this action.
[0146] However, this advantage function is calculated only from the values within a single moment t, resulting in a large deviation. Therefore, the Generalized Advantage Estimation (GAE) is used to balance the rewards in the near and far future and reduce the variance of the advantage function estimation. GAE is calculated from the data of multiple time slots, that is The final objective function is rewritten as:
[0147]
[0148] The update of the gradient policy is expressed as:
[0149]
[0150] where α is the step size parameter, i.e., the gradient of the objective function; finally, the local maximum corresponding to the objective function is found by the method of gradient ascent.
[0151] In step S4, the loss function corresponding to the evaluation network is the second loss function, that is where refers to the state value of the evaluation network in the state. At this time, the evaluation network is updated by the gradient descent method, that is where is the gradient of the loss function.
[0152] If t%F = 0, then the old policy network is updated, that is
[0153] In this embodiment, the update amplitude of the new policy network is also controlled, that is, the probability ratio of the new policy network relative to the old policy network is calculated and clipped by a clipping function; the calculation formula is:
[0154]
[0155] If the probability ratio is within the interval [1 - ε, 1 + ε], then the new policy is updated; otherwise, the update amplitude of the new policy network is readjusted until the probability ratio is within the interval [1 - ε, 1 + ε]; where ε is a hyperparameter representing the clipping amplitude.
[0156] Although importance sampling solves the problem of a large amount of sampled data, to prevent the policy update from being too drastic, the gap between the policy network responsible for sampling and the policy network responsible for decision-making cannot be too large. Therefore, the PPO algorithm uses the method of "Proximal Policy Optimization" to control the update amplitude of the policy. The probability ratio It reflects the degree of difference before and after the update of the new policy network. Since the new policy network before the update is synchronously copied to the old policy network, the probability ratio also represents the degree of difference between the updated new policy network and the old policy network.
[0157] Execute step S6.
[0158] In steps S4 and S5, the reward value r m (t) is calculated as:
[0159]
[0160] The reward function needs to have the effect of guiding the agent to act in the direction of minimizing the age of information. The larger the value of the reward function, the less the AoI increases. The design of a good state space should not contain too much redundant information, otherwise it is difficult for the algorithm to find truly useful features, resulting in difficulty in convergence.
[0161] Step S6: If t < T, then t = t + 1, and execute step S3; otherwise, end the iteration.
[0162] To adapt to the changes in different beam environments, this scheme proposes a multi-agent hierarchical DQN-PPO (MHDP) algorithm. Among them, the main network θ and the target network θ′ (which can also be called the Q network and the target Q network respectively), both have the same structure but different update frequencies, and the two together constitute the DQN network, that is, the deep Q network, belonging to the upper network; the new policy network the old policy network and the evaluation network constitute the proximal policy optimization PPO, belonging to the lower network. The state corresponds to the upper network, and the state corresponds to the lower network. Correspondingly, α m (t) is the action selected by the state ; p m (t) is the action selected by the state This invention mainly minimizes the age of information by optimizing the spectrum allocation strategy α and the transmit power allocation strategy p. Therefore, the action space contains both discrete actions and continuous actions, expressed as
[0163] a t = [[α 1 (t), p 1 (t)], …, [α m (t), p m (t)], …, [α M (t), p M (t)]].
[0164] Since it is necessary to optimize the mixed non - linear programming problem of the discrete spectrum allocation strategy α and the continuous power allocation strategy p simultaneously, it is very difficult to solve and obtain the global optimal solution using traditional optimization algorithms. Therefore, a solution method based on hierarchical multi - agent DRL is proposed. The link from each satellite to user k (ground satellite communication equipment) in beam m is regarded as an agent, and a hierarchical solution framework is adopted. The upper layer first uses DQN to optimize the resource block allocation, and the lower layer then uses proximal policy optimization (PPO) to optimize the power allocation within each resource block based on the result of spectrum allocation. Each agent independently learns the strategy and optimizes the global AoI through local observation and cooperation.
[0165] According to the disclosure and teaching of the above - mentioned specification, those skilled in the art to which the present invention pertains can also make changes and modifications to the above - mentioned embodiments. Therefore, the present invention is not limited to the specific embodiments disclosed and described above, and some modifications and changes to the present invention should also fall within the protection scope of the claims of the present invention. In addition, although some specific terms are used in this specification, these terms are only for convenience of description and do not constitute any limitation to the present invention.
Claims
1. A method for optimizing the information age of a satellite-ground fusion network based on multi-agent hierarchical deep reinforcement learning, wherein the satellite-ground fusion network includes a multi-beam low-orbit satellite LEO, the LEO includes M orthogonal frequency band resource blocks RB, and one RB is occupied by a ground cell base station; characterized in that: It includes the following steps: Step S1, initialize network parameters θ, θ′, ζ, ζ′ and They correspond to the main network, target network, new policy network, old policy network and evaluation network respectively; one RB corresponds to one intelligent agent; Step S2: Initialize the experience pool and Set the total training rounds T, the allocation frequency Ns, and the frequency F of updating the old strategy network; initialize t=1; Step S3: If t%Ns = 0, then execute Step S4; otherwise, execute Step S5; Step S4: For each agent's current state Use the greedy strategy to select action α m (t); m represents the mth agent; get data And deposit in; from Randomly select multiple samples from the α and update θ and θ′; m (t) represents the spectrum allocation strategy of the mth intelligent agent; execute step S6; Step S5: For each agent, input the lower layer observation state Go to the PPO new strategy network and output the action probability distribution; get the data And deposit in; from Randomly select multiple samples from the update ζ and p m (t) represents the power allocation strategy of the mth agent; If t%F = 0, then update the old policy network, i.e., ζ' = ζ; execute Step S6; Step S6: If t < T, then t = t + 1 and execute Step S3; otherwise, end the iteration.
2. The method for optimizing the information age of a satellite-ground fusion network based on multi-agent hierarchical deep reinforcement learning according to claim 1 is characterized in that: There are N satellite communication devices on the ground; one RB allows at most K ground satellite communication devices to access, where K < N; Satellite communication equipment sharing the same RB is denoted as The information sent is recorded as {W 1,m , …, W k,m , …, W K,m }; express The information sent by the next k-th satellite communication device; W k,m The public sub-signal of W k,m 's private child signal; All public sub-signals of are encoded to form a public stream, denoted as All K sub-signals of are encoded to form K private streams, and the kth private stream is recorded as right The information sent is encoded and recorded as in, for The transmission power, for The transmission power of The signal received by the kth ground satellite communication device is z m is Gaussian white noise that obeys normal distribution; express The channel between the kth ground satellite communication equipment and LEO in Among them, G s is the transmit antenna gain of LEO, G k is the transmitting antenna gain of the kth ground satellite communication device; c represents the speed of light, f c represents the carrier frequency, and d represents the height of LEO from the ground; represents the path loss in free space; δ k,m represents the Rayleigh small-scale fading distribution.
3. The method for optimizing the information age of a satellite-ground fusion network based on multi-agent hierarchical deep reinforcement learning according to claim 2 is characterized in that: The public stream The signal-to-interference-plus-noise ratio is Calculated as: The private flow The signal-to-interference-plus-noise ratios are Calculated as: p G represents the transmission power of the mth ground cell base station during communication; It represents the interference channel gain of the mth ground cell base station to the kth satellite communication equipment, that is, β k,m and Respectively represent the Rice large-scale fading and the Rice small-scale fading; σ 2 represents the Gaussian white noise power.
4. The method for optimizing the information age of a satellite-ground fusion network based on multi-agent hierarchical deep reinforcement learning according to claim 3 is characterized in that: In step S4, the state It is expressed as: In step S5, the state It is expressed as: in, is the spectrum allocation factor, Indicates that in time slot t, the nth satellite communication device is assigned to the mth RB for receiving satellite signals. It means it is not assigned and cannot communicate; Respectively and Multiply to confirm communication; It represents the power allocation of the mth RB to the nth satellite communication equipment in time slot t.
5. The method for optimizing the information age of a satellite-ground fusion network based on multi-agent hierarchical deep reinforcement learning according to claim 4 is characterized in that: The AoI of the kth ground satellite communication device in the mth RB at the tth time slot is denoted as The calculation process is: The constraint conditions include: in, Indicates the actual transmission rate of the public flow under the mth RB; Indicates the maximum transmission rate allocated to the public flow; Indicates the actual transmission rate of the kth private flow under the mth RB; represents the maximum transmission rate assigned to the private flow; σ N is the interference power; Δσ is the power threshold.
6. The method for optimizing the information age of a satellite-ground fusion network based on multi-agent hierarchical deep reinforcement learning according to claim 5 is characterized in that: Reward value r m (t) is calculated as:
7. The method for optimizing satellite-ground fusion network information age based on multi-agent hierarchical deep reinforcement learning according to claim 1 is characterized in that: In Step S4, use the gradient descent method to minimize the first loss function to update θ and θ'; The first loss function is expressed as: Among them, Q predict Represents the predicted Q value, Q target represents the target Q value, Indicates that in the target network θ′, it is the state Select an action α m (t+1), so maximum; γ is a hyperparameter, indicating the discount rate; r m (t+1) represents the state Select an action α m The reward value obtained after (t+1).
8. The method for optimizing the information age of a satellite-ground fusion network based on multi-agent hierarchical deep reinforcement learning according to claim 1 is characterized in that: In Step S5, perform importance sampling on randomly selected samples, which is expressed as: where x represents any one of the collected samples.
9. The method for optimizing the information age of a satellite-ground fusion network based on multi-agent hierarchical deep reinforcement learning according to claim 8 is characterized in that: In step S5, the update amplitude of the new policy network is also controlled, that is, the probability ratio r of the new policy network relative to the old policy network is calculated. ζ (t), and is clipped by the clipping function; the calculation formula is: If the probability ratio r ζ (t) is in the interval [1-ε, 1+ε], the new strategy is updated; otherwise, the update amplitude of the new strategy network is readjusted until the probability ratio r ζ (t) is in the interval [1-ε, 1+ε], where ε is a hyperparameter indicating the magnitude of the clipping.
10. The method for optimizing satellite-ground fusion network information age based on multi-agent hierarchical deep reinforcement learning according to claim 9 is characterized in that: In Step S5, also use the stochastic gradient ascent method to optimize the objective function; the objective function is expressed as: in, The corresponding function is λ is the cutoff parameter, γ is the attenuation factor, δ represents the time difference error function, min() represents the minimum value, clip(r ζ (t), 1-∈, 1+∈) means r ζ (t) is limited to [1-ε,1+ε].
Citation Information
Patent Citations
Air-space-ground network congestion control method based on deep reinforcement learning
CN119364423A
Unmanned aerial vehicle track adaptive optimization method based on information age
CN115696211A
Satellite-ground fusion network information age optimization method based on deep reinforcement learning
CN118157745A
Information age-oriented scheduling optimization method for related Internet of Things equipment of unmanned aerial vehicle relay Internet of Things system
CN118631820A
Space-air-ground integrated UAV-assisted IoT data collectioncollection method based on aoi
US20230239037A1