Low earth orbit satellite collaborative federated learning air aggregation system and joint scheduling method

By optimizing the beam hopping pattern through the federated learning aerial aggregation system coordinated by low-orbit satellites and the deep reinforcement learning algorithm, the model training and aggregation problems of large-scale equipment in different geographical areas are solved, the communication efficiency and training efficiency are improved, and resource utilization in dynamic environments is adapted.

CN120639149APending Publication Date: 2025-09-12XI AN JIAOTONG UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510791699.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing technologies cannot effectively support large-scale devices participating in model training and aggregation of low-orbit satellite collaborative federated learning in different geographical areas, resulting in communication bottlenecks and scheduling complexity.

Method used

A federated learning aerial aggregation system that collaborates with low-orbit satellites is adopted. Combining multi-beam low-orbit satellites and ground data processing centers, it realizes synchronous upload and physical layer aggregation of device terminals through beam hopping scheduling and aerial computing technology. It uses deep reinforcement learning algorithms to optimize beam hopping patterns and transmission power configurations, and constructs a joint scheduling method.

Benefits of technology

It improves the communication efficiency and training efficiency of wide-area federated learning tasks, reduces communication latency and bandwidth consumption, adapts to resource utilization in dynamic environments, and solves communication bottlenecks and scheduling complexity problems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120639149A_ABST
    Figure CN120639149A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of satellite communication and edge intelligent fusion, and relates to a low earth orbit satellite collaborative federated learning air aggregation system and a joint scheduling method, comprising: a plurality of static ground equipment terminals, a constellation network composed of a plurality of multi-beam low earth orbit (LEO) satellites, a ground data processing center, and a data processing center. The static ground equipment terminal is in communication connection with the multi-beam low-orbit satellite, and the satellite is in communication connection with the ground data processing center through a gateway; global model training is completed through a double-layer air aggregation mechanism of equipment terminal-LEO satellite-data processing center, an optimization model is constructed, a wave beam hopping mode, terminal transmitting power and a receiving end normalization factor are jointly optimized, a loss function of the training model is minimized, an air aggregation error is taken as a constraint, and the overall performance of the device terminal-LEO satellite-data processing center is improved. Adaptive scheduling is realized through deep reinforcement learning; according to the method, the multi-device cooperative training efficiency can be remarkably improved, and efficient federal learning of large-scale distributed devices is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of satellite communication and edge intelligence fusion, and relates to a federated learning aerial aggregation system and a joint scheduling method for low-orbit satellite collaboration. Background Art

[0002] Federated Learning (FL) is a distributed machine learning paradigm that allows devices to independently train models locally using private data and upload the updated model parameters to a server for aggregation, thereby completing iterative optimization of the global model. This mechanism avoids the centralized transmission of raw data, significantly reducing the communication burden while ensuring privacy protection. It has been widely used in mobile terminals, IoT edge devices, and intelligent sensing scenarios.

[0003] Traditional federated learning relies primarily on terrestrial cellular communications, which poses serious limitations in remote, wide-area scenarios or scenarios with weak terrestrial communication infrastructure (such as remote mountainous areas, oceans, and high-altitude platforms). Low Earth Orbit (LEO) satellites, with their low latency, wide coverage, and frequent visibility, are crucial for building a wide-area federated learning communication infrastructure. Implementing satellite-ground collaborative federated learning through LEO satellite relays can provide a unified learning platform for large-scale heterogeneous devices worldwide, improving system generalization and scalability.

[0004] In federated learning systems based on low-orbit satellites, the process of uploading model parameters remains a major bottleneck, particularly when there are many participating devices and limited communication links. This significantly impacts model aggregation efficiency and system convergence speed. To improve uplink communication efficiency, "over-the-air computing" (OTA) technology has been proposed in recent years. This allows multiple devices to transmit weighted model parameters in parallel over wireless channels. The receiving end achieves physical layer aggregation through signal superposition, significantly reducing communication latency and bandwidth consumption.

[0005] Furthermore, to improve spectrum utilization and system capacity, multi-beam low-orbit satellites typically employ beam hopping technology, dynamically activating some beams across different geographic regions to flexibly respond to regional resource demands. In federated learning systems, beam hopping can be used as an effective means of spatial resource scheduling, selecting some devices to participate in upload operations during each round, thereby controlling interference and improving aggregation efficiency.

[0006] However, although existing studies have modeled and optimized the applications of air computing or beam hopping, they are still unable to effectively support large-scale devices participating in model training and aggregation in different geographical areas, resulting in communication bottlenecks and scheduling complexity.

[0007] Therefore, there is an urgent need for an aerial aggregation system or method for collaborative federated learning of low-orbit satellites, which can efficiently implement model aggregation operations under the beam hopping scheduling mechanism and improve the performance and resource utilization efficiency of satellite-ground federated model training. Summary of the Invention

[0008] The technical solution adopted by the present invention to solve the technical problem is: a low-orbit satellite-coordinated federated learning aerial aggregation system, including:

[0009] Multiple stationary ground equipment terminals, used to collect data locally and perform local training based on the received global model, and generate and upload local model update parameters;

[0010] A LEO constellation composed of multiple multi-beam low-orbit satellites is used to receive local model parameters synchronously uploaded by the stationary ground equipment terminals in the coverage area according to a preset beam hopping pattern, and perform weighted superposition at the physical layer to achieve aerial aggregation;

[0011] A ground data processing center is configured to receive intermediate model parameters of the multi-beam low-orbit satellite through secondary air aggregation to complete a global model update;

[0012] The stationary ground equipment terminal is communicatively connected to the multi-beam low-orbit satellite, and the satellite is communicatively connected to the ground data processing center via a gateway.

[0013] Preferably, the stationary ground equipment terminals are distributed in different area blocks, and each area block has at most one equipment terminal;

[0014] For the area block without the stationary ground equipment terminal, a virtual terminal with a training data size of 0 is introduced;

[0015] Each of the stationary ground equipment terminals collects training data in real time in each round of federated learning, and the training data of the stationary ground equipment terminals that are not selected in a round of federated learning are cached to the next round.

[0016] Preferably, in the LEO constellation composed of multiple multi-beam low-orbit satellites, each satellite can cover multiple area blocks and transmit multiple beams simultaneously, and one area block may be covered by multiple satellites at the same time.

[0017] The present invention also discloses a joint scheduling method for a federated learning air aggregation system coordinated by low-orbit satellites. The joint scheduling method uses the above-mentioned federated learning air aggregation system and includes the following steps:

[0018] Step S1: Obtain the amount of training data for each ground device terminal in the current round, and the channel gains from the device terminal to the satellite and from the satellite to the data processing center;

[0019] Step S2: setting system optimization constraints;

[0020] Step S3: Construct a long-term optimization problem for joint scheduling with the objective function of maximizing the total amount of data involved in model training;

[0021] Step S4: Solve the optimization problem to obtain the optimal scheduling method.

[0022] Preferably, the constraints in step S2 include:

[0023] Beam constraint: Each device terminal is illuminated by at most one beam, and each device terminal can only be illuminated by the beam of one satellite at a time. Each satellite can activate at most V beams at a time. Each satellite can only select devices within its coverage area. The distance between devices i and j selected by two satellites must exceed the minimum interference distance to avoid inter-satellite interference.

[0024] Power constraint: The power of the device terminal and satellite must not exceed their maximum transmit power, and the transmit power of unselected devices is constrained to 0.

[0025] Mean squared error constraint: The mean squared error (MSE) of the global model aggregation does not exceed the set threshold.

[0026] Preferably, the optimization problem P0 in step S3 is:

[0027]

[0028] In formula (23), X n (t), b k (t), b n (t), η n (t), η(t) represent the variables to be optimized, X n (t) = {x n,1 (t),…,x n,k (t),…x n,C (t)} represents a zero-one vector, where x n,k (t) = 1 means that device k is selected by satellite n, otherwise x n,k (t) = 0, b k (t), b n (t) are the transmission power coefficients of device terminal k and satellite n, respectively, η n (t) and η(t) represent the receiving normalization factors of the satellite and ground station, respectively. C1 to C5 represent beam constraints, C6 and C7 represent power constraints, and C8 represents mean square error constraint.

[0029] More preferably, the specific steps in step S4 include:

[0030] Step S401: Formula (23) is reformulated as a Markov Decision Process (MDP). The Markov Decision Process includes a state space, an action space, a reward function, and a state transition function. The state of the state space includes the amount of training data of each ground device terminal in the current round, the channel gains from the device terminal to the satellite and from the satellite to the data processing center, and the actions of the action space include the beam hopping mode, the transmit power of the device terminal and the satellite, and the receiver normalization factors of the satellite and the data processing center.

[0031] Step S402: construct an immediate reward function with the goal of maximizing the total amount of data involved in model training;

[0032] Step S403: Using a deep reinforcement learning strategy training algorithm to perform strategy learning on the Markov decision process to obtain a joint strategy function for online scheduling;

[0033] Step S404: input the policy function according to the current state and output the corresponding joint scheduling action.

[0034] Preferably, the federated learning process includes the following steps: in each round of learning, the ground data processing center sends the global model to all satellites; each satellite selects a subset of device terminals within the corresponding coverage area, and the same device terminal cannot be selected by two satellites at the same time; each satellite illuminates the beam to the selected device terminal and sends the global model; each device terminal runs a local update algorithm using the training data of this round and the global model to generate an updated local model; each satellite locally integrates the local updates of the selected device terminals to obtain an intermediate model; the data processing center globally integrates the local models of all satellites to obtain an updated global model.

[0035] The beneficial effects of the present invention are:

[0036] 1. In response to the practical needs of widely distributed ground equipment and limited communication resources, this paper proposes a satellite-ground collaborative federated learning system architecture suitable for low-orbit satellite networks. It can effectively support large-scale equipment to participate in model training and aggregation in different geographical areas. By combining beam hopping scheduling and aerial aggregation mechanisms, this paper provides a systematic solution to the communication bottlenecks and scheduling complexity problems in wide-area federated learning tasks.

[0037] 2. This invention proposes a two-layer aerial aggregation mechanism: device terminal-LEO satellite-data processing center. The first step of physical layer aggregation is completed by synchronous uplink transmission of device terminals within the satellite beam coverage area. The aggregation result is then forwarded to the data processing center via satellite synchronous relay to complete the second step of fusion, significantly reducing communication latency and bandwidth overhead.

[0038] 3. The present invention models the joint scheduling problem as a Markov decision process and uses a deep reinforcement learning algorithm to automatically learn the optimal joint strategy of beam hopping mode, transmit power configuration and receive normalization factor, taking into account both training efficiency and communication quality, and is suitable for deployment in dynamic environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 Schematic diagram of a federated learning aerial aggregation system model for a low-orbit satellite-coordinated federated learning aerial aggregation system and a joint scheduling method of the present invention;

[0040] Figure 2 is a flowchart of the federated learning process of the present invention;

[0041] Figure 3 It is a flowchart of the joint scheduling method based on deep reinforcement learning of the present invention. DETAILED DESCRIPTION

[0042] The following will provide a clear and complete description of the relevant technologies in the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0043] refer to Figures 1 to 3 , this embodiment provides a low-orbit satellite collaborative federated learning air aggregation system and joint scheduling method, such as Figure 1 As shown in the figure, the federated learning aerial aggregation system includes a ground data processing center, a LEO constellation consisting of multiple multi-beam low-orbit satellites, and multiple stationary ground equipment terminals.

[0044] Consider a constellation consisting of N LEO satellites Serving an area consisting of S blocks area There are K device terminals in total, forming a device group For the ground part, consider that each area block has at most one device terminal, i.e. |k s |≤1, where k s Indicates the device terminal in area s. In the air, each satellite can cover C blocks and transmit up to V beams. The area covered by satellite n is denoted as The beam set transmitted by satellite n is denoted as According to the multiple coverage of LEO satellite communication system, a regional block may be covered by multiple satellites, i.e.

[0045] The federated online learning framework is described as follows:

[0046] Without loss of generality, this paper considers a total of T rounds of federated learning, which can be expressed as The duration of each round is τ. In round t:

[0047] 1) The data processing center sends the global model z(t) to all satellites

[0048] 2) Each satellite Select the corresponding coverage area A subset of the device terminals within In order to ensure that the weights of each terminal's local model are not repeatedly superimposed during model aggregation, it is stipulated that the same device terminal cannot be selected by two satellites at the same time, that is, The set of device terminals selected by all satellites is denoted as

[0049] 3) Each satellite A beam Illuminate the selected device terminal And send the global model z(t);

[0050] 4) Define the local dataset of each device terminal k in this round as

[0051] Represents all the training data used. Each device terminal Use this round's local dataset Run the local update algorithm with the global model z(t) to generate the updated local model z k (t+1);

[0052] 5) Each satellite For the selected device terminal The local update is used to perform local data integration and obtain z n (t+1), at the same time each device terminal Start collecting the next round of training data;

[0053] 6) The data processing center integrates the global data of the local models of all satellites to obtain the updated global model z(t+1).

[0054] The local loss function of device k in round t is defined as:

[0055]

[0056] Where z(t) is the model parameter that needs to be optimized, Representing a dataset The size of is the loss function for the input-output data pair (x,y) under the model parameters z(t).

[0057] The goals of federated learning are:

[0058]

[0059] where d is the dimension of the model parameter vector z(t), is the preprocessing scalar for device k, is the post-processing scalar of the ground station.

[0060] In order to facilitate processing and simplify the mapping relationship between devices and regional blocks, it is recorded that there is a device terminal in each regional block, and the amount of training data corresponding to the "virtual terminal" that does not actually exist is recorded as 0, that is, thus Can be used with One-to-one correspondence, the set of all device terminals covered by satellite n can be equivalent to The beam hopping pattern of satellite n in round t is recorded as:

[0061] X n (t) = {x n,1 (t),…,x n,k (t),…x n,C (t)} (3)

[0062] where x n,k (t)∈{0,1} indicates whether satellite n transmits beam v k Illuminate the device k it covers, so it satisfies

[0063] Although satellite n only transmits one beam to illuminate a selected device terminal k, the beam it transmits All beams can receive the local model signal transmitted by device k. Assume that the device terminal compensates for the Doppler frequency offset caused by satellite movement and the rainfall attenuation can be ignored. Let G k is the transmitting antenna gain of device k, G v is the receiving antenna gain of beam v. Therefore, the device Beam with satellite n The channel gain between is expressed as:

[0064]

[0065] where d k,v (t) is the distance between device k and beam v, and λ is the wavelength.

[0066] G k and Gv The values ​​of depend on the radiation patterns of the user and satellite antennas, respectively. Based on 3GPP TR 38.811, this embodiment gives the satellite antenna pattern as follows:

[0067]

[0068] where θ represents the off-axis angle, a is the radius of the antenna aperture, and J1(·) is a first-order Bessel function of the first kind.

[0069] The maximum antenna gain is given by:

[0070]

[0071] Where D′=2a is the diameter of the antenna aperture.

[0072] Likewise, the radiation pattern of the device terminal transmitting antenna is given by:

[0073]

[0074] Where θ is the off-axis angle between the main lobe direction of the device terminal antenna and the satellite beam, is the maximum transmission gain, when θ exceeds the main axis range θ main When , a gain attenuation ΔG(θ) is generated, and ΔG(θ) is modeled as follows:

[0075]

[0076] where k a is the attenuation factor, which controls the size of the initial attenuation, n a is the sharpness parameter, which controls how fast the gain decreases. It can be calculated by formula (6).

[0077] Similar to (7), this embodiment gives the radiation pattern of the ground station:

[0078]

[0079] in The maximum receiving gain of the ground station is calculated by formula (6). Considering that the transmission gain and receiving gain of the satellite antenna are consistent, and when the satellite uploads the intermediate model to the ground station, the antenna is aimed at the ground station. Therefore, the satellite The channel gain between the ground station and the ground station can be expressed as:

[0080]

[0081] where d n,g (t) is the distance between satellite n and the ground station.

[0082] The intermediate model parameters obtained by local aggregation of each satellite n are:

[0083]

[0084] where z k (t) is the local update of device k, is the post-processing scalar for satellite n. It defines the signal vector for each local update z k (t) is standardized to unit variance, i.e. In each time slot i∈{1,…,d}, devices Send a signal to the corresponding satellite n The signal received by satellite n under ideal conditions is calculated over the air:

[0085]

[0086] In order to simplify the symbol representation, the time slot index i of the symbol is omitted in this embodiment. During the uplink process, satellite n receives the signal from device k. signals. The signal actually received by satellite n can be expressed as:

[0087]

[0088] where b k (t) is the power control coefficient of the transmitter, is additive Gaussian white noise, η n (t) is the normalization factor for satellite n. The power limit for device k is:

[0089]

[0090] P k >0 indicates the maximum transmit power of the device.

[0091] The final model parameters obtained by global aggregation of the ground station are:

[0092]

[0093] is the post-processing scalar of satellite n. Using a similar method, the signal vector transmitted by the satellite is defined as:

[0094]

[0095] Right now for The unit variance standardization satisfies If the distortion of the equipment terminal-LEO link and the LEO-ground station link is ignored, the signal received by the ground station in each time slot under ideal conditions can be calculated over the air:

[0096]

[0097] The signal actually received by the ground station can be expressed as:

[0098]

[0099] where b n (t) is the transmission power control coefficient of satellite n, and η(t) is the normalization factor of the ground station. The power limit of satellite n is:

[0100]

[0101] P n >0 is the maximum transmit power of the satellite.

[0102] The ideal signal s(t) in equation (17) and the estimated signal in equation (18) The distortion between them is measured by the mean square error (MSE), which is defined as:

[0103]

[0104] For the local data set collected by device terminal k in round t definition:

[0105]

[0106] in is the training data collected by device terminal k in round t. That is, the training data of devices that were not selected in the previous round will be cached in this round. According to the federated learning framework proposed in this invention, the key factor affecting the performance of the global model z(t) is the sum of the local data volume of all device terminals participating in the training. It can be expressed as:

[0107]

[0108] Due to the limitations on the number of beams, transmit power, and MSE mentioned in the beam hopping and over-the-air computation sections, it is difficult to aggregate local models for all devices in a single pass. It is also important to note that beams from different satellites may serve adjacent cells in the same pass, leading to significant inter-satellite interference. Therefore, the beam hopping pattern for each satellite needs to be carefully planned. The optimization goal is to maximize the total amount of training data utilized by the local model. The long-term optimization problem can be modeled as follows:

[0109]

[0110] In the above optimization problem, X n (t), bk (t), b n (t), η n (t), η(t) are the variables to be optimized, X n (t) = {x n,1 (t),…,x n,k (t),…x n,C (t)} is a zero-one vector, where x n,k (t) = 1 means that device k is selected by satellite n, otherwise x n,k (t) = 0. b k (t), b n (t) are the transmission power coefficients of device terminal k and satellite n, respectively, η n (t) and η(t) are the receiving normalization factors of the satellite and ground station respectively. C1 will optimize the variable x n,k (t) is constrained to be a zero-one variable. C2 limits each device terminal to be selected by at most one satellite at a time. C3 stipulates that each satellite can activate at most V beams at a time, that is, each satellite can select at most V device terminals within its coverage area per round. C4 limits each satellite to only select devices within its coverage area. C5 requires that the distance between devices i and j selected by each of the two satellites exceeds the minimum interference distance to avoid inter-satellite interference. Where ω i,j =dist(i,j) represents the distance between devices i and j. C6 and C7 limit the power of the device terminal and the satellite to not exceed their maximum transmission power P respectively. k and P n , and at the same time constrain the transmission power of unselected devices to 0. C8 limits the MSE threshold ρ of the global model aggregation.

[0111] Equation (23) is clearly nonconvex and nonlinear, and involves both continuous and discrete variables. Therefore, this optimization problem is NP-hard, and it is difficult to find an optimal solution in polynomial time. Therefore, in the following sections, this embodiment will first reformulate Equation (23) as an MDP and then employ a deep reinforcement learning (DRL) algorithm to efficiently solve it.

[0112] In DRL, an agent interacts with a dynamic environment and takes actions based on the observed state at each round. The environment then changes, and rewards are returned as an evaluation of the actions taken. Through continuous interaction with the environment, the agent searches for the optimal strategy from its experience and learns the strategy with the highest long-term reward. In the optimization problem of this embodiment, the agent continuously makes decisions on beam scheduling and transmit power with the goal of maximizing the total amount of training data. The agent's state space, action space, and reward function are defined as follows:

[0113] State: During the decision-making process, the amount of local training data and channel conditions of the device will affect the scheduling decision. Therefore, this embodiment defines the state of round t as:

[0114]

[0115] in, h represents the amount of training data that device k has in the tth round of coverage by satellite n, which can be obtained by (21). n,k (t)=[h n1,k (t) h n2,k (t) … h nC,k (t)] T is the channel gain from the C beams of the t-th round satellite n to the device k it covers, which can be obtained by (4).

[0116] Action: As an intelligent agent, the ground station should make decisions on the beam pattern, the transmit power of the satellite and the device, and the normalization factors of the ground station and the satellite. Therefore, this embodiment defines the action at time step t as:

[0117] a(t)={X(t),p k (t), p n (t), η n (t),η(t)}, (25)

[0118] where X(t) is the set of beam pattern decisions, p k (t) and p n (t) represents the set of transmission power of the device and satellite respectively, η n (t) represents the set of normalization factors received by the satellite, and η(t) is the normalization factor received by the ground station.

[0119] Reward: In order to maximize the goal defined in Equation (23), the reward for round t is designed as follows:

[0120]

[0121] Due to the dynamic nature of the satellite network environment, the present invention proposes a joint scheduling algorithm based on model-free policy gradient, which repeatedly updates the parameterized strategy through gradient descent to learn the optimal strategy.

[0122] This implementation considers a diagonal Gaussian strategy applicable to the continuous action space, denoted as π θ (a(t)|s(t)), where θ is the vector of corresponding parameters. This stochastic policy can be approximated by a neural network called an actor network, and the parameters can be regarded as its weights and biases. Using the idea of ​​importance sampling, the agent samples actions from the parameterized policy at each time step, i.e., a(t)~π θ(·|s(t)), and the sampling uncertainty can be reflected by the entropy defined as follows:

[0123]

[0124] The entropy also indicates the exploration ability by evaluating the average degree of action selection. Specifically, the higher the entropy, the stronger the exploration ability of the policy.

[0125] At the same time, using the value function V φ (s(t)) estimates the long-term value and is approximated by a neural network with parameter φ (called the critic network).

[0126] As described above, the policies for beam scheduling, power, and normalization factor configuration are optimized by maximizing expected reward via gradient descent. However, in the dynamic environment of this embodiment, it is difficult for the agent to accurately predict the reward, and the policy gradient estimate has a large variance.

[0127] Among the various expressions used to adjust policy gradient updates, the advantage function exhibits relatively low variance. It encourages actions that are better than the policy's default behavior while discouraging actions that are below average. However, inaccurate estimates of the advantage function can produce bias and variance. To ensure effective and accurate estimates, this paper adopts a generalized advantage estimate widely used in model-free policy gradient algorithms. The exponentially weighted estimate of the advantage function can be given as:

[0128]

[0129] c v and c b are parameters that adjust the bias and variance. Specifically, c v This reduces the variance, but at the expense of introducing bias in the policy gradient estimate. b Helps to balance the bias-variance trade-off. b The value is usually lower than c v The value of .

[0130] In order to ensure the stability and steady improvement of the parameterized policy during the learning process, the objective function of the pruning agent based on the nearest policy optimization is defined as:

[0131]

[0132] Specifically, Measure the difference between the old and new policies. θ old represents the parameters of the old strategy before the update, ∈ is a hyperparameter that limits the change of strategy, The ratio Clipping to the interval [1-∈, 1+∈]. This clipping can ensure the appropriate change from the old policy to the new policy and further contribute to the stable improvement of policy updates. As mentioned above, the entropy of the policy can reflect the randomness of action selection. Therefore, this embodiment uses the coefficient c e Supplement the entropy reward to encourage full exploration and rewrite the agent objective function as:

[0133]

[0134] Obviously, the new objective function combines the advantages of bias-variance trade-off, stable improvement and sufficient exploration, which are helpful for learning the optimal beam scheduling and resource allocation strategies in the proposed joint scheduling framework.

[0135] At the end of each training episode, the state transitions stored in the experience replay buffer are sampled to update the parameterized policy and estimated value function, which are approximated by the actor and critic network. The parameterized policy can be updated by maximizing the clipped agent objective as follows:

[0136]

[0137] And the estimated value function can be updated by minimizing the mean square error between the estimated value function and the target value function:

[0138]

[0139] where V(t) is the target value calculated as the discount of the reward, i.e. The above updates can be achieved through a gradient descent optimizer. After multiple training and updates, the agent learns the optimal strategy. This implementation can obtain the optimal solution for the beam scheduling and resource allocation strategy with the highest cumulative reward.

[0140] In summary, the architecture proposed in this invention can greatly improve the training efficiency and communication quality of wide-area federated learning tasks.

[0141] It should be emphasized that the above are only preferred embodiments of the present invention and do not limit the present invention in any form. Any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention are still within the scope of the technical solution of the present invention.

Claims

1. Low-orbit satellite-coordinated federated learning aerial aggregation system, characterized by: include: Multiple stationary ground equipment terminals, used to collect data locally and perform local training based on the received global model, and generate and upload local model update parameters; A LEO constellation composed of multiple multi-beam low-orbit satellites is used to receive local model parameters synchronously uploaded by the stationary ground equipment terminals in the coverage area according to a preset beam hopping pattern, and perform weighted superposition at the physical layer to achieve aerial aggregation; A ground data processing center is configured to receive intermediate model parameters of the multi-beam low-orbit satellite through secondary air aggregation to complete a global model update; The stationary ground equipment terminal is communicatively connected to the multi-beam low-orbit satellite, and the satellite is communicatively connected to the ground data processing center via a gateway.

2. The low-orbit satellite-coordinated federated learning aerial aggregation system according to claim 1, characterized in that: The stationary ground equipment terminals are distributed in different area blocks, and each area block has at most one equipment terminal; For the area block without the stationary ground equipment terminal, a training data virtual terminal is introduced; Each of the stationary ground equipment terminals collects training data in real time in each round of federated learning, and the training data of the stationary ground equipment terminals that are not selected in a round of federated learning are cached to the next round.

3. The low-orbit satellite-coordinated federated learning aerial aggregation system according to claim 1, characterized in that: In the LEO constellation composed of multiple multi-beam low-orbit satellites, each satellite can cover multiple area blocks and transmit multiple beams simultaneously, and one area block may be covered by multiple satellites at the same time.

4. A joint scheduling method for a federated learning aerial aggregation system coordinated by low-orbit satellites, characterized in that: The joint scheduling method is used in the federated learning air aggregation system according to any one of claims 1 to 3, and the joint scheduling method comprises the following steps: Step S1: Obtain the amount of training data for each ground device terminal in the current round, and the channel gains from the device terminal to the satellite and from the satellite to the data processing center; Step S2: setting system optimization constraints; Step S3: Construct a long-term optimization problem for joint scheduling with the objective function of maximizing the total amount of data involved in model training; Step S4: Solve the optimization problem to obtain the optimal scheduling method.

5. The joint scheduling method of the low-orbit satellite coordinated federated learning air aggregation system according to claim 4 is characterized in that: The constraints in step S2 include: Beam constraint: Each device terminal is illuminated by at most one beam, and each device terminal can only be illuminated by the beam of one satellite at a time. Each satellite can only select devices within its coverage area. The distance between devices i and j selected by two satellites must exceed the minimum interference distance to avoid inter-satellite interference. Power constraint: The power of the device terminal and satellite must not exceed their maximum transmit power, and the transmit power of unselected devices is constrained to 0. Mean square error constraint: The mean square error of the global model aggregation does not exceed the set threshold.

6. The joint scheduling method of the low-orbit satellite coordinated federated learning air aggregation system according to claim 4 is characterized in that: The optimization problem P0 in step S3 is: In formula (23), X n (t), b k (t), b n (t), η n (t), η(t) represent the variables to be optimized, X n (t) = {x n,1 (t),…,x n,k (t),…x n,C (t)} represents a zero-one vector, where x n,k (t) = 1 means that device k is selected by satellite n, otherwise x n,k (t) = 0, b k (t), b n (t) are the transmission power coefficients of device terminal k and satellite n, respectively, η n (t) and η(t) represent the receiving normalization factors of the satellite and ground station, respectively. C1 to C5 represent beam constraints, C6 and C7 represent power constraints, and C8 represents mean square error constraint.

7. The joint scheduling method of the low-orbit satellite coordinated federated learning aerial aggregation system according to claim 6 is characterized in that: The specific steps in step S4 include: Step S401: Formula (23) is reformulated as a Markov decision process. The Markov decision process includes: a state space, an action space, a reward function, and a state transition function. The state of the state space includes: the amount of training data of each ground device terminal in the current round, the channel gain from the device terminal to the satellite and from the satellite to the data processing center. The action of the action space includes: the beam hopping mode, the transmit power of the device terminal and the satellite, and the receiving end normalization factor of the satellite and the data processing center. Step S402: construct an immediate reward function with the goal of maximizing the total amount of data involved in model training; Step S403: Using a deep reinforcement learning strategy training algorithm to perform strategy learning on the Markov decision process to obtain a joint strategy function for online scheduling; Step S404: input the policy function according to the current state and output the corresponding joint scheduling action.

8. The joint scheduling method of the low-orbit satellite coordinated federated learning aerial aggregation system according to claim 4, characterized in that: The federated learning process includes the following steps: in each round of learning, the ground data processing center sends the global model to all satellites; each satellite selects a subset of device terminals within the corresponding coverage area, and the same device terminal cannot be selected by two satellites at the same time; each satellite illuminates the selected device terminal with a beam and sends the global model; each device terminal runs a local update algorithm using the training data of this round and the global model to generate an updated local model; each satellite locally integrates the local updates of the selected device terminals to obtain an intermediate model; and the data processing center globally integrates the local models of all satellites to obtain an updated global model.

Citation Information

Cited By

  • Federal learning-based routing policy network training method and device

    CN121509301A