Low earth orbit satellite decentralized federated learning method and system based on dichotomy and multiplication
By employing a decentralized federated learning method that combines binary search and doubling, the model training and data transmission loads in low-Earth orbit satellite networks are distributed, optimizing latency and energy consumption. This solves the inefficiency problem caused by concentrated load in traditional federated learning and improves network stability and scalability.
Patent Information
- Application Number
- CN202411703794.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-26
- Publication Date
- 2026-01-06
- Estimated Expiration
- 2044-11-26
AI Technical Summary
Traditional federated learning in low-Earth orbit satellite networks suffers from the problem of concentrated load on servers, leading to increased overall latency and low network efficiency. Furthermore, existing decentralized methods do not adequately consider load and energy consumption issues.
A decentralized federated learning method based on binary search and doubling is adopted. Through the first round of model parameter exchange, local training, reduction and distribution and full collection stages, the model training and data transmission load are distributed. Particle swarm optimization algorithm is used to optimize latency and energy consumption to achieve load balancing.
It effectively solves the problem of load concentration on servers, improves the stability and scalability of low-Earth orbit satellite networks, reduces overall latency and energy consumption, and improves the efficiency of federated learning.
Smart Images

Figure CN119519817B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of federated learning technology, and specifically relates to a decentralized federated learning method and system for low-orbit satellites based on binary division and doubling. Background Technology
[0002] In recent years, fifth-generation mobile communication systems have matured and been put into practice, achieving high-speed, high-capacity communication between ground terminals. However, due to geographical limitations, the deployment of global internet coverage, the Internet of Things, and emergency communications in base station networks still faces challenges. To supplement terrestrial networks and achieve global coverage, low-Earth orbit (LEO) satellite networks, with their advantages of low cost, fast cycle time, and the ability to handle complex tasks using machine learning, have become a popular research direction for sixth-generation mobile communication systems. Considering data privacy issues, federated learning allows users to exchange only model parameters rather than raw data, thereby protecting privacy and reducing risks. However, as users of federated learning, LEO satellites are distributed in different orbits, forming a highly distributed architecture. However, traditional federated learning requires receiving the global model after all models are aggregated before training local models, leading to increased overall latency. The heterogeneity and computing power differences of LEO satellite networks further exacerbate this problem, reducing the efficiency of the entire federated learning process. In addition, during the model parameter exchange process, some intermediate nodes, especially aggregation servers, are overloaded, while other nodes are lightly loaded, resulting in load bottlenecks. Considering that LEO satellite networks also need to support other services, maintaining node load balance is crucial for network performance.
[0003] Most related research focuses on centralized federated learning architectures, where the federated learning server acts as a central node, and other federated learning users transmit data to the server for aggregation, leading to unbalanced network load. Although some studies have proposed decentralized federated learning methods, they mainly focus on energy or latency constraints, without fully considering load issues. Furthermore, the total energy consumption of each node along the routing path, especially the energy consumption of intermediate satellites, is often neglected. In existing decentralized federated learning methods for low-Earth orbit satellite networks, a comprehensive consideration of latency, load, and energy constraints remains insufficient. Summary of the Invention
[0004] The purpose of this invention is to provide a decentralized federated learning method and system for low-Earth orbit satellites based on binary division and multiplication, so as to solve the above-mentioned problems.
[0005] To achieve the above objectives, the present invention adopts the following technical solution:
[0006] In a first aspect, this invention provides a decentralized federated learning method for low-Earth orbit satellites based on binary division and doubling, including:
[0007] In the first round of federated learning, federated learning users obtain aggregated models by exchanging model parameters;
[0008] In the aggregation model, federated learning users train local models on their respective datasets;
[0009] After completing the local model training, the model enters the reduction and dispersion stage to obtain unique partial aggregate model parameters.
[0010] Based on the unique partial aggregated model parameters, each federated learning user sends their own partial aggregated model parameters along with the parameters to obtain the complete aggregated model parameters, which are then used for training a new local model, thus initiating a new round of federated learning.
[0011] Furthermore, the federated learning user obtains the aggregated model by exchanging model parameters, including:
[0012] The initial model parameters in the first round of federated learning are: Each user is assigned a unique index, represented as a permutation of integers. Instead Specifically, the federated learning round consists of three phases: local model training, reduction distribution, and full collection.
[0013] Furthermore, under the aggregation model, a local model is trained on the respective datasets of the federated learning users using the stochastic gradient descent algorithm, including:
[0014] After exchanging model parameters, federated learning users obtain complete aggregated model parameters and enter the local model training phase in a new federated learning round. In this phase, federated learning users train their local models on their respective datasets using the stochastic gradient descent algorithm, continuing for multiple training cycles.
[0015] Furthermore, after completing local model training, the process enters the reduction and distribution phase, obtaining unique partial aggregated model parameters. In this phase, the amount of data transmitted by each federated learning user decreases by a factor of two, i.e., "binary splitting," including:
[0016] After completing local model training, the model enters the reduction and distribution phase and continues to perform co-training. Each step of the operation, in each step, forms a total based on the index of the federated learning user. For user groups, it is represented as ,in The step index is represented; each pair of users exchanges their respective partial model parameters and aggregates them locally through a weighted summation operation; specifically, the partial model parameters transmitted by the federated learning users are divided by... and multiply by As the number of steps increases, the index gap between each pair of federated learning users doubles exponentially, and the size of the transferred data increases from... Gradually decrease to ,in Indicates the bit size of the corresponding data.
[0017] Furthermore, based on the unique partial aggregation model parameters, each federated learning user sends their own partial aggregation model parameters along with the complete aggregation model parameters. In this stage, the amount of data transmitted by each federated learning user increases by a factor of two, i.e., "doubling," including:
[0018] Finish After the steps, federated learning users will receive unique partial aggregation model parameters. The number of steps in the collection phase is equal to the number of steps in the reduction and distribution phase. For each federated learning user, the process is the reverse of the reduction and distribution phase. Specifically, after receiving the partial aggregation model parameters, each user sends them along with their own partial aggregation model parameters to the corresponding party. After the final step, all federated learning users obtain the complete aggregated model parameters, and the results will be used to train new local models, initiating a new round of federated learning.
[0019] Furthermore, between the reduction distribution and full collection phases, the indexes for each pair of federated learning users are different, defining pairwise associated indexes under various conditions:
[0020] First, define the symbolic function as...
[0021]
[0022] At this time, the corresponding party Represented as:
[0023]
[0024] Determine the delay, steps Transmitted model parameters The size is represented as:
[0025]
[0026] In the In a round of federated learning, federated learning users In the steps The latency in the process is defined as the time from the start of the federated learning process to the time from the corresponding party. The time when data reception ends; covering the completion steps. The entire time, expressed as In the first step of each federated learning round, the latency for federated learning users includes the training time incurred during the local model training phase. Divided into two parts: and ,in express Completed in the previous step Data reception latency or model training latency yes In the current step from The latency when receiving data; therefore, Equal to the higher of the two, expressed as:
[0027]
[0028] also,, Represented as:
[0029]
[0030] in, Represents the unit impulse function;
[0031] Indicates from arrive The transfer size is Data latency; in low-Earth orbit satellite networks, the total latency of wireless signal propagation during inter-satellite link transmission, and is used as... This indicates that transmission delay occurs across every transmission along the entire routing path; the number of intermediate satellites is assumed to be... ,but Represented as:
[0032] .
[0033] Furthermore, after the local model training phase, the index exceeds the maximum less than The user with the power of two sends the fully trained model parameters to an index lower than the maximum less than Users with a power of two; when the index exceeds the maximum less than When using a power-two user, assume Indicates receiving from user Data users Conversely, when the index is lower than the maximum less than When using a power-two user, Indicates the sender user When the index is lower than the maximum less than When using a power-two user, use Replace, and represent as:
[0034]
[0035] Among them, when , ,otherwise, Represented as:
[0036]
[0037] In addition, when At this time, it sends the complete trained model parameters to other users, without directly participating in the reduction distribution and full collection process. The overall latency at this point is:
[0038]
[0039] In centralized federated learning, latency is determined by the node with the longest latency in a single round. Decentralized federated learning focuses on the average latency of all federated learning users, aiming to minimize latency. The optimization problem is formulated as follows:
[0040]
[0041] in, This is represented as a list of federated learning users. and These represent the energy consumption of a single user and the total energy consumption of all nodes involved in the entire network, respectively. This represents the load balancing constraint value.
[0042] Secondly, this invention provides a decentralized federated learning system for low-Earth orbit satellites based on binary division and multiplication, comprising:
[0043] The aggregation model building module is used in the first round of federated learning, where federated learning users obtain the aggregation model by exchanging model parameters.
[0044] The local training module is used to allow federated learning users to train local models on their respective datasets under the aggregated model.
[0045] The reduction and dispersion module is used to obtain unique partial aggregated model parameters after completing local model training and entering the reduction and dispersion stage.
[0046] The full collection module is used to collect the model parameters based on unique partial aggregated model parameters. Each federated learning user sends their partial aggregated model parameters together to obtain the complete aggregated model parameters, which are then used for training a new local model and starting a new round of federated learning.
[0047] Thirdly, the present invention provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the low-Earth orbit satellite decentralized federated learning method based on binary division and multiplication.
[0048] Fourthly, the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the low-Earth orbit satellite decentralized federated learning method based on binary division and doubling.
[0049] Compared with the prior art, the present invention has the following technical effects:
[0050] This invention effectively solves the problem of concentrated server load in traditional federated learning by employing a decentralized federated learning method using binary search and doubling. In this decentralized architecture, each LEO satellite participates in model training and data transmission as a federated learning user, without relying on a central server for model parameter aggregation. This method not only distributes the load of data processing and model training but also avoids the phenomenon of a single server becoming a bottleneck, thereby improving the stability and scalability of the entire LEO satellite network.
[0051] By adopting a decentralized approach, each satellite undertakes a portion of the computing and communication tasks, thus distributing the load evenly across the network. This avoids the problems of excessive server load and vulnerability to failure inherent in traditional centralized architectures.
[0052] Decentralized architectures are easier to scale because new satellites can be seamlessly integrated into the network without requiring complex adjustments to the central server.
[0053] Due to the distributed load, the network's resilience to single points of failure is enhanced, and its overall stability is improved.
[0054] This invention derives in detail the recursive formula for latency and energy consumption at each step, and considers the energy consumption of a single satellite, the total network energy consumption, and load balancing constraints. By optimizing the arrangement of users in federated learning using the particle swarm optimization algorithm, the average latency is maximized (actually, it is minimized; the term "maximize" may be an oversight), thereby accelerating model convergence and improving the efficiency of federated learning.
[0055] By precisely calculating the latency at each step and taking into account the differences in inter-satellite links, this invention can more accurately assess and optimize latency in the federated learning process. This helps reduce unnecessary waiting time and improve overall learning efficiency.
[0056] In addition to considering latency, this invention also addresses energy consumption. By optimizing the transmission and aggregation methods of model parameters and rationally allocating computational tasks, the energy consumption of the entire network is effectively reduced.
[0057] In summary, this invention effectively addresses the problems of federated learning in low-Earth orbit (LEO) satellite networks by employing a decentralized architecture, optimizing latency and energy management, and implementing load balancing. This not only improves network stability and scalability but also accelerates model convergence and enhances the efficiency of federated learning. These technological advantages provide strong support for the widespread application of LEO satellite networks in remote sensing, communication, and navigation. Attached Figure Description
[0058] Figure 1 This is a schematic diagram of the low-orbit satellite network structure of the present invention.
[0059] Figure 2 This is a schematic diagram of the three stages of the present invention.
[0060] Figure 3 This is a logic block diagram of the present invention.
[0061] Figure 4 This is a schematic diagram of the process of the present invention using seven users as an example.
[0062] Figure 5 This invention addresses the latency of 10 federated learning users in all rounds of the proposed decentralized federated learning based on binary search and doubling.
[0063] Figure 6 For the time delay variation between different federated learning rounds.
[0064] Figure 7 The load distribution of satellites in the network in decentralized federated learning and centralized federated learning based on binary and doubling.
[0065] Figure 8 The comparison of latency for 10 users across federated learning rounds is shown, where the error bars represent the latency range for all 10 federated learning users in each round, and the circle markers represent the average latency.
[0066] Figure 9 This represents the relationship between image recognition accuracy and time.
[0067] Figure 10 This is a flowchart of the present invention. Detailed Implementation
[0068] The present invention will be further described below with reference to the accompanying drawings:
[0069] Example 1, please refer to Figure 10 Decentralized federated learning methods for low-Earth orbit satellites based on binary division and doubling include:
[0070] In the first round of federated learning, federated learning users obtain aggregated models by exchanging model parameters;
[0071] In the aggregation model, federated learning users train local models on their respective datasets;
[0072] After completing the local model training, the model enters the reduction and dispersion stage to obtain unique partial aggregate model parameters.
[0073] Based on the unique partial aggregated model parameters, each federated learning user sends their own partial aggregated model parameters along with the parameters to obtain the complete aggregated model parameters, which are then used for training a new local model, thus initiating a new round of federated learning.
[0074] Example 2: This invention provides a decentralized federated learning method for low-Earth orbit satellites based on binary division and doubling, specifically including:
[0075] This invention considers a method by A low-Earth orbit satellite network consisting of several low-Earth orbit satellites, such as Figure 1 As shown, these satellites provide seamless connectivity to ground equipment via ground-to-satellite links. Under specific routing algorithms, satellites can achieve single-hop or multi-hop routing paths through inter-satellite links between any two satellites, unaffected by changes in the satellite topology over time. The low-Earth orbit (LEO) satellite network consists of multiple orbits defined by six parameters: eccentricity, semi-major axis, inclination, ascending node longitude, perigee argument, and true anomaly. The eccentricity is almost zero, indicating that the orbit is nearly circular. In this case, the semi-major axis is the radius of the circle, equal to the average radius of the Earth. Add track height The network supports two types of inter-satellite links: inter-orbit and intra-orbit inter-satellite links, to support communication between low-Earth orbit satellites.
[0076] In a decentralized federated learning architecture in low-Earth orbit satellite networks, there are a total of A low-orbit satellite acts as a federated learning user, indicating that... In each round of federated learning, federated learning users base their learning on their respective local datasets. Train a local model. Then, according to a specific federated learning method, without compromising data privacy, the trained model parameters are... The model is then passed on to other users. As multiple rounds of federated learning iterate, the machine learning model gradually converges until it reaches the expected performance or completes all rounds. In federated learning, the classic federated averaging algorithm is widely used; its optimization objective is to minimize the empirical loss function, as shown in the following equation.
[0077]
[0078] in, This indicates the number of samples retrieved from the dataset. For machine learning model parameters, It is its empirical loss function. This formula can be further transformed into:
[0079]
[0080] in, These are represented as global model parameters, aggregated from the model parameters of all other federated learning users. The aggregate weights for each federated learning user are determined by the sample size contained in their dataset.
[0081] Clearly, in federated learning, the latency of the local model training process depends on factors such as dataset size, model parameter size, number of iterations, and the computing power of the low-Earth orbit satellites. Model parameter size and computing power can be integrated into... This represents the number of cycles required to process each sample in the dataset. Given the heterogeneity among low-Earth orbit satellites, the latency of the local model training process... It can be represented as:
[0082]
[0083] in, Number of training epochs for the machine learning model This indicates the operating frequency of low-Earth orbit satellites.
[0084] In low-Earth orbit (LEO) satellite networks, inter-satellite links play a crucial role in enabling communication between satellites in the same or different orbits, thereby improving the coverage of the service area. The establishment of inter-satellite links depends on several factors. Assume there exists an imaginary straight line used to determine the inter-satellite link, called the line of sight. From a geometric and geographical perspective, this line of sight must satisfy… ,in, This represents the Euclidean distance between two low-Earth orbit satellites. and Here are the orbital altitudes of the two satellites. The transmission rate of the inter-satellite link is equal to the theoretical maximum bit rate defined in information theory, i.e., the Shannon capacity, which is expressed as:
[0085]
[0086] in, Represents link bandwidth. Indicates the transmission power. Let be the noise power spectral density. In this case, the transmission delay depends entirely on the size of the transmitted object, as shown in the following equation:
[0087]
[0088] in, This represents the bit size of the transmitted object.
[0089] In addition, there is signal propagation delay, which refers to the time delay in the transmission of wireless signals from the transmitter to the receiver. This delay depends only on the distance of the direct path between the transmitter and the receiver, and is calculated using the following formula:
[0090]
[0091] in, It is expressed as the speed of light.
[0092] Due to frequent entry into Earth's shadow and limited onboard power capacity, energy management remains a key challenge for ensuring the long-term efficient operation of low-Earth orbit (LEO) satellites. In federated learning within LEO satellite networks, energy is typically allocated to two main components: model training and data transmission. The energy consumption for model training can be expressed as:
[0093]
[0094] in, This represents the energy consumption coefficient. The energy consumption for data transmission can be expressed as:
[0095]
[0096] Furthermore, to achieve load balancing in a low-Earth orbit (LEO) satellite network, the load on all inter-satellite link paths should be considered. The load for both federated learning users and intermediate nodes is represented by the sum of the data processed. Unlike range, standard deviation, and variance, the coefficient of variation (CV) is a dimensionless statistical measure suitable for describing the variability of data points relative to the mean, and is particularly useful for comparing the degree of variation across different datasets. In the context of measuring load balancing among related LEO satellites, the CV measures the uniformity of load distribution, and its expression is:
[0097]
[0098] in, Indicates the satellite's payload. and These represent the standard deviation and mean of the load in the network, respectively.
[0099] To improve efficiency and manageability, binary search and doubling are common parallel computing strategies. During computation, all nodes can be activated, thereby increasing the utilization of computing resources. Compared to centralized federated learning, decentralized federated learning has advantages in load balancing in LEO satellite networks because it does not rely on a central server to aggregate all local model parameters, thus avoiding the problem of single-point overload. This distributed characteristic effectively distributes the load and enhances load balancing in LEO satellite networks.
[0100] Considering these efficiency improvements, this invention proposes a decentralized federated learning method based on binary search and doubling in low-Earth orbit satellite networks. In this method, federated learning users exchange model parameters to obtain an aggregated model, which is then further trained until convergence. The key point is the aggregation of machine learning models among federated learning users.
[0101] For the power of two federated learning users, assuming the initial model parameters of the users in the first round of federated learning are... Each user is assigned a unique index, represented as a permutation of integers. To avoid confusion with the indexes of low-Earth orbit satellites, the assigned indexes are... Instead Specifically, the method proposed in the federated learning round can be divided into three stages: local model training, reduction distribution, and full collection, such as... Figure 2 As shown in the figure. Taking four users as an example, the model parameters are divided into four parts, from... arrive Each part is the same size. For example, Indicates the first part, and is composed of and The corresponding data aggregation in the process. To clarify the following description, two terms are defined here: the first is "partially aggregated model parameters," which refers to the complete machine learning model parameters aggregated by the federated learning partial users, such as... The second is "aggregated partial model parameters," which refers to a portion of the machine learning model parameters after aggregation from all federated learning users.
[0102] In the final step of the previous federated learning round, the federated learning user obtains the complete aggregated model parameters and enters the local model training phase in the new federated learning round. In this phase, the federated learning user trains its local model on its respective dataset using the stochastic gradient descent algorithm, continuing for multiple training epochs.
[0103] After local model training is complete, the process enters the reduction and distribution phase, and continues with the co-processing... Each step is an operation. In each step, a total is formed based on the index of the federated learning user. For user groups, it is represented as ,in This represents the step index. Each pair of users exchanges their respective partial model parameters and aggregates them locally through a weighted summation operation. Specifically, the partial model parameters transmitted by the federated learning users are divided by... and multiply by As the number of steps increases, the index gap between each pair of federated learning users doubles exponentially, and the size of the transferred data increases from... Gradually decrease to ,in Indicates the bit size of the corresponding data.
[0104] Finish After the steps, federated learning users will receive unique partially aggregated model parameters. The number of steps in the full collection phase is equal to the number of steps in the reduction-dispersion phase. For each federated learning user, the process is the reverse of the reduction-dispersion phase, but involves updating instead of aggregation. Specifically, after receiving the partially aggregated model parameters, each user sends them along with their own partially aggregated model parameters to the corresponding party. After this final step, all federated learning users obtain the complete aggregated model parameters. This result will then be used to train a new local model, initiating a new round of federated learning. Figure 3 The above process has been explained in detail.
[0105] Given that each federated learning round contains three identical phases, analyzing the latency within a single federated learning round is crucial. Between the reduction scatter and full collection phases, the indices for each pair of federated learning users are different; for example, the first step of the reduction scatter is... The steps of the full collection phase are as follows: To further evaluate the efficiency of the proposed method, it is necessary to define pairwise associative indices under various conditions. First, the symbolic function is defined as follows:
[0106]
[0107] At this time, the corresponding party It can be represented as:
[0108]
[0109] Whether aggregation is performed during the reduction and distribution phase or direct updates are performed during the full collection phase, the prerequisite is receiving data from the corresponding party in the previous step, which directly limits the latency for federated learning users. To determine the latency, the steps... Transmitted model parameters The size can be expressed as:
[0110]
[0111] First, in the In a round of federated learning, federated learning users In the steps The latency in the process is defined as the time from the start of the federated learning process to the time from the corresponding party. The time when data reception ended. This includes the steps completed. The entire time, expressed as Due to the continuity of the federated learning process, This always holds true. In the first step of each federated learning round, the latency for the federated learning user includes the training time incurred during the local model training phase. Generally, It can be divided into two parts: and ,in express Completed in the previous step The latency when receiving data or the latency of model training, and yes In the current step from The latency when receiving data. Therefore, Equal to the higher of the two can be expressed as:
[0112]
[0113] Furthermore, considering the model training latency in step 1, It can be represented as:
[0114]
[0115] in, This represents the unit impulse function.
[0116] also, Indicates from arrive The transfer size is Data latency. In low-Earth orbit satellite networks, the total latency of wireless signal propagation during inter-satellite link transmission should be considered, and used as... This is represented as follows. Furthermore, considering the store-and-forward mechanism in practice, transmission delay occurs across every transmission along the entire routing path. Assume the number of intermediate satellites is represented as... ,but It can be represented as:
[0117]
[0118] In practical applications, the number of users in federated learning often does not meet the requirement that the number equals a power of two in traditional binary search and doubling methods. Therefore, after the local model training phase, the index exceeds the maximum less than The user with the power of two sends the fully trained model parameters to an index lower than the maximum less than Users with a power of two. When the index exceeds the maximum less than... When using a power-two user, assume Indicates receiving from user Data users Conversely, when the index is lower than the maximum less than... When using a power-two user, Indicates the sender user . Figure 4 This shows the scenario when there are 7 users. In this case, the index is lower than the maximum less than... When using a power-two user, It can be used Replace, and represent as:
[0119]
[0120] Among them, when , ,otherwise, Represented as:
[0121]
[0122] In addition, when At this time, it sends the complete trained model parameters to other users, without directly participating in the reduction distribution and full collection process. The overall latency at this point is:
[0123]
[0124] In low-Earth orbit (LEO) satellite networks, energy consumption must be considered due to limited battery capacity, solar energy collection capabilities, and the impact of planetary eclipse on available collection time. Furthermore, node load is also a key concern for satellite networks. Uneven load can lead to single points of failure, and effectively balancing network load can significantly reduce network congestion, lower latency, and enhance user experience. Based on these considerations, when reducing latency to improve federated learning efficiency, both energy consumption and load balancing in LEO satellite networks should be taken into account. (Assuming the index exceeds the maximum less than...) The user with the power of two sends the fully trained model parameters to the nearest index lower than or equal to itself. For users with a power of two, the latency of federated learning users depends on... and The mapping relationship, i.e., the permutation. In centralized federated learning, latency is determined by the node with the longest latency in a round, while decentralized federated learning focuses on the average latency of all federated learning users. To minimize latency, the optimization problem can be formulated as follows:
[0125]
[0126] in, This is represented as a list of federated learning users. and These represent the energy consumption of a single user and the total energy consumption of all nodes involved in the entire network, respectively. This represents the load balancing constraint value.
[0127] Here, the optimization variable is the permutation of federated learning users, resulting in a computational complexity of O(n log n). .along with The increasing computational complexity presents a significant challenge. However, the output of the optimization function can be verified in polynomial time. Due to its polynomial-time verifiability, this problem is classified as a nondeterministic polynomial problem. Based on this, Particle Swarm Optimization (PSO), a metaheuristic algorithm, is employed to solve it. In this algorithm, each particle represents a possible solution and moves in the search space at a certain velocity. As iterations proceed, the particle swarm tracks any best solution found by a particle, thus obtaining an approximate solution. Specifically, the design of the particle velocity in each iteration is crucial to ensuring that each particle remains within the constraints. If the permutation consists of... A vector representing the particle's position, where each element corresponds to a federated learning user, could lead to duplicate positions due to random velocity, violating the uniqueness requirement of the permutation. Therefore, each particle's vector is designed to contain... The problem involves a set of elements that randomly change within a fixed range. Sort these elements by size to find a feasible solution. As particles move, they always satisfy the uniqueness requirement. Furthermore, if a particle violates energy consumption and load balancing constraints, its fitness value is penalized by subtracting a sufficiently large fixed value from its fitness value using a penalty function.
[0128] The simulation considers a Walker Delta constellation for a low-Earth orbit (LEO) satellite network, consisting of 24 LEO satellites. Eight satellites share the same orbit, and each satellite is equidistant along the orbit. Connections between any two satellites can be achieved via multi-hop routing using Dijkstra's algorithm. The LEO network topology is updated every 30 seconds within a snapshot. The computational power of the federated learning users is set to execute 50,000 iterations per sample at a frequency of 1 GHz. For wireless communication in inter-satellite links, the transmission power is 40 dBm, and the noise power spectral density is -88 dBm / Hz. Furthermore, the system's coefficient of performance (COP) is 10⁻²⁵.
[0129] Image recognition was used as the specific machine learning task, employing the MNIST dataset, which contains 10,000 images representing 10 classes of handwritten digits from 0 to 9, with 1,000 images of the same digit per class. All sample images are 28×28×1 pixels in size. The machine learning neural network for this dataset consists of a combination of convolutional layers with varying kernel numbers, max-pooling layers, and rectified linear layers. Furthermore, mini-batch stochastic gradient descent with momentum was used for local training with a learning rate of 0.001, 10 training epochs, and a batch size of 512 samples. Before each training iteration, the dataset was randomly shuffled, with 30% used for testing and 70% for training. The test set was further divided equally into a validation set and a test set for the global model. The entire federated learning process concluded after 100 epochs.
[0130] Furthermore, appropriate parameters must be set to achieve optimal results in the particle swarm optimization algorithm. To improve the exploratory nature of the search space, the number of particles is set to 100. During velocity updates, to balance exploration and development, both the self-awareness coefficient and the swarm awareness coefficient are set to 1.5. The solution set is obtained by sorting randomly assigned elements within a fixed range. To simplify computation, the position of each particle is a scalar within that range, restricted to [−2, 2]. To ensure robustness, the minimum and maximum velocities are both set to -1 and 1, respectively, representing half the position constraints. The number of iterations is set to 200. To avoid premature convergence, the inertia coefficient is set to 0.9. If a particle does not meet the energy consumption or load balancing constraints, its fitness is reduced by 10⁶ using a penalty function.
[0131] Figure 5 The graph shows the latency for 10 federated learning users across all rounds of the proposed binary and doubling-based decentralized federated learning approach. The bar chart represents the average latency for each user, and the error bars indicate the latency range. As shown, the average latency for each user is approximately the same, around 997 ms. This is due to... and This is due to the interdependencies in computation between them. Because there are differences in latency between users before a training round, the latency for all users and rounds fluctuates between approximately 850 ms and 1150 ms.
[0132] Figure 6 This illustrates the latency variations across different federated learning rounds. The solid blue line represents the average latency per round, while the shaded area represents the latency range for all federated learning users within each round. Notably, the latency generally fluctuates between 900 ms and 1100 ms, although occasional spikes and troughs occur in some rounds. Despite these fluctuations, the average latency for all federated learning users remains relatively stable at approximately 1000 ms per round.
[0133] Figure 7 This diagram illustrates the load distribution of satellites in a network during decentralized federated learning based on binary search and doubling, as well as centralized federated learning. The color intensity of each box corresponds to the load experienced by the respective satellite, with higher values represented by warmer colors and lower values by cooler colors. Furthermore, each value in the box represents the ratio of the number of bits received to the total number of bits in the local model parameters. Zero load indicates that the corresponding satellite did not participate in the federated learning process. Specifically, the left-hand diagram shows a relatively balanced load among satellites under decentralized federated learning. In contrast, the right-hand diagram shows the load distribution in centralized federated learning, which clearly exhibits an uneven load distribution. In particular, satellite 1 in orbit acts as the server for federated learning, with a load nine times the size of the local model parameters, while many other satellites have relatively small loads. This imbalance stems from the star topology of centralized federated learning, where the server accumulates a large load, creating a bottleneck.
[0134] Figure 8 This paper compares the latency of 10 users across federated learning rounds, with error bars representing the latency range of all 10 users in each round and circles indicating the average latency. In addition to results obtained from the particle swarm optimization algorithm, fixed and randomized methods are also introduced for comparison. In the fixed method, the arrangement of federated learning users remains constant throughout the federated learning process. In the randomized method, the order of federated learning users is randomly assigned. Regardless of the method used, the optimal method exhibits smaller fluctuations in average latency across all federated learning rounds. The average latency of the fixed method fluctuates slightly in the initial rounds but tends to stabilize in subsequent rounds. In contrast, the average latency of the randomized method shows significant fluctuations. Furthermore, in any given federated learning round, the latency range among federated learning users in the optimal method is smaller than that in the randomized method, while the latency range in the fixed method is almost zero. In absolute values, the optimal result achieves lower latency than other methods, thus improving the efficiency of federated learning.
[0135] Figure 9 The relationship between image recognition accuracy and time is shown. The accuracy trends of the fixed and random methods are similar, converging to 0.92 after 1083.1 s and 1098.8 s, respectively. In contrast, the optimal method achieves high accuracy close to 997.0 s. Therefore, this invention accelerates global model convergence and improves federated learning efficiency in low-Earth orbit satellite networks.
[0136] 1. This invention addresses the problem of concentrated load on servers in traditional federated learning for low-Earth orbit satellite networks. It also considers latency management and uses binary search and doubling methods to form a decentralized architecture, which distributes the load of data processing and model training, thus avoiding the bottleneck phenomenon of a single server.
[0137] 2. This invention takes into account the different inter-satellite links, derives the specific delay recursion formula and energy consumption for each step, and considers the energy consumption of a single satellite, the total energy consumption of the network and load balancing constraints. It uses the particle swarm optimization algorithm to maximize the average delay, accelerate model convergence, and improve federated learning efficiency.
[0138] In another embodiment of the present invention, a decentralized federated learning system for low-Earth orbit satellites based on binary division and doubling is provided, which can be used to implement the above-mentioned decentralized federated learning method for low-Earth orbit satellites based on binary division and doubling. Specifically, the system includes:
[0139] The aggregation model building module is used in the first round of federated learning, where federated learning users obtain the aggregation model by exchanging model parameters.
[0140] The local training module is used to allow federated learning users to train local models on their respective datasets under the aggregated model.
[0141] The reduction and dispersion module is used to obtain unique partial aggregated model parameters after completing local model training and entering the reduction and dispersion stage.
[0142] The full collection module is used to collect the model parameters based on unique partial aggregated model parameters. Each federated learning user sends their partial aggregated model parameters together to obtain the complete aggregated model parameters, which are then used for training a new local model and starting a new round of federated learning.
[0143] The module division in this embodiment of the invention is illustrative and represents only one logical functional division. In actual implementation, other division methods may be used. Furthermore, the functional modules in the various embodiments of the invention can be integrated into a single processor, exist as separate physical entities, or be integrated into a single module. The integrated modules described above can be implemented in hardware or as software functional modules.
[0144] In another embodiment of the present invention, a computer device is provided, comprising a processor and a memory. The memory stores a computer program, which includes program instructions. The processor executes the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing and control core of the terminal, suitable for implementing one or more instructions, specifically suitable for loading and executing one or more instructions from the computer storage medium to achieve a corresponding method flow or function. The processor described in this embodiment can be used for operation based on a binary and multiplicative decentralized federated learning method for low-Earth orbit satellites.
[0145] In another embodiment of the present invention, a storage medium is provided, specifically a computer-readable storage medium (Memory), which is a memory device in a computer device used to store programs and data. It is understood that the computer-readable storage medium here can include both the built-in storage medium in the computer device and extended storage media supported by the computer device. The computer-readable storage medium provides storage space that stores the terminal's operating system. Furthermore, the storage space also stores one or more instructions suitable for loading and execution by a processor. These instructions can be one or more computer programs (including program code). It should be noted that the computer-readable storage medium here can be high-speed RAM or non-volatile memory, such as at least one disk storage device. The processor can load and execute one or more instructions stored in the computer-readable storage medium to implement the corresponding steps of the low-Earth orbit satellite decentralized federated learning method based on binary division and doubling in the above embodiments.
[0146] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0147] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0148] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0149] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0150] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.
Claims
1. A low earth orbit satellite decentralized federated learning method based on dichotomy and multiplication, characterized in that, Comprise: In the first round of federated learning, the federated learning user obtains an aggregated model by exchanging model parameters, wherein the federated learning user is for each low-orbit satellite; Under the aggregated model, the federated learning user trains a local model on the respective data set; After completing the local model training, enter the reduction and dispersion stage to obtain a unique partial aggregated model parameter; In the full collection stage, based on the unique partial aggregated model parameter, each federated learning user sends its own partial aggregated model parameter together to obtain the complete aggregated model parameter, which is used for new local model training and starts a new round of federated learning; After completing the local model training, enter the reduction and dispersion stage to obtain a unique partial aggregated model parameter, comprising: After the local model training is completed, a regulation dispersion phase is entered, and the federated learning is continued in a step-by-step operation, in each step, a total of combinations of users are formed according to the indexes of the federated learning users , wherein denotes the step index; each pair of users exchanges respective partial model parameters and locally aggregates through a weighted summation operation; as the step increases, the index gap between each pair of federated learning users doubles exponentially, and the size of the transmitted data gradually decreases from to , wherein denotes the bit size of the corresponding data; Based on the unique partial aggregated model parameter, each federated learning user sends its own partial aggregated model parameter together to obtain the complete aggregated model parameter, comprising: complete After the step, the federal learning user will obtain the unique partial aggregation model parameters, the number of steps of the full collection stage is equal to the number of steps of the reduction dispersion stage, for each federal learning user, the process is opposite to the reduction dispersion stage, after receiving the unique partial aggregation model parameters, each user sends it together with its own partial aggregation model parameters to the corresponding party After the last step, all federal learning users obtain the complete aggregation model parameters, and the results will be used for new local model training to start a new round of federal learning.
2. The bipartition and multiplication based low earth orbit satellite decentralized federated learning method according to claim 1, characterized in that, The federated learning user obtains an aggregated model by exchanging model parameters, comprising: The initial model parameters in the first round of federated learning are Each user is assigned a unique index, denoted as an integer permutation, and the assigned index Rather than denoted.
3. The bipartition and multiplication based low earth orbit satellite decentralized federated learning method according to claim 2, characterized in that, Under the aggregated model, a local model is trained on the respective data set of the federated learning user using a stochastic gradient descent algorithm, comprising: After exchanging model parameters, the federated learning user obtains the complete aggregated model parameter and enters the local model training stage in the new federated learning round, in which the federated learning user trains its local model on the respective data set using a stochastic gradient descent algorithm, and continues for multiple training periods.
4. The bipartition and multiplication based low earth orbit satellite decentralized federated learning method according to claim 1, characterized in that, Between the reduction and dispersion and the full collection stage, the index of each pair of federated learning users is different, and the index of the pair is defined under various conditions: First, define the symbol function as At this time, the counterpart is represented as: determining the latency, step part of the model parameters transmitted the size of the In the In a round of federated learning, federated learning users In the steps The latency in the process is defined as the time from the start of the federated learning process to the time from the corresponding party. The time when data reception ends; covering the completion steps. The entire time, expressed as In the first step of each federated learning round, the latency for federated learning users includes the training time incurred during the local model training phase. Divided into two parts: and ,in express Completed in the previous step Data reception latency or model training latency yes In the current step from The latency when receiving data; therefore, Equal to the higher of the two, expressed as: Furthermore, is represented as: wherein denotes the unit impulse function; denotes the time delay of transmitting data with a size of to a size of data; in a low-orbit satellite network, the total time delay of wireless signal propagation in the process of inter-satellite link transmission is denoted by ; the transmission time delay occurs including each transmission on the entire routing path; assuming that the number of intermediate satellites is denoted by , then is denoted as: 。 5. The bipartition and multiplication based low earth orbit satellite decentralized federated learning method according to claim 4, characterized in that, After the local model training phase, the index exceeds the maximum less than The user with the power of two sends the fully trained model parameters to an index lower than the maximum less than Users with a power of two; when the index exceeds the maximum less than When using a power-two user, assume Indicates receiving from user Data users ; Conversely, when the index is lower than the largest power of two user, denotes the sender user ; when the index is lower than the largest power of two user, is replaced by and is denoted as: wherein, when , , otherwise, is represented as: In addition, when the trained complete model parameters are sent to other users, and the user does not directly participate in the regulation dispersion and full collection process, at this time the overall time delay is: In centralized federated learning, the delay is determined by the node with the longest time consumption in a round, and in decentralized federated learning, the average delay of all federated learning users is focused on, and the delay is minimized, and the optimization problem is expressed as follows: wherein, denotes the permutation of the federated learning users, denotes the load balancing constraint value. denotes the individual user energy consumption and the total energy consumption of the involved nodes in the whole network, respectively; denotes the load balancing constraint value.
6. A low earth orbit satellite decentralized federated learning system based on dichotomy and doubling, characterized in that, Comprise: An aggregated model construction module is used for the federated learning user to obtain an aggregated model by exchanging model parameters in the first round of federated learning; A local training module is used for the federated learning user to train a local model on the respective data set under the aggregated model; A reduction and dispersion module is used for entering the reduction and dispersion stage to obtain a unique partial aggregated model parameter after completing the local model training; A full collection module is used for each federated learning user to send its own partial aggregated model parameter together based on the unique partial aggregated model parameter to obtain the complete aggregated model parameter, which is used for new local model training and starts a new round of federated learning; After completing the local model training, enter the reduction and dispersion stage to obtain a unique partial aggregated model parameter, comprising: After the local model training is completed, the regulation dispersion phase is entered, and the federated learning is continued in each step, and a total of combinations of users, denoted as , where represents the step index; each pair of users exchanges their respective partial model parameters and locally aggregates through a weighted summation operation; as the step increases, the index gap between each pair of federated learning users doubles exponentially, and the size of the transmitted data gradually decreases from to , where represents the bit size of the corresponding data; Based on the unique partial aggregated model parameter, each federated learning user sends its own partial aggregated model parameter together to obtain the complete aggregated model parameter, comprising: complete After the step, the federal learning user will obtain the unique partial aggregation model parameters, the number of steps of the full collection phase is equal to the number of steps of the reduction dispersion phase, for each federal learning user, the process is opposite to the reduction dispersion phase, after receiving the unique partial aggregation model parameters, each user sends it together with its own partial aggregation model parameters to the corresponding party After the last step, all federal learning users obtain the complete aggregation model parameters, and the results will be used for new local model training to start a new round of federal learning.
7. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The processor executes the computer program to implement the steps of the low-orbit satellite decentralized federated learning method based on dichotomy and multiplication according to any one of claims 1 to 5.
8. A computer-readable storage medium storing a computer program, the computer-readable storage medium comprising: The computer program is executed by the processor to implement the steps of the low-orbit satellite decentralized federated learning method based on dichotomy and multiplication according to any one of claims 1 to 5.