Intermittent communication-oriented continuous computing federated learning optimization method

By introducing a continuous computation federated learning framework and a two-layer joint optimization algorithm in the context of intermittent satellite communication, the problem of slow training speed of federated learning models in this scenario is solved, thereby improving the model convergence speed and test accuracy.

CN122496085APending Publication Date: 2026-07-31EAST CHINA NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
EAST CHINA NORMAL UNIV
Filing Date
2026-05-13
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

In existing intermittent satellite communication scenarios, federated learning models are slow to train, and the local resources of unscheduled devices are idle, resulting in low resource utilization and decreased learning performance.

Method used

A continuous computation federated learning framework (CoCoFL) is proposed, in which unscheduled devices continuously train locally and are processed in parallel with the global aggregation of scheduled devices. A two-layer joint optimization algorithm is combined to optimize device scheduling and local training rounds, and equal bandwidth allocation and frequency division multiple access are used to optimize communication.

Benefits of technology

It improves model convergence speed and test accuracy in intermittent satellite communication scenarios, making full use of local computing and data resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122496085A_ABST
    Figure CN122496085A_ABST
Patent Text Reader

Abstract

This invention discloses a continuous computation federated learning optimization method for intermittent satellite communication scenarios. A continuous computation federated learning framework is proposed, and its convergence is analyzed and the problem is modeled. Equipment scheduling and local training rounds are jointly optimized using a convex function difference algorithm and Gibbs sampling, which accelerates global model training and improves test accuracy in intermittent satellite communication scenarios. This invention solves the problem of slow training speed of traditional federated learning models in intermittent satellite communication scenarios, enabling ground equipment to perform tasks such as environmental monitoring and disaster prediction more quickly and effectively using collected data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of satellite communications, and more particularly to a continuous computation federated learning optimization method for intermittent communications. Background Technology

[0002] Intermittent communication is a typical characteristic of satellite-to-ground links, and federated learning optimization methods targeting this characteristic urgently need research. In recent years, low-Earth orbit satellite constellations have developed rapidly, and have the capability to provide global communication coverage for ground equipment deployed in remote areas. The data collected by ground equipment is characterized by its large volume and highly dispersed distribution, and can be used to support applications such as environmental monitoring and disaster prediction. Traditional methods require transmitting this data to cloud servers for processing, but the transmission overhead of satellite-to-ground links is extremely high, making it difficult to meet practical needs. Federated learning (FL Federated Learning), as a distributed learning paradigm that does not require sharing raw data, can collaboratively train a global model without aggregating raw data, thereby significantly reducing communication overhead.

[0003] Existing research has explored satellite-assisted federated learning, reducing communication overhead and mitigating the impact of data heterogeneity by considering satellite mobility and constellation topology. However, most existing work assumes continuous connectivity between satellites and ground equipment, neglecting scenarios with intermittent satellite-to-ground links. Under conditions of intermittent satellite-to-ground link connectivity and limited bandwidth, only some ground equipment can complete local model uploads and participate in global aggregation within the visible window, while non-participating equipment remains idle, failing to fully utilize its local computing and data resources. Simultaneously, data heterogeneity among equipment will further lead to uneven local training latency and may cause the global model to favor specific devices, thereby reducing learning performance.

[0004] Therefore, those skilled in the art are dedicated to developing a continuous computation federated learning optimization method for intermittent communication. A continuous computation-based federated learning framework, CoCoFL (Continuous Computing based Federated Learning), is proposed, specifically for intermittent satellite-to-ground link scenarios. In this framework, devices not scheduled for global aggregation continuously perform local training and temporarily cache local updates, transmitting them when the satellite passes over again. The local update process of unscheduled devices is processed in parallel with the global aggregation of scheduled devices, thereby fully utilizing local computing and data resources. Furthermore, this invention performs convergence analysis and, based on this, develops a two-level optimization algorithm within the CoCoFL framework to simultaneously optimize device scheduling and local training rounds, thereby improving model convergence speed and simultaneously enhancing test accuracy. Summary of the Invention

[0005] In view of the above-mentioned deficiencies of the prior art, the technical problem to be solved by the present invention is the slow training speed of federated learning models in the existing intermittent satellite communication scenario.

[0006] To achieve the above objectives, the present invention provides a continuous computation federated learning optimization method for intermittent communication, including a continuous local training mechanism for unscheduled devices, which processes the local update process of unscheduled devices and the global aggregation of scheduled devices in parallel.

[0007] Furthermore, we jointly optimize equipment scheduling strategies and local training round allocation.

[0008] Furthermore, the uplink communication from ground equipment to the satellite adopts a frequency division multiple access method with equal bandwidth allocation.

[0009] Furthermore, the downlink broadcast rate from satellite to ground equipment is determined by the worst-case channel.

[0010] Furthermore, the goal of federated learning is to minimize the weighted sum of the local loss functions of all devices.

[0011] Furthermore, devices with a large proportion of local datasets are prioritized for global aggregation.

[0012] Furthermore, this includes a two-layer joint optimization algorithm.

[0013] Furthermore, the two-layer joint optimization algorithm employs Gibbs sampling optimization for equipment scheduling in the outer layer.

[0014] Furthermore, in the dual-layer joint optimization algorithm, the inner layer uses a convex function difference algorithm to optimize the local training rounds.

[0015] Furthermore, this includes convergence analysis and problem modeling, joint optimization of equipment scheduling and local training rounds through convex function difference algorithm and Gibbs sampling method.

[0016] Existing research on federated learning using satellite-assisted ground equipment primarily focuses on scenarios with continuous satellite-to-ground links, neglecting the intermittent link scenario. In continuous connection scenarios, all devices can participate in global aggregation at any time; however, under intermittent link conditions, ground equipment can only communicate with the satellite within a limited visible window, and only some devices can complete local model uploads within this window. Unscheduled devices and their local resources are forced into an idle state, thus reducing the model's convergence speed. This invention proposes a continuous computational federated learning framework (CoCoFL) for intermittent satellite-to-ground links. The core innovation of this framework lies in its design of a continuous local training mechanism for unscheduled devices to address the issue of idle resources in intermittent links. This mechanism processes the local update process of unscheduled devices in parallel with the global aggregation of scheduled devices, thereby overcoming the technical bottleneck of low resource utilization in traditional federated learning under intermittent connection scenarios and achieving full utilization of local computing and data resources. In the proposed architecture, unscheduled devices can continue local training, thus fully utilizing local computing and data resources in intermittent connection scenarios.

[0017] The heterogeneity of data between existing devices further leads to uneven local training latency and may cause the global model to favor specific devices, thereby reducing learning performance. Furthermore, how to rationally schedule devices and allocate local training rounds under intermittent connectivity to maximize resource utilization and model convergence speed is also a critical problem that urgently needs to be solved. This invention proposes a two-level joint optimization algorithm based on convergence analysis. The core innovation of this algorithm lies in revealing the influence of device scheduling and local training rounds on model convergence speed through theoretical analysis, and constructing a joint optimization framework accordingly to optimize device scheduling strategies and local training round allocation. This maximizes model convergence speed while ensuring convergence and improving test accuracy. This invention improves model convergence speed and training-test accuracy by jointly optimizing device scheduling strategies and local training round allocation.

[0018] Compared with the prior art, the present invention has the following obvious substantive features and significant advantages: 1. This invention improves the model convergence speed and enhances test accuracy.

[0019] 2. This invention solves the problem of slow training speed of traditional federated learning models in existing intermittent satellite communication scenarios, accelerates global model training in intermittent satellite communication scenarios, improves test accuracy, and enables ground equipment to perform tasks such as environmental monitoring and disaster prediction more quickly and effectively using the collected data.

[0020] The following will further explain the concept, specific structure, and technical effects of the present invention in conjunction with the accompanying drawings, so as to fully understand the purpose, features, and effects of the present invention. Attached Figure Description

[0021] Figure 1 This is a diagram illustrating the continuous computation federated learning process; Figure 2 This is a comparison of the test accuracy results of each algorithm; Figure 3 This is a comparison of the number of communication rounds required for each algorithm to achieve 85% test accuracy. Figure 4 It represents the scheduling probability of each device. Detailed Implementation

[0022] The following description, with reference to the accompanying drawings, illustrates several preferred embodiments of the present invention to make its technical content clearer and easier to understand. The present invention can be embodied in many different forms, and the scope of protection of the present invention is not limited to the embodiments mentioned herein.

[0023] In the accompanying drawings, components with the same structure are indicated by the same numerical designation, and components with similar structures or functions are indicated by similar numerical designations. The dimensions and thicknesses of each component shown in the drawings are arbitrary, and the present invention does not limit the dimensions and thicknesses of each component. To make the illustrations clearer, the thickness of some components has been appropriately exaggerated in the drawings.

[0024] This invention addresses the problem of slow training speed of traditional federated learning models in intermittent satellite communication scenarios. The steps involve proposing a continuous computation federated learning framework, performing convergence analysis and problem modeling, and jointly optimizing device scheduling and local training rounds using convex function difference algorithms and Gibbs sampling. This accelerates global model training in intermittent satellite communication scenarios while improving testing accuracy, enabling ground equipment to perform tasks such as environmental monitoring and disaster prediction more quickly and effectively using collected data.

[0025] like Figure 1 As shown in the diagram, the solid lines on the right represent the communication periods (visible windows) between the satellite and ground equipment, while the dashed lines indicate the non-communication periods (invisible windows). Each communication round includes one visible window and one invisible window. Mesh-like filled blocks correspond to the communication phase, and dot-like filled blocks correspond to the computation phase. This represents the global model after the k-th round of aggregation. Indicates user equipment With global model The local model obtained after local training of the initial model.

[0026] This invention considers a federated learning scenario where multiple ground devices are deployed in remote areas, and low-Earth orbit (LEO) satellites sequentially act as parameter servers. At most one LEO satellite assumes the parameter server role within each visible window. As the satellite moves along its orbit, the parameter server role is transferred to the next satellite entering the area. Let the set of ground devices be denoted as . ,in This represents the total number of devices. Each device... Holding local datasets Its size is The total number of samples across all devices is The goal of federated learning is to minimize the weighted sum of the local loss functions of all devices. The global loss function is defined as:

[0027] in Indicates equipment The proportion of local datasets, For equipment The local loss function is specifically defined as:

[0028] In the formula The model represents the data sample. The loss function value on.

[0029] like Figure 1 As shown, the first The duration of round communication is from to This corresponds to the time interval between two consecutive satellite transits. The interval is... Indicates the visible window, interval Indicates an invisible window. Round communication is performed according to the following steps: (a) Equipment Scheduling and Model Upload: In the visible window Internally, define the obsolescence coefficient. If the equipment In the If the wheel is scheduled, then ;otherwise The scheduled device uploads its local model update. .

[0030] (II) Global Model Update: Satellites update the global model using the following aggregation formula:

[0031] Define the set of users to be scheduled as This incorporates the global model from the previous round. To suppress model fluctuations when there are fewer scheduling devices.

[0032] (iii) Global model broadcast: at the end of the visible window Before, the updated global model The broadcast is sent to the scheduled equipment and simultaneously relayed to the next satellite via inter-satellite links to ensure the continuity of the learning process.

[0033] (iv) Local model training: Scheduled devices train based on the latest global model, while unscheduled devices continue training based on the outdated model. Both types of devices must be trained on the [missing information - likely a specific date or time]. End of round Update completed.

[0034] Scheduled device: Receive Execute after Local training rotation: in Unscheduled devices: based on the first The old local model obtained in the round Continue training: in For from the first From the beginning of the round to the number The cumulative number of local iteration rounds at the end of the round is expressed as:

[0035] For uplink communication from ground equipment to a low-Earth orbit satellite, a frequency division multiple access (FDMA) scheme with equal bandwidth allocation is used. The uplink transmission delay of device i in the k-th round is denoted as . The downlink broadcast rate is determined by the worst-case channel: in, Broadcast bandwidth (Hz) The satellite transmit power is (W). The free-space channel gain is... In the formula The speed of light in a vacuum (m / s) The carrier frequency (Hz) Let be the distance (m) between device i and the satellite. and These are the antenna gains (dimensionless) of the equipment and the satellite, respectively. Let be the noise power spectral density (W / Hz). Based on this, the downlink transmission delay of device i is expressed as:

[0036] in This represents the amount of data (in bits) in the global model. For local model updates, device i executes in the k-th round. The computation latency of round-robin local training is expressed as: in The computational effort (FLOPs) required to process a single data sample. The computing power of device i (FLOPs / s).

[0037] In the assumption -smooth、 Under the four conditions of strong convexity, bounded gradients, and bounded local and global gradient deviations, the proposed continuously computed federated learning algorithm can be derived through convergence analysis, showing the learning rate... satisfy When, its upper bound on convergence is:

[0038] in, It is a constant. Defined as all The maximum value, and the value ranges from 0 to 1. The specific expression is:

[0039] In the formula This is a scheduling indicator variable; it is 1 when scheduled and 0 when not scheduled. As shown in the convergence expression, decreasing... This can improve convergence speed and accuracy. Under the assumption that device scheduling is independent, minimizing Equivalent to maximizing ,in achievable The range of values ​​is At the same time, in order to ensure , must meet in , It is a very small positive number.

[0040] The above analysis shows that, to accelerate learning convergence, under conditions of limited communication resources, devices with a larger proportion of local datasets should be prioritized for global aggregation. Meanwhile, devices with excessively large or small accumulated local training epochs should not be prioritized: the former carries outdated information due to stale models, while the latter fails to effectively update its model due to insufficient local learning. Therefore, a reasonable scheduling strategy needs to strike a balance between the data proportion of devices and the number of local training epochs to achieve faster convergence speed and higher final accuracy.

[0041] To this end, the present invention constructs a problem of jointly optimizing device scheduling and local training rounds to accelerate global model convergence, as detailed below.

[0042]

[0043] in, and To optimize the variables, let represent the device scheduling scheme and the number of local training rounds, respectively. The objective function is derived from the expression in the convergence analysis. C1 guarantees that the sum of the total computation delay and communication delay of each device in each round does not exceed the length of that round. C2 ensures that all communication is completed within the visible window. C3 and C4 jointly constrain the upper bound of the cumulative number of local training rounds for each device. C5 defines the scheduling decision as a binary variable. C6 specifies that the number of local training rounds is an integer. 'a' and 'b' are constants obtained through parameter estimation experiments.

[0044] To address the aforementioned nonlinear mixed integer optimization problem, this invention proposes a two-layer joint optimization algorithm. The outer layer uses Gibbs sampling to optimize equipment scheduling, while the inner layer uses a convex function difference algorithm to optimize the local training rounds.

[0045] (I) Inner layer optimization: Convex function difference algorithm

[0046] Under the condition of fixed equipment scheduling, the local training rounds $E_k^i$ are first relaxed into continuous variables to reduce the difficulty of the solution. Then, the local rounds for unscheduled and scheduled equipment are optimized separately. For unscheduled equipment, which does not generate communication latency, the local rounds are constrained by the total time of each round and the upper bound of the cumulative rounds. The feasible range is as follows: To fully utilize computing resources, the maximum feasible number of rounds for unscheduled devices is: For the scheduling device, after fixing the schedule and relaxing the integer constraints, the optimization subproblem is formulated as follows:

[0047] Among them, the visible window constraint C8 is independent of the local round and is used to determine the feasibility of the scheduling scheme: if it is not satisfied, the current scheduling scheme is not feasible; if it is satisfied, the remaining constraints are solved. Since there are non-convex terms in the constraints, this invention linearizes them into a convex approximation form:

[0048] The above convex approximation problem is solved iteratively using the convex function difference algorithm until convergence, where This is the solution from the previous iteration.

[0049] (ii) Outer layer optimization: Gibbs sampling

[0050] This invention employs the Gibbs sampling method to optimize device scheduling. In the t-th sampling, one of three operations—add, remove, or swap—is randomly selected based on the current scheduling. Generate candidate schedules For candidate scheduling, first determine the local round number of unscheduled devices and scheduled devices using the method described above. If the visibility window constraint is not satisfied, discard the candidate scheme and proceed to the next sampling; otherwise, calculate the candidate target value and calculate the acceptance probability using the following formula: in The difference between the current target value and the candidate target value. Temperature is a parameter used to control the balance between exploration and utilization. It is expressed as a probability. Accept the candidate solution; otherwise, retain the current solution. After T samplings, output the final device schedule. and local rounds As an optimization parameter, it is used to accelerate the convergence of the global model.

[0051] The instances were assigned using the Fashion-MNIST dataset in a non-independent, identically distributed manner and trained using the VGG-11 model. Comparison algorithms included: DSA (scheduling in descending order of data volume, prioritizing devices with larger data volumes), SAS (scheduling prioritizing devices with higher obsolescence coefficients), and FedAvg (random scheduling, unscheduled devices do not participate in continuous computation). The effectiveness of the proposed joint optimization strategy was verified by comparing the accuracy and scheduling probability.

[0052] Figure 2 The test accuracy comparison results of various algorithms are presented. Experiments show that the convergence speed of the proposed CoCoFL optimization algorithm is significantly faster than FedAvg and SAS. CoCoFL's performance improvement over FedAvg stems from its continuous computation mechanism, which allows the local model to incorporate more knowledge. Compared to SAS, CoCoFL prioritizes scheduling devices with larger datasets while considering model staleness, thus ensuring more effective information participates in global model updates. Notably, although CoCoFL's learning speed is slightly slower than DSA in the early stages of training, it achieves higher test accuracy in subsequent rounds. This is because DSA consistently schedules devices with larger datasets, causing the converged model to deviate from the information from devices with smaller datasets, ultimately affecting global performance. Figure 3 It can be seen that the number of communication rounds required for CoCoFL to achieve 85% test accuracy is reduced by 45.8%, 54.1%, and 54.5% compared to DSA, SAS, and FedAvg, respectively.

[0053] Figure 4The scheduling probabilities of each device are presented, further revealing the source of CoCoFL's performance improvement. DSA only selects devices with larger datasets, completely ignoring devices with smaller datasets; FedAvg uses random scheduling, and SAS uses round-robin to maintain a low obsolescence coefficient. Neither of these considers dataset size, resulting in roughly uniform scheduling probabilities for devices with different dataset sizes. In contrast, the CoCoFL proposed in this invention prioritizes scheduling devices with larger datasets while still allocating lower scheduling probabilities to devices with smaller datasets, thereby accelerating convergence while simultaneously improving overall test accuracy.

[0054] The preferred embodiments of the present invention have been described in detail above. It should be understood that those skilled in the art can make numerous modifications and variations based on the concept of the present invention without creative effort. Therefore, all technical solutions that can be obtained by those skilled in the art based on the concept of the present invention through logical analysis, reasoning, or limited experimentation on the basis of existing technology should be within the scope of protection defined by the claims.

Claims

1. A continuous computation federated learning optimization method for intermittent communication, characterized in that, This includes a mechanism for continuous local training of unscheduled devices, which processes the local update process of unscheduled devices in parallel with the global aggregation of scheduled devices.

2. The continuous computation federated learning optimization method for intermittent communication as described in claim 1, which jointly optimizes the device scheduling strategy and the local training round allocation.

3. The continuous computation federated learning optimization method for intermittent communication as described in claim 1 adopts a frequency division multiple access method with equal bandwidth allocation for uplink communication from ground equipment to satellite.

4. In the continuous computation federated learning optimization method for intermittent communication as described in claim 1, the downlink broadcast rate from satellite to ground equipment is determined by the worst-case channel.

5. The continuous computation federated learning optimization method for intermittent communication as described in claim 1, wherein the goal of federated learning is to minimize the weighted sum of the local loss functions of all devices.

6. The continuous computation federated learning optimization method for intermittent communication as described in claim 1 prioritizes scheduling devices with a large proportion of local datasets to participate in global aggregation.

7. The continuous computation federated learning optimization method for intermittent communication as described in claim 1, including a two-layer joint optimization algorithm.

8. The continuous computation federated learning optimization method for intermittent communication as described in claim 7, wherein the two-layer joint optimization algorithm employs Gibbs sampling optimization for device scheduling in the outer layer.

9. The continuous computation federated learning optimization method for intermittent communication as described in claim 7, wherein the two-layer joint optimization algorithm employs a convex function difference algorithm in the inner layer to optimize the local training rounds.

10. The continuous computation federated learning optimization method for intermittent communication as described in claim 1, comprising convergence analysis and problem modeling, and joint optimization of device scheduling and local training rounds through convex function difference algorithm and Gibbs sampling method.