Model training method, device and system based on satellite communication

Through distributed training and dynamic weight allocation, satellite computing power and local data resources are used to solve the problem of limited communication bandwidth between satellites and ground stations, the model training efficiency and quality are improved, and the model generalization ability is enhanced.

CN120541516APending Publication Date: 2025-08-26SHANTOU UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510500916.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

The communication bandwidth between satellites and ground stations is limited and the data transmission speed is slow, resulting in low efficiency in model parameter updates and data interactions. The satellite's energy and computing capabilities are limited, making it difficult to undertake large-scale model training tasks, and the model is prone to convergence difficulties in complex resource-constrained environments.

Method used

The distributed training method is adopted. After receiving the global model and training round, the satellite conducts local training, generates local models and synthetic data sets, and dynamic weight allocation and iterative optimization are performed through the ground station. It uses satellite computing capabilities and local data resources to generate synthetic data sets to enrich training samples and improve model training efficiency and quality.

Benefits of technology

Through distributed computing and dynamic weight adjustment, communication costs and time overhead are significantly reduced, model training efficiency and quality are improved, model generalization capabilities are enhanced, and model training problems are solved in environments of constrained satellite resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120541516A_ABST
    Figure CN120541516A_ABST
Patent Text Reader

Abstract

The method is mainly applied to the technical field of satellite communication and machine learning. The invention provides a model training method, device and system based on satellite communication. The method comprises the following steps: receiving a global model and a current training round from a ground station; training the global model by using the local data set; when the number of times of training the global model reaches the current training round, outputting a local model and updating the number of rounds; generating a synthetic data set based on the local data set; and sending the local model, the update round number and the synthetic data set to a ground station. According to the technical scheme, the satellite data can be used for model training, and meanwhile the model training efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of satellite communications and machine learning technology, and in particular to a model training method, device and system based on satellite communications. Background Art

[0002] With the deep integration of aerospace technology and artificial intelligence, deploying large machine learning models on satellites has become a new area of ​​exploration. While in orbit, satellites collect vast amounts of Earth observation data, including meteorological, geographical, and oceanographic information. Using large machine learning models directly on the satellite side to process and analyze this data in real time would greatly improve data utilization efficiency and response speed, bringing immense value to fields such as scientific research, disaster warning, and resource management. However, this process faces numerous technical challenges. First, the limited communication bandwidth and slow data transmission speed between satellites and ground stations lead to inefficient model parameter updates and data exchange. Second, training large machine learning models typically requires massive computing resources and high costs, while satellites have limited energy and computing power, making them difficult to handle large-scale model training tasks. Finally, large machine learning models are prone to convergence difficulties in the complex and resource-constrained environment of satellites, compromising model accuracy and stability. Summary of the Invention

[0003] The present invention provides a model training method, device and system based on satellite communication, which can utilize satellite data for model training and improve the efficiency of model training.

[0004] The present invention provides a model training method based on satellite communication, the method comprising: Receive the global model and current training round from the ground station; Training the global model using the local dataset; When the number of times the global model is trained reaches the current training round, outputting the local model and updating the round number; Based on the local dataset, generating a synthetic dataset; The local model, the update round number, and the synthetic data set are sent to the ground station.

[0005] Optionally, generating a synthetic dataset based on the local dataset includes: Convert each image in the local dataset into an image block; determining a discreteness between each of the image blocks and the local dataset; Based on the discreteness, each of the image blocks is screened, and the screened image blocks are classified to obtain the synthetic data set.

[0006] Optionally, determining the discreteness between each of the image blocks and the local dataset includes: The image block outputs a predicted distribution value through a pre-trained prediction model; The difference between the true distribution value of the label corresponding to the image block and the predicted distribution value is used as the discreteness.

[0007] Optionally, the satellite communication-based model training method further includes: Based on the predicted distribution value and the true distribution value corresponding to each of the image blocks, a data quality score value of each image block in the synthetic data set is determined and then accumulated to obtain a score value of the synthetic data set.

[0008] The present invention also provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the satellite communication-based model training method as described in any one of the above items is implemented.

[0009] The present invention also provides a model training method based on satellite communication, the method comprising: Sending the current global model to each satellite respectively, so that each satellite performs localized model training according to the current global model; Obtaining a local model and a synthetic data set obtained by completing localization model training for each of the satellites; Determining a weight value of a local model corresponding to each synthetic data set based on the score value of each synthetic data set; Based on the weight value of each local model, aggregating all the local models to obtain an updated global model; Using the data of each synthetic data set as training data, training the updated global model to obtain a trained global model; After using the trained local model as the current global model and updating the training round, returning to the step of sending the current global model to each satellite respectively so that each satellite performs localized model training according to the current global model; When the number of training rounds is greater than a preset threshold, the current global model is used as the optimal global model.

[0010] Optionally, determining a weight value of a local model corresponding to each synthetic data set based on a score value of each synthetic data set includes: Determining the staleness of the global model corresponding to the current training round based on a preset functional relationship between a preset threshold and the training round; The sum of the score value and the obsolescence is determined according to a first preset weight value corresponding to the score value and a second preset weight value corresponding to the obsolescence and is used as the weight value of the local model.

[0011] Optionally, the aggregating all the local models based on the weight value of each local model to obtain an updated global model includes: When receiving a plurality of local models sent by the satellites within a preset period, sorting each of the local models to form a sorting queue, and updating the global model one by one based on each local model in the sorting queue; When the preset period ends, all the local models are aggregated based on the weight value of each local model to obtain the updated global model.

[0012] Optionally, the aggregating all the local models based on the weight value of each local model to obtain an updated global model further includes: When the number of the received local models is greater than the upper limit of the sorting queue, all the local models in the sorting queue are aggregated based on the weight value of each local model to obtain the updated global model.

[0013] The present invention also provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the satellite communication-based model training method as described in any one of the above items is implemented.

[0014] The present invention also provides a model training system based on satellite communication, the system comprising a plurality of ground stations and a plurality of satellites; The ground station is used to send the global model and the current training round to the target satellite; The satellite is configured to train the global model using a local data set until the number of training rounds reaches the current training round, output a local model, and update the number of rounds; The satellite is configured to generate a synthetic dataset based on the local dataset; The satellite is configured to transmit the local model, the update round number, and the synthetic data set to a target ground station; The ground station is configured to determine a weight value of a local model corresponding to each synthetic data set based on a score value of each synthetic data set; The ground station is configured to aggregate all the local models according to a weight value of each local model to obtain an updated global model, and train the updated global model according to each synthetic data set to obtain a trained global model; The ground station is used to update the training round, and when the updated training round is greater than a threshold, the trained global model is used as the optimal global model.

[0015] Optionally, the satellite communication-based model training system is further used to: receiving, via each of said ground stations, a complete synthetic data set; Determining a weight value for each of the synthetic data sets, wherein the weight value for the synthetic data set is equal to the proportion of the amount of data of the synthetic data set in the total amount of data of all synthetic data sets; Determine the loss function of the global model on each of the synthetic datasets respectively; Based on the weight value of each synthetic data set, all the loss functions are accumulated to obtain the minimized global loss function of the global model.

[0016] Optionally, the satellite communication-based model training system is further used to: Determining the trajectory of each of the ground stations and the trajectory of each of the satellites based on a preset geocentric inertial coordinate system; Determining the angle value between each ground station and the target satellite based on the angle function relationship between the trajectory of the ground station and the trajectory of the satellite; When the angle between the ground station and the target satellite is within a preset value range, the ground station is controlled to connect with the target satellite and communicate with each other.

[0017] The present invention has at least the following beneficial effects: First, the satellite receives the global model and the current training round from the ground station and trains the global model using its local dataset. This distributed training approach allows the satellite to process local data in parallel, avoiding the need to transmit large amounts of raw data to the ground station, thereby reducing communication costs and data transmission time. Second, after reaching the specified number of training rounds, the satellite outputs the local model and the updated round number, and generates a synthetic dataset based on the local data. The generation of the synthetic dataset further enriches the training samples, compensates for any deficiencies in the local data, and enhances the model's generalization capabilities. Finally, the local model, updated round number, and synthetic dataset are sent to the ground station for subsequent updating and optimization of the global model. In this way, this solution fully utilizes the satellite's computing power and local data resources, while also improving the efficiency and quality of model training with the assistance of the synthetic dataset. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The accompanying drawings are used to provide a further understanding of the technical solution of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the technical solution of the present invention and do not constitute a limitation to the technical solution of the present invention.

[0019] Figure 1 It is a structural diagram of a model training system based on satellite communication; Figure 2 It is a flowchart of the steps of a satellite communication-based model training method; Figure 3 is a flowchart of step S104 in a satellite communication-based model training method; Figure 4 This is a flowchart of the steps of another satellite communication-based model training method; Figure 5 It is a schematic diagram of the interactive process between a ground station and multiple satellites in a model training system based on satellite communication; Figure 6 It is a schematic diagram of the FL framework of a model training system based on satellite communication; Figure 7 It is a flow chart of the steps in the operation of a satellite communication-based model training system; Figure 8 This is a comparison chart of the effects of the model training method based on satellite communication and the model training method based on existing technology. DETAILED DESCRIPTION

[0020] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0021] Researchers have found that traditional machine learning based on satellite imagery often requires downloading image data from satellites to ground stations for model training. However, since LEO satellites only communicate with ground stations a few times per day, the duration of each communication is very limited. Furthermore, the transmission of high-resolution imagery requires significant bandwidth and storage capacity, increasing latency and energy consumption, making it impossible for satellites to fully transmit their image data to ground stations in a single communication. Furthermore, high-resolution satellite imagery may infringe on privacy and be subject to regulatory restrictions. Therefore, downloading high-resolution satellite imagery to ground stations is not always feasible. Distributed model training across LEO constellations is also not feasible. The high-speed motion and relative position fluctuations of satellites complicate the establishment and maintenance of inter-satellite links (ISLs). The complex pointing, pointing, and tracking (PAT) systems required to maintain stable links further burden the system, leading to increased communication complexity and latency. Therefore, LEO satellites currently cannot implement inter-satellite communication for distributed training.

[0022] To address these challenges, federated learning (FL), a new paradigm for distributed training on large-scale edge devices, is being applied to machine learning training tasks in satellite constellations. FL is a new paradigm for distributed training on large-scale edge devices. FL requires a central server to coordinate all participating edge devices and aggregate their local model parameters. This approach protects data privacy, increases the model's adaptability to heterogeneous data distribution, and reduces communication costs between the central server and edge devices. Typically, for FL methods on LEO satellite networks, satellites train local models using their own datasets, while ground stations act as central servers to aggregate the satellites' local models. However, in current FL frameworks, LEO satellites act as clients, contributing their local model parameters to the ground stations acting as servers while storing the raw satellite data onboard the satellites. For example, existing FL methods like FedAvg are ineffective for satellite network computations because the links between LEO satellites and ground stations are limited in number and duration, resulting in slow synchronization of the FL global model while waiting for updates from other models. Limited bandwidth resources and the sparse and heterogeneous connectivity of satellites make the transmission of high-resolution imagery difficult. In addition, the data distribution differences among different satellite clients affect the generalization ability of the global model.

[0023] To address these issues, dataset distillation technology has been introduced based on the existing FL for satellite communications. This technology can help synthesize and compress high-resolution raw image datasets, thereby facilitating satellite transmission of information-rich synthetic image data.

[0024] Dataset distillation was first proposed by [ 1 ]. Its core idea is to generate a set of small but information-rich synthetic datasets to replace the original large-scale datasets for model training, thereby achieving similar model capabilities. This method not only reduces the amount of data transmitted, but also improves the model's generalization ability while protecting data privacy.

[0025] Because federated learning requires extensive communication to transmit models when training distributed models while preserving data privacy, model parameters are typically much larger than the small number of data points after distillation. This approach can significantly reduce the cost of uploading data from the client back to the server, and the synthesized dataset can retain the characteristic information of the entire original dataset. In the context of federated learning, the application of dataset distillation technology has been further expanded. The FedDistill method, first introduced dataset distillation into the federated learning framework, reduces the communication overhead between the client and the server by performing dataset distillation locally on the client. This method demonstrates significant performance improvements under non-IID data distributions. Some research on federated learning transmits locally generated synthetic datasets from the client to the server instead of model updates. Jin et al. proposed a multi-round dataset distillation algorithm that, through multiple iterative distillations, further compresses the data while preserving key information. This allows for the transmission of more valuable data within limited communication bandwidth, improving the feasibility and efficiency of federated learning in specialized scenarios.

[0026] Dong et al. showed that dataset distillation can provide privacy protection against accidental data leakage. They theoretically analyzed the connection between dataset distillation and differential privacy. Furthermore, they empirically verified that synthetic datasets are irreversible to the original data in terms of L2 and LPIPS similarity metrics. Chen et al. applied dataset distillation to generate high-dimensional data and provide different privacy guarantees for private data sharing with low memory and computational costs. This demonstrates that synthetic data can be transmitted while preserving data privacy.

[0027] The main problems in existing FL systems are: (1) When the connection density between satellites and ground stations is too sparse, other satellites will spend a lot of time waiting; when the connection density between satellites and ground stations is too dense, a large number of satellites will be idle, resulting in a waste of computing resources. (2) Most dataset distillation methods require high computing power of the equipment, and the limited computing power of LEO satellites makes it impossible to use these distillation methods normally. (3) The obsolete heterogeneity of satellite models and the heterogeneous distribution of satellite data make model convergence difficult.

[0028] To address these issues, the present invention provides a satellite communication-based model training method, device, and system that can utilize satellite data for model training while improving the efficiency of model training. The following are various embodiments of the technical solution of this application: Please refer to Figure 1 , Figure 1 It is a structural diagram of a model training system based on satellite communication.

[0029] The present invention also provides a model training system based on satellite communication, which includes multiple ground stations and multiple satellites.

[0030] The ground station is used to send the global model and the current training round to the target satellite; Satellites are used to train the global model using local datasets until the number of training rounds reaches the current training round, and then output the local model and update the round number; Satellite, used to generate synthetic datasets based on local datasets; Satellites, used to send local models, update rounds, and synthetic datasets to target ground stations; The ground station is used to determine the weight value of the local model corresponding to each synthetic data set based on the score value of each synthetic data set; The ground station is used to aggregate all local models according to the weight value of each local model to obtain an updated global model, and train the updated global model according to each synthetic data set to obtain a trained global model; The ground station is used to update the training rounds. When the updated training rounds are greater than the threshold, the trained global model is used as the optimal global model.

[0031] It will be appreciated that in this embodiment, the ground station transmits the global model and training rounds to the satellite. The satellite then uses local data for training and generates local models and synthetic datasets, while outputting the updated round number. This process fully utilizes the satellite's local computing power, avoids large amounts of data transmission, and significantly reduces communication costs and time overhead. The ground station assigns weights to the local models based on the scores of the synthetic dataset. This dynamic weight allocation differentiates the contributions of different satellites based on data quality and model performance, further optimizing the model aggregation process. The ground station then aggregates the local models based on the weights and trains the global model using the synthetic dataset, further improving model performance. Training rounds are iteratively updated until a preset threshold is reached, ultimately resulting in the optimal global model. Through distributed computing, dynamic weight adjustment, and iterative optimization, the entire solution effectively utilizes satellite data and significantly improves the efficiency and quality of model training.

[0032] In some embodiments, a satellite communication-based model training system is further configured to: All synthetic data sets are received through each ground station; a weight value of each synthetic data set is determined, wherein the weight value of the synthetic data set is equal to the proportion of the data volume of the synthetic data set in the total data volume of all synthetic data sets; a loss function of the global model on each synthetic data set is determined respectively; based on the weight value of each synthetic data set, all loss functions are accumulated to obtain the minimized global loss function of the global model.

[0033] Specifically, assume that a FL system has a constellation of K satellites and G ground stations. Satellite k∈K={1,…K} has a private dataset D k In this FL system, the ground station and satellite play the roles of central server and client respectively. Therefore, the ground station needs to coordinate with the satellite to train a global model w∈R d , by minimizing the global objective function as follows: (1) Wherein, subscript k represents the satellite subscript, and K represents the set of satellites contained in the FL system; is the size of the entire dataset, so the meaning of formula (1) is to use the number of records of each satellite's private dataset to the total number of records of all satellite datasets as the weight of each satellite, and to minimize the global loss function by weighted summing the loss function of the global model on the local dataset of each satellite. Local objective function Representation and Dataset The associated loss function is: (2) Where l(w,x) is the training loss for the data x and the model parameters w. The communication latency between ground stations is lower than the satellite-to-ground station communication latency. Therefore, multiple ground stations can function as FL servers.

[0034] In some embodiments, a satellite communication-based model training system is further configured to: Based on the preset geocentric inertial coordinate system, the trajectory of each ground station and the trajectory of each satellite are determined; based on the angle function relationship between the trajectory of the ground station and the trajectory of the satellite, the angle value between each ground station and the target satellite is determined; when the angle value between the ground station and the target satellite is within the preset value range, the ground station and the target satellite are controlled to connect and communicate with each other.

[0035] Specifically, the geocentric inertial coordinate system refers to a coordinate system that remains stationary or in uniform linear motion (no acceleration) in space. Its origin is at the center of the earth, the Z axis coincides with the earth's rotation axis, and the X and Y axes are located in the equatorial plane, pointing to the prime meridian and the meridian of 90° east longitude, respectively.

[0036] In the geocentric inertial coordinate system, assume that the satellite k∈K ={1,... ,K} and the ground station g∈G={K + 1,..., K+ G} have satellite trajectories r k (t) and ground station r g (t), where t is a continuous wall-clock time. When satellite k is at the minimum elevation angle α min The link between satellite k and ground station g is feasible when it is visible from ground station g, that is: (3) in, Represents the trajectory difference between the ground station trajectory and the satellite trajectory and the angle between the ground station trajectory, which is recorded as , which represents the angle between the satellite and the ground station. Therefore, the above formula (3) means that the angle between the satellite and the ground station is less than or equal to , that is, to meet the minimum elevation angle Satellite k is visible to the ground station. The position and trajectory of the ground station and satellites allow accurate prediction of future connectivity. When a satellite is visible to the ground station, the satellite can establish a connection with the ground station for communication.

[0037] In a specific application scenario, the orbital edge computing simulation platform cote, which provides orbital edge computing and runtime service simulation, integrates orbital mechanics modeling with Earth rotation dynamics algorithms to dynamically analyze the relative positions of satellites and the ground and simulate the communication schedule between satellite constellations and ground stations. The experiment constructed a heterogeneous network consisting of 100 low-orbit satellites and 10 distributed ground terminals. By simulating and collecting satellite-to-ground communication frequency statistics over a 24-hour period, including constellation-level global link up / down events and individual satellite-to-ground node interaction time series data, the communication schedule for each satellite and ground station was derived.

[0038] Please refer to Figure 2 , Figure 2 It is a flowchart of the steps of the model training method based on satellite communication.

[0039] This embodiment provides a satellite communication-based model training method including: S101: Receive the global model and current training round from the ground station.

[0040] S102: Use the local data set to train the global model.

[0041] S103: When the number of times the global model is trained reaches the current training round, the local model is output and the number of rounds is updated.

[0042] S104: Generate a synthetic dataset based on the local dataset.

[0043] S105: Send the local model, update round number, and synthetic data set to the ground station.

[0044] It will be appreciated that in this embodiment, first, the satellite receives the global model and the current training round from the ground station and trains the global model using a local dataset. This distributed training approach allows the satellite to process local data in parallel, avoiding the need to centrally transmit large amounts of raw data to the ground station, thereby reducing communication costs and data transmission time. Secondly, after reaching the specified number of training rounds, the satellite outputs the local model and the updated round number, and generates a synthetic dataset based on the local data. The generation of the synthetic dataset further enriches the training samples, compensates for any deficiencies in the local data, and enhances the model's generalization capabilities. Finally, the local model, updated round number, and synthetic dataset are sent to the ground station for subsequent updating and optimization of the global model. In this way, the solution fully utilizes the satellite's computing power and local data resources, while improving the efficiency and quality of model training with the assistance of the synthetic dataset.

[0045] In some embodiments, the process of a satellite using a local dataset to train a localized model on a global model is as follows: (1) Receive the global model and the current training round t from the ground station, and record the number of update rounds as τ; (2) Use the local dataset to train the local update model to the current training round (the satellite can switch between working and idle states); (3) Satellites distill local datasets to obtain synthetic datasets; (4) Push the local model, synthetic dataset, and update round number τ to the ground station.

[0046] In each training round, when the satellite first establishes a connection with the ground station, it receives the latest global model from the ground station. After receiving the global model, the satellite will use the global model to initialize the local model, and then perform SGD (Stochastic Gradient Descent) steps on the satellite's local dataset to train the local model. The above process is represented as follows: (4) Where η is the learning rate, the superscript j is the local training index, indicating the round of local training, and the subscript k is the satellite index. Represents the dataset D from point j k Select a mini-batch of size B in , is the loss function with a random mini-batch X.

[0047] Please refer to Figure 3 , Figure 3 The present invention is a flowchart of step S104 in the satellite communication-based model training method.

[0048] In some embodiments, step S104 includes: S201: Convert each image in the local dataset into an image block.

[0049] S202: Determine the discreteness between each image block and the local dataset.

[0050] S203 : Screen each image block based on the discreteness, and classify the screened image blocks to obtain a synthetic data set.

[0051] Specifically, after the satellite completes the localization model training, it needs to k Each image in is smoothly cropped to Crop parts to form smaller rectangular image blocks, and each small image is recorded as , measured by KL divergence Image and local dataset D k The distribution gap is obtained, and the N image blocks with the most similar distribution are added to the synthetic dataset S k middle.

[0052] In some embodiments, the method of screening each image block includes: For each individual image block , the algorithm calculates the output of the pre-trained model and the true label y i The KL divergence between .

[0053] The divergence is determined by the information value measure, A measure of the agreement between the model's predictions and the actual class labels.

[0054] Will The negative KL divergence is assigned by formula (6), which is as follows:

[0055] This step aims to identify the image patches that minimize the KL divergence, thereby selecting the image patches that are most similar to the true distribution of the satellite local dataset.

[0056] Select the one with the highest The first N image blocks with the same value are selected, because these image blocks are considered to be the most representative of category C. In order to save the transmission bandwidth between satellite and ground, one synthetic image is finally distilled from a single category C, for a total of N images.

[0057] In some embodiments, step S202 includes: The image block outputs a predicted distribution value through a pre-trained prediction model; the difference between the true distribution value and the predicted distribution value of the label corresponding to the image block is used as the discreteness.

[0058] In some embodiments, the discreteness between each image block and the local dataset is calculated as follows: (5) Formula (5) is the calculation formula for KL divergence. KL divergence is an indicator that measures the difference between one probability distribution and another probability distribution. In this context, it is used to quantify the predicted distribution and the true distribution y i The difference between them. and denote the probability of the jth category in the true distribution and the predicted distribution, respectively.

[0059] It's understandable that a smaller KL divergence indicates that the model's predicted distribution for the sample is closer to the true distribution, meaning the model fits the sample better. Therefore, the smaller the KL divergence, the greater the sample's contribution to the model. By using an exponential function to attenuate the KL divergence, samples with smaller KL divergences contribute more to the weight.

[0060] In some embodiments, a satellite communication-based model training method further includes: Based on the predicted distribution value and the true distribution value corresponding to each image block, the data quality score value of each image block in the synthetic data set is determined and then accumulated to obtain the score value of the synthetic data set.

[0061] Optionally, the score of the synthetic dataset can be calculated using the following formula: (7) in, Denotes the dataset S belonging to satellite k k A sample in ξ represents the true distribution of sample ξ, Represents the model's predicted distribution for sample ξ.

[0062] It can be understood that the score value α of the synthetic dataset k To reflect the data quality of satellite k. Accumulate the corresponding data set S of satellite k through formula (7) k The data quality is measured by the contribution of each sample ξ in . The score reflects the importance of satellite data in the model. The larger the score, the greater the contribution of the satellite data to the model.

[0063] In the above embodiment, the dataset distillation method deployed on the satellite can obtain a representative synthetic image set with low computing power cost by randomly cropping and KL divergence scoring the images.

[0064] In some embodiments, a satellite device includes a memory and a processor, the memory stores a computer program, and the processor implements the satellite communication-based model training method of any of the above embodiments when executing the computer program.

[0065] It can be understood that the contents of the above method embodiments are all applicable to the electronic device embodiment. The functions specifically implemented by the electronic device embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0066] Please refer to Figure 4 , Figure 4 This is a flowchart of the steps of another satellite communication-based model training method.

[0067] The present invention also provides a satellite communication-based model training method comprising: S301. Send the current global model to each satellite respectively, so that each satellite performs localized model training according to the current global model.

[0068] S302: Obtain the local model and synthetic data set obtained by completing localization model training for each satellite.

[0069] S303: Based on the score value of each synthetic data set, determine the weight value of the local model corresponding to each synthetic data set.

[0070] S304: Based on the weight value of each local model, aggregate all local models to obtain an updated global model.

[0071] S305: Use the data of each synthetic data set as training data to train the updated global model to obtain a trained global model.

[0072] S306: After using the trained local model as the current global model and updating the training round, return to step S301.

[0073] S307: When the number of training rounds is greater than a preset threshold, the current global model is used as the optimal global model.

[0074] As will be appreciated, in this embodiment, the ground station first transmits the current global model to each satellite. The satellite then performs localized model training based on local data and generates a synthetic dataset. This process fully utilizes the satellite's computing resources, enables parallel training, and significantly improves training efficiency. Secondly, the ground station assigns weights to each local model based on the synthetic dataset's score. This dynamic weight assignment differentiates the contributions of different satellites based on data quality and model performance, further optimizing the model aggregation process. Subsequently, an updated global model is obtained through weighted aggregation, and further trained using the synthetic dataset to further improve model performance. Finally, through iterative optimization, the global model is continuously updated and the training process is repeated until a preset number of rounds is reached, ultimately resulting in the optimal global model. Through distributed computing, dynamic weight adjustment, and iterative optimization, this entire solution effectively utilizes satellite data and significantly improves the efficiency and quality of model training.

[0075] In some embodiments, step S303 includes: Based on the preset functional relationship between the preset threshold and the training round, the staleness of the global model corresponding to the current training round is determined; according to the first preset weight value corresponding to the score value and the second preset weight value corresponding to the staleness, the sum of the score value and the staleness is determined and used as the weight value of the local model.

[0076] In some embodiments, staleness is calculated as follows: (8) Among them, β k Staleness weight coefficient, T is the model version of the ground station server model at the current time point, t k Submit the version of the model for the k-th satellite client.

[0077] A specific method of determining the sum of the score value and the staleness according to the first preset weight value corresponding to the score value and the second preset weight value corresponding to the staleness and using the sum as the weight value of the local model is as follows: (9) Among them, λ∈(0,1) is an adaptive hyperparameter, w t is the global model of the t-th round of training, w t+1 is the global model trained in the t+1th round.

[0078] It is understandable that when the distribution of the synthetic dataset matches that of the original satellite data, α k Increase. For α k For satellites k with low scores, the impact of bad data and models on global data is reduced by reducing λ. When updating the global model, greater staleness will lead to greater errors. For satellites with large staleness (T−tk ), where λ can be increased to mitigate the error caused by staleness. When t=τ, β k It should be 1, indicating that the satellite has just obtained a new global model in the last round of communication, and directly transmits the updated local model to the ground station to update the global model; when T−t k When β increases, k will decrease, then according to formula (9), λ will decrease, indicating that the satellite model is updated and transmitted to the ground station for update. At this time, the ground station has updated T−t k wheel.

[0079] In a specific embodiment, since training the optimal global model requires T max (preset threshold) cycles, when t∈Tmax cycles, first subtract the number of training rounds t uploaded by the satellite from the number of cycles T k The staleness β of the model uploaded by the satellite is calculated by the staleness function, and then the γ of this round is updated by formula (9). The ground station receives the locally trained models of the satellite in sequence. , and update the global model by weighted average using formula (10). After completing the hybrid weighted aggregation, the ground station uses the synthetic data set provided by all satellites to perform gradient updates according to formula (11). After the ground station updates the global model, it coordinates the satellites to conduct the next round of training and sends the latest global model. Formulas (9) and (10) are: (10) (11) It can be understood that this embodiment introduces a staleness function to adaptively handle delayed updates, thereby alleviating the negative impact of the asynchrony of the satellite-to-ground link and retaining synthetic distilled samples to both meet the storage constraints of onboard equipment and ensure the knowledge fidelity during long-term model evolution.

[0080] In some embodiments, step S304 includes: When local models sent by multiple satellites are received within a preset period, each local model is sorted to form a sorting queue, and the global model is updated one by one based on each local model in the sorting queue.

[0081] When the preset period ends, all local models are aggregated based on the weight value of each local model to obtain an updated global model.

[0082] In some embodiments, step S304 further includes: When the number of received local models is greater than the upper limit of the sorting queue, all local models in the sorting queue are aggregated based on the weight value of each local model to obtain an updated global model.

[0083] Please refer to Figure 5 , Figure 5 This is a schematic diagram of the interaction process between the ground station and multiple satellite communications in the satellite communication-based model training system.

[0084] The aggregation method in the first embodiment is: by presetting a time period length T e ( Figure 5 Middle T e =3 cycles), at the ground station, each time the cycle length T e When , all the models sent by satellites in the model sorting queue will be aggregated once, such as Figure 5 Assume that the time period length is T e cycles, i∈R represents the length of the cycle, and the global model is at time t*T e w t , then when satellite k is in the preset period [t*T e ,(t+1)T e ] is connected to the ground station, then the ground station will send w t Give it to satellite k, and then satellite k uses its own local dataset D k To train the local update model, after satellite k has trained the local update model, in the preset period [t*T e ,(t+1)T e ] connect to the ground station for the second time, send the model to the ground station to receive and store in sequence, if (t+1)T e At the time when it is not yet time to aggregate the model sent by the satellite k, due to the preset cycle length T e When the cycle is up, the ground station will aggregate the models sent by all stored satellites to train and update the global model. The above triggering process of training the global model can be expressed as follows: (12) Where i represents the time interval [t*T e ,(t+1)T e ], t represents the number of rounds of global model training at the ground station, T e It is a pre-set time period length, B represents the satellite index set sent to the ground station for local update model in this round, a(B) represents the aggregation index of asynchronous periodic aggregation. When a(B)=1, it means that the set period length is reached, and the ground station will start to aggregate the local updates sent by all satellites to train the global model, otherwise no update aggregation will be performed.

[0085] It's understandable that when satellite-ground station connections are sparse, periodic aggregation can significantly improve the efficiency of ground station global model updates. By setting a set periodic trigger to periodically update the ground station and aggregate all received local updates, this approach saves update time compared to traditional asynchronous aggregation schemes (where the satellite sends local update models to the ground station, which then updates the global model one by one in the order received), thereby alleviating the challenge of long update wait times caused by sparse satellite connections.

[0086] The aggregation method in the second specific embodiment is as follows: a buffer queue with a length of V is set on the ground station to receive and store the local update models sent by the satellite. When the number of models sent by the satellite is greater than V, the ground station will perform an aggregation to update the global model. The specific triggering process is as follows: (13) Where a(B) represents the aggregation index for asynchronous buffered aggregation, B represents the index set of satellites that send local update models to the ground station, |B| represents the number of satellites sending local update models to the ground station, that is, the number of local update models received by the ground station, and V represents the length of the buffer queue set by the ground station. When the number of local update models sent by satellites is greater than or equal to the buffer queue length, that is, a(B) = 1, the ground station aggregates and updates the global model; otherwise, no aggregate update is performed.

[0087] It's understandable that when the connection between satellites and ground stations is too dense, satellites continuously send local models to the ground station. When the number of models is greater than or equal to the ground station's buffer queue length, the ground station promptly aggregates and updates the global model to account for the temporal heterogeneity of satellite connections. This prevents a large number of satellites from being idle when the connection between satellites and ground stations is dense, thus avoiding wasted computing resources and long satellite wait times. The set buffer queue length V limits the maximum possible staleness of the models sent by the satellites. Keeping the staleness of the models aggregated by the ground station within the range [0, V] helps mitigate the negative impact of excessively stale models on global model optimization.

[0088] The polymerization method in the third embodiment is: for example, Figure 3 As shown (cycle length T e =2, buffer queue V=3), the specific execution process is described as follows: When the satellite establishes a connection with the ground station for the first time and the connection density is sparse, in round t, the satellite receives the global update w sent by the ground station t-1 , and use the local dataset D k To train and update the local update model w new, and then sent to the ground station during the second connection. After receiving the local update model sent by the satellite, the ground station will store it in the buffer queue V (the buffer queue is not full). The ground station will update the model in the preset period [t*T e ,(t+1)T e ] will continue to select local models in order to continuously update the global model until (t+1)T e When the predefined cycle length is reached, the ground station aggregates all local models in the buffer queue to update the global model, and then sends the global model to the satellite when it connects to the satellite next time; when the satellite connection density is dense, the ground station may not reach the predefined cycle length T e When the number of local models received is greater than or equal to the buffer queue V, the ground station will directly perform an aggregation to update the global model according to the buffer aggregation strategy. In asynchronous dynamic aggregation, whether the cycle length is reached or the buffer queue is full, the ground station will be triggered to aggregate and update the global model. The specific process can be expressed as follows: (14) Among them, B represents the satellite index set returned to the local model in this round, V is the buffer queue length of the buffer, a(B) is the aggregation index, and a(B)=1 means that the ground station can perform aggregation to update the global model.

[0089] It is understandable that the hybrid weighted aggregation method deployed in ground combat can enhance the impact of high-information data and improve the model convergence speed.

[0090] In some embodiments, a ground station device includes a memory and a processor, the memory stores a computer program, and the processor implements the satellite communication-based model training method of any one of the above embodiments when executing the computer program.

[0091] It can be understood that the contents of the above method embodiments are all applicable to the electronic device embodiment. The functions specifically implemented by the electronic device embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0092] Optionally, the present invention also provides a specific operation mode of a satellite communication-based model training system in an actual scenario: In this example, MNIST

[27] is a classic benchmark dataset in the field of machine learning. Its core consists of handwritten digits from 0 to 9, collected from samples of 250 different writers. FashionMNIST

[28] is known for its more difficult clothing classification task, covering 10 categories of fashion items such as T-shirts, jackets, and sandals. Both datasets use a standardized data scale design: the training set contains 60,000 samples and the test set contains 10,000 samples, each sample is a 28×28 pixel grayscale image. Both datasets contain 60,000 training samples and 10,000 test samples.

[0093] A satellite communication network model was constructed based on the Cote simulation platform

[26] . The system consists of 100 low-orbit satellites and 10 distributed ground stations. Two circular ground station clusters are deployed on the ground side, one at 87° north latitude and the other at 87° south latitude, each containing five terminal nodes. The space segment uses a polar orbit (inclination 97.3°), and satellite orbit data is generated using the PlanetLab constellation parameters. A cache queue with a capacity of 25 units is set in the algorithm design, and a time window threshold of 100 seconds is set. The experimental data comes from the satellite link status time series data generated by the simulation, which contains complete connection records for 9000 consecutive observation cycles.

[0094] Each satellite is provided with a corresponding amount of image data based on the number of data categories Ck contained in the satellite. Under the IID data distribution, set Ck=10, so that each satellite receives data from all categories. Under the Non-IID data distribution, set Ck=5, so that each satellite receives data from half of the categories. Under each data distribution, each satellite is randomly provided with the same amount of image data from each category for training.

[0095] For the satellite communication-based model training system, the connection time between the satellite and the ground station was used as a benchmark. The global model training performance during the first 9000 seconds of the dynamic connection between the satellite and the ground station was observed. An MLP model was used as the training model, with a learning rate of 0.001 and no learning rate decay or early stopping. After each communication, the satellite completed five rounds of local training.

[0096] Please refer to Figure 6 , Figure 6 It is a schematic diagram of the FL framework of a model training system based on satellite communication.

[0097] First, the satellite receives the global model and update round number from the ground station through the coordinator. Then, the satellite uses the local dataset to train the local update model to the target round number. Then, the satellite distills its own dataset to obtain a synthetic dataset. Finally, the coordinator pushes the local updated model, synthetic data, and update round number to the ground station.

[0098] Please refer to Figure 7 , Figure 7 It is a flow chart of the steps during the operation of the satellite communication based model training system.

[0099] First, the satellite is required to perform dataset distillation after completing the local training of the current round. In round t, the satellite receives the global update w sent by the ground station. t-1 And use the local dataset D k To train and update the local update model When the satellite completes local training, it needs to perform local data set D k Each image in is smoothly cropped to form a smaller rectangular image block, and each small image is recorded as , measured by KL divergence Image and local dataset D k The distribution gap is obtained, and the image block with the most similar distribution is added to the set S k After the satellite completes the dataset distillation, it uses formula (7) to obtain αk for all images in S. αk reflects the data quality of satellite k. The smaller the KL divergence (the greater the amount of information), the more significant the contribution. During the second connection, the satellite sets the local model parameters ,Synthetic dataset S k and the data quality score αk are sent to the ground station.

[0100] After the ground station receives the data from satellite k, it first obtains βk by formula (8), and then obtains the mixed weight γk by formula (9). Then, after updating the local model and the mixed weight γk, it stores them in the buffer queue V (the buffer queue is not full). According to the dynamic aggregation strategy of ADA, it waits for the current preset cycle to end or the length of the buffer queue V to meet the preset requirements, and then performs weighted model aggregation. Then, all the synthetic data D received by the ground station are used. syn Perform gradient update.

[0101] It can be understood that by pre-setting a preset cycle length, asynchronous aggregation is performed before the preset cycle is reached, that is, a local model is received to update the global model once. If the preset cycle length is reached, the global model is aggregated and updated once, thereby shortening the long update waiting time caused by sparse satellite connections; and in the case of dense satellite connections, the asynchronous buffer aggregation strategy is used to fully utilize a large number of satellites that are idle when the satellite connections are dense, thereby fully utilizing the satellite's computing resources.

[0102] Please refer to Figure 8 , Figure 8 This is a comparison chart of the effects of the model training method based on satellite communication and the model training method based on existing technology.

[0103] like Figure 8 As shown in Figure 1, the model training method based on satellite communication is FL-SA, and the existing model training methods include AFL, Asyperiod, ADA, FedBuff, GMM, and FL-M3D.

[0104] The satellite communication-based model training method (FL-SA) of this technical solution can perform aggregate updates of the global model faster than FedAvg and FedBuff, saves a lot of time for ground stations to wait for satellite synchronization compared to AsyPeriod periodic synchronization updates, and improves the convergence speed of the model.

[0105] This technical solution's satellite communication-based model training method (FL-SA) uses a hybrid weighted aggregation method by calculating staleness scores and synthetic data quality scores. This significantly reduces the impact of highly stale models, low-quality synthetic data, and lagging models on the optimization of the global model. After aggregation, gradient compensation is performed using the synthetic dataset, accelerating global model optimization and effectively improving the model's prediction accuracy, including performance under both IID and non-IID data distributions. This significantly enhances model performance and generalization, enabling the model to better handle a variety of data distributions and complex real-world situations.

[0106] Those skilled in the art will appreciate that all or some of the steps and devices in the methods disclosed above can be implemented as software, firmware, hardware, or any suitable combination thereof. Some or all of the physical components may be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on computer-readable media, which may include computer storage media (or non-transitory media) and communication media (or transient media). As is well known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disks (DVDs) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. As is well known to those skilled in the art, communication media typically embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.

[0107] An embodiment of the present application also provides a computer-readable storage medium, which stores a program executable by a processor. When the program executable by the processor is executed by the processor, it is used to implement the satellite communication-based model training method as described in any one of the above specific embodiments.

[0108] An embodiment of the present application also discloses a computer program product, including a computer program or computer instructions, which are stored in a computer-readable storage medium. A processor of a computer device reads the computer program or computer instructions from the computer-readable storage medium, and the processor executes the computer program or computer instructions, so that the computer device executes the satellite communication-based model training method as described in any of the previous embodiments.

[0109] It can be understood that the contents of the above method embodiments are all applicable to the present storage medium embodiment, the functions specifically implemented by the present storage medium embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0110] The terms "first", "second", "third", "fourth" etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can, for example, be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, device, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices. It should be understood that in the present application, "at least one (item)" refers to one or more, and "a plurality of" refers to two or more.

[0111] In the several embodiments provided in this application, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0112] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0113] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0114] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0115] Although the description of the present application has been quite detailed and specifically describes several embodiments, it is not intended to be limited to any of these details or embodiments or any particular embodiment, but should be considered to provide a broad possible interpretation of these claims by reference to the appended claims, taking into account the prior art, so as to effectively cover the intended scope of the present application. In addition, the above description of the present application is based on the embodiments foreseen by the inventors, which is intended to provide a useful description, and those non-substantial changes to the present application that have not yet been foreseen may still represent equivalent changes to the present application.

Claims

1. A model training method based on satellite communication, characterized in that: The method comprises: Receive the global model and current training round from the ground station; Training the global model using the local dataset; When the number of times the global model is trained reaches the current training round, outputting the local model and updating the round number; Based on the local dataset, generating a synthetic dataset; The local model, the update round number, and the synthetic data set are sent to the ground station.

2. A satellite communication-based model training method according to claim 1, characterized in that: The generating a synthetic dataset based on the local dataset includes: Convert each image in the local dataset into an image block; determining a discreteness between each of the image blocks and the local dataset; Based on the discreteness, each of the image blocks is screened, and the screened image blocks are classified to obtain the synthetic data set.

3. A satellite communication-based model training method according to claim 2, characterized in that: The determining of the discreteness between each of the image blocks and the local data set includes: The image block outputs a predicted distribution value through a pre-trained prediction model; The difference between the true distribution value of the label corresponding to the image block and the predicted distribution value is used as the discreteness.

4. A satellite device, characterized in that: The satellite device includes a memory and a processor, the memory stores a computer program, and the processor implements the satellite communication-based model training method described in any one of claims 1 to 3 when executing the computer program.

5. A model training method based on satellite communication, characterized in that: The method comprises: Sending the current global model to each satellite respectively, so that each satellite performs localized model training according to the current global model; Obtaining a local model and a synthetic data set obtained by completing localization model training for each of the satellites; Determining a weight value of a local model corresponding to each synthetic data set based on the score value of each synthetic data set; Based on the weight value of each local model, aggregating all the local models to obtain an updated global model; Using the data of each synthetic data set as training data, training the updated global model to obtain a trained global model; After using the trained local model as the current global model and updating the training round, returning to the step of sending the current global model to each satellite respectively so that each satellite performs localized model training according to the current global model; When the number of training rounds is greater than a preset threshold, the current global model is used as the optimal global model.

6. A satellite communication-based model training method according to claim 5, characterized in that: The step of aggregating all the local models based on the weight value of each local model to obtain an updated global model includes: When receiving a plurality of local models sent by the satellites within a preset period, sorting each of the local models to form a sorting queue, and updating the global model one by one based on each local model in the sorting queue; When the preset period ends, all the local models are aggregated based on the weight value of each local model to obtain the updated global model.

7. A satellite communication-based model training method according to claim 6, characterized in that: The method of aggregating all the local models based on the weight value of each local model to obtain an updated global model further includes: When the number of the received local models is greater than the upper limit of the sorting queue, all the local models in the sorting queue are aggregated based on the weight value of each local model to obtain the updated global model.

8. A ground station device, characterized in that: The ground station device includes a memory and a processor, the memory stores a computer program, and the processor implements the satellite communication-based model training method described in any one of claims 5 to 7 when executing the computer program.

9. A model training system based on satellite communication, characterized in that: The system includes a plurality of ground stations and a plurality of satellites; The ground station is used to send the global model and the current training round to the target satellite; The satellite is configured to train the global model using a local data set until the number of training rounds reaches the current training round, output a local model, and update the number of rounds; The satellite is configured to generate a synthetic dataset based on the local dataset; The satellite is configured to transmit the local model, the update round number, and the synthetic data set to a target ground station; The ground station is configured to determine a weight value of a local model corresponding to each synthetic data set based on a score value of each synthetic data set; The ground station is configured to aggregate all the local models according to a weight value of each local model to obtain an updated global model, and train the updated global model according to each synthetic data set to obtain a trained global model; The ground station is used to update the training round, and when the updated training round is greater than a threshold, the trained global model is used as the optimal global model.

10. A satellite communication-based model training system according to claim 9, characterized in that: The system is also used to: Determining the trajectory of each of the ground stations and the trajectory of each of the satellites based on a preset geocentric inertial coordinate system; Determining the angle value between each ground station and the target satellite based on the angle function relationship between the trajectory of the ground station and the trajectory of the satellite; When the angle between the ground station and the target satellite is within a preset value range, the ground station is controlled to connect with the target satellite and communicate with each other.

Citation Information

Cited By

  • Two-way personalized federal learning method, system and device for spatial information network

    CN122226761A