A Satellite Hopping Beam Method Based on Federated Learning
By training local model gradients locally on the satellite and performing asynchronous federal aggregation, the problems of high user data security risks and excessive complexity in system synchronous communication in beam hopping technology are solved, and the user data security and computing efficiency are improved to ensure the stable operation of the satellite communication system.
Patent Information
- Application Number
- CN202510293980.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-03-12
AI Technical Summary
In the existing beam hopping technology, the user data security risks and the complexity of the system synchronous communication are too high, resulting in security and efficiency problems of satellite communication systems.
Using a federated learning method, by training local model gradients locally on the satellite and performing asynchronous federated aggregation, new global model gradients are obtained, the complexity of system synchronous communication is reduced, and user data security is ensured.
It realizes that while ensuring user data security, it reduces the complexity of system synchronous communication, improves computing efficiency, and ensures the effective utilization and long-term stable operation of satellite communication system resources.
Smart Images

Figure CN119834871B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of satellite communication technology, and particularly to a satellite hopping beam method based on federated learning. Background Art
[0002] With the advent of the information age, people's demand for efficient and convenient information acquisition has been increasing day by day, and communication networks have become an important factor affecting economic and social development. In satellite communication, users can only obtain services when they are located within the area irradiated by the satellite beam, otherwise communication cannot be completed. However, due to the uneven distribution of ground users and the rapid movement of low-earth orbit satellites around the earth, the beam coverage area changes frequently, often resulting in a waste of satellite resources. To solve this problem, by reducing the beam irradiation range and appropriately increasing the number of beams, the limited on-board resources can be concentrated and applied to a smaller area, so as to irradiate different areas at different times, achieving the same coverage effect as a wide beam, while avoiding resource waste. Among them, the technology that determines to irradiate a specific smaller area at a specific time is called the hopping beam technology.
[0003] Although the hopping beam technology improves the flexibility and efficiency of the satellite communication system, it also brings new challenges of how to optimize beam allocation to ensure the long-term efficient operation of the system. After determining the beam coverage area, it is still necessary to further optimize the allocation of limited on-board resources, fully considering performance indicators such as system throughput and latency. In addition, the resource allocation strategy not only needs to meet the immediate user needs and system load, but also needs to estimate the impact of decisions on future system performance to ensure the continuous optimization of communication quality.
[0004] As a powerful extension of the ground network, the satellite communication network is an important communication means for building difficult areas and emergency disaster relief. Among them, low-earth orbit multi-beam satellites are an important part of the satellite communication network, and the research on the hopping beam and resource allocation algorithms of low-earth orbit multi-beam satellites is of great significance.
[0005] However, although the hopping beam technology plays an important role in improving the resource utilization rate of the satellite communication system, the key challenges in its development are still issues such as the difficulty in ensuring user data security and the high system complexity caused by synchronous communication. In current research, the hopping beam technology relies on frequent communication between satellites and ground stations to transmit data such as user requirements and channel states. Most of this transmission is raw data without processing, which is easily attacked and has security risks.
[0006] Currently, the existing technologies mainly focus on aspects such as adaptive scheduling and resource allocation algorithms, cross-satellite collaborative optimization, and the application of AI large model algorithms. Their goal is to improve the resource utilization rate and communication efficiency of satellite systems. However, in many current studies, resource scheduling still focuses on the ground control center. Each time the strategy is updated, a large amount of data exchange is required between the satellite and the ground station, resulting in a high security vulnerability of the central node of the system. Once the central node is attacked, the key data of the entire system will face the risk of leakage, which urgently needs to be solved. Summary of the Invention
[0007] This application provides a satellite hopping beam method based on federated learning to solve the problems of high user data security risk and excessive system synchronization communication complexity in the existing hopping beam technology. It realizes reducing the system synchronization communication complexity while ensuring user data security, improving the computing efficiency, and ensuring the effective utilization and long-term stable operation of satellite communication system resources.
[0008] The first aspect embodiment of this application provides a satellite hopping beam method based on federated learning, including the following steps:
[0009] Obtain the local model gradients of each satellite;
[0010] Estimate the quality of the local model gradients of each satellite, and asynchronously federate and aggregate multiple gradients that meet the preset quality compliance conditions to obtain new global model gradients;
[0011] Calculate the similarity between the local model gradients of each satellite and the new global model gradients, and when the similarity is greater than the preset threshold, determine that the new global model is in a preset convergence state.
[0012] According to an embodiment of this application, the obtaining of the local model gradients of each satellite includes:
[0013] Obtain the global model parameters of the central node;
[0014] Initialize the global model parameters, and based on the initialized global model parameters, train the hopping beam strategy locally for each satellite, and obtain the local model gradients of each satellite after the training is completed.
[0015] According to an embodiment of this application, the training of the hopping beam strategy locally for each satellite based on the initialized global model includes:
[0016] Obtain the current system state information of each satellite within a preset time step, where the current system state information includes at least one of user requirements, interference conditions, beam overlap regions, and channel states;
[0017] Based on the current system state information and the current policy network of each satellite, determine the target execution actions corresponding to each satellite, update the current system state information of each satellite after each satellite executes the corresponding target execution actions, and calculate the rewards obtained by each satellite;
[0018] Record the quadruple composed of the current system state information, the target execution actions, the rewards, and the updated current system state information, and calculate the advantage function based on the quadruple;
[0019] According to the value of the advantage function in the new update period, determine the update amplitude of the policy network and / or the parameters of the policy network, and update the parameters of the value network based on minimizing the mean square error loss function to obtain the updated parameters of the value network.
[0020] According to an embodiment of the present application, the quality estimation of the local model gradients of each satellite and the asynchronous federated aggregation of multiple gradients that meet the preset quality compliance conditions include:
[0021] Based on a preset quality index calculation formula, calculate the comprehensive quality index value of the local model gradients of each satellite;
[0022] If the comprehensive quality index value is greater than the preset index value, determine that the local model gradient meets the preset quality compliance conditions, and obtain the number of local model gradients that meet the preset quality compliance conditions;
[0023] If the number of local model gradients that meet the preset quality compliance conditions reaches the preset quantity, perform asynchronous federated aggregation on multiple gradients that meet the preset quality compliance conditions based on a preset aggregation formula;
[0024] Wherein, the preset quality index calculation formula is:
[0025] ;
[0026] Wherein, is the comprehensive quality index value of the i-th satellite node, is the time step of this asynchronous federated aggregation; is the staleness of the i-th satellite node, is a constant controlling the staleness decay rate, is the gradient of the policy network parameters of the i-th satellite node and the gradient of the value network parameters the time elapsed since the calculation; is and the cosine similarity of, is and The cosine similarity is the gradient of the global policy network parameters is the gradient of the global value network parameters is the Euclidean norm of the gradient of the local policy network parameters is the Euclidean norm of the gradient of the global policy network parameters is the Euclidean norm of the gradient of the local value network parameters is the Euclidean norm of the gradient of the global value network parameters
[0027] According to an embodiment of the present application, after calculating the similarity between the local model gradient of each satellite and the new global model gradient, it further includes:
[0028] If the similarity is less than or equal to the preset threshold, then broadcast the new global model gradient to each satellite;
[0029] Update the local model parameters of each satellite according to the new global model gradient.
[0030] According to an embodiment of the present application, after determining that the new global model is in a preset convergence state, it further includes:
[0031] Calculate the inversion success rate of the new global model;
[0032] Verify whether the new global model meets the preset training termination condition according to the inversion success rate;
[0033] If the new global model meets the preset training termination condition, terminate the training of the new global model, otherwise, execute the step of broadcasting the new global model gradient to each satellite.
[0034] According to a satellite hopping beam method based on federated learning in an embodiment of the present application, the quality of the local model gradient of each satellite is estimated, and asynchronous federated aggregation is performed on multiple gradients that meet the preset quality standard conditions to obtain a new global model gradient; calculate the similarity between the local model gradient of each satellite and the new global model gradient, and when the similarity is greater than the preset threshold, determine that the new global model is in a preset convergence state. Thereby, the problems of high user data security risk and too high system synchronous communication complexity in the existing hopping beam technology are solved, realizing reducing the system synchronous communication complexity while ensuring user data security, improving the computing efficiency, and ensuring the effective utilization and long-term stable operation of satellite communication system resources.
[0035] An embodiment of the second aspect of the present application provides a satellite hopping beam device based on federated learning, including:
[0036] An acquisition module, configured to acquire the local model gradients of each satellite;
[0037] An aggregation module, configured to perform quality estimation on the local model gradients of each satellite, and asynchronously federate and aggregate multiple gradients that meet the preset quality compliance conditions to obtain a new global model gradient;
[0038] An optimization module, configured to calculate the similarity between the local model gradients of each satellite and the new global model gradient, and determine that the new global model is in a preset convergence state when the similarity is greater than a preset threshold.
[0039] According to an embodiment of the present application, the acquisition module is configured to:
[0040] Acquire the global model parameters of the central node;
[0041] Initialize the global model parameters, and perform hopping beam strategy training locally for each satellite based on the initialized global model parameters, and obtain the local model gradients of each satellite after the training is completed.
[0042] According to an embodiment of the present application, the acquisition module is configured to:
[0043] Acquire the current system state information of each satellite within a preset time step, where the current system state information includes at least one of user requirements, interference conditions, beam overlap regions, and channel states;
[0044] Based on the current system state information and the current policy network of each satellite, determine the target execution actions corresponding to each satellite, update the current system state information of each satellite after each satellite executes the corresponding target execution actions, and calculate the rewards obtained by each satellite;
[0045] Record a quadruple composed of the current system state information, target execution actions, rewards, and updated current system state information, and calculate the advantage function based on the quadruple;
[0046] Determine the update amplitude of the policy network and / or the parameters of the policy network according to the value of the advantage function in a new update cycle, and update the parameters of the value network based on minimizing the mean square error loss function to obtain updated value network parameters.
[0047] According to an embodiment of the present application, the aggregation module is configured to:
[0048] Calculate the comprehensive quality index value of the local model gradients of each satellite based on a preset quality index calculation formula;
[0049] If the comprehensive quality index value is greater than the preset index value, it is determined that the local model gradient meets the preset quality compliance condition, and the number of local model gradients that meet the preset quality compliance condition is obtained;
[0050] If the number of local model gradients that meet the preset quality compliance condition reaches the preset quantity, asynchronous federated aggregation is performed on multiple gradients that meet the preset quality compliance condition based on a preset aggregation formula;
[0051] Among them, the preset quality index calculation formula is:
[0052] ;
[0053] Among them, is the comprehensive quality index value of the i-th satellite node, is the time step of this asynchronous federated aggregation; is the staleness of the i-th satellite node, is a constant controlling the decay rate of staleness, is the gradient of the policy network parameters of the i-th satellite node and the gradient of the value network parameters the time elapsed since the calculation; is and the cosine similarity of, is and the cosine similarity of, is the gradient of the global policy network parameters, is the gradient of the global value network parameters, is the Euclidean norm of the gradient of the local policy network parameters, is the Euclidean norm of the gradient of the global policy network parameters, is the Euclidean norm of the gradient of the local value network parameters, is the Euclidean norm of the gradient of the global value network parameters.
[0054] According to an embodiment of the present application, after calculating the similarity between the local model gradient of each satellite and the new global model gradient, the optimization module is further configured to:
[0055] If the similarity is less than or equal to the preset threshold, the new global model gradient is broadcast to each satellite;
[0056] Update the local model parameters of each satellite according to the new global model gradient.
[0057] According to an embodiment of the present application, after determining that the new global model is in a preset convergence state, the optimization module is further configured to:
[0058] Calculate the inversion success rate of the new global model;
[0059] Verify whether the new global model meets the preset training termination condition according to the inversion success rate;
[0060] If the new global model meets the preset training termination condition, terminate the training of the new global model; otherwise, execute the step of broadcasting the gradient of the new global model to each satellite.
[0061] A satellite hopping beam device based on federated learning according to an embodiment of the present application estimates the quality of the local model gradients of each satellite, and performs asynchronous federated aggregation on multiple gradients that meet the preset quality standard conditions to obtain a new global model gradient; calculates the similarity between the local model gradient of each satellite and the new global model gradient, and determines that the new global model is in a preset convergence state when the similarity is greater than a preset threshold. Thus, the problems of high user data security risks and excessive system synchronous communication complexity in the existing hopping beam technology are solved, and while ensuring user data security, the system synchronous communication complexity is reduced, the computing efficiency is improved, and the effective utilization and long-term stable operation of satellite communication system resources are ensured.
[0062] An embodiment of the third aspect of the present application provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the processor executes the program to implement a satellite hopping beam method based on federated learning as described in the above embodiment.
[0063] An embodiment of the fourth aspect of the present application provides a computer-readable storage medium, on which a computer program is stored, and the program is executed by a processor to be used to implement a satellite hopping beam method based on federated learning as described in the above embodiment.
[0064] Advantages of the present invention:
[0065] (1) Compared with the centralized architecture, the system synchronous communication complexity is reduced: The traditional centralized method of the hopping beam technology requires the central node to communicate frequently with all devices and perform large-scale global optimization. However, the federated learning adopted by the present invention reduces the need for global synchronization through the mechanism of local training and asynchronous uploading of model parameters, and avoids the system complexity brought by synchronous large-scale communication.
[0066] (2) Effectively protected the security of user data: Since the demand and location information data of ground users are of high importance, if the data is to be uploaded to the central node for further processing, it increases the possibility of user data leakage during transmission. However, the characteristics of federated learning adopted in the present invention do not require uploading the original data to the central server.
[0067] (3) Improved the computing efficiency: Federated learning allows each device to independently perform local model training, which means that the system can make full use of distributed computing resources, reduce the computing load of the central node, and effectively improve the computing efficiency when the scale of satellite networking is large.
[0068] Additional aspects and advantages of the present application will be partly given in the following description, partly will become obvious from the following description, or will be understood through the practice of the present application. Brief Description of the Drawings
[0069] The above-mentioned and / or additional aspects and advantages of the present application will become obvious and easy to understand from the following description of the embodiments in conjunction with the drawings, wherein:
[0070] Figure 1 It is a flowchart of a satellite hopping beam method based on federated learning according to an embodiment of the present application;
[0071] Figure 2 It is a flowchart of model training of a satellite local according to an embodiment of the present application;
[0072] Figure 3 It is a flowchart of asynchronous federated aggregation according to an embodiment of the present application;
[0073] Figure 4 It is a schematic flowchart of a satellite hopping beam method based on federated learning according to an embodiment of the present application;
[0074] Figure 5 It is a schematic block diagram of a satellite hopping beam device based on federated learning according to an embodiment of the present application;
[0075] Figure 6 It is a schematic structural diagram of an electronic device according to an embodiment of the present application. Detailed Embodiments
[0076] The embodiments of the present application will be described in detail below. The examples of the embodiments are shown in the drawings, wherein the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described below with reference to the drawings are exemplary and are intended to explain the present application and should not be construed as a limitation of the present application.
[0077] The following describes a satellite hopping beam method based on federated learning according to an embodiment of the present application. Aiming at the problems of high user data security risk and high system synchronization communication complexity mentioned in the above background technology, the present application provides a satellite hopping beam method based on federated learning, which solves the problems of high user data security risk and high system synchronization communication complexity in the existing hopping beam technology, optimizes the algorithm for hopping beam resource allocation by using federated learning to enhance the protection of key data, reduces the complexity of system synchronization operations, adapts to the actual operation requirements of hopping beam technology, and improves the feasibility of hopping beam resource allocation technology.
[0078] Specifically, Figure 1 It is a schematic flowchart of a satellite hopping beam method based on federated learning provided by an embodiment of the present application.
[0079] As Figure 1 shown, the satellite hopping beam method based on federated learning includes the following steps:
[0080] In step S101, obtain the local model gradients of each satellite.
[0081] Furthermore, in some embodiments, obtaining the local model gradients of each satellite includes: obtaining the global model parameters of the central node; initializing the global model parameters, and based on the initialized global model parameters, training the hopping beam strategy locally for each satellite, and obtaining the local model gradients of each satellite after the training is completed.
[0082] Specifically, at the beginning of the system, initialize the global model parameters and training data of the central node, that is, the parameters of the policy network and the parameters of the value network , and broadcast the initialized parameters of the policy network and the value network to all participating satellite nodes, so that all satellite nodes can train based on the initialized global model parameters.
[0083] Exemplarily, the specific initialization method in the embodiments of the present application can adopt Xavier initialization:
[0084] (1)
[0085] Among them, is the initial parameter of the policy network, is the initial parameter of the value network, is the number of input neurons of the policy network and the value network.
[0086] Furthermore, the initial training data set can be expressed as , where is at The system state at a time step, where the system state includes information such as user requirements, channel state, beam allocation state, and system load within the time step.
[0087] Furthermore, in some embodiments, based on the initialized global model, beam hopping strategy training is performed locally for each satellite, including: obtaining the current system state information of each satellite within a preset time step, where the current system state information includes at least one of user requirements, interference situation, beam overlap area, and channel state; determining the target execution action corresponding to each satellite based on the current system state information and the current policy network of each satellite, and updating the current system state information of each satellite after each satellite executes the corresponding target execution action, and calculating the reward obtained by each satellite; recording the quadruple composed of the current system state information, the target execution action, the reward, and the updated current system state information, and calculating the advantage function based on the quadruple; determining the update amplitude of the policy network and / or the parameters of the policy network according to the value of the advantage function in the new update period, and updating the parameters of the value network based on minimizing the mean square error loss function to obtain the updated value network parameters.
[0088] Among them, the preset time step can be a time step preset by those skilled in the art, and no specific limitation is made here.
[0089] Specifically, the specific policy training process of the embodiments of the present application adopts the Proximal Policy Optimization (PPO) algorithm. The PPO algorithm has the advantages of not relying on an experience replay buffer and being more suitable for continuous action environments compared to deep reinforcement learning algorithms; at the same time, compared to other policy-based reinforcement learning algorithms, such as Deep Deterministic Policy Gradient and Actor-Critic algorithms, the PPO algorithm has a unique clipping mechanism, which can reduce the occurrence of policy over-update problems.
[0090] The process of performing beam hopping strategy training locally for each satellite based on the initialized global model in the embodiments of the present application may specifically include the following steps:
[0091] First, obtain the current system state of each satellite at the time step including user requirements, interference situation, beam overlap area, and channel state.
[0092] Secondly, based on the current policy network each satellite selects the corresponding target execution action from the current policy network that is, determines the beam allocation and hopping decision.
[0093] Further, each satellite executes the corresponding target execution action After that, the system enters a new state , and each satellite obtains a reward according to the corresponding target execution action .
[0094] Among them, the reward function is designed as follows:
[0095] (2)
[0096] Among them, is the throughput reward, is the user demand satisfaction, is the inter-satellite interference penalty, is the delay penalty, is the weight of the throughput reward, is the weight of the user demand satisfaction, is the weight of the inter-satellite interference penalty, is the weight of the delay penalty.
[0097] Further, each satellite records its executed action and the state change as a quadruple , for subsequent local model gradient calculation, thus obtaining a series of quadruples of state-action-reward-next state for forming a decision trajectory , and stores the decision trajectory obtained from each interaction with the environment into the experience pool .
[0098] Among them, the calculation formula of the advantage function is as follows:
[0099] (3)
[0100] Among them, is the advantage function, is the action value function, representing the expected return after executing the action , is the output of the value network, representing the state value function.
[0101] It can be understood that, assuming that satellite A selects and executes a certain action according to the current state in a certain update cycle, and then records the corresponding quadruple and calculates the advantage function. If in the new update cycle, the value of the advantage function decreases, it indicates that the newly updated local model parameters may lead to performance degradation. Therefore, it is necessary to use a clipping mechanism to suppress the update of the satellite local model. The clipping mechanism will limit the amplitude of parameter update to ensure that the local model of the satellite will not degrade in performance due to excessive update.
[0102] Based on this, in the embodiment of the present application, by adding a clipping term to the policy loss function, the change range of the probability ratio is restricted, preventing the satellite local model parameter update amplitude from being too large, thereby avoiding performance degradation. The policy loss function is as follows:
[0103] (4)
[0104] wherein, is the policy loss function, is the new policy to be updated, is the old policy to be updated, is the advantage function; the clip function is used to constrain the ratio of the new and old policies within to achieve restricting the update amplitude of the policy, is a small constant used to control the update amplitude, usually between 0.1 and 0.2.
[0105] Furthermore, the policy network of each satellite local model updates the policy network parameters using the gradient ascent method based on the advantage function. The specific formula is as follows:
[0106] (5)
[0107] wherein, are the updated policy network parameters, is the learning rate, which can be a preset constant, represents the gradient of the loss function with respect to the policy network parameters of the satellite local model taken.
[0108] Finally, the value network of each satellite local model updates the value network parameters by minimizing the mean squared error loss function to better predict the state value. The specific formula is as follows:
[0109] (6)
[0110] wherein, is the value loss function, is the experience pool, is the output of the value network, representing the state value function for estimating the system state value; γ is the discount factor used to control the influence of future estimated rewards, usually taking values between [0,1]; is the decision trajectory the number of time steps in; is the immediate reward at the new time step, obtained from the reward function, i.e., formula (2). are the updated value network parameters, is the learning rate, Denote the value loss function for the value network parameters of the satellite local model to calculate the gradient.
[0111] To facilitate a clearer and more intuitive understanding by those skilled in the art of the process of each satellite performing hop beam strategy training locally in the embodiments of the present application, the following is combined with Figure 2 for detailed description.
[0112] As Figure 2 shown, the process of each satellite performing hop beam strategy training locally includes the following steps:
[0113] S201, Obtain the current system state.
[0114] S202, Select an action according to the current policy network and execute it.
[0115] S203, Update the system state.
[0116] S204, Calculate the reward.
[0117] S205, Record the quadruple (state, action, reward, new state).
[0118] S206, Calculate the advantage function.
[0119] S207, Update the policy network parameters and trigger the clipping mechanism.
[0120] S208, Update the value network parameters.
[0121] Furthermore, each satellite calculates its respective local model gradients, including the gradient of the updated policy network loss function with respect to the policy network parameters and the gradient of the updated value network loss function with respect to the value network parameters , Since the results of each satellite at each time step are different, for the convenience of those skilled in the art to understand, these two gradient results can be denoted as and , respectively, and upload the local model gradients of each satellite to the central node.
[0122] In step S102, perform quality estimation on the local model gradients of each satellite, and asynchronously federate and aggregate multiple gradients that meet the preset quality compliance conditions to obtain new global model gradients.
[0123] Further, in some embodiments, quality estimation is performed on the local model gradients of each satellite, and asynchronous federated aggregation is performed on multiple gradients that meet the preset quality compliance conditions, including: calculating the comprehensive quality index value of the local model gradient of each satellite based on a preset quality index calculation formula; if the comprehensive quality index value is greater than the preset index value, determining that the local model gradient meets the preset quality compliance conditions, and obtaining the number of local model gradients that meet the preset quality compliance conditions; if the number of local model gradients that meet the preset quality compliance conditions reaches a preset quantity, performing asynchronous federated aggregation on multiple gradients that meet the preset quality compliance conditions based on a preset aggregation formula.
[0124] Among them, the preset quality compliance conditions can be quality compliance conditions preset by those skilled in the art, such as the comprehensive quality index value exceeding a quality threshold; the preset quantity can be a quantity preset by those skilled in the art, and no specific limitation is made here.
[0125] Specifically, after the central node receives the local model gradients of a preset number of satellites that meet the preset quality compliance conditions, asynchronous federated aggregation is performed to calculate a new global model. Among them, the quality estimation in the embodiments of the present application is calculated based on the cosine similarity and staleness of the gradients, and the final comprehensive quality index is represented by the sum of the gradient cosine similarity and staleness, which facilitates subsequent allowing the local model gradients whose comprehensive quality index values meet the preset quality compliance conditions to participate in the subsequent weighted asynchronous federated aggregation.
[0126] Among them, the preset quality index calculation formula is:
[0127] ;
[0128] Among them, is the comprehensive quality index value of the i-th satellite node, is the time step of this asynchronous federated aggregation; is the staleness of the i-th satellite node, is a constant controlling the staleness decay rate, is the gradient of the policy network parameters of the i-th satellite node and the gradient of the value network parameters the time elapsed since the calculation; is and the cosine similarity of, is and the cosine similarity of, is the gradient of the global policy network parameters, is the gradient of the global value network parameters, is the Euclidean norm of the gradient of the local policy network parameters, is the Euclidean norm of the gradient of the global policy network parameters, is the Euclidean norm of the gradient of the local value network parameters, is the Euclidean norm of the gradient of the global value network parameters.
[0129] Furthermore, the flag indicating that the local model gradients of each satellite can participate in asynchronous federated aggregation finally can be that the comprehensive quality index value of the i-th satellite node exceeds a preset quality threshold , where the asynchronous federated aggregation method is as follows:
[0130] (8)
[0131] where, is the gradient of the new global policy network parameters, is the gradient of the new global value network parameters, is the weight of asynchronous federated aggregation.
[0132] To facilitate those skilled in the art to more clearly and intuitively understand the process of asynchronous federated aggregation in the embodiments of the present application, the following will be combined with Figure 3 to be described in detail.
[0133] As Figure 3 shown, the process of this asynchronous federated aggregation includes the following steps:
[0134] S301, Receive satellite local model gradient information.
[0135] S302, Perform quality screening on the local model gradients.
[0136] S303, Determine whether K gradients with qualified quality are received. If so, execute S304, otherwise, execute S301.
[0137] S304, Calculate the new global model parameters after aggregation.
[0138] S305, Calculate the new global model gradient.
[0139] In step S103, calculate the similarity between the local model gradient of each satellite and the new global model gradient, and when the similarity is greater than a preset threshold, determine that the new global model is in a preset convergence state.
[0140] Among them, the preset threshold can be a threshold preset by those skilled in the art, and no specific limitation is made here.
[0141] Specifically, the similarity between the gradient of the new global model parameters and the gradient of the local model parameters can be used to evaluate the convergence degree of the model. The specific determination is as follows:
[0142] (10)
[0143] Among them, and are the similarities in formula (7), is and 's cosine similarity, is and 's cosine similarity, is the similarity threshold, that is, the preset threshold, usually taking values between . When the average similarity is higher than the preset threshold, it is determined that the new global model is in the preset convergence state.
[0144] Furthermore, in some embodiments, after calculating the similarity between the local model gradient of each satellite and the new global model gradient, it further includes: if the similarity is less than or equal to the preset threshold, broadcasting the new global model gradient to each satellite; and updating the local model parameters of each satellite according to the new global model gradient.
[0145] Specifically, if the average similarity calculated according to formula (10) is less than or equal to the preset threshold, the central node broadcasts the updated new global model to all satellites, and each satellite receives the updated new global model gradient and and then updates its local local model parameters by gradient.
[0146] (9)
[0147] Among them, is the local policy network parameter updated based on the new global model gradient, is the local value network parameter updated based on the new global model gradient, is the learning rate, which is consistent with formula (5) and is used to control the update step size.
[0148] Furthermore, in some embodiments, after determining that the new global model is in the preset convergence state, it further includes: calculating the inversion success rate of the new global model; verifying whether the new global model meets the preset training termination condition according to the inversion success rate; if the new global model meets the preset training termination condition, terminating the training of the new global model, otherwise, executing the step of broadcasting the new global model gradient to each satellite.
[0149] Among them, the preset training termination condition may be that the safety performance of the new global model no longer improves.
[0150] Specifically, the inversion success rate refers to the success rate of measuring the attacker's inference of the original user data from the machine learning model through the model inversion attack. Simulate the inversion attack to crack the transmitted user information, and evaluate whether the optimization method has reached convergence through the inversion success rate. Among them, the calculation formula of the inversion success rate is as follows:
[0151] (11)
[0152] Among them, is the inversion success rate at time step t, is the set of original data successfully recovered by the simulated attacker through the inversion attack at time step t, is the set of all original user data involved in the system at time step t, is the ratio of the number of successfully inverted data entries to the total number of data entries.
[0153] Furthermore, the inversion success rate is used to determine whether the security performance of the new global model has been improved. The condition for the embodiment of this application to determine that there is no longer an improvement is:
[0154] (12)
[0155] Among them, is the inversion success rate at time step ; is the inversion success rate at time step ; is the current time step, is the set inspection period. It means that if there is no change in the security performance in the past consecutive n periods, it is considered that the security performance of the new global model has reached its peak; is a small constant used to determine whether the change in security performance is significant. Usually, it can be set to a very small value, such as 0.01, and no specific limitation is made here.
[0156] To facilitate those skilled in the art to more clearly and intuitively understand a satellite hopping beam method based on federated learning in the embodiment of this application, the following will be combined with Figure 4 for detailed description.
[0157] As Figure 4 shown, the process of asynchronous federated aggregation includes the following steps:
[0158] S401, start.
[0159] S402, initialize the model parameters.
[0160] S403, satellite local model training.
[0161] S404, Calculate and upload the local gradient.
[0162] S405, Perform asynchronous federated aggregation.
[0163] S406, Determine whether the similarity between the local model parameters of each satellite and the new global model parameters reaches a preset threshold. If so, execute S409; otherwise, execute S407.
[0164] S407, Broadcast the new global gradient.
[0165] S408, All satellites are updated according to the new global gradient.
[0166] S409, Determine whether the inversion success rate has improved. If so, execute S407; otherwise, execute S410.
[0167] S410, Determine that the new global model has converged and terminate the training.
[0168] S411, End.
[0169] Thus, the present invention adopts the method of federated learning, enabling each satellite to train the hopping beam strategy on its own on-board computing unit, transmitting only the gradient calculated from the model parameters at an appropriate time, and using the central computing node to optimize the overall model to avoid the transmission of raw data. At the same time, the mechanism of K-asynchronous federated learning is adopted to prevent the central computing node from processing too much data at once and reduce the complexity of data processing.
[0170] According to a satellite hopping beam method based on federated learning in an embodiment of the present application, the quality of the local model gradient of each satellite is estimated, and asynchronous federated aggregation is performed on multiple gradients that meet the preset quality compliance conditions to obtain a new global model gradient; the similarity between the local model gradient of each satellite and the new global model gradient is calculated, and when the similarity is greater than the preset threshold, it is determined that the new global model is in a preset convergence state. Thus, the problems of high user data security risk and excessive system synchronous communication complexity in the existing hopping beam technology are solved, realizing the reduction of system synchronous communication complexity while ensuring user data security, improving computing efficiency, and ensuring the effective utilization and long-term stable operation of satellite communication system resources.
[0171] Next, a satellite hopping beam device based on federated learning proposed in an embodiment of the present application is described with reference to the accompanying drawings.
[0172] Figure 5 It is a block diagram of a satellite hopping beam device based on federated learning in an embodiment of the present application.
[0173] As Figure 5As shown, the satellite hopping beam device 10 based on federated learning includes: an acquisition module 100, an aggregation module 200, and an optimization module 300.
[0174] Among them, the acquisition module 100 is used to obtain the local model gradients of each satellite; the aggregation module 200 is used to estimate the quality of the local model gradients of each satellite, and asynchronously federate and aggregate multiple gradients that meet the preset quality compliance conditions to obtain new global model gradients; the optimization module 300 is used to calculate the similarity between the local model gradient of each satellite and the new global model gradient, and when the similarity is greater than the preset threshold, determine that the new global model is in the preset convergence state.
[0175] Further, in some embodiments, the acquisition module 100 is used to: obtain the global model parameters of the central node; initialize the global model parameters, and based on the initialized global model parameters, perform hopping beam strategy training locally on each satellite, and obtain the local model gradients of each satellite after the training is completed.
[0176] Further, in some embodiments, the acquisition module 100 is used to: obtain the current system state information of each satellite within a preset time step, where the current system state information includes at least one of user requirements, interference conditions, beam overlap regions, and channel states; based on the current system state information and the current policy network of each satellite, determine the target execution actions corresponding to each satellite, and after each satellite executes the corresponding target execution actions, update the current system state information of each satellite, and calculate the rewards obtained by each satellite; record the quadruple composed of the current system state information, target execution actions, rewards, and updated current system state information, and calculate the advantage function based on the quadruple; according to the value of the advantage function in the new update cycle, determine the update amplitude of the policy network and / or the parameters of the policy network, and update the parameters of the value network based on minimizing the mean square error loss function to obtain the updated parameters of the value network.
[0177] Further, in some embodiments, the aggregation module 200 is used to: calculate the comprehensive quality index value of the local model gradient of each satellite based on a preset quality index calculation formula; if the comprehensive quality index value is greater than the preset index value, determine that the local model gradient meets the preset quality compliance conditions, and obtain the number of local model gradients that meet the preset quality compliance conditions; if the number of local model gradients that meet the preset quality compliance conditions reaches the preset quantity, asynchronously federate and aggregate multiple gradients that meet the preset quality compliance conditions based on a preset aggregation formula; where the preset quality index calculation formula is:
[0178] ;
[0179] Among them, is the comprehensive quality index value of the i-th satellite node, is the time step of this asynchronous federated aggregation; is the staleness of the i-th satellite node, is a constant that controls the decay rate of staleness, is the gradient of the policy network parameters of the i-th satellite node and the gradient of the value network parameters is the time elapsed since the calculation; is and is the cosine similarity of, is and is the cosine similarity of, is the gradient of the global policy network parameters, is the gradient of the global value network parameters, is the Euclidean norm of the gradient of the local policy network parameters, is the Euclidean norm of the gradient of the global policy network parameters, is the Euclidean norm of the gradient of the local value network parameters, is the Euclidean norm of the gradient of the global value network parameters.
[0180] Further, in some embodiments, after calculating the similarity between the local model gradient of each satellite and the new global model gradient, the optimization module 300 is further configured to: if the similarity is less than or equal to a preset threshold, broadcast the new global model gradient to each satellite; perform gradient update on the local model parameters of each satellite according to the new global model gradient.
[0181] Further, in some embodiments, after determining that the new global model is in a preset convergence state, the optimization module 300 is further configured to: calculate the inversion success rate of the new global model; verify whether the new global model meets the preset training termination condition according to the inversion success rate; if the new global model meets the preset training termination condition, terminate the training of the new global model, otherwise, perform the step of broadcasting the new global model gradient to each satellite.
[0182] It should be noted that the foregoing explanation of the embodiments of a satellite hopping beam method based on federated learning also applies to a satellite hopping beam device based on federated learning in this embodiment, and will not be repeated here.
[0183] A satellite hopping beam device based on federated learning according to an embodiment of the present application estimates the quality of the local model gradients of each satellite, and asynchronously federally aggregates multiple gradients that meet the preset quality compliance conditions to obtain new global model gradients; calculates the similarity between the local model gradient of each satellite and the new global model gradient, and determines that the new global model is in a preset convergence state when the similarity is greater than a preset threshold. Thereby, the problems of high user data security risks and excessive system synchronous communication complexity in the existing hopping beam technology are solved, and while ensuring user data security, the system synchronous communication complexity is reduced, the computing efficiency is improved, and the effective utilization and long-term stable operation of satellite communication system resources are ensured.
[0184] Figure 6 The structural schematic diagram of the electronic device provided by the embodiment of the present application. The electronic device may include:
[0185] A memory 601, a processor 602, and a computer program stored on the memory 601 and executable on the processor 602.
[0186] When the processor 602 executes the program, it implements a satellite hopping beam method based on federated learning provided in the above embodiment.
[0187] Further, the electronic device further includes:
[0188] A communication interface 603 for communication between the memory 601 and the processor 602.
[0189] The memory 601 is used to store a computer program executable on the processor 602.
[0190] The memory 601 may include a high-speed RAM memory, and may also include a non-volatile memory, such as at least one disk memory.
[0191] If the memory 601, the processor 602, and the communication interface 603 are implemented independently, the communication interface 603, the memory 601, and the processor 602 may be interconnected through a bus and communicate with each other. The bus may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 6It is represented only by a thick line, but it does not mean that there is only one bus or one type of bus.
[0192] Optionally, in a specific implementation, if the memory 601, the processor 602, and the communication interface 603 are integrated on a single chip, the memory 601, the processor 602, and the communication interface 603 can communicate with each other through an internal interface.
[0193] The processor 602 may be a central processing unit (CPU for short), or an application specific integrated circuit (ASIC for short), or one or more integrated circuits configured to implement the embodiments of the present application.
[0194] The embodiments of the present application also provide a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements a satellite hopping beam method based on federated learning as described above.
[0195] In the description of this specification, the descriptions with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or N embodiments or examples. In addition, without conflict, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0196] In addition, the terms "first" and "second" are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present application, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically defined.
[0197] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
Claims
1. A satellite hopping beam method based on federated learning, characterized in that, Including the following steps: Obtain the local model gradients of each satellite; Perform quality estimation on the local model gradients of each satellite, and asynchronously federate and aggregate multiple gradients that meet the preset quality compliance conditions to obtain new global model gradients; Calculate the similarity between the local model gradients of each satellite and the new global model gradients, and when the similarity is greater than the preset threshold, determine that the new global model is in a preset convergence state; Among them, the performing quality estimation on the local model gradients of each satellite and asynchronously federating and aggregating multiple gradients that meet the preset quality compliance conditions includes: calculating the comprehensive quality index value of the local model gradients of each satellite based on a preset quality index calculation formula; if the comprehensive quality index value is greater than the preset index value, determine that the local model gradient meets the preset quality compliance conditions, and obtain the number of local model gradients that meet the preset quality compliance conditions; if the number of local model gradients that meet the preset quality compliance conditions reaches a preset quantity, asynchronously federate and aggregate multiple gradients that meet the preset quality compliance conditions based on a preset aggregation formula; Among them, the preset quality index calculation formula is: ; wherein, is the comprehensive quality index value of the i-th satellite node; is the time step of this asynchronous federated aggregation; is the staleness of the i-th satellite node; is a constant that controls the decay rate of staleness; is the gradient of the policy network parameters of the i-th satellite node and the gradient of the value network parameters since the calculation; is and 's cosine similarity; is and 's cosine similarity; is the gradient of the global policy network parameters; is the gradient of the global value network parameters; is the Euclidean norm of the gradient of the local policy network parameters; is the Euclidean norm of the gradient of the global policy network parameters; is the Euclidean norm of the gradient of the local value network parameters; is the Euclidean norm of the gradient of the global value network parameters.
2. The satellite hopping beam method based on federated learning according to claim 1, wherein The obtaining the local model gradients of each satellite includes: Obtain the global model parameters of the central node; Initialize the global model parameters, and perform beam hopping strategy training on each satellite locally based on the initialized global model parameters, and obtain the local model gradients of each satellite after the training is completed.
3. The satellite hopping beamforming method based on federated learning according to claim 2, wherein The performing beam hopping strategy training on each satellite locally based on the initialized global model includes: Obtain the current system state information of each satellite within a preset time step, where the current system state information includes at least one of user requirements, interference conditions, beam overlap regions, and channel states; Based on the current system state information and the current policy network of each satellite, determine the target execution actions corresponding to each satellite, update the current system state information of each satellite after each satellite executes the corresponding target execution actions, and calculate the rewards obtained by each satellite; Record the quadruple composed of the current system state information, target execution actions, rewards, and updated current system state information, and calculate the advantage function based on the quadruple; Determine the update amplitude of the policy network and / or the parameters of the policy network according to the value of the advantage function in a new update cycle, and update the parameters of the value network based on minimizing the mean square error loss function to obtain the updated parameters of the value network.
4. The satellite hopping beam method based on federated learning according to claim 1, wherein After calculating the similarity between the local model gradients of each satellite and the new global model gradients, it further includes: If the similarity is less than or equal to the preset threshold, broadcast the new global model gradients to each satellite; Update the local model parameters of each satellite according to the new global model gradients.
5. The satellite hopping beam method based on federated learning according to claim 1, wherein After determining that the new global model is in a preset convergence state, it further includes: Calculate the inversion success rate of the new global model; Verify whether the new global model meets the preset training termination condition according to the inversion success rate; If the new global model meets the preset training termination condition, terminate the training of the new global model; otherwise, execute the step of broadcasting the gradient of the new global model to each satellite.
6. A satellite hopping beam device based on federated learning, characterized in that, Including: An acquisition module, configured to acquire the local model gradients of each satellite; An aggregation module, configured to perform quality estimation on the local model gradients of each satellite, and perform asynchronous federated aggregation on multiple gradients that meet the preset quality compliance conditions to obtain a new global model gradient; An optimization module, configured to calculate the similarity between the local model gradient of each satellite and the new global model gradient, and determine that the new global model is in a preset convergence state when the similarity is greater than a preset threshold; Wherein, the aggregation module is configured to: calculate the comprehensive quality index value of the local model gradient of each satellite based on a preset quality index calculation formula; if the comprehensive quality index value is greater than a preset index value, determine that the local model gradient meets the preset quality compliance condition, and obtain the number of local model gradients that meet the preset quality compliance condition; if the number of local model gradients that meet the preset quality compliance condition reaches a preset quantity, perform asynchronous federated aggregation on multiple gradients that meet the preset quality compliance condition based on a preset aggregation formula; Wherein, the preset quality index calculation formula is: ; wherein, is the comprehensive quality index value of the i-th satellite node, is the time step of this asynchronous federated aggregation; is the obsolescence of the i-th satellite node, is a constant controlling the decay rate of obsolescence, is the gradient of the policy network parameters of the i-th satellite node and the gradient of the value network parameters is the time elapsed since the calculation; is and 's cosine similarity, is and 's cosine similarity, is the gradient of the global policy network parameters, is the gradient of the global value network parameters, is the Euclidean norm of the gradient of the local policy network parameters, is the Euclidean norm of the gradient of the global policy network parameters, is the Euclidean norm of the gradient of the local value network parameters, is the Euclidean norm of the gradient of the global value network parameters.
7. The satellite hopping beam device based on federated learning according to claim 6, characterized in that, The acquisition module is configured to: Acquire the global model parameters of the central node; Initialize the global model parameters, and perform beam hopping strategy training locally on each satellite based on the initialized global model parameters, and obtain the local model gradients of each satellite after the training is completed.
8. An electronic device, characterized in that, Including: A memory, a processor, and a computer program stored on the memory and executable on the processor, and the processor executes the program to implement a satellite beam hopping method based on federated learning according to any one of claims 1-5.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by the processor to be used to implement a satellite beam hopping method based on federated learning according to any one of claims 1-5.