Federal learning energy consumption optimization method based on reinforcement learning in low earth orbit satellite network

By building a dynamic architecture in a low-orbit satellite network and using reinforcement learning optimization parameter distribution and uploading processes, the problem of high federal learning energy consumption in the low-orbit satellite network is solved, and balanced optimization of energy consumption and improvement of energy efficiency is achieved.

CN120295771APending Publication Date: 2025-07-11XI AN JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510348680.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

Existing research has failed to effectively optimize the overall energy consumption and energy consumption balance of the federal learning process in low-orbit satellite networks, ignoring the energy consumption optimization opportunities of intermediate nodes, resulting in high overall energy consumption.

Method used

Build a dynamic architecture of a low-orbit satellite network, use reinforcement learning methods to optimize the parameter distribution and upload process, train the policy network through the soft actor-critic algorithm SAC, dynamically select server nodes and transmission paths, combine multi-path routing algorithm to reduce redundant energy consumption, and balance real-time energy consumption with network load.

Benefits of technology

It significantly reduces the overall energy consumption during the federal learning process, achieves balanced optimization of energy consumption, and improves energy efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120295771A_ABST
    Figure CN120295771A_ABST
Patent Text Reader

Abstract

The invention provides a federated learning energy consumption optimization method based on reinforcement learning in a low earth orbit satellite network, and the method comprises the steps: constructing a dynamic network architecture composed of low earth orbit satellites, and initializing a federated learning framework; under a federated learning framework, establishing a two-stage energy consumption model of federated learning parameter distribution PDP and uploading PUP; energy consumption optimization in the two-stage energy consumption model is converted into a Markov decision process; before each round of federal learning starts, dynamically selecting a new server node and a transmission path; according to a path selection result, executing intermediate parameter aggregation at a first cross node, and updating multi-hop routing energy consumption hierarchical division; and iteratively optimizing global model parameters and a routing strategy, and balancing instant energy consumption reduction and network load balance until a specified federal learning round is completed. According to the method, a first path aggregation point in the federal learning parameter uploading process is used as an intermediate aggregation node, aggregation is carried out in advance, and finally only one subsequent path is reserved. According to the method, the total energy consumption in the federal learning process is remarkably reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of federated learning energy consumption optimization, and in particular relates to a federated learning energy consumption optimization method based on reinforcement learning in a low-orbit satellite network. Background Art

[0002] As a core component of the sixth-generation wireless communication system, the low-orbit satellite network can provide low-latency communication services and has significant cost-effectiveness due to its low orbital altitude, easy mass production and high-frequency transmission. The rapid development of artificial intelligence technology has significantly improved the intelligence level of low-orbit satellite networks. Complex machine learning models have been deployed on onboard equipment to make up for the computing power and energy reserve limitations of ground equipment in specific tasks. In the distributed machine learning framework, federated learning technology has attracted widespread attention due to its unique localized training mechanism. This technology avoids central server data collection by training local data sets, and is particularly suitable for distributed low-orbit satellite network architecture. However, the existing energy supply system is difficult to ensure the continuous high-performance operation of low-orbit satellites, making energy efficiency management a key technical bottleneck that needs to be solved urgently: on the one hand, satellites as federated learning clients need to consume a lot of computing resources for model training; on the other hand, network intermediate nodes generate additional energy consumption during parameter routing and forwarding. Therefore, in the implementation of federated learning, it is necessary to systematically evaluate the energy consumption of all participating nodes.

[0003] Regarding the energy consumption optimization problem of low-orbit satellites, existing research mainly reduces energy consumption through data allocation and power control, thereby optimizing the energy consumption of the federated learning process. At the same time, some scholars have proposed to reduce the energy consumption of low-orbit satellites by optimizing the aggregation frequency or increasing the compression ratio to improve efficiency. However, existing research usually does not consider the optimization of overall energy consumption and energy consumption balance in the federated learning process, and ignores the energy consumption optimization opportunities brought by the early aggregation of intermediate nodes in the low-orbit satellite network transmission, resulting in the problem of high overall energy consumption in the federated learning process. Summary of the invention

[0004] The purpose of the present invention is to provide a federated learning energy consumption optimization method based on reinforcement learning in a low-orbit satellite network to solve the above problems.

[0005] To achieve the above object, the present invention adopts the following technical solutions: In a first aspect, the present invention provides a method for optimizing energy consumption of federated learning based on reinforcement learning in a low-orbit satellite network, comprising: Build a dynamic network architecture consisting of low-orbit satellites, select satellites as federated learning clients, and the remaining satellites as potential server nodes to initialize the federated learning framework; Under the federated learning framework, a two-stage energy consumption model of federated learning parameter distribution PDP and upload PUP is established; Transform the energy consumption optimization in the two - stage energy consumption model into a Markov decision process, adopt the Soft Actor - Critic (SAC) algorithm to train the policy network, configure a dual - objective critic network and an experience replay mechanism; Before each round of federated learning starts, generate actions through the trained reinforcement learning agent based on the current server state, dynamically select new server nodes and transmission paths; according to the path selection result, perform intermediate parameter aggregation at the first cross - node, and update the multi - hop routing energy consumption level division; iteratively optimize the global model parameters and routing strategies to balance immediate energy consumption reduction and network load balancing until the specified number of federated learning rounds is completed.

[0006] Furthermore, for constructing a dynamic network architecture composed of low - earth orbit satellites, select satellites as federated learning clients and the remaining satellites as potential server nodes, and initialize the federated learning framework, including: Select federated learning client nodes and initial server nodes in the low - earth orbit satellite network, initialize the global model parameters and broadcast them to the clients through a multi - path routing algorithm; calculate the dynamic distance between satellites based on celestial mechanics laws and construct an inter - satellite communication topology; Federated learning clients train the model according to the local dataset, update the parameters using the gradient descent algorithm, and aggregate the weights through the FedAvg rule, where the weight distribution is proportional to the amount of client data.

[0007] Furthermore, when aggregating weights using the FedAvg rule, the global model parameters are expressed as follows:

[0008] where, is the number of federated learning rounds, extracts the number of samples contained in the data; Gradient descent is used as the training algorithm for federated learning clients, and the machine learning model parameters are expressed as:

[0009] where, represents the loss function; represents the learning rate in the training, represents the extraction of the gradient sign.

[0010] Furthermore, under the federated learning framework, establish a two - stage energy consumption model for federated learning parameter distribution (PDP) and upload (PUP), including: Calculate the local training energy consumption, which is related to the processor frequency, the number of training rounds, and the sample processing rate; Model the signal attenuation of the inter - satellite link based on the Friis formula and derive the transmission energy consumption under adaptive power control; Design the first cross-node intermediate aggregation mechanism to reduce the energy consumption of redundant transmission paths through a hierarchical grouping strategy.

[0011] Furthermore, the inter-satellite link signal attenuation modeling based on the Friis formula to derive the transmission energy consumption under adaptive power control includes: Low Earth Orbit (LEO) satellites communicate with each other through inter-satellite links under the constraint of line-of-sight visibility. When two satellites are visible to each other, there is an inter-satellite link, denoted as , where is the straight-line distance between the two satellites; represents the signal power at the receiver, which is expressed by the Friis transmission formula as:

[0012] where, is the transmission power at the transmitter, and represent the effective aperture areas of the transmitter and receiver respectively, represents the signal wavelength; The transmission power is expressed as:

[0013] where, and represent the standard transmission power and distance respectively.

[0014] Furthermore, the design of the first cross-node intermediate aggregation mechanism to reduce the energy consumption of redundant transmission paths through a hierarchical grouping strategy includes: The energy consumption in the PDP is determined by the deterministic routing path as follows:

[0015] where represents the set of single-hop paths of all federated learning clients in the PDP; at the first routing cross-node, path selection eliminates the energy consumption along redundant paths. For the final routing topology, the total energy consumption is expressed as:

[0016] where, and ; and represent the levels of the transmission source and destination respectively.

[0017] Furthermore, the conversion of the energy consumption optimization in the two-stage energy consumption model into a Markov decision process, adopting the Soft Actor-Critic (SAC) algorithm training strategy network, configuring a dual-objective critic network and an experience replay mechanism, includes: State space: One-hot encoding is used to represent the server location in the previous round of federated learning; Hybrid action space: It includes discrete actions: server selection in the current round and continuous actions: PDP / PUP path selection vectors of each client; Reward function: Integrate the total energy consumption and the energy consumption balance index, and introduce a discount factor to achieve the long-term optimization goal; The Soft Actor-Critic algorithm is used to minimize the total energy consumption of FL rounds under given constraints. An actor network and two critic networks are used. The actor network determines the strategy, and the critic networks evaluate its value. Both critic networks contain target networks.

[0018] In a second aspect, the present invention provides a federated learning energy consumption optimization system based on reinforcement learning in a low-earth orbit satellite network, including: An initialization module for constructing a dynamic network architecture composed of low-earth orbit satellites, selecting satellites as federated learning clients, and the remaining satellites as potential server nodes, and initializing the federated learning framework; An energy consumption model construction module for establishing a two-stage energy consumption model of parameter distribution PDP and upload PUP in federated learning under the federated learning framework; An energy consumption optimization module for converting the energy consumption optimization in the two-stage energy consumption model into a Markov decision process, training a policy network using the Soft Actor-Critic algorithm SAC, configuring a dual-objective critic network and an experience replay mechanism; An optimization output module for generating actions through a trained reinforcement learning agent based on the current server state before the start of each round of federated learning, dynamically selecting new server nodes and transmission paths; according to the path selection results, performing intermediate parameter aggregation at the first cross node, and updating the multi-hop routing energy consumption level division; iteratively optimizing the global model parameters and routing strategies to balance immediate energy consumption reduction and network load balancing until the specified number of federated learning rounds is completed.

[0019] In a third aspect, the present invention provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the federated learning energy consumption optimization method in the low-earth orbit satellite network are implemented.

[0020] In a fourth aspect, the present invention provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps of the federated learning energy consumption optimization method in the low-earth orbit satellite network are implemented.

[0021] Compared with the prior art, the present invention has the following technical effects: The present invention takes into account the multi-hop routing characteristics of a low-Earth orbit satellite network, uses the first path aggregation point in the process of uploading federated learning parameters as an intermediate aggregation node, and performs aggregation in advance, finally retaining only one subsequent path. Compared with other strategies, this method significantly reduces the overall energy consumption in the federated learning process.

[0022] The present invention combines the multi-round characteristics of federated learning to construct a Markov process model for energy consumption, including observation, action, state transition, and reward space. By constructing a hybrid action space to improve the training efficiency, and using the deep reinforcement learning method to minimize the overall energy consumption and energy consumption balance in the multi-round process. Compared with other strategies, this method obtains better results from an overall perspective and provides effective support for the actual deployment of the reinforcement learning agent. Brief Description of the Drawings

[0023] Figure 1 It is a schematic diagram of a low-Earth orbit satellite network.

[0024] Figure 2 It is a schematic diagram of the transmission of multiple routing paths.

[0025] Figure 3 It is a flow chart of energy consumption optimization aggregation under multi-path routing.

[0026] Figure 4 It is a topological structure diagram of a Markov decision process.

[0027] Figure 5 It is a schematic diagram of the detailed information of the agent operation.

[0028] Figure 6 It is a schematic diagram of the energy consumption of relevant satellites involved in federated learning in a single round of federated learning process.

[0029] Figure 7 It is the action selection sequence of the deep reinforcement learning agent in the first few rounds of the federated learning process when there are 5 optional paths.

[0030] Figure 8 For when , the relationship between the cumulative energy consumption, cumulative negative reward, and average energy consumption balance value of the proposed method and the delay.

[0031] Figure 9 For when , the relationship between the accuracy of image recognition and the cumulative negative reward.

[0032] Figure 10 It is a flow chart of the present invention. Detailed Embodiments

[0033] The present invention is further described below with reference to the accompanying drawings: Example 1, please refer to Figure 10, the present invention provides an energy consumption optimization method for federated learning based on reinforcement learning in a low-earth orbit satellite network, including: Construct a dynamic network architecture composed of low-earth orbit satellites, select satellites as federated learning clients, and the remaining satellites as potential server nodes, and initialize the federated learning framework; Under the federated learning framework, establish a two-stage energy consumption model for federated learning parameter distribution PDP and upload PUP; Convert the energy consumption optimization in the two-stage energy consumption model into a Markov decision process, use the Soft Actor-Critic algorithm SAC to train the policy network, and configure a dual-objective critic network and an experience replay mechanism; Before each round of federated learning starts, generate actions through the trained reinforcement learning agent based on the current server state, dynamically select new server nodes and transmission paths; according to the path selection result, perform intermediate parameter aggregation at the first cross-node, and update the multi-hop routing energy consumption level division; iteratively optimize the global model parameters and routing strategies to balance immediate energy consumption reduction and network load balancing until the specified number of federated learning rounds is completed.

[0034] Example 2, please refer to Figures 1 to 9 , the present invention provides an energy consumption optimization method for federated learning based on reinforcement learning in a low-earth orbit satellite network, specifically including: As Figure 1 shown, the low-earth orbit satellite network consists of satellites, and these satellites can communicate with maritime, ground, and air devices in real time through a specific routing algorithm based on ground-satellite links and inter-satellite links. According to the laws of celestial mechanics, the velocity of a satellite is given by the formula (unit: m / s), where is the geocentric gravitational constant, is the average radius of the earth.

[0035] To ensure data privacy, the federated learning framework deploys different satellites as federated learning clients and a central server. The federated learning clients are denoted as , and each client has its own dataset, collectively referred to as . According to specific applications, the federated learning clients are equipped with appropriate machine learning models. According to the federated learning mechanism, the server initially broadcasts the global model parameters to all federated learning clients. After receiving the global model, the federated learning clients train local models on their respective datasets, and then send the updated model parameters

[0036] The aggregation algorithm adopted on the server is crucial for federated learning because it directly affects the loss function and overall performance of the federated learning system. The present invention uses the widely applied and recognized FedAvg algorithm, where the aggregation weights assigned to each federated learning client are proportional to the number of samples in its dataset, and the global model parameters are expressed as follows:

[0037] where, is the number of federated learning rounds, extracts the number of samples contained in the data.

[0038] With the explosive growth of the number of devices in recent years, the amount of data has also increased significantly. This provides opportunities for machine learning to solve more complex tasks. As a data-driven technology, machine learning analyzes data through training algorithms applied to datasets to reveal the inherent patterns and laws in classification, prediction, and decision-making. Generally, a dataset is divided into three parts: a training set, a validation set, and a test set. The training set is used for model training, the validation set is used for hyperparameter tuning during the training process, and the final performance of the model is evaluated through the test set. Considering that the goal is to minimize the empirical loss function, gradient descent is used as the training algorithm for federated learning clients. The machine learning model parameters are expressed as:

[0039] where, represents the loss function. represents the learning rate during the training, represents the extraction of the gradient sign.

[0040] For low-earth orbit satellites, limited on-board resources and power constraints exacerbate the energy demand. Fine management of energy consumption has become the key to ensuring sustainable and efficient operation. In this case, the energy consumption of machine learning model training can be expressed as:

[0041] where, is the number of epochs for local model training, represents the energy consumption coefficient. and respectively represent the number of samples processed per unit time and the processor frequency.

[0042] Due to the high maturity and reliability of microwave communication, it remains the most widely used communication method. Low-earth orbit satellites communicate with each other through inter-satellite links under the constraint of line-of-sight visibility. When two satellites are visible to each other, there is an inter-satellite link, expressed as where is the straight-line distance between the two satellites.

[0043] Represents the signal power at the receiver. Since there are no obvious obstacles between two satellites in space, the signal only experiences free space attenuation and does not fade. Therefore, It can be expressed by the Friis transmission formula as:

[0044] Where, is the transmission power at the transmitter, and represent the effective aperture areas of the transmitter and the receiver respectively, represents the signal wavelength. As the distance increases, will decrease and the transmission efficiency will ultimately decrease. Assuming adaptive power control is implemented at the transmitter to maintain a constant transmission rate during transmission, it can be achieved by relying on channel state information estimation through pilot signals or feedback links. Therefore, the transmission power can be expressed as:

[0045] Where, and represent the standard transmission power and distance respectively.

[0046] The energy consumption during the transmission process is mainly at the transmitter and depends on its transmission power and duration, which can be expressed as:

[0047] Where is the transmission duration.

[0048] The multipath routing algorithm is a key technology to improve the data transmission efficiency and network reliability of low Earth orbit satellite networks. In practice, due to limited resources, a single path may not be able to meet the required performance. By providing multiple available paths, the stability and flexibility of data transmission are improved. In addition, multi-hop networks, especially low Earth orbit satellite networks, promote more efficient communication between federated learning clients and servers because the server is both the routing source in the parameter distribution phase and the destination in the parameter upload phase. Considering the aggregation mechanism, the present invention considers an innovative aggregation method under the multipath routing algorithm to reduce energy consumption. Among them, several first cross-nodes in multiple routing paths perform intermediate aggregation to reduce the subsequent transmission energy consumption, as Figure 2 shown.

[0049] The first cross-nodes in the overlapping paths in PUP are classified into groups according to their hop counts from the server. Due to the multipath routing algorithm, this value may vary, and the maximum number of layers is taken here. Then, these sets are assigned levels, where , sorted by hop count. Specifically, the server is considered to belong to layer . The first cross-node at the layer is denoted as , where . At this time, the federated learning clients are considered to belong to layer, where Figure 3 . Although they may also act as the first cross-nodes of other layers. The specific process of the proposed method is summarized in the algorithm and described as

[0050] Since the data transmission sources are all federated learning servers, the energy consumption in PDP is only determined by the deterministic routing path, as follows:

[0051] where represents the set of single-hop paths of all federated learning clients in PDP. The redundant energy consumption of overlapping paths cannot be directly reduced like broadcast to multicast, because although there are overlapping paths in PUP, all FL clients arrive at the same node at different times. The path selection at the first routing cross-node only eliminates the energy consumption along the redundant paths. For the final routing topology, the total energy consumption can be expressed as:

[0052] where, and . and represent the levels of the transmission source and the destination respectively. Since data can only be transmitted from a higher level to a lower level (client to server), corresponds to any level greater than .

[0053] In the federated learning framework with a multi-path routing algorithm, the path selection in PDP and PUP can be independent. The various transmission paths in the two phases provide the potential for further improving energy efficiency. However, some challenges arise in the process of achieving this goal. As the number of FL clients increases, the number of available routing paths per FL round grows exponentially. Considering multiple rounds of FL in practice, the computational complexity becomes huge and unacceptable. In addition, the selection of the server not only affects the current routing topology in the FL round but also the topology in subsequent rounds, which requires a comprehensive global optimization method for multiple decisions. Moreover, due to the position change of low-earth orbit satellites, the physical topology may change drastically, increasing the uncertainty of decision-making.

[0054] Due to the rapid development of artificial intelligence in recent years, different from supervised learning or unsupervised learning, reinforcement learning enables an agent to make decisions by following a policy through interaction with the environment, which maximizes the long-term cumulative reward of events. As a branch of reinforcement learning, deep reinforcement learning has been widely applied in research to solve highly complex sequential decision-making problems. Deep reinforcement learning uses deep neural networks as function approximators to handle non-linear, non-stationary, and high-dimensional tasks. Specifically, the reinforcement learning agent selects and executes an action based on the current observation. Then, the agent receives a reward, and the state of the environment is transformed into a new state for the next action. After being trained by the learning algorithm, the agent achieves the optimal long-term reward through sequential decision-making. Therefore, reinforcement learning is very suitable for federated learning under low-earth orbit satellite networks, which can select the best routing path from alternative options and choose an aggregation server to achieve the best long-term results. Considering the importance of long-term trajectories, the goal of deep reinforcement learning is to maximize the discounted reward, expressed as:

[0055] where represents the discount factor, is the reward at the -th episode of the trajectory. In particular, represents the current reward. As the basis of reinforcement learning, the Markov decision process includes observation, action, transition probability, and reward space, represented by respectively. Therefore, the optimization problem is to minimize the overall energy consumption and energy consumption balance of all federated learning rounds, expressed as:

[0056] where represents the normalized average energy consumption of the -th round of federated learning. represents the degree of energy consumption balance in this round of federated learning, where the numerator is the standard deviation of the total energy consumption and the denominator is the mean. and represent their weights respectively. and represent the path selection at the parameter distribution and upload nodes respectively. represents the server of the -th round of federated learning.

[0057] As Figure 4 shown, in the deep reinforcement learning environment, an episode is represented as a round of federated learning, and the Markov decision process can be represented as follows: Although the topology is a factor affecting energy consumption, its information is often too large to be shared with relevant LEO satellites in a timely manner in a dynamic LEO satellite network. However, the server can be identified at negligible cost. Therefore, the observation space is the server of the previous round of federated learning , which is responsible for the final aggregation of the previous round. Specifically, the observation space can be expressed as .

[0058] The action space includes the routing path selection of each FL client in the PDP and PUP processes, that is and . Each group contains a total of elements representing federated learning clients, with alternative routing paths. In addition, the server of the current federated learning round is also part of the action space, which affects the actual routing topology and performs the final aggregation in the current round. Therefore, the action space can be expressed as .

[0059] Obviously, according to the federated learning process, the previous server in the observation space is actually the server of the previous round in the action space. The transition probability space can be expressed as . Specifically, it is 1 when the action contains , and 0 otherwise.

[0060] The goal of the minimization problem is to minimize the total energy consumption. Therefore, at this time , which is interpreted as the cumulative reward. The immediate reward is obtained after each round of the federated learning process and is equal to the negative total energy consumption in the current FL round, expressed as .

[0061] Obviously, the observation space is discrete. However, the significant mutual influence of different servers on the routing topology, subsequent actions, and rewards weakens the convergence of the deep learning network in the deep reinforcement learning agent. Therefore, the observation space scalar is one-hot encoded into a vector containing only one element as 1 and the other elements as 0. The position index with a value of 1 corresponds to the selected server. Compared with directly using integers, it clearly reflects the independence of discrete values and helps with deep reinforcement learning policy learning and value function optimization.

[0062] Although the action space is also discrete, its scale is too large, growing exponentially with the number of federated learning clients and posing challenges to high-performance hardware. Therefore, the present invention adopts a hybrid action space composed of a discrete space and a continuous space. The discrete space includes , representing the server in the current round. It is also one-hot encoded. On the other hand, the continuous space includes , represented by a vector containing elements, where each element ranges from 0 to K. The value is rounded to determine the actual path selection. In particular, when the value of the element is exactly 0, the first path is still selected. This hybrid design avoids the curse of dimensionality and reduces the complexity of the action space.

[0063] Based on the above, the Soft Actor-Critic algorithm is adopted to minimize the total energy consumption of FL rounds under given constraints. It supports a hybrid action space, minimizing the cumulative reward while maximizing the policy entropy. This method effectively balances exploration and exploitation in a complex environment. In the agent, an actor network and two critic networks are used. The actor network determines the policy, while the critic networks evaluate its value, thus improving stability. Both critic networks contain target networks to mitigate drastic value changes and prevent overestimation. Figure 5 Details of the agent's operation are shown. After training with multi-episode experience replay, the agent can perform the best action according to different observations without further exploration, thus obtaining the best cumulative reward.

[0064] The low Earth orbit satellite constellation considered in the simulation consists of 40 satellites at an altitude of approximately 800 km. The K-shortest routing algorithm is applied for inter-satellite communication within the constellation, providing a total of K alternative paths for each satellite. Considering the dynamic topology of the low Earth orbit satellite network, it is assumed that the satellite positions remain stationary within a fixed time period to obtain a snapshot topology. This time period is set to 30 seconds. Each FL client processes a dataset sample at a processing frequency of Hz with a period of . The standard transmit power of the satellite is dBm, and the reference distance is km. For high-reliability inter-satellite link communication, the transmitter and receiver antenna gains are set to 30 and 20 dBi respectively. The center frequency of the wireless signal is within the Ka band, specifically 28 GHz, and the signal bandwidth is 100 MHz. The noise power spectral density at the receiver is -97.5 dBm / Hz. The energy consumption coefficient is .

[0065] The specific machine learning task of federated learning uses the well-known MNIST dataset for image recognition. The MNIST dataset consists of a total of 10,000 grayscale handwritten digit images from 0 - 9, with an equal number of samples for each digit. The size of each image is 28x28. During the federated learning process, the dataset is divided into 10 clients, and each client has 1,000 samples corresponding to a specific handwritten digit. For local model training, each client further splits its dataset into a training set and a test set, allocating 30% for training and the remainder for testing. In addition to the input and output layers, the local model for image recognition also includes convolutional layers, max pooling layers, and a rectified linear unit (ReLU) activation function. The learning algorithm uses mini-batch stochastic gradient descent with momentum. For global model evaluation, the test sets of all clients are combined into the test set of the server. The FL process is carried out for 300 rounds.

[0066] To minimize the cumulative reward, the hyperparameters of the deep reinforcement learning agent are crucial for achieving optimal performance. Each critic network consists of two input layers: one for discrete actions and the other for continuous actions. These are connected into a single layer, followed by two fully connected layers, each with a fully connected layer and a ReLU activation function. Additionally, the actor network outputs the probability distribution of discrete actions as well as the mean and standard deviation of continuous actions through two layers similar to the critic network. The Adam algorithm is used for training the critic and actor networks. The learning rate is set to 0.01, and the offset is 。According to the requirements of the optimization objective, the discount factor is set to 1. In addition, the entropy weight is also trained using the Adam algorithm. During target updates, the weight update rate of the critic network parameters is set to 0.001. The batch size of experience replay is equal to 64, and the buffer size is set to 10,000.

[0067] Figure 6 Shows the satellite energy consumption related to federated learning during a single round of the federated learning process. Each axis on the radar chart represents the energy consumption of a satellite, with the outermost circle corresponding to the highest energy consumption value of 10 and the center representing the lowest value of 0. Among them, red represents the federated learning server satellite, while blue represents the federated learning client satellite. The area enclosed by the satellite energy consumption reflects the total energy consumption of this round of the federated learning process. Obviously, compared with the traditional federated learning method, the energy consumption of a total of eight satellites has decreased, while the energy consumption of other satellites remains unchanged, and the total energy consumption has decreased from 64.4 to 54.4, highlighting the improvement in energy efficiency of the proposed scheme.

[0068] Figure 7The figure shows the action selection sequence of the deep reinforcement learning agent in the first few rounds of the federated learning process when there are 5 optional paths. The solid and dashed lines respectively represent the selected routing paths during the parameter distribution and parameter upload processes. The different colors in the background of the figure correspond to the optional servers indicated in the color bar. It can be seen that except for the special case in the first round where the result is affected by the initial value, the transmission path selection shows periodic changes, alternating every two rounds. For each federated learning client, the path selection is also different. In addition, since the source (the server in the previous round) and the destination (the server in this round) are different in each round of the federated learning process, for the same federated learning client, there are also differences in the path selection during the parameter distribution and parameter upload processes.

[0069] Figure 8 Shows when the relationship between the cumulative energy consumption, cumulative negative reward, and average energy consumption balance value of the proposed method and the delay. Among them, the cumulative energy consumption and cumulative negative reward are plotted on the left Y-axis, while the average energy consumption balance value corresponds to the right. Due to the different action sequences selected in the federated learning rounds, different values correspond to different curves. Specifically, due to the periodicity of the policy of the deep learning agent, which causes the agent to return the same server selection result, the cumulative energy consumption and negative reward almost increase linearly. In addition, the average energy consumption balance value tends to gradually stabilize after initial fluctuations, which is because of the influence of the initial server value. At the same time, the delay varies with the different values of, and this change is reflected in the different lengths of the curve endpoints on the X-axis, which is mainly due to the different transmission paths executed, resulting in different routing topologies, thus affecting the transmission delay.

[0070] Figure 9 Shows when the relationship between the accuracy of image recognition and the cumulative negative reward. The random scheme performs random path selection and server selection in each round of the federated learning process. In the fixed scheme, the server and routing path remain unchanged throughout the federated learning process. The results show that the final recognition accuracy of different strategies can all approach 0.98, but their cumulative negative rewards are different. Specifically, at convergence, the cumulative rewards corresponding to different strategies are 329.7, 386.6, and 458.8 respectively. In addition, their delays are 123.7 s, 112.7 s, and 114.1 s respectively. Although the delays of the two comparison schemes are lower, in terms of the cumulative negative reward, the proposed method achieves a better balance between energy consumption and energy consumption balance, and the overall performance is better than other methods, further demonstrating the advantage of this method in energy efficiency.

Claims

1. An energy consumption optimization method for federated learning based on reinforcement learning in a low-earth orbit satellite network, characterized in that Including: Construct a dynamic network architecture composed of low-earth orbit satellites, select satellites as federated learning clients, and the remaining satellites as potential server nodes, and initialize the federated learning framework; Under the federated learning framework, establish a two-stage energy consumption model for federated learning parameter distribution PDP and upload PUP; Convert the energy consumption optimization in the two-stage energy consumption model into a Markov decision process, use the Soft Actor-Critic algorithm SAC to train the policy network, and configure a dual-objective critic network and an experience replay mechanism; Before each round of federated learning starts, generate actions through the trained reinforcement learning agent based on the current server state, and dynamically select new server nodes and transmission paths; according to the path selection result, perform intermediate parameter aggregation at the first cross-node, and update the multi-hop routing energy consumption level division; iteratively optimize the global model parameters and routing strategies to balance immediate energy consumption reduction and network load balancing until the specified number of federated learning rounds is completed.

2. The method for optimizing the energy consumption of federated learning based on reinforcement learning in a low-earth orbit satellite network according to claim 1, wherein The construction of a dynamic network architecture composed of low-earth orbit satellites, selecting satellites as federated learning clients, and the remaining satellites as potential server nodes, and initializing the federated learning framework includes: Select federated learning client nodes and initial server nodes in the low-earth orbit satellite network, initialize the global model parameters and broadcast them to the clients through the multi-path routing algorithm; calculate the dynamic distance between satellites based on celestial mechanics laws, and construct the inter-satellite communication topology; The federated learning client trains the model according to the local dataset, updates the parameters using the gradient descent algorithm, and aggregates the weights through the FedAvg rule. The weight allocation is proportional to the client data volume.

3. The energy consumption optimization method for federated learning based on reinforcement learning in a low-earth orbit satellite network according to claim 2, wherein When aggregating weights by the FedAvg rule, the global model parameters are expressed as follows: Among them, is the number of rounds of federated learning, extracts the number of samples contained in the data; Gradient descent is used as the training algorithm for the federated learning client, and the machine learning model parameters are expressed as: Among them, represents the loss function; represents the learning rate in the training, represents extracting the gradient sign.

4. The method for optimizing the energy consumption of federated learning based on reinforcement learning in a low-earth orbit satellite network according to claim 1, wherein Under the federated learning framework, the establishment of a two-stage energy consumption model for federated learning parameter distribution PDP and upload PUP includes: Calculation of local training energy consumption, related to processor frequency, number of training rounds, and sample processing rate; Modeling of inter-satellite link signal attenuation based on Friis formula, and derivation of transmission energy consumption under adaptive power control; Design the intermediate aggregation mechanism at the first cross-node, and reduce the energy consumption of redundant transmission paths through the hierarchical grouping strategy.

5. The method for optimizing the energy consumption of federated learning based on reinforcement learning in a low-earth orbit satellite network according to claim 4, characterized in that, The modeling of inter-satellite link signal attenuation based on Friis formula, and the derivation of transmission energy consumption under adaptive power control includes: Low Earth Orbit (LEO) satellites communicate with each other via inter-satellite links under the constraint of line-of-sight visibility. When two satellites are visible to each other, there is an inter-satellite link, denoted as , where is the straight-line distance between the two satellites; represents the signal power at the receiver, which is expressed by the Friis transmission formula as: wherein, is the transmission power at the transmitter, and respectively represent the effective aperture areas of the transmitter and the receiver, represents the signal wavelength; The transmission power is expressed as: Among them, and respectively represent the standard transmission power and distance.

6. The method for optimizing the energy consumption of federated learning based on reinforcement learning in a low-earth orbit satellite network according to claim 4, wherein The design of the intermediate aggregation mechanism at the first cross-node, and the reduction of the energy consumption of redundant transmission paths through the hierarchical grouping strategy includes: The energy consumption in PDP is determined by the deterministic routing path, as follows: Among them represents the set of single-hop paths of all federated learning clients in the PDP; the path selection at the first routing crossover node eliminates the energy consumption along redundant paths. For the final routing topology, the total energy consumption is expressed as: Among them, and ; and respectively represent the levels of the transmission source and the destination.

7. The method for optimizing energy consumption of federated learning based on reinforcement learning in a low-earth orbit satellite network according to claim 1, characterized in that, Converting the energy consumption optimization in the two-stage energy consumption model into a Markov decision process, using the Soft Actor-Critic algorithm SAC to train the policy network, and configuring a dual-objective critic network and an experience replay mechanism includes: State space: Use one-hot encoding to represent the server location in the previous round of federated learning; Hybrid action space: Includes discrete actions: server selection in the current round and continuous actions: PDP / PUP path selection vectors for each client; Reward function: Integrate the total energy consumption and the energy consumption balance index, and introduce a discount factor to achieve long-term optimization goals; The soft actor-critic algorithm is adopted to minimize the total energy consumption of FL rounds under given constraints. An actor network and two critic networks are used. The actor network determines the strategy, and the critic networks evaluate its value. Both critic networks include target networks.

8. An energy consumption optimization system for federated learning based on reinforcement learning in a low-earth orbit satellite network, characterized in that, It includes: An initialization module for constructing a dynamic network architecture composed of low-earth orbit satellites, selecting satellites as federated learning clients, and the remaining satellites as potential server nodes to initialize the federated learning framework; An energy consumption model construction module for establishing a two-stage energy consumption model of federated learning parameter distribution PDP and upload PUP under the federated learning framework; An energy consumption optimization module for converting the energy consumption optimization in the two-stage energy consumption model into a Markov decision process, training a policy network using the soft actor-critic algorithm SAC, configuring a dual-objective critic network and an experience replay mechanism; An optimization output module for generating actions through the trained reinforcement learning agent based on the current server state before each round of federated learning, dynamically selecting new server nodes and transmission paths; according to the path selection result, performing intermediate parameter aggregation at the first cross node to update the multi-hop routing energy consumption level division; iteratively optimizing the global model parameters and routing strategy to balance immediate energy consumption reduction and network load balancing until the specified number of federated learning rounds is completed.

9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method for optimizing the energy consumption of federated learning based on reinforcement learning in the low-earth orbit satellite network according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method for optimizing the energy consumption of federated learning based on reinforcement learning in the low-earth orbit satellite network according to any one of claims 1 to 7.