Intelligent edge selection and resource allocation method for DT-FL satellite health management

By constructing an edge selection model and a joint resource allocation optimization model and combining it with the MMP-DQN algorithm, the problems of edge selection and resource allocation in satellite health management are solved, the precise allocation of computing resources and the optimization of training delay and energy consumption are achieved, and the flexibility and efficiency of satellite health management are improved.

CN120750404APending Publication Date: 2025-10-03TAIYUAN UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511126153.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-12
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing satellite health management technologies have problems such as high data privacy and security risks, severe system response delays, and tight communication resources. They are unable to meet the requirements of real-time and efficient communication. In addition, the DT-FL framework faces the dual challenges of edge selection and resource allocation in satellite health management.

Method used

An intelligent edge selection and resource allocation method is adopted. By constructing a joint optimization model of edge selection and resource allocation, combined with the Markov decision process of multi-agent hybrid action space, the MMP-DQN algorithm is used to learn the approximately optimal edge selection and resource allocation strategy to optimize the allocation of computing and communication resources.

Benefits of technology

It improves the flexibility and efficiency of satellite health management, achieves precise allocation of computing resources, reduces training latency and energy consumption, makes personalized decisions suitable for complex scenarios, and improves model accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120750404A_ABST
    Figure CN120750404A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent edge selection and resource allocation method for DT-FL satellite health management, and the method comprises the steps: carrying out the clustering of satellites in each iteration, and constructing an edge selection model; then, a calculation and communication model is constructed on the basis of the edge selection model, an edge selection and resource allocation joint optimization model with the purpose of minimizing cumulative training time delay and energy consumption is established, and the edge selection and resource allocation joint optimization model is converted into a partially observable Markov decision process; finally, an MMP-DQN algorithm is provided for learning an approximately optimal edge selection and resource allocation strategy. According to the invention, LEO satellites are supported to participate in federated learning through local training or construction of digital twinborn training, and the training flexibility and federated learning efficiency are improved. According to the method, under the condition that calculation, storage and communication resources are limited, the optimal balance between model precision improvement and training time delay energy consumption can be realized through scientific resource scheduling; according to the method, large-scale state space can be processed, and training individuation is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of satellite communications and computing, and in particular to an intelligent edge selection and resource allocation method for DT-FL satellite health management. Background Art

[0002] Satellite systems are crucial to modern infrastructure and are widely used in fields such as communications and navigation. The deployment of large-scale low-Earth orbit satellite constellations, such as Starlink, has led to increasingly intensive utilization of near-Earth orbit resources. With the rapid increase in the number of satellites, the risk of system failure has also increased significantly, making satellite health management increasingly important. However, existing satellite health management technologies present numerous challenges. For example, traditional centralized architectures employ centralized data storage, which poses significant privacy and security risks. Once vulnerabilities in the storage link occur, sensitive information is exposed to the threat of leakage. Furthermore, traditional architectures suffer from hierarchical redundancy, requiring information to pass through multiple layers, resulting in significant system response delays and making it difficult to meet the real-time requirements of health management. Satellite health management requires uploading massive amounts of data, placing significant pressure on communication resources, leading to link congestion and reduced data transmission efficiency. Consequently, traditional centralized health management architectures are no longer able to meet current needs.

[0003] Digital twins are currently widely used in distributed machine learning. When combined with a federated learning framework, the Digital Twin-Federated Learning (DT-FL) framework offers three core advantages: data privacy protection, with digital twins providing simulated data to prevent original data leakage; enhanced system reliability, with real-time monitoring and prediction of physical entity states; and optimized resource utilization, with federated learning resources allocated based on real-time information. This framework is widely used in areas such as the Industrial Internet, intelligent driving, and high-speed mobile networks. By leveraging digital twins and reinforcement learning techniques, it optimizes device feature capture, environmental simulation, and node configuration, achieving parameter optimization, reducing training latency and communication costs, and supporting intelligent applications and security assurance in various fields. Applying this framework to satellite health management offers three benefits: first, it can safeguard data privacy and address the issue of insufficient computing resources per satellite through local training and digital twin-assisted training; second, it can simplify system complexity using the three-tiered architecture of federated learning; and third, it can alleviate transmission pressure on constellation communication links through a parameterized update upload mechanism. However, while addressing these challenges, the DT-FL framework also presents two challenges: (1) Edge selection problem: Determine whether the Low Earth Orbit (LEO) satellite chooses to use the digital twin established in the edge server (ES) deployed on the satellite base station for model training in this round of global update to solve the problem of insufficient or faulty computing resources on a single satellite.

[0004] (2) Resource allocation problem: Under the constraints of limited computing, storage, and communication resources, how much computing resources should be allocated to each LEO satellite digital twin for model training to speed up training and reduce resource overhead.

[0005] In summary, an intelligent edge selection and resource allocation method for DT-FL satellite health management is urgently needed to meet the efficient operation requirements of LEO satellite clusters under the DT-FL framework, achieve accurate edge twin construction selection and optimal allocation of limited resources, effectively cope with the dual challenges and improve the practicality and reliability of the overall framework. Summary of the Invention

[0006] In order to solve the above problems, the purpose of the present invention is to provide an intelligent edge selection and resource allocation method for DT-FL satellite health management.

[0007] The present invention provides an intelligent edge selection and resource allocation method for DT-FL satellite health management, comprising: Step S1: During each round of federated learning global update iteration, LEO satellites are clustered under each edge server ES by selecting access satellites and setting a fixed time interval to ensure that the accessible satellites of each edge server ES remain unchanged during the entire time interval of a single iteration.

[0008] Step S2: Construct an edge selection model to determine whether each LEO satellite chooses to build a corresponding twin in ES and perform local model training in ES or whether the LEO satellite performs local model training on board.

[0009] Step S3: Construct a multi-stage DT-FL computational model and calculate the latency and energy consumption of each of the following computational stages. These stages include the LEO satellite building and maintaining a digital twin in the ES, local computation using the LEO satellite's own real-time data for model training, edge computing using data generated by the digital twin for model training in the ES, aggregation of local model parameters in the ES (referred to as local aggregated computation), and global aggregation of local model parameters in the Central Cloud (CC).

[0010] Step S4: Construct a multi-communication phase DT-FL communication model and calculate the latency and energy consumption for each phase. The communication phase includes inter-satellite communication, LEO satellite-to-ES communication, ES-to-CC communication, and CC-to-ES communication for global parameter transmission and onboard training with the LEO satellite. The latency and energy consumption associated with global parameter transmission from the central cloud are negligible because the data volume is much smaller than the previous three phases.

[0011] Step S5: Calculate the total system latency and total energy consumption based on the latency and energy consumption of the computation and communication phases constructed in Steps S3 and S4. Based on the total system latency and total energy consumption and related constraints, establish a joint optimization model for edge selection and resource allocation with the goal of minimizing cumulative training latency and energy consumption.

[0012] Step S6: Convert the joint optimization model constructed in step S5 into a multi-agent hybrid action space Markov decision (POMDP) ​​problem model. Specifically, the POMDP model of the system can be constructed by constructing the state space, action space, state transition probability and reward function deployed on a single ES agent.

[0013] Step S7: Design a deep reinforcement learning (MMP-DQN) algorithm capable of solving the aforementioned POMDP problem model. This algorithm is used to learn near-optimal edge selection and resource allocation strategies, minimizing the cumulative latency and energy consumption of federated learning. The MMP-DQN algorithm design includes determining the policy network and Q network structures deployed in the ES agent, as well as the parameter update strategy for these two networks; and determining the fusion network structure deployed in the CC and the fusion network parameter update strategy.

[0014] Furthermore, the specific operations of step S1 are as follows: The terminal layer LEO satellite set is represented as ,in, is the number of LEO satellites in the terminal layer. The set of edge layer ES is expressed as ,in, is the number of edge layers ES.

[0015] Choose to leave The nearest visible satellite is used as the access satellite , will be connected to the satellite The path hop count is The satellites that jump to the cluster In this case, when the satellite distances are different and the number of access satellite hops is the same, they are divided into corresponding clusters according to the principle of shortest path distance. The LEO satellite clusters under , is the total number of satellites in the cluster. The LEO satellite clusters under each ES are disjoint.

[0016] Furthermore, in step S2, an edge selection model needs to be constructed: Using variables express The corresponding LEO satellite cluster LEO satellites Is it Build and maintain its digital twin in the local machine learning platform for local model training. Indicates satellite exist Build and maintain its digital twin model in the local environment and use it for local model training. Indicates absence Building and maintaining satellites Data twin model of satellite Perform local model training on the satellite. Definition ,in, Indicates satellite Whether to perform local model training on the satellite and transfer the training parameters to Local polymerization is performed in

[0017] The total number of LEO satellite digital twins that each ES can maintain simultaneously is expressed as , then for any The number of LEO satellite digital twins maintained by Zhongtong must meet Given the real-time requirements for the twin model, the LEO satellite twins constructed and maintained by each ES are created before the start of the global iteration during the federated learning process. During the federated learning training, the twin model in each global iteration of the ES maintains the satellite twin model based on whether it participates in the federated learning training.

[0018] Furthermore, the calculation model of DT-FL in step S3 is constructed as follows: In the In the global iteration, create The required delay of the digital twin is expressed as , twins are all in When created. Among them, The number of CPU cycles required to process one byte of data for ES; for No. The amount of data used to create its digital twin in the first round of global iterations; for Available CPU frequencies. Create The computational energy consumption of the digital twin is expressed as .in, is the computing power of ES. From this we can see that The computational energy consumption for creating a digital twin of a LEO satellite is ,and The energy consumption for maintaining the LEO digital twin is .in, Indicates the power of ES to maintain a single digital twin.

[0019] The computational latency of training a local model on-board a satellite is expressed as: , Digital twins in The training delay is expressed as: .in, The number of CPU cycles required to train one byte of data, For the The size of the data that the local model participates in training in the global iteration, for The frequency of local training cycles, For the Round global iteration ES is assigned to The CPU cycle frequency of the digital twin. Correspondingly, The computational energy consumption of training the local model on board is ,in, express The on-board power calculation of The computational energy consumption of training the local model of the digital twin is , from this we can know The total computational energy consumption of training the local model using satellite twins is .

[0020] The time delay for aggregating local model parameters is .in, for Managed satellite clusters The model parameter size, The number of CPU cycles required to aggregate one byte of data. Correspondingly, The aggregate energy consumption of aggregating local model parameters is ,in, express Aggregate power.

[0021] CC aggregation The calculation delay of the uploaded local model parameters is .in, for The size of the uploaded local model parameters, Indicates the CPU cycle frequency of the CC aggregation model. Correspondingly, the aggregate energy consumption of the CC aggregation local model parameters ,in, Indicates the aggregate power of CC.

[0022] Furthermore, the communication model of DT-FL in step S4 is constructed as follows: Intersatellite transmission, and room, and The transmission delays between CC and GPIO are: ;

[0023] ; .

[0024] The following is a detailed description of the parameters in the formula: The data transmission rate between satellites is .in, represents the channel width of the intersatellite link, Represents the signal-to-noise ratio (SNR) of intersatellite communication. and The data transmission rate between .in, express arrive The channel width, express and The signal-to-noise ratio between . The data transmission rate between CC is .in, express The channel bandwidth, express and the signal-to-noise ratio between CC.

[0025] Correspondingly, the energy consumption of intersatellite transmission is .in, is the transmission power of the intersatellite link. and The transmission energy consumption is .in, yes transmission power. The transmission power between CC is .in, for transmission power.

[0026] Furthermore, the joint optimization model of edge selection and resource allocation in step S5 is constructed as follows: In the DT-FL In the global iteration, The total delay required to complete one round of iterative training is: .

[0027] The local aggregation phase of the model parameters adopts a semi-asynchronous aggregation method, and the maximum time limit for ES to receive the local model parameters of the LEO satellites in the cluster is .when Receiving time exceeds hour, No longer waiting for the unacquired local model parameters in the corresponding cluster, the received local model parameters are directly aggregated. Satellites in the same cluster are trained and learned simultaneously in each round of iteration, so The total delay to obtain all local model parameters in the cluster is determined by the LEO satellite that completes the local model training the slowest, which is expressed as Considering that the maximum time limit of semi-asynchronous aggregation is , The total delay caused by clustering the LEO satellite set is: .

[0028] The total delay is expressed as: .

[0029] The global aggregation phase of the model parameters adopts synchronous aggregation mode, which shows that the system completes the first The total delay caused by the global iteration of the round is expressed as: .

[0030] In the DT-FL In the global iteration, 、 The total energy consumption generated by and CC are: ; ; .

[0031] From the above, we can see that the system completes the The total energy consumption generated by the global iteration is expressed as: .

[0032] Based on the above analysis of training delay and energy consumption, After rounds of global iterations, the joint optimization model of edge selection and resource allocation can be expressed as:

[0033] in, is the set of edge selection results of LEO satellites; It is a collection of computing resource allocation results in ES; and are the coefficients of delay and energy consumption respectively. Constraints: : Make sure the model is After rounds of global iterations, it converges to the target loss , is the convergence threshold.

[0034] : LEO satellites can only choose to be trained locally or with the assistance of digital twins in ES.

[0035] :The total number of LEO satellite digital twins maintained by a single ES does not exceed .

[0036] :Ensure that the CPU frequency allocated by ES to the LEO satellite digital twin is between.

[0037] Furthermore, the POMDP problem model in step S6 is constructed as follows: The POMDP problem model consists of the following four tuples Indicates that, represents the global state space; express The joint action space of the agents in the ES; express The joint state transition probability of the agents in the ES; express The unified reward function of the intelligent agent in ES. The quadruple of the agent To explain: (1) State space :refer to What the agent observes with the help of the digital twin and Resource conditions and status information, etc. Specifically including: The collection of available computing resources ; Communication resource collection ; Current training loss function of the federated learning global model ; Training delay of federated learning in the current round and energy consumption In the Round global iteration, system state It can be expressed as: .

[0038] (2) Action Space : The action space of the agent can be defined as ,in . express Down Edge selection strategy for LEO satellites. . express Zhongwei The computing resource allocation strategy for the digital twin of LEO satellites is proposed. , is the minimum CPU frequency that ES can allocate to LEO satellite digital twin training, The maximum CPU frequency that ES can allocate to LEO satellite digital twin training.

[0039] (3) State transition probability : It gives the current state and actions Transfer to a new state The probability of .

[0040] (4) Reward Function :The instantaneous reward function should pay attention to both the training delay and energy consumption of federated learning and the global model loss. The instantaneous reward of the round iteration is .in, and are the weights of federated training latency and energy consumption and model loss, ; Furthermore, the MMP-DQN algorithm in step S7 is constructed as follows: make express The action value function of the agent in the state and , is a set of discrete actions, is a set of continuous actions. In the global iteration, for each discrete action , the corresponding optimal continuous action Can be regarded as observations When When the maximum value is taken, the optimal continuous action can be obtained Therefore, the policy network can be used To approximate the mapping function from observation to the agent's optimal continuous action, that is, .

[0041] Here, we use a Parameterized Q-network Actions are used to approximate the action-value function. The optimal discrete action is selected based on the value of the discrete action-value function. That is, the discrete action corresponding to the largest discrete value function can be expressed as: .

[0042] In the forward propagation of the Q network, a multi-channel forward propagation method is used. For each action , the state and the action parameter vector Perform a forward propagation as input, where yes Middle dimensional standard basis vectors. Therefore, is a joint action parameter vector, where each of are all set to 0. This makes all error gradients 0, i.e. , and eliminate irrelevant action parameters in the input layer The influence of the network weights, so that Depends only on .

[0043] In the centralized training phase, the QMIX algorithm is used to solve the multi-agent problem. The CC fusion network will fuse all agents and perform nonlinear combination of the Q functions of each agent. The fusion network is based on the global state S and the Q function of each agent. Function as input, and get the global Q function: .in represents a fusion network; For M agents Nonlinear combinations of functions.

[0044] Since the local action Q value is derived based on the observation value of the agent, we can Expressed in general form .in, Represents the parameters of the fusion network. In order to implement the learning process in a distributed manner, The gradient of Items, each of which is generated only based on the individual observations of each agent. Therefore, the global Q value can be decomposed into Independent items ,in express In order to make the global Q value decomposable, the second-order partial derivative of any two observations Should be approximately zero.

[0045] The global loss function can be expressed as Among them, the target value , is the penalty parameter for non-zero second-order partial derivatives. At this time, the global gradient Can be decomposed into the local gradient of each agent The summation form of .

[0046] During training, each agent first calculates its local gradient As a contribution to the global parameter update, after receiving the gradients of all agents, the central cloud CC performs global parameter update through gradient aggregation, i.e. .in, Represents the learning rate.

[0047] In addition, the policy network and Q Network The parameters are trained alternately by minimizing their respective loss functions. The loss functions of the Q network and the deterministic policy network can be expressed as: ; .

[0048] The network parameters of the Q network and the policy network are updated by the gradient descent method, which is specifically expressed as: , .in, and are the learning rates of the Q network and the policy network, respectively.

[0049] In the DT-FL satellite health management system, LEO satellites in the terminal layer complete local training using local data sets and their own computing resources, or upload twin data to satellite base stations equipped with ES in the edge layer, and use the computing resources of ES for auxiliary training in the digital twin space; LEO satellites that have completed local training upload model parameters to the nearest ES through intersatellite links and downlinks. The ES aggregates the local model within the communication range with the model in the digital twin space; the ES uploads the local model parameters to the central cloud (CC) in the cloud layer, which is equipped with a large number of parameter servers; after receiving the parameters, the CC performs global aggregation and sends the updated global model to the LEO satellite and ES. All parties repeat the iteration until the model loss reaches the target value.

[0050] The MMP-DQN algorithm is a multi-agent algorithm that combines the MP-DQN algorithm and the QMIX algorithm. Based on the P-DQN algorithm, the MP-DQN algorithm separates action parameters through multiple propagations, enabling a single Q network with fully connected layers to implement separate action parameter inputs. This allows it to handle mixed action space problems and outperforms P-DQN. The QMIX algorithm uses a "centralized training, independent execution" approach to achieve collaborative learning, combining the Q values ​​of individual agents into a global Q value through a fusion network. The MMP-DQN algorithm leverages the QMIX algorithm to meet the requirements of multi-agent environments. It distributes the established MP-DQN algorithm on a single agent to solve the multi-agent mixed action space problem and learn near-optimal edge selection and resource allocation strategies.

[0051] In this intelligent edge selection and resource allocation method for DT-FL satellite health management, an edge selection model is constructed to determine whether each LEO satellite chooses to train its model in the ES during the current global update. Secondly, computation and communication models are constructed based on the edge selection model. A joint optimization model for edge selection and resource allocation is established to minimize cumulative training latency and energy consumption, and this model is converted into a partially observable Markov decision process. Finally, a multi-agent multi-channel parameterized deep Q-network (MMP-DQN) algorithm is proposed to learn near-optimal edge selection and resource allocation strategies. The advantages of this invention are: (1) Edge selection gives LEO satellites dynamic decision-making capabilities, allowing them to flexibly choose local training to participate in federated learning or build digital twins to participate in federated learning, thereby improving training flexibility and federated learning efficiency.

[0052] (2) The resource allocation mechanism precisely allocates ES resources for LEO satellite digital twin training. Under the condition of limited computing, storage, and communication resources, scientific resource scheduling is used to achieve the optimal balance between model accuracy improvement and training latency and energy consumption. (3) The MMP-DQN algorithm solves the problem of multi-agent mixed action space, learns edge selection and resource allocation strategies, is applicable to complex scenarios, can make decisions in parallel, process large-scale state spaces, achieve personalization, and improve model accuracy.

[0053] Based on the following detailed description of specific embodiments of the present invention in conjunction with the accompanying drawings, those skilled in the art will become more aware of the above and other objects, advantages and features of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiment below. The accompanying drawings are for illustration purposes only and are not to be considered as limiting the present invention. The same reference symbols are used throughout the drawings to represent the same components. In the drawings: Figure 1 1. It is a flowchart of an intelligent edge selection and resource allocation method for DT-FL satellite health management according to an embodiment of the present invention; Figure 2 3. It is a schematic diagram of the framework of the intelligent edge selection and resource allocation method for DT-FL satellite health management according to an embodiment of the present invention. DETAILED DESCRIPTION

[0055] The embodiments of the present invention are described below with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are only used to illustrate the present invention and are not intended to limit the present invention.

[0056] The present invention provides an intelligent edge selection and resource allocation method for DT-FL satellite health management, and the specific operation process is as follows: steps S1 to S7.

[0057] Step S1: During each round of federated learning global update iteration, LEO satellites are clustered under each ES by selecting access satellites and inter-connected satellites, and a fixed time interval is set to ensure that the accessible satellites of each ES remain unchanged during the entire time interval of a single iteration.

[0058] Step S2: Construct an edge selection model to determine whether each LEO satellite chooses to build a corresponding twin in ES and perform local model training in ES or whether the LEO satellite performs local model training on board.

[0059] Step S3: Construct a multi-stage DT-FL computational model and calculate the latency and energy consumption of each of the following computational stages. These computational stages include the LEO satellite building and maintaining a digital twin in the ES, local computation using its own real-time data for model training, edge computing using the ES for model training using the digital twin space, aggregation of local model parameters by the ES (referred to as local aggregated computation), and global aggregation of local model parameters by the Central Cloud (CC).

[0060] Step S4: Construct a multi-phase DT-FL communication model and calculate the latency and energy consumption for each of the following communication phases. The communication model includes four components: inter-satellite communication, LEO satellite-to-ES communication, ES-to-CC communication, and CC-to-ES communication for global parameter transmission and onboard training with the LEO satellite. The latency and energy consumption associated with global parameter transmission from the central cloud are negligible because the data volume is much smaller than the previous three components.

[0061] Step S5: Calculate the total system latency and total energy consumption based on the latency and energy consumption of the computation and communication phases constructed in Steps S3 and S4. Based on the total system latency and total energy consumption and related constraints, establish a joint optimization model for edge selection and resource allocation with the goal of minimizing cumulative training latency and energy consumption.

[0062] Step S6: Convert the joint optimization model constructed in step S5 into a multi-agent hybrid action space Markov decision (POMDP) ​​problem. Specifically, the POMDP model of the system can be constructed by constructing the state space, action space, state transition probability and reward function deployed on a single ES agent.

[0063] Step S7: Construct a deep reinforcement learning (MMP-DQN) algorithm capable of solving the aforementioned POMDP problem. This algorithm is used to learn near-optimal edge selection and resource allocation strategies, minimizing the cumulative latency and energy consumption of federated learning. The MMP-DQN algorithm design involves determining the policy network and Q network structures deployed in the agent in the ES, as well as the parameter update strategies for these two networks; and determining the fusion network structure and parameter update strategy deployed in the CC. Each step is described in detail below.

[0064] Step S1: In an optional embodiment of the present invention, the terminal layer is composed of a set of LEO satellites of the Walker constellation, represented as ,in, is the number of LEO satellites in the terminal layer. The edge layer consists of Starlink gateways equipped with ES within 50 to 60 degrees in the north-south dimension, expressed as ,in, is the number of edge layers ES. The cloud layer consists of central clouds (CC) located near the equator.

[0065] Choose to leave The nearest visible satellite is used as the access satellite , will be connected to the satellite The path hop count is The satellites that jump to the cluster In this case, when the satellite distances are different and the number of access satellite hops is the same, they are divided into corresponding clusters according to the principle of shortest path distance. The LEO satellite clusters under , is the total number of satellites in the cluster, and the LEO satellite clusters under each ES are disjoint. Finally, a fixed time gap is set to ensure that the number of accessible satellites for each ES remains unchanged throughout the time gap of a single iteration.

[0066] Step S2: Construct edge selection model, using variables express The corresponding LEO satellite cluster LEO satellites Is it Build and maintain its digital twin in the local machine learning platform for local model training. When it is equal to 1, exist Build and maintain its digital twin model in the local environment and use it for local model training. When it is equal to 0, Indicates absence Building and maintaining satellites Data twin model of satellite Perform local model training on the satellite. Definition , Indicates satellite Whether to perform local model training on the satellite and transfer the training parameters to Local polymerization is performed in

[0067] The total number of LEO satellite digital twins that each ES can maintain simultaneously is expressed as , then for any The number of LEO satellite digital twins maintained by Zhongtong must meet Given the real-time requirements for the twin model, the LEO satellite twins constructed and maintained by each ES are created before the start of the global iteration during the federated learning process. During the federated learning training, the twin model in each global iteration of the ES maintains the satellite twin model based on whether it participates in the federated learning training.

[0068] Step S3: Construct the computational model of DT-FL in stages.

[0069] exist In the global iteration, the time delay and energy consumption of the LEO satellite in the ES stage of building and maintaining the digital twin are calculated (the energy consumption of maintaining the LEO digital twin is calculated). The required delay of the digital twin is expressed as , twins are all in When created. Among them, The number of CPU cycles required to process one byte of data for ES; for No. The amount of data used to create its digital twin in the first round of global iterations; for Available CPU frequencies. Create The computational energy consumption of the digital twin is expressed as .in, for Available computing power. From this, we can see that The computational energy consumption for creating a digital twin of a LEO satellite is ,and The energy consumption for maintaining the LEO digital twin is .in, Indicates the power of ES to maintain a single digital twin.

[0070] Calculate the latency and energy consumption of the LEO satellite using its own real-time data for model training and the model training phase with the help of the ES digital twin space. The computational latency of training a local training model on board is expressed as: , Digital twins in The training delay is expressed as: .in, The number of CPU cycles required to train one byte of data, For the The size of the data that the local model participates in training in the global iteration, for The frequency of local training cycles, For the Round global iteration ES is assigned to The CPU cycle frequency of the digital twin. Correspondingly, The computational energy consumption of training the local model on board is .in, express The computing power of Computational energy consumption of training local models for digital twins , from this we can know The total computational energy consumption of training the local model using satellite twins is .

[0071] Calculate the latency and energy consumption of the ES aggregation local model parameter stage. The time delay for aggregating local model parameters is: .in, for Managed satellite clusters The size of the model parameters. The number of CPU cycles required to aggregate one byte of data. Correspondingly, The aggregate energy consumption of aggregating local model parameters is .in, express Aggregate power.

[0072] Calculate the time delay and energy consumption of CC aggregation local model parameter stage. The calculation delay of the uploaded local model parameters is: .in, for The size of the uploaded local model parameters, Indicates the CPU cycle frequency of the CC aggregation model. Correspondingly, the aggregate energy consumption of the CC aggregation local model parameters .in, Indicates the aggregate power of CC.

[0073] Step S4: Construct the communication model of DT-FL in stages.

[0074] Calculate the delay of inter-satellite transmission, LEO satellite to ES communication, and ES to CC communication. arrive as well as The transmission delays to CC are: ; ; .

[0075] The following is a detailed description of the parameters in the formula: The data transmission rate between satellites is .in, represents the channel width of the intersatellite link, Represents the signal-to-noise ratio (SNR) of intersatellite communication. and The data transmission rate between .in, express arrive The channel width, express and The signal-to-noise ratio between . The data transmission rate between CC is .in, express The channel bandwidth, express and the signal-to-noise ratio between CC.

[0076] Correspondingly, the energy consumption of intersatellite transmission is .in, is the transmission power of the intersatellite link. and The transmission energy consumption is .in, yes transmission power. The transmission power between CC is .in, for transmission power.

[0077] Step S5: Use system latency, energy consumption, and related constraints to establish a joint optimization model for edge selection and resource allocation with the goal of minimizing cumulative training latency and energy consumption.

[0078] In the DT-FL In the global iteration, The total delay required to complete one round of iterative training is: .

[0079] The local aggregation phase of the model parameters adopts a semi-asynchronous aggregation method, and the maximum time limit for ES to receive the local model parameters of the LEO satellites in the cluster is .when Receiving time exceeds hour, No longer waiting for the unacquired local model parameters in the corresponding cluster, the received local model parameters are directly aggregated. Satellites in the same cluster are trained and learned simultaneously in each round of iteration, so The total delay to obtain all local model parameters in the cluster is determined by the LEO satellite that completes the local model training the slowest, which is expressed as Considering that the maximum time limit of semi-asynchronous aggregation is , The total delay caused by clustering the LEO satellite set is: .

[0080] The total delay is expressed as: .

[0081] The global aggregation phase of the model parameters adopts synchronous aggregation mode, which shows that the system completes the first The total delay caused by the global iteration of the round is expressed as: .

[0082] In the DT-FL In the global iteration, 、 The total energy consumption generated by and CC are: ; ; .

[0083] From the above, we can see that the system completes the The total energy consumption generated by the global iteration is expressed as: .

[0084] Based on the above analysis of training delay and energy consumption, After rounds of global iterations, the joint optimization model of edge selection and resource allocation can be expressed as:

[0085] in, is the set of edge selection results of LEO satellites; It is a collection of computing resource allocation results in ES; and are the coefficients of delay and energy consumption respectively. Constraints: : Make sure the model is After rounds of global iterations, it converges to the target loss , is the convergence threshold.

[0086] : LEO satellites can only choose to be trained locally or with the assistance of digital twins in ES.

[0087] :The total number of LEO satellite digital twins maintained by a single ES does not exceed .

[0088] :Ensure that the CPU frequency allocated by ES to the LEO satellite digital twin is between.

[0089] Step S6: Construct the POMDP model of the system by constructing the state space, action space, state transition probability and reward function deployed on a single agent of ES.

[0090] The POMDP problem model consists of the following four tuples Indicates that, represents the global state space; express The joint action space of the agents in the ES; express The joint state transition probability of the agents in the ES; express The unified reward function of the intelligent agent in ES. The quadruple of the agent in the : (1) State space :refer to What the agent observes with the help of the digital twin and Resource conditions and status information, etc. Specifically including: The collection of available computing resources ; Communication resource collection ; Current training loss function of the federated learning global model ; Training delay of federated learning in the current round and energy consumption In the Round global iteration, system state It can be expressed as: .

[0091] (2) Action Space : The action space of the agent can be defined as ,in . express Down Edge selection strategy for LEO satellites. . express Zhongwei The computing resource allocation strategy for the digital twin of LEO satellites is proposed. , is the minimum CPU frequency that ES can allocate to LEO satellite digital twin training, The maximum CPU frequency that ES can allocate to LEO satellite digital twin training.

[0092] (3) State transition probability : It gives the current state and actions Transfer to a new state The probability of .

[0093] (4) Reward Function :The instantaneous reward function should pay attention to both the training delay and energy consumption of federated learning and the global model loss. The instantaneous reward of the round iteration is .in, and are the weights of federated training latency and energy consumption and model loss, .

[0094] Step S7: Construct an MMP-DQN algorithm to learn the approximately optimal edge selection and resource allocation strategy to minimize the cumulative delay and energy consumption of federated learning.

[0095] make express The action value function of the agent in the state and , is a set of discrete actions, A collection of continuous actions.

[0096] In the In the global iteration, for each discrete action , the corresponding optimal continuous action Can be regarded as observations When When the maximum value is taken, the optimal continuous action can be obtained Therefore, the policy network can be used To approximate the mapping function from observation to the agent's optimal continuous action, that is, .

[0097] Here, we use a Parameterized Q-network Actions are used to approximate the action-value function. The optimal discrete action is selected based on the value of the discrete action-value function. That is, the discrete action corresponding to the largest discrete value function can be expressed as: .

[0098] In the forward propagation of the Q network, a multi-channel forward propagation method is used. For each action , the state and the action parameter vector Perform a forward propagation as input, where yes Middle dimensional standard basis vectors. Therefore, is a joint action parameter vector, where each of are all set to 0. This makes all error gradients 0, i.e. , and eliminate irrelevant action parameters in the input layer The influence of the network weights, so that Depends only on .

[0099] In the centralized training phase, the QMIX algorithm is used to solve the multi-agent problem. The CC fusion network will fuse all agents and perform nonlinear combination of the Q functions of each agent. The fusion network is based on the global state S and the Q function of each agent. Function as input, and get the global Q function: .in represents a fusion network; For M agents Nonlinear combinations of functions.

[0100] Since the local action Q value is derived based on the observation value of the agent, we can Expressed in general form .in, Represents the parameters of the fusion network. In order to implement the learning process in a distributed manner, The gradient of Items, each of which is generated only based on the individual observations of each agent. Therefore, the global Q value can be decomposed into Independent items ,in express In order to make the global Q value decomposable, the second-order partial derivative of any two observations Should be approximately zero.

[0101] The global loss function can be expressed as Among them, the target value , is the penalty parameter for non-zero second-order partial derivatives. At this time, the global gradient Can be decomposed into the local gradient of each agent The summation form of .

[0102] During training, each agent first calculates its local gradient As a contribution to the global parameter update, after receiving the gradients of all agents, the central cloud CC performs global parameter update through gradient aggregation, i.e. .in, Represents the learning rate.

[0103] In addition, the policy network and Q Network The parameters are trained alternately by minimizing their respective loss functions. The loss functions of the Q network and the deterministic policy network can be expressed as: ; .

[0104] The network parameters of the Q network and the policy network are updated by the gradient descent method, which is specifically expressed as: , .in, and are the learning rates of the Q network and the policy network, respectively.

[0105] The present invention proposes an intelligent edge selection and resource allocation method for DT-FL satellite health management, which can be applied to any DT-FL satellite health management system. In addition, the system includes at least one satellite base station deploying an ES and one LEO satellite.

[0106] LEO satellites in the terminal layer complete local training using local datasets and their own computing resources, or upload twin data to satellite base stations equipped with ESs in the edge layer, using the ES's computing resources for auxiliary training in the digital twin space. LEO satellites that have completed local training upload model parameters to the nearest ES via intersatellite links and downlinks. The ES aggregates the local model within the communication range with the model in the digital twin space. The ES uploads the local model parameters to the central cloud (CC) in the cloud layer, which is equipped with a large number of parameter servers. Upon receiving the parameters, the CC performs global aggregation and sends the updated global model to the LEO satellites and ESs. All parties repeat the iteration until the model loss reaches the target value.

[0107] It should be noted that this implementation does not restrict the local model parameters (e.g., weights, gradients, etc.) or the content of local raw data sets (e.g., images, text, audio, etc.) uploaded by LEO satellites. Furthermore, the local raw data sets of each satellite can overlap or be completely different. Regarding health management models, while there are no restrictions on specific types (rules, machine learning, deep learning, etc.), all LEO satellites must deploy the same model and adhere to consistent data interfaces, algorithm protocols, and evaluation standards to ensure system compatibility, stability, and efficient operation and maintenance.

[0108] When aggregating local models within the ES communication range and models in the digital twin space, or local models within the CC communication range, model parameter preprocessing is required to ensure the quality and efficiency of model aggregation. This includes data standardization, outlier processing, and missing value filling. However, this solution does not limit the specific preprocessing methods. Appropriate preprocessing methods and techniques can be flexibly selected based on actual needs to adapt to different models and data characteristics.

[0109] This paper proposes an intelligent edge selection and resource allocation method for DT-FL satellite health management. Based on the deployment of intelligent agents in the ES, the MMP-DQN algorithm uses the MP-DQN algorithm and the QMIX algorithm to work in tandem. The MP-DQN algorithm is distributed across each agent in the ES, utilizing its decision network and Q network to select continuous and discrete actions, effectively solving the multi-agent mixed action space problem. Simultaneously, the QMIX algorithm's fusion network, deployed in the central cloud (CC), integrates information from each agent, enabling collaborative optimization of strategies in a multi-agent environment. This allows the learning of near-optimal edge selection and resource allocation strategies.

[0110] An embodiment of the present invention further provides an intelligent edge selection and resource allocation device for DT-FL satellite health management, which is used to execute the intelligent edge selection and resource allocation method for DT-FL satellite health management of the above embodiment.

[0111] An embodiment of the present invention further provides a computer-readable storage medium for storing program code, wherein the program code is used to execute the intelligent edge selection and resource allocation method for DT-FL satellite health management of the above embodiment.

[0112] An embodiment of the present invention further provides a computing device, comprising a processor and a memory: the memory is used to store program code and transmit the program code to the processor; the processor is used to execute the intelligent edge selection and resource allocation method for DT-FL satellite health management of the above embodiment according to instructions in the program code.

[0113] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiment. All technical solutions based on the concept of the present invention are within the scope of protection of the present invention. It should be noted that for those skilled in the art, various improvements and modifications that do not depart from the principles of the present invention should also be considered within the scope of protection of the present invention.

Claims

1. An intelligent edge selection and resource allocation method for DT-FL satellite health management, characterized by: The method comprises: S1, in each round of federated learning global update iteration, LEO satellites are clustered under each edge server ES and a fixed time interval is set to ensure that the number of accessible satellites of each edge server ES remains unchanged during the entire time interval of a single iteration; S2, builds an edge selection model to determine whether each LEO satellite chooses to build a corresponding twin in ES and perform model training in ES or whether the LEO satellite performs local model training on board; S3: Build a multi-stage DT-FL computational model and calculate the latency and energy consumption of each computational stage. The computational stages include the LEO satellite building and maintaining a digital twin in the ES, local computation in which the LEO satellite uses its own real-time data for model training, edge computation in which the ES uses data generated by the digital twin for model training, local aggregation computation in which the ES aggregates local model parameters, and global aggregation computation in which the central cloud CC aggregates local model parameters. S4: Build a multi-communication phase DT-FL communication model and calculate the delay and energy consumption of each communication phase. The communication phase includes four parts: inter-satellite link communication, communication between LEO satellite and ES, communication between ES and CC, and communication from the central cloud to the ES and LEO satellite. The central cloud sends global parameters. Compared with the first three, the data transmission volume is very small, so the related delay and energy consumption are negligible. S5, calculating the total delay and total energy consumption of the system based on the delay and energy consumption of the computing phase and the communication phase constructed in step S3 and step S4; based on the total delay and total energy consumption of the system and related constraints, establishing an edge selection and resource allocation joint optimization model with the goal of minimizing the cumulative training delay and energy consumption; S6, converting the joint optimization model described in step S5 into a Markov decision POMDP problem model in a multi-agent hybrid action space; S7. Design a deep reinforcement learning MMP-DQN algorithm that can solve the POMDP problem model, which is used to learn the approximately optimal edge selection and resource allocation strategy to minimize the cumulative delay and energy consumption of federated learning.

2. The method according to claim 1, characterized in that The total delay includes: the total delay of LEO satellites, the total delay of LEO satellite clusters under ES, the total delay of ES, the total delay of ES sets, and the total delay of CCs; The total energy consumption includes: the total energy consumption of LEO satellites, the total energy consumption generated by ES, and the total energy consumption of CC.

3. The method according to claim 1, characterized in that In step S1, the terminal layer LEO satellite set is represented as , the set representation of the edge layer ES ; is the number of edge layer ES; Choose to leave The nearest visible satellite is used as the access satellite , will be connected to the satellite The path hop count is The satellites that jump to the cluster In this case, when the satellite distances are different and the number of access satellite hops is the same, they are divided into corresponding clusters according to the principle of shortest path distance. The LEO satellite clusters under , is the total number of satellites in the cluster. The LEO satellite clusters under each ES are disjoint.

4. The method according to claim 1, wherein In step S2, use the variable express The corresponding LEO satellite cluster LEO satellites Is it Build and maintain its digital twin for local model training; When it is equal to 1, exist Build and maintain its digital twin model in the cloud and use it for local model training; When it is equal to 0, Indicates absence Building and maintaining satellites Data twin model of satellite Perform local model training on the satellite; express Whether to perform local model training on the satellite and transfer the training parameters to Local aggregation is performed in the following formula: , The total number of LEO satellite digital twins that each ES can maintain simultaneously is expressed as , then for any The number of LEO satellite digital twins maintained by Zhongtong must meet ; In view of the real-time requirements of the twin model, the LEO satellite twin built and maintained by each ES is created before the start of the global iteration during the federated learning process. During the federated learning training, the twin model in each round of global iteration of the ES maintains the satellite twin model based on whether it participates in the federated learning training.

5. The method according to claim 2, characterized in that The calculation model of DT-FL in step S3 is constructed as follows: exist In the global iteration, create The required delay of the digital twin is expressed as , twins are all in When created; among them, The number of CPU cycles required to process one byte of data for ES; for No. The amount of data used to create its digital twin in the first round of global iterations; for Available CPU frequencies; create The computational energy consumption of the digital twin is expressed as ;in, for Available computing power; hence, The computational energy consumption for creating a digital twin of a LEO satellite is ;and The energy consumption for maintaining the LEO digital twin is ;in, Indicates the power of ES to maintain a single digital twin; The computational latency of training a local model on-board a satellite is expressed as: , Digital twins in The training delay is expressed as: ;in, The number of CPU cycles required to train one byte of data, is the data size of the local model participating in the training in the kth round of global iteration, for The frequency of local training cycles, Assigned to the kth round of global iteration The CPU cycle frequency of the digital twin; correspondingly, The computational energy consumption of training the local training model on board is in, express The computing power of The computational energy consumption of training the local model of the digital twin is , from this we can know The total computational energy consumption of training the local model using satellite twins is ; The time delay for aggregating local model parameters is: ;in, for Managed satellite clusters The model parameter size, The number of CPU cycles required to aggregate one byte of data; correspondingly, The aggregate energy consumption of aggregating local model parameters is ;in, express Aggregate power; CC aggregation The calculation delay of the uploaded local model parameters is: ;in, for The size of the uploaded local model parameters, Indicates the CPU cycle frequency of the CC aggregation model; correspondingly, the aggregate energy consumption of the CC aggregation local model parameters ,in, Indicates the aggregate power of CC.

6. The method according to claim 2, characterized in that The communication model of DT-FL in step S4 is constructed as follows: Intersatellite transmission, and room, and The transmission delays between CC and GPIO are: ; ; ; The data transmission rate between satellites is: ;in, represents the channel width of the intersatellite link, represents the signal-to-noise ratio (SNR) of intersatellite communication; and The data transmission rate between ;in, express arrive The channel width, express and The signal-to-noise ratio between The data transmission rate between CC is ;in, express The channel bandwidth, express and the signal-to-noise ratio between CC; Correspondingly, the energy consumption of intersatellite transmission is ;in, is the transmission power of the intersatellite link; and The transmission energy consumption is ;in, yes The transmission power; The transmission power between CCs is ;in, for transmission power.

7. The method according to claim 2, characterized in that The joint optimization model of edge selection and resource allocation in step S5 is constructed as follows: In the DT-FL In the global iteration, The total delay required to complete one round of iterative training is: ; The local aggregation phase of the model parameters adopts a semi-asynchronous aggregation method, and the maximum time limit for ES to receive the local model parameters of the LEO satellites in the cluster is ;when Receiving time exceeds hour, No longer waiting for the unacquired local model parameters in the corresponding cluster, the received local model parameters are directly aggregated; the satellites in the same cluster are trained and learned simultaneously in each round of iteration, so The total delay to obtain all local model parameters in the cluster is determined by the LEO satellite that completes the local model training the slowest, which is expressed as ; Considering the maximum time limit of semi-asynchronous aggregation , The total delay caused by clustering the LEO satellite set is: ; The total delay is expressed as: ; The global aggregation stage of the model parameters adopts synchronous aggregation. It can be seen that the total delay caused by the system completing the kth round of global iteration is expressed as: ; In the DT-FL In the global iteration, 、 The total energy consumption generated by and CC are: ; ; ; From the above, we can see that the system completes the The total energy consumption generated by the global iteration is expressed as: ; Based on the above analysis of training delay and energy consumption, After rounds of global iterations, the joint optimization model of edge selection and resource allocation can be expressed as: ; in, is the set of edge selection results of LEO satellites; It is a collection of computing resource allocation results in ES; and are the coefficients of delay and energy consumption respectively; the constraints are as follows: : Ensure that the model converges to the target loss after K rounds of global updates , is the convergence threshold; : LEO satellites can only choose to train locally or use digital twins in ES to assist training; :The total number of LEO satellite digital twins maintained by a single ES does not exceed ; :Ensure that the CPU frequency allocated by ES to the LEO satellite digital twin is between.

8. The method according to claim 2, characterized in that The POMDP model in step S6 is constructed as follows: The POMDP model consists of the following four tuples Indicates that, represents the global state space; express The joint action space of the agents in the ES; express The joint state transition probability of the agents in the ES; represents the unified reward function of the agents in B ES; Below The quadruple of the agent in the : (1) State space :refer to What the agent observes with the help of the digital twin and Resource status and information; specifically including: The collection of available computing resources ; Communication resource collection ; Current training loss function of the federated learning global model ; Training delay of federated learning in the current round and energy consumption ; In the kth round of global iteration, the system state It can be expressed as: ; (2) Action Space : The action space of the agent can be defined as ,in , Indicates Down Edge selection strategy for LEO satellites; ; express Zhongwei Computational resource allocation strategy for digital twins of LEO satellites; , is the minimum CPU frequency that ES can allocate to LEO satellite digital twin training, The maximum CPU frequency that ES can allocate to the training of LEO satellite digital twins; (3) State transition probability : It gives the current state and actions Transfer to a new state The probability of ; (4) Reward Function :The instantaneous reward function should pay attention to the federated learning training delay and energy consumption as well as the global model loss. The instantaneous reward of the kth iteration is ;in, and are the weights of federated training latency and energy consumption and model loss, .

9. The method according to claim 2, characterized in that The MMP-DQN algorithm in step S7 is constructed as follows: make express The action value function of the agent in the state and , is a set of discrete actions, A collection of continuous actions; In the kth global iteration, for each discrete action , the corresponding optimal continuous action Can be regarded as observations function; when When the maximum value is taken, the optimal continuous action can be obtained ; Therefore, using the policy network To approximate the mapping function from observation to the agent's optimal continuous action, that is, ; Use a Parameterized Q-network Actions are used to approximate the action value function; the optimal discrete action is selected according to the value of the discrete action value function, that is, the discrete action corresponding to the largest discrete value function can be expressed as: ; In the forward propagation of the Q network, a multi-channel forward propagation method is adopted; for each action , the state and the action parameter vector Perform a forward propagation as input, where yes The standard basis vectors for the i-th dimension in ; therefore, is a joint action parameter vector, where each of are all set to 0, which makes all error gradients 0, that is, , and eliminate irrelevant action parameters in the input layer The influence of the network weights, so that Depends only on ; In the centralized training phase, the QMIX algorithm is used to solve the multi-agent problem; the CC fusion network will fuse all agents and perform nonlinear combination of the Q functions of each agent; the fusion network uses the global state S and the Q functions of each agent to calculate the Q function of each agent. Function as input, and get the global Q function: ;in represents a fusion network; For M agents Nonlinear combinations of functions; Since the local action Q value is derived based on the observation value of the agent, Expressed in general form ;in, represents the parameters of the fusion network; in order to implement the learning process in a distributed manner, The gradient of Each item is generated only based on the individual observation value of each agent, so the global Q value can be decomposed into Independent items ,in express In order to make the global Q value decomposable, the second-order partial derivatives of any two observations should be approximately zero; The global loss function can be expressed as ; Among them, the target value , is the penalty parameter for non-zero second-order partial derivatives; at this time the global gradient Can be decomposed into the local gradient of each agent The summation form of ; During training, each agent first calculates its local gradient As a contribution to the global parameter update, after receiving the gradients of all agents, the central cloud CC performs the global parameter update through gradient aggregation, i.e. ;in, represents the learning rate; In addition, the policy network and Q Network The parameters are trained alternately by minimizing their respective loss functions; the loss functions of the Q network and the deterministic policy network can be expressed as: ; ; The network parameters of the Q network and the policy network are updated by the gradient descent method, which is specifically expressed as: , ;in, and are the learning rates of the Q network and the policy network, respectively.