Satellite edge calculation unloading method and system based on deep reinforcement learning model, and electronic equipment
By combining deep reinforcement learning and generating diffusion models, the problems of low stability and poor environmental adaptability of satellite edge computing offload decisions in the prior art are solved, and higher quality and stable offload decisions are achieved.
Patent Information
- Application Number
- CN202510249273.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-07-01
AI Technical Summary
The existing satellite edge computing offloading method based on deep reinforcement learning has problems such as low stability, poor environmental adaptability and poor decision-making quality in complex and dynamic LEO inter-satellite network environments.
Using a method based on deep reinforcement learning and generating diffusion models, the satellite network state is collected through deep reinforcement learning models, empirical tuples are generated, and random sampling is used to train the unload decision strategy, and the optimal unload decision is generated by generating the diffusion model.
It improves the stability and environmental adaptability of satellite edge computing offload decisions, significantly improves decision quality, and can better cope with complex and dynamic LEO inter-satellite network environments.
Smart Images

Figure CN120234112A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of satellite edge computing offloading, and particularly relates to a satellite edge computing offloading method, system and electronic device based on a deep reinforcement learning model. Background Art
[0002] The low Earth orbit (LEO) satellite network provides seamless and flexible access services for users with the characteristics of low latency, high bandwidth and global coverage, and has become one of the key elements for constructing the sixth generation of mobile communication technology (6G). With the successive launch of low-orbit satellites equipped with edge computing servers, such as Tianzhi-1 and Tianxuan constellations, providing on-orbit edge computing services for remote Internet users, satellite edge computing (SEC) technology has emerged as the times require and is regarded as one of the most promising computing architectures in the future 6G communication network. On the one hand, by extending the rich computing resources of ground cloud servers to the low-earth orbit edge, it can effectively realize the in-situ processing of IoRT user data, relieve the dependence of satellite networks on ground networks, reduce task transmission latency while optimizing bandwidth occupancy; on the other hand, satellite edge processing reduces the need to transmit sensitive data to external cloud servers, which helps to protect user privacy and improve data security.
[0003] However, with the development trend of "communication-sensing-computation" integration, the resource heterogeneity shown by current LEO satellite constellation nodes will bring new challenges to the SEC computing offloading decision-making task. In addition, the high-speed motion mode of satellites and the complex inter-satellite network environment will further increase the difficulty of this task. Due to its performance advantages in complex, dynamic and uncertain network scenarios, currently, deep reinforcement learning (DRL) strategies have become the mainstream method for optimizing satellite computing offloading tasks. However, there are still some limitations when existing DRL-based methods are applied to LEO scenarios. On the one hand, the inter-satellite environment is complex, and DRL algorithms usually need to interact with a large number of environmental samples before they can accurately understand the environmental state. The difficulty of detailed data collection and manual content creation between satellites limits the performance of existing methods. On the other hand, DRL algorithms are limited in the high-dimensional state and action space between satellites, which may lead to poor decision-making quality.
[0004] Recently, generative artificial intelligence (GenAI) technology has emerged as a major advancement in the field of artificial intelligence. Different from traditional discriminative AI models that only focus on and process existing data, applying GenAI to LEO satellite network scenarios can not only effectively generate high-quality data related to historical operation trajectories, solving the problem of difficult detailed data collection between satellites, but also accurately learn and adapt to various unstructured data sources, making it more flexible in handling various applications. Combining GenAI with DRL, GenAI generates richer data and scenario simulations based on existing data and reinforcement learning strategies, and deep reinforcement learning makes more reasonable offloading decisions by interacting with more comprehensive environmental data. The two complement each other and are more suitable for dealing with complex and dynamic LEO inter-satellite network environments. Common GenAI models include Transformers, generative adversarial networks (GANs), variational autoencoders (VAEs), flow-based generative models, energy-based generative models, and generative diffusion models (GDMs). Compared with other models, GDM stands out with its unique data generation method and the ability to model complex data distributions, showing higher flexibility and adaptability, and is particularly suitable for complex and dynamic satellite network environments.
[0005] In addition, the sources of computing tasks in satellite networks are complex, mainly including in-node tasks and offloading tasks. In-node tasks refer to tasks generated by the satellite itself, such as attitude control, orbit calculation, and sensor data processing. Due to restrictions such as data privacy, these tasks must be completed locally on the satellite. Offloading tasks come from ground users, and these tasks usually have the characteristics of large data volume, high computational intensity, and high sensitivity to latency. Processing such tasks on a single satellite node with limited resources may lead to high latency and the risk of resource exhaustion. Therefore, they can be offloaded to other satellite nodes through inter-satellite cooperation to make full use of the constellation resources to achieve higher computing performance. More complexly, the distribution of in-node tasks usually shows dynamic and unpredictable changes. This uncertainty may lead to instability of the system state, thereby weakening the task scheduling efficiency and affecting network performance. Summary of the Invention
[0006] In view of the above technical problems existing in the prior art, the present invention provides a satellite edge computing offloading method, system, and electronic device based on reinforcement learning and diffusion models, which overcomes the problems of low stability, poor environmental adaptability, and low decision-making quality existing in traditional computing offloading methods.
[0007] The technical solutions adopted by the present invention are as follows:
[0008] A satellite edge computing offloading method based on a deep reinforcement learning model, comprising the following steps:
[0009] Step 1: The deep reinforcement learning model collects the current satellite network state, obtains the reward returned by the environment and the state at the next moment;
[0010] Step 2: Store the information obtained during the interaction between the deep reinforcement learning model and the environment in Step 1 in the form of experience tuples;
[0011] Step 3: Randomly sample the experience tuples in Step 2 for training the offloading decision-making policy in the deep reinforcement learning model; during the training process, adjust the neural network weights in the deep reinforcement learning model until the offloading decision-making policy converges;
[0012] Step 4: The deep reinforcement learning model makes an optimal offloading decision according to the offloading decision-making policy in combination with the current satellite network environment and task requirements.
[0013] Furthermore, the deep reinforcement learning DRL architecture adopts the Soft-Actor-Critic (SAC) framework. The architecture is shown in Figure 3 .
[0014] Furthermore, the core of the policy network of the SAC framework is the Generative Diffusion Model (GDM). The Generative Diffusion Model is shown in Figure 4 .
[0015] The Generative Diffusion Model encodes the observation s effectively by inputting the current state s. The diffusion process can effectively capture the dependencies in the observation and action spaces and improve the performance of the DRL task.
[0016] Furthermore, the specific process of generating the offloading decision-making policy by the Generative Diffusion Model (GDM) in Step 1 is the same as the process of generating the offloading decision-making policy by the Generative Diffusion Model in Step 4.
[0017] Furthermore, in Step 3, the training process is as follows:
[0018] (1) Initialize the parameters of the policy network θ, value network ψ, value target network and Soft-Q network φ;
[0019] (2) The deep reinforcement learning model obtains the current environment state s; generates a preliminary decision ρ based on the initial policy, obtains the reward r and the environment state s' at the next moment;
[0020] (3) The replay buffer stores the experience tuple (s, ρ, s', r);
[0021] (4) In each round of the update training process, a batch of experiences are randomly sampled from the replay buffer for training, where K is the batch size;
[0022] (5) According to the maximum entropy reinforcement learning process, the state value si Estimated:
[0023]
[0024] Wherein, represents the mathematical expectation, represents the policy network θ policy function, ρ new is the input s i and the action generated by the policy network after ρ i and ρ new are sent to the Soft-Q network to generate the state-action value Q φ (s i ,ρ new ); represents the conditional probability distribution that ρ i follows the policy network when the state s new is used as the input;
[0025] (6) Update the value network ψ through the loss function:
[0026]
[0027] (7) Obtain the estimate of the action-state value for (s i ,ρ i ) according to the Soft Bellman equation:
[0028]
[0029] Wherein, r i is the currently obtained reward, γ is the discount factor, s i+1 is the state of the next moment environment, is the output after the value target network inputs s i+1 ; The value target network periodically updates the parameters by soft-copying from the value network ψ to calculate the loss of the Soft-Q network φ:
[0030]
[0031] Wherein, α represents the learning rate;
[0032] (8) Update the policy network θ through the loss function:
[0033] Loss θ = logπθ(ρ i |s i ) - Q φ (s i ,ρ i ).
[0034] Further, in Step 3, during the training process, the long-term stability problem is transformed into a slot-by-slot online optimization task through the Lyapunov optimization framework, and the objective function is:
[0035]
[0036] where B is the upper bound constant of the queue backlog, Q v (τ) is the task queue backlog of node v, and y v (τ) is the change in the queue backlog. W is the trade-off parameter between delay and stability, and t(τ) is the task delay.
[0037] Further, in Step 4, the offloading decision strategy is generated by the generative diffusion model in the deep reinforcement learning model, specifically as follows:
[0038] In each denoising step t, use the deep neural network to infer and scale the denoising distribution tanh(∈ θ (x t ,t,s)), where ∈ θ (x t ,t,s) is the denoising noise generated by the deep neural network. Here, θ represents the parameters of the neural network, x t is the noise distribution at the denoising step t, t is the current denoising step, and s is the current environmental state;
[0039] Calculate the mean of the reverse transition distribution :
[0040]
[0041] where α t = 1 - β t , and β t is the noise variance at the denoising step t, and α t represents the noise retention ratio at the denoising step t;
[0042] Obtain the distribution x t-1 several times using the following update rule until the final output x0 of the diffusion process is obtained:
[0043]
[0044] where represents the random noise of the standard normal distribution;
[0045] Finally, apply the Softmax function to x0 to obtain the offloading decision ρ:
[0046]
[0047] The present invention also discloses a satellite edge computing offloading system based on a deep reinforcement learning model for performing the above method, including the following modules:
[0048] State collection module: The deep reinforcement learning model collects the current satellite network state to obtain the reward returned by the environment and the state at the next moment.
[0049] Storage module: Stores the information obtained by the deep reinforcement learning model and the environment during the interaction in the form of experience tuples.
[0050] Training module: Randomly samples experience tuples for training the offloading decision-making policy in the deep reinforcement learning model; during the training process, adjusts the neural network weights in the deep reinforcement learning model until the offloading decision-making policy converges.
[0051] Decision module: The deep reinforcement learning model makes an optimal offloading decision according to the offloading decision-making policy in combination with the current satellite network environment and task requirements.
[0052] The present invention also discloses an electronic device, including a processor and a memory. The memory is used to store a program. When the program is called and executed by the processor, the processor executes the above method or system.
[0053] Compared with the prior art, the present invention has the following advantages:
[0054] 1. Train the offloading decision-making policy through DRL and make offloading decisions in the way of model decision-making. The trained offloading decision-making policy can make appropriate decisions that conform to the current network environment by using the experience obtained from historical learning.
[0055] 2. In the preferred solution of the present invention, through the Lyapunov optimization framework, the complex edge computing offloading problem can be transformed into an online optimization problem for each time slot in the case where the complete task arrival information is usually difficult to predict. While optimizing the decision delay of each time slot, the present invention can ensure the long-term stability of the satellite edge computing system.
[0056] 3. In the preferred solution of the present invention, GDM is used to generate the optimal offloading decision to improve the generalization ability and dynamic adaptability of the model. GDM can create new high-quality content related to the context, thus reducing the need for exhaustive data collection and manual content creation. At the same time, using GDM to generate high-quality decisions can significantly improve the performance of DRL in the LEO satellite network. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 It is a flowchart of a satellite edge computing offloading method based on a reinforcement learning model according to a preferred embodiment of the present invention.
[0058] Figure 2It is a schematic diagram of the satellite edge computing offloading task.
[0059] Figure 3 It is a schematic diagram for the generation diffusion model GDM to generate the optimal decision.
[0060] Figure 4 It is a schematic diagram of the deep reinforcement learning DRL framework.
[0061] Figure 5 This is a block diagram of a satellite edge computing offloading system based on a reinforcement learning model according to a preferred embodiment of the present invention. Specific implementation manners
[0062] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions of the present invention will be described in detail below. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other implementation manners obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope protected by the present invention.
[0063] As Figure 1 shown, this embodiment provides a satellite edge computing offloading method based on a deep reinforcement learning model, and its general steps are as follows:
[0064] 1. The deep reinforcement learning model collects the current satellite network state and obtains the reward and the state at the next moment returned by the environment; in this embodiment, the deep reinforcement learning model is implemented using the soft actor-critic framework, and the soft actor-critic framework uses the generation diffusion model as the policy network; the specific process of the generation diffusion model (GDM) generating the offloading decision strategy in this step is the same as the process of the offloading decision strategy being generated by the generation diffusion model in step four, and refer to the description of step four below.
[0065] 2. Store the information obtained during the above interaction process between the deep reinforcement learning model and the environment in the form of experience tuples;
[0066] 3. In the training stage, randomly sample a batch of experience tuples from the database for training the offloading decision strategy in the deep reinforcement learning model; during the training process, continuously train and adjust the neural network weights of the deep reinforcement learning model until the offloading decision strategy converges;
[0067] 4. In the decision-making stage, the deep reinforcement learning model makes the optimal offloading decision according to the offloading decision strategy in combination with the current satellite network environment and task requirements.
[0068] Figure 2 It is a schematic diagram of the satellite edge computing offloading task. As Figure 2As shown in the figure, it includes LEO satellites, GEO satellites, and ground users with computing capabilities and capable of cross-layer communication. Among them, both ground users and satellites generate internal node tasks, and such computing tasks will be processed by the ground users or the satellites themselves that generate the internal node tasks. In addition, ground users also generate offloading tasks, and such computing tasks will be offloaded to the satellite network through access satellites and computed jointly by multiple LEO satellites and GEO satellites. After the task computation is completed, the results will be returned to the users.
[0069] Figure 3 It is a schematic diagram of the DRL framework. In step one, the DRL network structure consists of six main modules: the policy network θ, the value network ψ, the value target network the Soft-Q network φ, the replay buffer, and the environment. The deep reinforcement learning model continuously interacts with the environment during training, takes the environmental state as the input of the algorithm, and outputs gradually optimal network parameters.
[0070] The training process in step three is as follows:
[0071] (1) Initialize the parameters of the policy network θ, the value network ψ, the value target network and the Soft-Q network φ;
[0072] (2) The deep reinforcement learning model obtains the current environmental state s; generates a preliminary decision ρ based on the initial policy, obtains the reward r and the environmental state s' at the next moment;
[0073] (3) The replay buffer stores the experience tuple (s, ρ, s', r);
[0074] (4) In each round of the update training process, a batch of experiences will be randomly sampled from the replay buffer for training, where K is the batch size.
[0075] (5) According to the maximum entropy reinforcement learning process, the state value s i can be estimated:
[0076]
[0077] where denotes the mathematical expectation, denotes the policy function of the policy network θ, ρ new is the action generated by the policy network after inputting s i . s i and ρ new are sent to the Soft-Q network to generate the state-action value Q φ (s i , ρ new ). represents when the state si When used as input, ρ new follows the conditional probability distribution of the policy network;
[0078] (6) Update the value network ψ through the loss function:
[0079]
[0080] (7) Obtain an estimate of the action-state value for (s i , ρ i ) according to the Soft Bellman equation:
[0081]
[0082] where r i is the currently obtained reward, γ is the discount factor, s i+1 is the state of the next moment's environment, is the output after the value target network inputs s i+1 . The value target network periodically updates the parameters by soft-copying from the value network ψ to calculate the loss of the Soft-Q network φ:
[0083]
[0084] where α represents the learning rate;
[0085] (8) Update the policy network θ through the loss function:
[0086] Loss θ = Lossπθ(ρ i |s i ) - Q φ (s i , ρ i ).
[0087] In step three, considering the high-dynamic nature of the inter-satellite environment, the training process of the offloading decision strategy needs to consider the adaptive adjustment ability to the changing environment to ensure the long-term stable operation of the satellite network. By introducing the Lyapunov optimization process in the training process, the long-term stability problem of the offloading decision strategy can be transformed into a per-slot optimization adaptive adjustment process, and the objective function is:
[0088]
[0089] where B is the upper bound constant of the queue backlog, Q v (τ) is the task queue backlog of node v, y v (τ) is the change in the queue backlog, W is the trade-off parameter between delay and stability, and t(τ) is the task delay.
[0090] Figure 4 Generate a schematic diagram of the optimal decision for GDM. In step 4, an example of using GDM to generate the optimal decision is as follows:
[0091] (1) In each denoising step t, use the deep neural network to infer and scale the denoising distribution tanh(∈ θ (x t , t, s)),
[0092] ∈ θ (x t , t, s) is a denoising noise generated by the deep neural network, where θ represents the parameters of the neural network, and x t is the noise distribution at the denoising step t, t is the current denoising step, and s is the current environmental state;
[0093] (2) Calculate the mean of the reverse jump distribution :
[0094]
[0095] Among them, represents the Gaussian distribution, I represents the identity matrix, α t = 1 - β t , β t is the noise variance at the denoising step t, α t represents the noise retention ratio at the denoising step t, is the cumulative product of α at the previous denoising step k , is the computable deterministic variance magnitude;
[0096] (3) Repeatedly use the following update rule to obtain the distribution x t-1 , until the final output x0 of the diffusion process is obtained:
[0097]
[0098] Among them, represents the random noise of the standard normal distribution;
[0099] (4) Finally, apply the Softmax function to x0 to obtain the offloading decision ρ:
[0100]
[0101] Among them, represents the final output of the diffusion process of satellite node i, and n represents the number of satellite nodes.
[0102] Based on the above scheme design, the present invention can solve the routing problem of the low-earth orbit satellite network. The offloading decision-making strategy trained by the method provided by the present invention can make appropriate offloading decisions according to the current environment, and can solve the problems of low stability, poor environmental adaptability, and low decision-making quality existing in the traditional deep reinforcement learning model.
[0103] As Figure 5 shown, this embodiment discloses a satellite edge computing offloading system based on a deep reinforcement learning model for executing the above method, including the following modules:
[0104] State collection module: The deep reinforcement learning model collects the current satellite network state, obtains the reward returned by the environment and the state at the next moment;
[0105] Storage module: Stores the information obtained by the deep reinforcement learning model and the environment during the interaction in the form of experience tuples;
[0106] Training module: Randomly samples experience tuples for training the offloading decision-making strategy in the deep reinforcement learning model; during the training process, adjusts the neural network weights in the deep reinforcement learning model until the offloading decision-making strategy converges;
[0107] Decision module: The deep reinforcement learning model makes an optimal offloading decision according to the offloading decision-making strategy in combination with the current satellite network environment and task requirements.
[0108] For other contents of this embodiment, reference can be made to the above method embodiment.
[0109] A preferred embodiment of the present invention also discloses an electronic device, including: a processor and a memory, and the memory is used to store a program. When the program is called and executed by the processor, the processor executes the above method or system.
[0110] In summary, the present invention has the beneficial effect of optimizing the computing offloading decision of the satellite network. The present invention uses a method based on deep reinforcement and diffusion models to solve the computing offloading problem of the satellite network. By creating new high-quality content related to the context through the diffusion model, it can effectively reduce the need for exhaustive data collection and manual content creation required in the reinforcement learning training process. Secondly, by using the diffusion model to generate high-quality decisions, the performance of the offloading decision model in the satellite network can be significantly improved. Finally, considering the complexity of the inter-satellite dynamic environment, the present invention introduces the Lyapunov optimization framework. By transforming the offloading problem into an online optimization task for each time slot, the long-term stability of the satellite edge computing system can be maintained under the condition that the task distribution and environmental state are difficult to predict.
[0111] The foregoing has shown and described the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments, and what is described in the above embodiments and the specification is only to illustrate the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements fall within the scope of the present invention claimed. The scope of the present invention claimed is defined by the appended claims and their equivalents.
Claims
1. A satellite edge computing offloading method based on a deep reinforcement learning model, characterized in that: The following steps are involved: Step 1: The deep reinforcement learning model collects the current satellite network status and obtains the reward and next-moment status returned by the environment; Step 2: Store the information obtained by the deep reinforcement learning model and the environment during the interaction in step 1 in the form of experience tuples; Step 3: Randomly sample the experience tuples in step 2 to train the unloading decision strategy in the deep reinforcement learning model; During the training process, the neural network weights in the deep reinforcement learning model are adjusted until the unloading decision strategy converges; Step 4: The deep reinforcement learning model makes the optimal offloading decision based on the offloading decision strategy combined with the current satellite network environment and mission requirements.
2. The method according to claim 1, characterized in that In step 1, the deep reinforcement learning model is implemented using a soft actor-critic framework.
3. The method according to claim 2, characterized in that The soft actor-critic framework uses a generative diffusion model as the policy network.
4. The method according to any one of claims 1 to 3, characterized in that: In step three, the training process is as follows: (1) Initialize the strategy network θ, value network ψ, and value target network and Soft-Q network φ parameters; (2) The deep reinforcement learning model obtains the current environment state s; generates a preliminary decision ρ based on the initial strategy, obtains the reward r and the next moment environment state s′; (3) The replay buffer stores the experience tuple (s,ρ,s′,r); (4) Each round of training update randomly extracts a batch of experience from the replay buffer. For training, where K is the batch size; (5) According to the maximum entropy reinforcement learning process, the state value s i It is estimated that: in, represents the mathematical expectation, represents the policy network θ policy function, ρ new is the input i Actions generated by the post-policy network; i and ρ new Sent to the Soft-Q network to generate the state action value Q φ (s i ,ρ new ); Represents when the state s i As input, ρ new Follow the conditional probability distribution of the policy network; (6) Update the value network ψ through the loss function: (7) According to the Soft Bellman equation, we can obtain the i ,ρ i )’s action state value estimation: Among them, r i is the current reward, γ is the discount factor, s i+1 It is the state of the environment at the next moment. is the value target network input s i+1 Output after; Value target network The loss of the Soft-Q network φ is computed by periodically updating the parameters by soft copying from the value network ψ: Among them, α represents the learning rate; (8) Update the policy network θ through the loss function: Loss θ =logπ θ (r i |s i )-Q φ (s i ,r i )。 5. The method according to claim 4, characterized in that In step 3, the training process transforms the long-term stability problem into a time-slot online optimization task through the Lyapunov optimization framework, and the objective function is: Among them, B is the upper bound constant of the queue backlog, Q v (τ) is the backlog of the task queue of node v, y v (τ) is the change in queue backlog, W is the trade-off parameter between latency and stability, and t(τ) is the task latency.
6. The method according to claim 1, characterized in that In step 4, the unloading decision strategy is generated by the generative diffusion model in the deep reinforcement learning model, as follows: In each denoising step t, a deep neural network is used to infer and scale the denoising distribution tanh(∈ θ (x t ,t,s)),∈ θ (x t ,t,s) is a denoised noise generated by a deep neural network, where θ represents the parameters of the neural network, x t is the noise distribution at denoising step t, t is the current denoising step number, and s is the current environment state; Calculate the reverse transition distribution The mean of : Among them, α t =1-β t , β t is the noise variance at denoising step t, α t represents the noise retention ratio at denoising step t; Use the following update rule several times to obtain the distribution x t-1 , until the final output x0 of the diffusion process is obtained: in, represents random noise from a standard normal distribution; Finally, the Softmax function is applied to x0 to obtain the unloading decision ρ:
7. A satellite edge computing offloading system based on a deep reinforcement learning model, used to execute the method according to any one of claims 1 to 6, characterized in that: Includes the following modules: State collection module: The deep reinforcement learning model collects the current satellite network state and obtains the reward returned by the environment and the state at the next moment; Storage module: stores the information obtained by the deep reinforcement learning model and the environment during the interaction in the form of experience tuples; Training module: randomly sampling experience tuples to train the unloading decision strategy in the deep reinforcement learning model; During the training process, the neural network weights in the deep reinforcement learning model are adjusted until the unloading decision strategy converges; Decision-making module: The deep reinforcement learning model makes the optimal offloading decision based on the offloading decision strategy combined with the current satellite network environment and mission requirements.
8. An electronic device, characterized in that: include: A processor and a memory, the memory being used to store a program. When the program is called and executed by the processor, the processor executes the method according to any one of claims 1 to 6 or the system according to claim 7.