A Computing Offloading Method and Terminal for 6G Computing Power Network

By building computing power network models and deep learning algorithms, the flexibility and resource management problems of unloading strategies in multi-satellite collaboration environments are solved, and efficient computing unloading and user experience improvement are achieved.

CN119729631BActive Publication Date: 2025-07-11NANJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510233101.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-07-11
Estimated Expiration
2045-02-28

AI Technical Summary

Technical Problem

The existing technology is difficult to flexibly adjust the offloading strategy in a multi-satellite dynamic collaboration environment, fails to effectively model satellite mobility and its coverage time changes, and fails to comprehensively optimize system delay and energy consumption, making it difficult to meet the needs of 6G computing power network for efficient resource management.

Method used

Build a computing power network model based on user local equipment, low-Earth orbit satellites and remote cloud data centers. Use Shannon formula to calculate the transmission rate of the communication path, combine the delay and energy consumption models, establish a global optimization objective function, and use the deep learning model to generate the optimal unloading action to achieve flexible allocation of user computing tasks.

Benefits of technology

Implement efficient computing offloading in 6G computing power network, improve user computing task completion efficiency and resource utilization, balance delay and energy consumption, and improve user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119729631B_ABST
    Figure CN119729631B_ABST
Patent Text Reader

Abstract

The present invention discloses a computing offloading method and a terminal for a 6G computing power network in the field of computing power networks and computing offloading methods. The method includes: constructing a computing power network model based on user local devices, low-earth orbit satellites, and remote cloud data centers; determining the offloading ratio of user computing tasks; establishing a communication model and calculating the transmission rates of different communication paths according to the Shannon formula; calculating the latency and energy consumption for task processing according to the characteristics of user computing tasks and the transmission rates of different communication paths; establishing a latency and energy consumption model; taking the minimization of the user service quality perception cost in the process of the user interacting with the computing power network environment as the goal, establishing a global optimization objective function, solving the optimal solution of the global optimization objective function, and generating the optimal offloading actions for each user. The present invention realizes optimized computing offloading in a 6G computing power network, significantly improving the completion efficiency of user computing tasks, user experience, and system energy efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computing power networks and computing offloading methods, and particularly relates to a computing offloading method and a terminal for a 6G computing power network. Background Art

[0002] The rapid development of 6G communication technology in today's society has led to an increase in the demand for seamless communication coverage and efficient computing globally. As an important architecture for future 6G communication development, the Space-Ground Integrated Network (SGIN) can not only achieve communication coverage in remote areas (such as deserts, oceans, forests, and mountains) by integrating satellite and terrestrial networks, but also enhance communication resilience in extreme environments such as natural disasters. However, with the continuous progress of 5G and artificial intelligence technologies, applications in fields such as computer vision, natural language processing, autonomous driving, and intelligent wearable devices are becoming increasingly widespread, bringing a large number of computationally intensive tasks, which pose higher requirements for resources, storage, and energy consumption. When processing such tasks, users expect to have low-latency and high-efficiency computing support, especially users located in remote areas outside the coverage of terrestrial cellular networks. Therefore, the Space-Ground Integrated Computing Power Network (SGICPN) has emerged, which can meet the computing resource requirements of users in complex scenarios.

[0003] Currently, there are already some research methods for computing task offloading on the market, such as methods based on heuristic algorithms and methods of single-agent reinforcement learning. These methods have improved the efficiency of computing resource allocation to a certain extent. However, these methods have various deficiencies. Firstly, many methods are difficult to flexibly adjust offloading strategies in a multi-satellite dynamic cooperation environment. Secondly, the lack of effective modeling of satellite mobility and its coverage time changes limits their effectiveness in practical applications. Thirdly, in a multi-agent environment, traditional solutions fail to comprehensively optimize system latency and energy consumption and are difficult to meet the requirements of the computing power network for efficient resource management. In addition, the current research on the collaborative computing of satellite nodes, edge computing nodes, and remote cloud data centers is still relatively preliminary. Many methods only consider the computing offloading between local and a single satellite in a multi-satellite cooperation scenario and fail to effectively combine the computing power resources of remote cloud data centers, making it difficult to meet the requirements of large-scale computing tasks. Summary of the Invention

[0004] The purpose of the present invention is to overcome the deficiencies in the prior art and provide a computing offloading method and a terminal for a 6G computing power network, which can significantly improve the completion efficiency of user computing tasks and the user experience.

[0005] To achieve the above object, the present invention is implemented by the following technical solutions:

[0006] In a first aspect, the present invention provides a computing offloading method for a 6G computing power network, including:

[0007] Construct a computing power network model based on user local devices, low-Earth orbit satellites, and remote cloud data centers, and define the communication methods between user local devices, remote cloud data centers, and low-Earth orbit satellites according to the constructed computing power network model;

[0008] According to the offloading decision of the user, determine the offloading ratio of the user's computing task to be offloaded to the user local device, low-Earth orbit satellite, and remote cloud data center for computing;

[0009] Establish a communication model between the user local device, remote cloud data center, and low-Earth orbit satellite, and calculate the transmission rate of different communication paths according to Shannon's formula;

[0010] According to the characteristics of the user's computing task and the transmission rate of different communication paths, calculate the latency and energy consumption of the user local device, low-Earth orbit satellite, and remote cloud data center for processing the user's computing task respectively in combination with the communication method;

[0011] Combine the offloading ratio, latency, and energy consumption to establish a latency model and an energy consumption model;

[0012] Based on the remaining battery state of the user local device, assign weights to the latency model and the energy consumption model with the goal of reducing latency or energy consumption;

[0013] Establish a global optimization objective function with the goal of minimizing the user service quality perception cost during the interaction between the user and the computing power network environment;

[0014] Solve the optimal solution of the global optimization objective function to generate the optimal offloading action for each user.

[0015] Further, the computing power network model includes a satellite group network composed of N user local devices and M low-Earth orbit satellites, and a remote cloud data center;

[0016] The communication methods between the user local device, remote cloud data center, and low-Earth orbit satellite include:

[0017] Within a time interval T, there are a constant number of low-Earth orbit satellites that can communicate with the user local device, and the user local device directly utilizes the computing power resources of the low-Earth orbit satellites within its communication coverage;

[0018] The low Earth orbit satellite processes user computing tasks as a computing node or forwards user computing tasks to a remote cloud data center as a relay node; where N and are both integers greater than or equal to 1, <M, the calculation formula for the time interval T is as follows:

[0019] ;

[0020] ;

[0021] ;

[0022] In the formula: represents the central angle of the Earth corresponding to the coverage area of the low Earth orbit satellite; R e represents the radius of the Earth; H represents the orbital altitude of the low Earth orbit satellite; represents the elevation angle between the user's local device and the low Earth orbit satellite; L represents the orbital arc length passed by the low Earth orbit satellite during the coverage time; represents the running speed of the low Earth orbit satellite.

[0023] Furthermore, the offloading decision of the user is defined as:

[0024] ;

[0025] In the formula: represents the proportion of the user computing task offloaded to the user's local device for computing; represents the proportion of the user computing task offloaded to the low Earth orbit satellite for computing; represents the proportion of the user computing task offloaded to the remote cloud data center for computing; , and are subject to the following relational expressions:

[0026] { a n ( t ) + b n ( t ) + c n ( t ) = 1 a n ( t ), b n ( t ), c n ( t ) ∈ [ 0 , 1 ] .

[0027] Furthermore, the communication model between the user's local device and the low Earth orbit satellite is defined as:

[0028] ;

[0029] In the formula: represents the data transmission rate of the ground-air communication between the user's local device of the nth user and the mth low Earth orbit satellite; represents the transmission bandwidth between the user's local device and the low Earth orbit satellite; denotes the number of users who choose to offload their user computing tasks to the \(m\)th low-earth orbit satellite. Users offloading to the same low-earth orbit satellite share the bandwidth resources in an equal-sharing manner; denotes the transmission rate of the user's local device of the \(n\)th user; denotes the channel gain between the user's local device of the \(n\)th user and the \(m\)th low-earth orbit satellite at time \(t\); denotes the channel noise power;

[0030] The communication model between the low-earth orbit satellite and the remote cloud data center is defined as:

[0031] ;

[0032] In the formula: denotes the data transmission rate between the \(m\)th low-earth orbit satellite and the remote cloud data center; denotes the congestion coefficient; denotes the transmission bandwidth between the low-earth orbit satellite and the remote cloud data center; denotes the number of users who choose to offload their user computing tasks to the remote cloud data center; denotes the transmission rate of the \(m\)th low-earth orbit satellite; denotes the channel gain between the \(m\)th low-earth orbit satellite and the remote cloud data center at time \(t\).

[0033] Furthermore, the characteristics of the user computing task are defined as:

[0034] ;

[0035] In the formula: denotes the size of the data volume of the user computing task, in bytes; denotes the amount of computation required for CPU computing in the data volume, in CPU cycles; denotes the amount of computation required for GPU computing in the data volume, in GPU floating-point operation times; denotes the maximum tolerable delay time of the user; denotes the maximum tolerable energy consumption of the user;

[0036] Calculating the delay and energy consumption of the user's local device, the low-earth orbit satellite, and the remote cloud data center for processing the user computing task respectively according to the characteristics of the user computing task and the transmission rates of different communication paths, and combining the communication methods, includes:

[0037] When the user computing task is offloaded to the user's local device, only the computing delay of the user computing task is considered, and the formula for calculating the delay is as follows:

[0038] ;

[0039] Where: Denotes the delay required to complete the user's computing task when the user's computing task is offloaded to the local device of the nth user for computing at time t; Denotes the CPU computing power of the local device of the nth user; Denotes the GPU computing power of the local device of the nth user;

[0040] The calculation formula for energy consumption is as follows:

[0041] ;

[0042] Where: Denotes the energy consumption required to complete the user's computing task when the user's computing task is offloaded to the local device of the nth user for computing at time t; Denotes the capacitance parameter of the local device of the nth user; Denotes the dynamic power consumption of the GPU of the local device of the nth user;

[0043] When the user's computing task is offloaded to a low-Earth orbit satellite, the satellite-ground link propagation delay, data transmission delay, and computing delay of the user's computing task need to be considered. The calculation formula for the delay is as follows:

[0044] ;

[0045] Where: Denotes the delay required to complete the user's computing task when the user's computing task is offloaded to a low-Earth orbit satellite for computing at time t; Denotes the data transmission delay from the local device of the nth user to the mth low-Earth orbit satellite; Denotes the computing delay of the user's computing task; Denotes the CPU computing power allocated by the mth low-Earth orbit satellite to the user's computing task; Denotes the GPU computing power allocated by the mth low-Earth orbit satellite to the user's computing task; d represents the distance between the low-Earth orbit satellite and the local device of the user, ; Denotes the speed of light; Denotes the round-trip propagation delay from the Earth to the low-Earth orbit satellite;

[0046] The calculation formula for energy consumption is as follows:

[0047] ;

[0048] Where: Denotes the energy consumption required to complete the user's computing task when the user's computing task is offloaded to a low-earth orbit satellite for computing at time t; Denotes the transmission energy consumption of the user's local device of the nth user; Denotes the capacitance parameter of the mth low-earth orbit satellite; Denotes the GPU dynamic power consumption of the mth low-earth orbit satellite;

[0049] When the user's computing task is offloaded to a remote cloud data center, it is necessary to consider the data transmission delay when the data is relayed by a low-earth orbit satellite and sent to the remote cloud data center. The formula for the delay is as follows:

[0050] ;

[0051] In the formula: Denotes the delay required to complete the user's computing task when the user's computing task is offloaded to a remote cloud data center for computing at time t; Denotes the data transmission delay when the data is relayed by the mth low-earth orbit satellite and sent to the remote cloud data center; Denotes the round-trip propagation delay from the earth via the low-earth orbit satellite to the remote cloud data center;

[0052] The formula for the energy consumption is as follows:

[0053] ;

[0054] In the formula: Denotes the energy consumption required to complete the user's computing task when the user's computing task is offloaded to a remote cloud data center for computing at time t; Denotes the transmission energy consumption of the mth low-earth orbit satellite; Combining the offloading ratio, delay, and energy consumption, a delay model and an energy consumption model are established as follows:

[0055] ;

[0056] .

[0057] Furthermore, the user service quality perception cost is as follows:

[0058] ;

[0059] In the formula: Denotes the weight of the delay model ; Denotes the weight of the energy consumption model ; ω L , ω E ∈[0, 1] , and According to the remaining power of different user local devices, and The values of are set as follows:

[0060] ;

[0061] In the formula: represents the current battery power percentage of the user local device.

[0062] Furthermore, aiming to minimize the user service quality perception cost in the process of user interaction with the computing power network environment, a global optimization objective function is established, including:

[0063] By accumulating the user service quality perception cost at each moment in the process of user interaction with the computing power network environment, the global optimization objective function is as follows:

[0064] ;

[0065] In the formula: represents the termination moment of user interaction with the computing power network environment; represents The user service quality perception cost at time.

[0066] Furthermore, solving the optimal solution of the global optimization objective function to generate the optimal offloading action for each user, including:

[0067] Define a state set to describe all possible states of each user at time t; The state includes the user's computing task information, offloading decision, and the resource allocation and offloading situation of low Earth orbit satellites. The state set at time t is as follows:

[0068] ;

[0069] Among them,

[0070] ;

[0071] In the formula: represents the computing task information of each user at time t; represents the computing task information of the nth user at time t; represents the current offloading decision of each user at time t; represents the offloading decision of the nth user at time t; represents the resource allocation situation of each low Earth orbit satellite at time t; Indicates the resource allocation of the m-th low Earth orbit satellite at time t; Indicates the load conditions of each low Earth orbit satellite at time t; Indicates the load condition of the m-th low Earth orbit satellite at time t;

[0072] Wherein,

[0073] { I n ( t ) = { Task n , ε n } Ω m ( t ) = { A CPU m ( t ), A GPU m ( t ), F max LEO , TP max LEO , num n } A CPU m ( t ) = [ f 1 m , LEO , f 2 m , LEO , … , f num n m , LEO ] A GPU m ( t ) = [ TP 1 m , LEO , TP 2 m , LEO , … , TP num n m , LEO ] Ψ m ( t ) = { U CPU m ( t ), U GPU m ( t ), N queue m ( t )} ;

[0074] In the formula: Indicates the current battery power percentage of the user's local device of the n-th user; And Respectively indicate the CPU computing resources and GPU computing resources allocated by the m-th low Earth orbit satellite to the user; Indicates the number of users unloaded to the m-th low Earth orbit satellite; Indicates the m-th low Earth orbit satellite allocated to the CPU computing power of the user computing task of the n-th user; Indicates the m-th low Earth orbit satellite allocated to the GPU computing power of the user computing task of the n-th user; Indicates the current CPU usage rate of the m-th low Earth orbit satellite; Indicates the current GPU usage rate of the m-th low Earth orbit satellite; Indicates the number of user computing tasks currently queued in the m-th low Earth orbit satellite;

[0075] Define the offloading action set, which is used to describe the offloading actions taken by each user in the state at time t. The offloading action set at time t Is shown in the following formula:

[0076] ;

[0077] Wherein,

[0078] ;

[0079] In the formula: Indicates the offloading action of the n-th user at time t; Indicates the target low Earth orbit satellite for offloading;

[0080] The offloading actions taken by each user, the resource allocation of each low Earth orbit satellite, and the load conditions of each low Earth orbit satellite are constrained by the following conditional expressions:

[0081] ;

[0082] ;

[0083] ;

[0084] ;

[0085] ;

[0086] ;

[0087] In the formula: represents the delay required for a single user to complete the user computing task; represents the energy consumption required for a single user to complete the user computing task; represents the maximum CPU computing power of the low Earth orbit satellite; represents the maximum GPU computing power of the low Earth orbit satellite;

[0088] Define the reward function by combining the state set and the offloading action set, which is used to describe the immediate reward obtained by each user taking the offloading action in the said state , and the reward function is defined as:

[0089]

[0090] In the formula: represents the penalty factor;

[0091] Train the deep learning model using the said state set, offloading action set and reward function to generate the optimal offloading action for each user.

[0092] Furthermore, the training of the deep learning model using the said state set, offloading action set and reward function includes:

[0093] Initialize the state set , offloading action set , experience replay pool , batch size , discount factor , soft update coefficient ;

[0094] Initialize the Actor network , and make it consistent with the action sets of each user under the initial state set ; Among them, under the initial state set , each user only selects to offload the user computing task to the user's local device for computing, and the bandwidth resources, CPU and GPU computing resources of the low Earth orbit satellite and the remote cloud data center are not allocated; ​

[0095] Initialize two Critic networks and and the corresponding target networks, where the target networks include a target Actor network and two target Critic networks ;

[0096] Among them, the Actor network is used to generate an offloading action set according to the current state set ; , ; and are the Q-values calculated by the current two Critic networks, indicating the long-term rewards that can be obtained by taking the offloading action set under the state set ; represents the offloading action generated by the target Actor network; and respectively represent the Q-values calculated by the two target Critic networks;

[0097] Initialize the number of training rounds, the number of training steps per round, and the delayed update step size Z. The number of training steps is the total number of time steps executed in one round of training;

[0098] At the beginning of each round of training, reset the state set to the initial state set of the environment , and let ;

[0099] In each time step, first each user selects an initial offloading action set according to the policy of the current Actor network . Add Gaussian noise to the initial offloading action set to obtain the actual offloading action set :

[0100] ;

[0101] In the formula: represents Gaussian noise with a mean of 0 and a variance of ;

[0102] Then apply the actual offloading action set to the computing power network environment, and record the current state set , the offloading action set , the immediate reward , the next state set and the boolean variable The value of , is false if the current state set is the terminal state set; otherwise, it is true, indicating that there is a next state set;

[0103] Store the experience generated at each time step in the form into the experience replay pool , and update the current state set, making ; Allocate priorities to each experience according to the temporal difference TD-error, and select the most important experience for training the Actor and Critic networks. The TD-error of each experience is calculated from the value calculated by the current Critic network and the target value, as shown in the following formula:

[0104] ;

[0105] where, ;

[0106] In the formula: represents the temporal difference TD-error of the j-th experience; represents the target value obtained based on the offloading action with added Gaussian noise; and are the values of the next state set calculated by the target network, representing future rewards; min represents taking the minimum value to avoid overestimation of the value; is the discount factor, used to balance the current immediate reward and future rewards, representing the weight of future rewards;

[0107] Calculate the sampling probability of each experience according to the TD-error, as shown in the following formula:

[0108] ;

[0109] In the formula: represents the sampling probability of the j-th experience; represents the hyperparameter of prioritized experience replay, controlling the influence of TD-error on the sampling probability; k represents the number of all experiences in the experience replay pool ; represents the sum of the absolute values of the temporal difference TD-errors of all experiences in the experience replay pool after the

[0110] After sampling B pieces of data from the experience replay pool according to the priority, optimize the loss function of the Critic network, and use the following loss function to update the parameters of the Critic network and such that the output by the Critic network is closer to the target value :

[0111] ;

[0112] In the formula: represents the average over the batch size B, reducing the impact of a single piece of data on the update; represents the importance sampling weight, used to correct the bias brought by priority sampling;

[0113] If the current training step is an integer multiple of the delayed update step size Z, then use the following policy gradient method to update the parameters of the Actor network to select the optimal offloading action to maximize the value calculated by the Critic network:

[0114] ;

[0115] In the formula: represents the policy gradient of the Actor network; represents the expectation; represents sampling from the experience replay pool ; represents the gradient with respect to the offloading action set; represents the derivative of the Critic network with respect to the offloading action set, used to guide the Actor network to learn better offloading actions; represents the gradient with respect to the parameters of the Actor network ; represents the gradient representation of the offloading action set with respect to the parameter ;

[0116] After each round of training is completed, update the parameters of the target network based on the following soft update formula and so that it gradually approaches the parameters of the current Actor network and Critic network , and , achieving smooth adjustment of the target network parameters;

[0117] ;

[0118] When all training rounds are completed, a trained deep learning model is obtained.

[0119] In a second aspect, the present invention provides an electronic terminal, including a processor and a memory connected to the processor. A computer program is stored in the memory. When the computer program is executed by the processor, the steps of the above-mentioned computing offloading method for a 6G computing power network are executed.

[0120] Compared with the prior art, the beneficial effects achieved by the present invention are as follows:

[0121] By specifically analyzing the characteristics of user computing tasks, modeling communication models, latency models, and energy consumption models, the present invention realizes efficient computing offloading in a 6G computing power network. Especially in remote areas where ground base stations cannot provide coverage, users can complete their computing tasks through the collaborative work of low-earth orbit satellites and remote cloud data centers, thus significantly improving the completion efficiency of user computing tasks, as well as the flexibility of the computing power network and the utilization rate of computing resources. At the same time, with the user service quality perception cost as the optimization goal and comprehensively considering key indicators such as the remaining battery power of the user's local device, the present invention can effectively balance the latency and energy consumption during the offloading process of user computing tasks, greatly enhancing the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0122] Figure 1 is a flowchart of the computing offloading method for a 6G computing power network provided in Embodiment 1 of the present invention;

[0123] Figure 2 is a computing power network model diagram provided in Embodiment 1 of the present invention;

[0124] Figure 3 is a satellite coverage model diagram provided in Embodiment 1 of the present invention;

[0125] Figure 4 is an algorithm block diagram of the reinforcement learning algorithm PER-MATD3 that combines the prioritized experience replay mechanism and the multi-agent twin-delayed deep deterministic policy gradient provided in Embodiment 1 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0126] The technical solutions of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific features in the embodiments of the present application and the embodiments are detailed descriptions of the technical solutions of the present application, rather than limitations on the technical solutions of the present application. Without conflict, the technical features in the embodiments of the present application and the embodiments can be combined with each other.

[0127] Embodiment 1:

[0128] Figure 1It is a flowchart of a computing offloading method for a 6G computing power network in Embodiment 1 of the present invention. This flowchart only shows the logical order of the method described in this embodiment. On the premise of non-conflict, in other possible embodiments of the present invention, the steps shown or described can be completed in a different Figure 1 order from that shown. Refer to Figure 1 , the method of this embodiment specifically includes the following steps:

[0129] Construct a computing power network model based on user local devices, low Earth orbit satellites, and remote cloud data centers, and define the communication methods between user local devices, remote cloud data centers, and low Earth orbit satellites according to the constructed computing power network model;

[0130] According to the offloading decision of the user, determine the offloading ratio of the user's computing tasks to be offloaded to the user local device, low Earth orbit satellite, and remote cloud data center for computing;

[0131] Establish a communication model between user local devices, remote cloud data centers, and low Earth orbit satellites, and calculate the transmission rates of different communication paths according to Shannon's formula;

[0132] According to the characteristics of the user's computing tasks and the transmission rates of different communication paths, calculate the latency and energy consumption of the user local device, low Earth orbit satellite, and remote cloud data center for processing the user's computing tasks respectively in combination with the communication methods;

[0133] Combine the offloading ratio, latency, and energy consumption to establish a latency model and an energy consumption model;

[0134] Based on the remaining battery status of the user local device, assign weights to the latency model and the energy consumption model with the goal of reducing latency or energy consumption;

[0135] With the goal of minimizing the user service quality perception cost during the interaction between the user and the computing power network environment, establish a global optimization objective function;

[0136] Solve the optimal solution of the global optimization objective function to generate the optimal offloading actions for each user.

[0137] Specifically, the computing power network model is as Figure 2 shown, including a satellite group network composed of N user local devices and M low Earth orbit satellites, and a remote cloud data center;

[0138] The communication methods between the user local device and the low Earth orbit satellite, and between the low Earth orbit satellite and the remote cloud data center include:

[0139] Within the time interval T, there is a constant A number of low Earth orbit satellites can communicate with user local devices, and the user local devices can directly utilize the computing power resources of the low Earth orbit satellites within their communication coverage; the low Earth orbit satellites act as computing nodes to process the user's computing tasks or as relay nodes to forward the user's computing tasks to a remote cloud data center; where N = {1, 2, …, n}, , <M. The satellite coverage model is as Figure 3 shown. In the figure, R e represents the radius of the Earth; H represents the orbital altitude of the low Earth orbit satellite; d represents the distance between the low Earth orbit satellite and the user local device; represents the elevation angle between the user local device and the low Earth orbit satellite. When > 0, the user local device can communicate with the low Earth orbit satellite; represents the central angle of the Earth corresponding to the coverage area of the low Earth orbit satellite. According to the relevant parameters provided by the satellite coverage model, the calculation time interval T is calculated as shown in the following formula:

[0140] ;

[0141] ;

[0142] ;

[0143] In the formula: L represents the orbital arc length passed by the low Earth orbit satellite during the coverage time; represents the running speed of the low Earth orbit satellite. As Figure 2 shown, some users are located in remote areas such as deserts, oceans, forests or mountainous areas, etc. These areas usually lack ground base station coverage and cannot directly communicate with the remote cloud data center. In this case, users need to rely on communicable low Earth orbit satellites for relay to offload user computing tasks to the remote cloud data center for calculation. The low Earth orbit satellites act as bridges. Through efficient communication links, even in an environment with imperfect ground infrastructure, users can still access powerful remote computing resources to complete user computing tasks.

[0144] Each user can choose to offload the user computing task to the user local device, the low Earth orbit satellite or the remote cloud data center for calculation. The offloading decision of the user is defined as:

[0145] ;

[0146] In the formula: represents the proportion of the user computing task offloaded to the user local device for calculation; represents the proportion of the user computing task offloaded to the low Earth orbit satellite for calculation; represents the proportion of the user's computing tasks offloaded to the remote cloud data center for computing; and and are constrained by the following relational expressions:

[0147] { a n ( t ) + b n ( t ) + c n ( t ) = 1 a n ( t ), b n ( t ), c n ( t ) ∈ [ 0 , 1 ] .

[0148] A single user can determine the offloading proportion of their user computing tasks according to their needs. For example, if the user computing tasks are fully offloaded to the user's local device for computing, the offloading proportion of the user computing tasks is 100% allocated to the user's local device. At this time takes a value of 100%, and take values of 0; if the user computing tasks are split into multiple parts, they can be offloaded to the user's local device, low Earth orbit satellites, and remote cloud data centers proportionally. For example, 30% of a single user's computing tasks are allocated to the user's local device for computing, 40% are allocated to low Earth orbit satellites for computing, and 30% are allocated to the remote cloud data center for computing. At this time , and take values of 30%, 40%, and 30% respectively. This flexible offloading method supports the multi-mode collaborative computing of user computing tasks among the user's local device, low Earth orbit satellites, and remote data centers, thus improving the overall utilization rate of computing resources.

[0149] According to Shannon's formula, calculate the data transmission rate of the ground-air communication between the user's local device and the low Earth orbit satellite and the data transmission rate

[0150] of the m-th low Earth orbit satellite and the remote cloud data center as follows:

[0151] In the formula: represents the data transmission rate of the ground-air communication between the user's local device of the n-th user and the m-th low Earth orbit satellite; represents the transmission bandwidth between the user's local device and the low Earth orbit satellite; represents the number of users who choose to offload their user computing tasks to the m-th low Earth orbit satellite. Users offloading to the same low Earth orbit satellite share the bandwidth resources equally; represents the transmission rate of the user's local device of the n-th user; represents the channel gain between the user's local device of the n-th user and the m-th low Earth orbit satellite at time t; represents the channel noise power.

[0152] The communication model between low-earth orbit satellites and remote cloud data centers is defined as:

[0153] ;

[0154] Where: represents the data transfer rate between the m-th low-earth orbit satellite and the remote cloud data center; represents the congestion coefficient, ; represents the transmission bandwidth between the low-earth orbit satellite and the remote cloud data center; represents the number of users who choose to offload their computing tasks to the remote cloud data center; represents the transmission rate of the m-th low-earth orbit satellite; represents the channel gain between the m-th low-earth orbit satellite and the remote cloud data center at time t.

[0155] After obtaining the transmission rates and of different communication paths, then analyze the characteristics of the user computing task. The characteristics of the user computing task include the data volume, CPU and GPU computing volume, the maximum latency acceptable to the user, and the energy consumption limit. The characteristics of the user computing task are defined as:

[0156] ;

[0157] Where: represents the size of the data volume of the user computing task, in bytes; represents the amount of computing required for CPU computing in the data volume, in CPU cycles; represents the amount of computing required for GPU computing in the data volume, in GPU floating-point operation times; represents the longest latency time tolerated by the user; represents the maximum energy consumption tolerated by the user.

[0158] According to the characteristics of the user computing task, combined with the communication method between the user's local device and the low-earth orbit satellite, and the communication method between the low-earth orbit satellite and the remote cloud data center, calculate the latency and energy consumption of the user's local device, low-earth orbit satellite, and remote cloud data center for processing the user computing task respectively.

[0159] When the user computing task is offloaded to the user's local device, no additional data transmission is required, so the data transmission latency is not considered, and only the computing latency of the user computing task is considered. The formula for calculating the latency is shown as follows:

[0160] ;

[0161] In the formula: represents the delay required to complete the user computing task when the user computing task is offloaded to the user's local device of the nth user at time t; represents the CPU computing power of the user's local device of the nth user; represents the GPU computing power of the user's local device of the nth user;

[0162] The calculation formula of energy consumption is shown in the following formula:

[0163] ;

[0164] In the formula: represents the energy consumption required to complete the user computing task when the user computing task is offloaded to the user's local device of the nth user at time t; represents the capacitance parameter of the user's local device of the nth user; represents the dynamic power consumption of the GPU of the user's local device of the nth user;

[0165] When the user computing task is offloaded to a low Earth orbit satellite, it is necessary to comprehensively consider the round-trip propagation delay, data transmission delay, and computing delay of the user computing task from the Earth to the low Earth orbit satellite. Since the computing result of the user computing task is much smaller than the input data, the data transmission delay only considers the delay from the user's local device of the nth user to the mth low Earth orbit satellite, and ignores the delay of returning the computing result from the mth low Earth orbit satellite to the user's local device. The calculation formula of the delay is shown in the following formula:

[0166] ;

[0167] In the formula: represents the delay required to complete the user computing task when the user computing task is offloaded to the low Earth orbit satellite for computing at time t; represents the data transmission delay of data from the user's local device of the nth user to the mth low Earth orbit satellite; represents the computing delay of the user computing task; represents the CPU computing power allocated by the mth low Earth orbit satellite to the user computing task; represents the GPU computing power allocated by the mth low Earth orbit satellite to the user computing task; d represents the distance between the low Earth orbit satellite and the user's local device;

[0168] ;

[0169] represents the speed of light; Denotes the round-trip propagation delay from the Earth to a low-Earth orbit satellite;

[0170] The calculation formula for energy consumption is shown as follows:

[0171] ;

[0172] In the formula: Denotes the energy consumption required to complete the user's computing task when the user's computing task is offloaded to the m-th low-Earth orbit satellite for computing at time t; Denotes the transmission energy consumption of the user's local device of the n-th user; Denotes the capacitance parameter of the m-th low-Earth orbit satellite; Denotes the dynamic power consumption of the GPU of the m-th low-Earth orbit satellite;

[0173] When the user's computing task is offloaded to a remote cloud data center, due to the lack of a ground base station in the user's area and the inability to directly communicate with the remote cloud data center, the user needs to relay through a communicable low-Earth orbit satellite to offload the user's computing task to the remote cloud data center. In this process, it is necessary to consider the data transmission delay of the user's computing task relayed through the low-Earth orbit satellite to the remote cloud data center. Since the remote cloud data center has rich computing resources, the computing delay of the user's computing task is negligible. The calculation formula for the delay is shown as follows:

[0174] ;

[0175] In the formula: Denotes the delay required to complete the user's computing task when the user's computing task is offloaded to the remote cloud data center for computing at time t; Denotes the data transmission delay of the data relayed by the m-th low-Earth orbit satellite and sent to the remote cloud data center; Denotes the round-trip propagation delay from the Earth via the low-Earth orbit satellite to the remote cloud data center;

[0176] The calculation formula for energy consumption is shown as follows:

[0177] ;

[0178] In the formula: Denotes the energy consumption required to complete the user's computing task when the user's computing task is offloaded to the remote cloud data center for computing at time t; Denotes the transmission energy consumption of the m-th low-Earth orbit satellite.

[0179] Combined with the offloading ratio, delay and energy consumption, a delay model and an energy consumption model are established as follows:

[0180] ;

[0181] ;

[0182] wherein: represents the total delay required to complete the user computing tasks of N users; represents the total energy consumption required to complete the user computing tasks of N users.

[0183] Based on the remaining battery level of the user's local device, with the goal of reducing latency or energy consumption, weights are assigned to the latency model and the energy consumption model; and the user quality of service perception cost is defined, and the user quality of service perception cost at time t is shown as follows:

[0184] ;

[0185] wherein: represents the weight of the latency model ; represents the weight of the energy consumption model ; ω L , ω E ∈[0, 1] , and ; According to the remaining battery levels of different user local devices, and are set as shown in the following formula:

[0186] ;

[0187] wherein: represents the current battery percentage of the user's local device.

[0188] Through this weight assignment, when the battery of the user's local device is fully charged, the offloading strategy will tend to reduce the computing latency to provide a faster response; while when the battery level is low, reducing energy consumption will be given priority to extend the usage duration of the user's local device. This flexible strategy can ensure that the user's experience requirements are always met according to different battery levels.

[0189] The service quality perception cost during the interaction between the user and the computing power network environment is the cumulative value of the user service quality perception costs at multiple moments. By aggregating the user service quality perception costs at each moment during the interaction between the user and the environment, the service quality perception cost is obtained. With the goal of minimizing the user service quality perception cost, a global optimization objective function is established, and the optimal solution of the global optimization objective function is solved. A deep learning model is used to generate the optimal offloading action for each user, that is, the optimal offloading strategy. The deep learning model combines the prioritized experience replay mechanism and the reinforcement learning algorithm PER-MATD3 (Prioritized Experience Replay - Multi-Agent Twin Delayed Deep Deterministic Policy Gradient) of multi-agent twin delayed deep deterministic policy gradient.

[0190] By accumulating the user service quality perception costs at each moment during the interaction between the user and the environment, the global optimization objective function is as shown in the following formula:

[0191] ;

[0192] In the formula: represents the termination moment of the interaction between the user and the computing power network environment; represents the user service quality perception cost at the moment of

[0193] In the multi-user computing task offloading scenario, if the optimization objective is established only by combining the latency and energy consumption of a single user, the agents in the reinforcement learning algorithm of multi-agent twin delayed deep deterministic policy gradient may tend to select the offloading strategy that is optimal for a single user. For example, the user computing tasks are concentratedly offloaded to the remote cloud data center. Although the remote cloud data center has strong computing power and the latency and energy consumption of a single user may be reduced, as the user computing tasks are concentratedly offloaded, the resources of the remote cloud data center are gradually overloaded, resulting in a decline in the overall performance of the computing power network, and the offloading effect of the user computing tasks of other users will also become worse. This locally optimal strategy not only reduces the utilization efficiency of the computing power network resources but also may increase the service quality perception cost of some users.

[0194] Therefore, in the present invention, by calculating the total latency and total energy consumption of all users, the optimization objective of the multi-agent deep reinforcement learning algorithm is defined by combining the total latency and total energy consumption of all users, guiding each agent to jointly explore a reasonable offloading strategy. Through global optimization, the agents can balance the resource allocation among the user local devices, low-earth orbit satellites, and remote cloud data centers, making the overall user service quality perception cost of the computing power network reach the optimal, avoiding the problem of uneven resource utilization, and improving the overall performance of the computing power network.

[0195] The process of solving the global optimization objective function is transformed into a Markov decision process, the state set, action set, and reward function are defined, and the optimization objective is linked to the user's offloading action through the reward function. By transforming the process of solving the global optimization objective function into a Markov decision process, the present invention can dynamically adapt to changes in the user's computing tasks and communication conditions, realize real-time adjustment of the offloading strategy, and ensure excellent performance in a changing computing power network environment.

[0196] First, define the state set to describe all possible states of each user at time t; the state includes the user's computing task information, offloading decision, and the resource allocation and offloading situation of low-earth orbit satellites, and the state set at time t is shown as follows:

[0197] ;

[0198] where

[0199] ;

[0200] In the formula: represents the user's computing task information of each user at time t; represents the user's computing task information of the nth user at time t; represents the current offloading decision of each user at time t; represents the offloading decision of the nth user at time t; represents the resource allocation situation of each low-earth orbit satellite at time t; represents the resource allocation situation of the mth low-earth orbit satellite at time t; represents the load situation of each low-earth orbit satellite at time t; represents the load situation of the mth low-earth orbit satellite at time t;

[0201] where

[0202] { I n ( t ) = { Task n , ε n } Ω m ( t ) = { A CPU m ( t ), A GPU m ( t ), F max LEO , TP max LEO , num n } A CPU m ( t ) = [ f 1 m , LEO , f 2 m , LEO , … , f num n m , LEO ] A GPU m ( t ) = [ TP 1 m , LEO , TP 2 m , LEO , … , TP num n m , LEO ] Ψ m ( t ) = { U CPU m ( t ), U GPU m ( t ), N queue m ( t )} ;

[0203] In the formula: represents the current battery power percentage of the user's local device of the nth user; and respectively represent the CPU computing resources and GPU computing resources allocated by the mth low-earth orbit satellite to the user; represents the number of users offloaded to the mth low-earth orbit satellite; represents the mth low-earth orbit satellite allocated to the th user's CPU computing power for the user's computing task; represents the GPU computing power of the m-th low Earth orbit satellite assigned to the user computing task of the n-th user; represents the current CPU utilization rate of the m-th low Earth orbit satellite; represents the current GPU utilization rate of the m-th low Earth orbit satellite; represents the number of user computing tasks currently queued in the m-th low Earth orbit satellite;

[0204] Next, define the offloading action set, which is used to describe the offloading actions taken by each user in the state at time t. The offloading action set at time t is shown as follows:

[0205] ;

[0206] where,

[0207] ;

[0208] In the formula: represents the offloading action of the n-th user at time t; represents the target low Earth orbit satellite for offloading;

[0209] The offloading actions taken by each user, the resource allocation of each low Earth orbit satellite, and the load of each low Earth orbit satellite are constrained by the following conditional expressions:

[0210] ;

[0211] ;

[0212] ;

[0213] ;

[0214] ;

[0215] ;

[0216] In the formula: represents the delay required for a single user to complete the user computing task; represents the energy consumption required for a single user to complete the user computing task; represents the maximum CPU computing power of the low Earth orbit satellite; represents the maximum GPU computing power of the low Earth orbit satellite; The conditional expressions and ensure that the delay and energy consumption of the offloading actions selected by the user are less than or equal to the maximum delay time And the maximum energy consumption ; Conditional And Ensure that the CPU computing power allocated by the low-earth orbit satellite to the user's computing tasks And the GPU computing power Is less than or equal to the maximum CPU computing power of the low-earth orbit satellite And the maximum GPU computing power ; Conditional And Ensure that the sum of the CPU computing power allocated by the low-earth orbit satellite to each user And the sum of the GPU computing power Is less than or equal to the sum of the maximum CPU computing power of the low-earth orbit satellite And the sum of the maximum GPU computing power ;

[0217] Define the reward function by combining the state set and the offloading action set, which is used to describe the immediate reward obtained by each user taking the offloading action in the state. The reward function is defined as:

[0218] ;

[0219] In the formula: Represents the penalty factor. When the offloading action of the user causes any of the following situations, take <0: When the user's computing task exceeds the longest delay time tolerated by the user, exceeds the maximum energy consumption limit, or fails to successfully match the target low-earth orbit satellite; otherwise, >0.

[0220] Train the deep learning model using the state set, offloading action set, and reward function, including:

[0221] Adopt the reinforcement learning algorithm of multi-agent double-delay deep deterministic policy gradient, sample data from the experience replay pool according to the priority through the prioritized experience replay mechanism, optimize the Q-value evaluation accuracy based on the double Critic network, guide the policy update of the Actor network with the target Q-value, and synchronize the target network parameters combined with the soft update mechanism, so that the reward function gradually increases, thereby reducing the user's perceived cost of service quality and generating the optimal offloading action for each user.

[0222] Such as Figure 4As shown in the figure, it is a block diagram of the reinforcement learning algorithm PER-MATD3, which combines the priority experience replay mechanism and the multi-agent double-delay deep deterministic policy gradient. In the figure, the agent is the decision maker in the reinforcement learning system, responsible for observing the state of the environment, selecting the unloading action, and adjusting its unloading strategy according to the environmental feedback. In this embodiment, the agent is the user, and the agent simulates the unloading decision of each user to decide whether to unload the user's computing task to the user's local device, low-Earth orbit satellite or remote cloud data center. Specifically, first initialize the state set , Uninstall action set , Experience Replay Pool , batch size , discount factor , soft update coefficient ;

[0223] Initializing the Actor Network , so that it is consistent with the initial state set The action set of each user Guaranteed consistency; among them, the initial state set In this case, each user only chooses to offload the user computing task to the user's local device for computing, and the bandwidth resources, CPU and GPU computing resources of the low-Earth orbit satellite and remote cloud data center are not allocated;

[0224] Initialize two Critic networks and And the corresponding target network, the target network includes the target Actor network And two target critic networks ;

[0225] Among them, Actor network Used to collect data based on the current state Generate uninstall action set , ; and is the Q value calculated by the two Critic networks, indicating that in the state set Take the following uninstall action set the long-term rewards that can be obtained; Represents the target Actor network generated unloading action; and Respectively represent the Q values ​​calculated by the two target Critic networks;

[0226] Initialize the number of training rounds episode, the number of training steps per round step, and the delay update step Z, where the number of training steps step is the total number of time steps executed in one round of training;

[0227] At the start of each round of training, reset the state set as the initial state set of the environment , and let ;

[0228] In each time step, first each user selects an initial set of offloading actions according to the policy of the current Actor network , and obtains the actual set of offloading actions by adding Gaussian noise to the initial set of offloading actions :

[0229] ;

[0230] In the formula: represents Gaussian noise with a mean of 0 and a variance of ;

[0231] Then apply the actual set of offloading actions to the computing power network environment, and record the current state set , the set of offloading actions , the immediate reward , the next state set and the value of the boolean variable . If the current state set is the termination state set , takes the value false; otherwise, takes the value true, indicating that there is a next state set;

[0232] Store the experience generated in each time step in the form of in the experience replay pool , and update the current state set, letting ; Allocate priorities to each experience according to the temporal difference error TD-error, and select the most important experience for training the Actor and Critic networks. The TD-error of each experience is calculated from the difference between the value calculated by the current Critic network and the target value, as shown in the following formula:

[0233] ;

[0234] Among them, ;

[0235] In the formula: represents the temporal difference error TD-error of the j-th experience; represents the target obtained based on the offloading actions with added Gaussian noise Value; and is the value of the next state set calculated by the target network, representing future rewards; min represents taking the minimum value to avoid overestimation of the value; is the discount factor, used to balance the current immediate reward and future rewards, representing the weight of future rewards;

[0236] The sampling probability of each experience is calculated according to the TD-error, as shown in the following formula:

[0237] ;

[0238] In the formula: represents the sampling probability of the j-th experience; represents the hyperparameter of the prioritized experience replay, controlling the influence of the TD-error on the sampling probability; k represents the number of all experiences in the experience replay pool ; represents the sum of the TD-error absolute values of all experiences in the experience replay pool after the power operation;

[0239] After sampling B data from the experience replay pool according to the priority, the loss function of the Critic network is optimized, and the parameters of the Critic network are updated through the following loss function to make the and output by the Critic network closer to the target value :

[0240] ;

[0241] In the formula: represents averaging over the batch size B to reduce the influence of a single data on the update; represents the importance sampling weight, used to correct the bias brought by the prioritized sampling;

[0242] ;

[0243] In the formula: represents the policy gradient of the Actor network; represents the expectation; represents sampling from the experience replay pool ; represents the gradient with respect to the offloading action set; Denotes the derivative of the Critic network with respect to the set of offloading actions, which is used to guide the Actor network to learn better offloading actions; Denotes the parameters of the Actor network Gradient; Denotes the gradient representation of the set of offloading actions with respect to the parameter ;

[0244] After each round of training is completed, update the parameters of the target network based on the following soft update formula and , so that it gradually approaches the parameters of the current Actor network and Critic network , and , realizing smooth adjustment of the target network parameters;

[0245] ;

[0246] After all training rounds are completed, a trained deep learning model is obtained. The offloading actions generated by the Actor network can maximize the value of the Critic network in different states, thereby gradually increasing the reward function , reducing the user's perceived cost of service quality , achieving minimization of the user's perceived cost of service quality in the process of interacting with the computing power network environment , and outputting the optimal offloading actions for each user.

[0247] Through analyzing the characteristics of user computing tasks, efficiently modeling communication, delay, and energy consumption models, the embodiments of the present invention achieve optimized computing offloading in the 6G computing power network, significantly improving the efficiency of user computing task completion, user experience, and system energy efficiency, and realizing dynamic adaptability and global optimal strategies through multi-mode collaborative computing and multi-agent deep reinforcement learning algorithms.

[0248] Embodiment 2:

[0249] The embodiments of the present invention also provide an electronic terminal, including a processor and a memory connected to the processor. A computer program is stored in the memory. When the computer program is executed by the processor, the steps of the computing offloading method for the 6G computing power network described in Embodiment 1 above are executed.

[0250] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0251] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0252] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0253] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0254] The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.

Claims

1. A computing offloading method for a 6G computing power network, characterized in that, Including: Construct a computing power network model based on the user's local device, low Earth orbit satellite, and remote cloud data center, and define the communication methods between the user's local device, remote cloud data center, and low Earth orbit satellite according to the constructed computing power network model; According to the user's offloading decision, determine the offloading ratio of the user's computing task to the user's local device, low Earth orbit satellite, and remote cloud data center for computing; Establish a communication model between the user's local device, remote cloud data center, and low Earth orbit satellite, and calculate the transmission rate of different communication paths according to Shannon's formula; According to the characteristics of the user's computing task and the transmission rate of different communication paths, calculate the latency and energy consumption of the user's local device, low Earth orbit satellite, and remote cloud data center for processing the user's computing task respectively in combination with the communication method; Combine the offloading ratio, latency, and energy consumption to establish a latency model and an energy consumption model; Based on the remaining battery level of the user's local device, assign weights to the latency model and the energy consumption model with the goal of reducing latency or energy consumption; With the goal of minimizing the user service quality perception cost in the process of the user interacting with the computing power network environment, establish a global optimization objective function; Solve the optimal solution of the global optimization objective function to generate the optimal offloading action for each user; Among them, the computing power network model includes a satellite group network composed of N user local devices and M low Earth orbit satellites, and a remote cloud data center; The communication methods between the user's local device, remote cloud data center, and low Earth orbit satellite include: Within the time interval T, there is a constant number of low-Earth orbit satellites that can communicate with the user's local device, and the user's local device directly utilizes the computing power resources of the low-Earth orbit satellites within its communication coverage area; The low Earth orbit satellite acts as a computing node to process the user's computing task, or acts as a relay node to forward the user's computing task to the remote cloud data center; where N and are both integers greater than or equal to 1, <M, the calculation formula for the time interval T is as follows: ; ; ; Where: L represents the orbital arc length passed by the low-earth orbit satellite during the coverage time; represents the operating speed of the low-earth orbit satellite; R e represents the radius of the earth; H represents the orbital altitude of the low-earth orbit satellite; represents the central angle of the earth corresponding to the coverage area of the low-earth orbit satellite; represents the elevation angle between the user's local device and the low-earth orbit satellite; The uninstallation decision of the user is defined as: ; Wherein: represents the proportion of the user's computing tasks unloaded to the user's local device for computing; represents the proportion of the user's computing tasks unloaded to the low-Earth orbit satellite for computing; represents the proportion of the user's computing tasks unloaded to the remote cloud data center for computing; 、 and are subject to the following relational expressions: ; The characteristics of the user computing task are defined as: ; Wherein: represents the size of the data volume of the user's computing task, in bytes; represents the amount of computation required for CPU computation in the data volume, in CPU cycles; represents the amount of computation required for GPU computation in the data volume, in GPU floating-point operation counts; represents the longest latency time tolerated by the user; represents the maximum energy consumption tolerated by the user.

2. The computing offloading method for a 6G computing power network according to claim 1, wherein The communication model between the user's local device and the low Earth orbit satellite is defined as: ; Wherein: represents the data transmission rate of the terrestrial-air communication between the user local device of the nth user and the mth low-earth orbit satellite; represents the transmission bandwidth between the user local device and the low-earth orbit satellite; represents the number of users who choose to offload their user computing tasks to the mth low-earth orbit satellite, and the users offloading to the same low-earth orbit satellite share the bandwidth resources in an equal-sharing manner; represents the transmission rate of the user local device of the nth user; represents the channel gain between the user local device of the nth user and the mth low-earth orbit satellite at time t; represents the channel noise power; The communication model between the low Earth orbit satellite and the remote cloud data center is defined as: ; Wherein: represents the data transmission rate between the m-th low-earth orbit satellite and the remote cloud data center; represents the congestion coefficient; represents the transmission bandwidth between the low-earth orbit satellite and the remote cloud data center; represents the number of users who choose to offload their computing tasks to the remote cloud data center; represents the transmission rate of the m-th low-earth orbit satellite; represents the channel gain between the m-th low-earth orbit satellite and the remote cloud data center at time t.

3. The computing offloading method for a 6G computing power network according to claim 2, wherein, The calculating the latency and energy consumption of the user's local device, low Earth orbit satellite, and remote cloud data center for processing the user's computing task respectively according to the characteristics of the user's computing task and the transmission rate of different communication paths in combination with the communication method includes: When the user's computing task is offloaded to the user's local device, only consider the computing latency of the user's computing task, and the formula for calculating the latency is shown as follows: ; In the formula: represents the latency required to complete the user's computing task when the user's computing task at time t is offloaded to the local device of the nth user for computing; represents the CPU computing power of the local device of the nth user; represents the GPU computing power of the local device of the nth user; The formula for calculating the energy consumption is shown as follows: ; Wherein: represents the energy consumption required to complete the user's computing task when the user's computing task at time t is offloaded to the local device of the nth user for computing; represents the capacitance parameter of the local device of the nth user; represents the dynamic power consumption of the GPU of the local device of the nth user; When the user's computing task is offloaded to the low Earth orbit satellite, it is necessary to consider the space-ground link propagation latency, data transmission latency, and the computing latency of the user's computing task. The formula for calculating the latency is shown as follows: ; Wherein: represents the time delay required to complete the user's computing task when the user's computing task is offloaded to a low-earth orbit satellite at time t; represents the data transmission delay from the user's local device of the nth user to the mth low-earth orbit satellite; represents the computing delay of the user's computing task; represents the CPU computing power allocated by the mth low-earth orbit satellite to the user's computing task; represents the GPU computing power allocated by the mth low-earth orbit satellite to the user's computing task; d represents the distance between the low-earth orbit satellite and the user's local device, , represents the speed of light; represents the round-trip propagation delay from the earth to the low-earth orbit satellite; The formula for calculating the energy consumption is shown as follows: ; In the formula: represents the energy consumption required to complete the user's computing task when the user's computing task is offloaded to the m-th low-Earth orbit satellite at time t; represents the transmission energy consumption of the user's local device of the n-th user; represents the capacitance parameter of the m-th low-Earth orbit satellite; represents the dynamic power consumption of the GPU of the m-th low-Earth orbit satellite; When the user's computing task is offloaded to the remote cloud data center, it is necessary to consider the data transmission latency of the data sent to the remote cloud data center through the relay of the low Earth orbit satellite. The formula for calculating the latency is shown as follows: ; Wherein: represents the latency required to complete the user's computing task when the user's computing task is offloaded to the remote cloud data center at time t; represents the data transmission latency when the data is relayed by the m-th low-earth orbit satellite and sent to the remote cloud data center; represents the round-trip propagation latency from the earth via the low-earth orbit satellite to the remote cloud data center; The formula for calculating the energy consumption is shown as follows: ; Wherein: represents the energy consumption required to complete the user's computing task when the user's computing task is offloaded to the remote cloud data center for computing at time t; represents the transmission energy consumption of the m-th low-earth orbit satellite; Combining the offloading ratio, delay, and energy consumption, a delay model and an energy consumption model are established as shown in the following formula: ; 。 4. The computing offloading method for the 6G computing power network according to claim 3, characterized in that The perceived cost of user service quality As shown in the following formula: ; wherein: represents the weight of the time delay model ; represents the weight of the energy consumption model ; , and ; according to the remaining power of different user local devices, and are set as shown in the following formula: ; In the formula: represents the current battery power percentage of the user's local device.

5. The computing offloading method for a 6G computing power network according to claim 4, wherein, With the goal of minimizing the user service quality perception cost in the process of the user interacting with the computing power network environment, establishing a global optimization objective function includes: By accumulating the user service quality perception cost at each moment in the process of the user interacting with the computing power network environment, the global optimization objective function is obtained as shown in the following formula: ; In the formula: represents the termination moment when the user interacts with the computing power network environment; represents the perceived cost of the user's service quality at the moment.

6. The computing offloading method for a 6G computing power network according to claim 5, wherein Solve the optimal solution of the global optimization objective function to generate the optimal offloading actions for each user, including: Define the state set, which is used to describe all possible states of each user at time t; the states include the user's computing task information, offloading decision, and the resource allocation and offloading situation of low Earth orbit satellites. The state set at time t is shown as follows: ; wherein, ; Wherein: represents the user computing task information of each user at time t; represents the user computing task information of the nth user at time t; represents the current offloading decision of each user at time t; represents the offloading decision of the nth user at time t; represents the resource allocation of each low-earth orbit satellite at time t; represents the resource allocation of the mth low-earth orbit satellite at time t; represents the load of each low-earth orbit satellite at time t; represents the load of the mth low-earth orbit satellite at time t; wherein, ; In the formula: represents the current battery power percentage of the user's local device of the nth user; and respectively represent the CPU computing resources and GPU computing resources allocated by the mth low Earth orbit satellite to the user; represents the number of users offloaded to the mth low Earth orbit satellite; represents the CPU computing power of the user computing task allocated by the mth low Earth orbit satellite to the th user; represents the GPU computing power of the user computing task allocated by the mth low Earth orbit satellite to the th user; represents the current CPU utilization rate of the mth low Earth orbit satellite; represents the current GPU utilization rate of the mth low Earth orbit satellite; represents the number of user computing tasks currently queued in the mth low Earth orbit satellite; Define the offloading action set, which is used to describe the offloading actions taken by each user in the state at time t. The offloading action set at time t As shown in the following formula: ; wherein, ; In the formula: represents the offloading action of the nth user at time t; represents the target low Earth orbit satellite for offloading; The offloading actions taken by each user, the resource allocation of each low-earth orbit satellite, and the load of each low-earth orbit satellite are constrained by the following conditional expressions: ; ; ; ; ; ; Wherein: represents the latency required for a single user to complete a user computing task; represents the energy consumption required for a single user to complete a user computing task; represents the maximum CPU computing power of a low-earth orbit satellite; represents the maximum GPU computing power of a low-earth orbit satellite; Define a reward function by combining the state set and the offloading action set, which is used to describe the immediate reward obtained by each user for taking the offloading action in the state , and the reward function is defined as: ; In the formula: represents the penalty factor; Use the state set, offloading action set, and reward function to train the deep learning model to generate the optimal offloading actions for each user.

7. The computing offloading method for a 6G computing power network according to claim 6, wherein The training of the deep learning model using the state set, offloading action set, and reward function includes: Initial state set , Uninstallation action set , Experience replay pool , Batch size , Discount factor , Soft update coefficient ; Initialize the Actor network to ensure consistency with the offloading actions of each user under the initial state set ; among them, under the initial state set , each user only selects to offload the user computing task to the user's local device for computing, and the bandwidth resources, CPU, and GPU computing resources of the low-earth orbit satellite and the remote cloud data center are not allocated; Initialize two Critic networks and the corresponding target networks, where the target networks include a target Actor network and two target Critic networks ; Among them, the Actor network is used to generate an offloading action set according to the current state set ; , ; and are the Q-values calculated by the current two Critic networks, indicating the long-term rewards that can be obtained by taking the offloading action set under the state set ; represents the offloading action set generated by the target Actor network; and respectively represent the Q-values calculated by the two target Critic networks; Initialize the number of training rounds, the number of steps per round of training, and the delayed update step size Z. The number of training steps is the total number of time steps executed in one round of training; At the start of each round of training, reset the set of states as the set of initial states of the environment , and let ; At each time step, first each user selects an initial set of offloading actions according to the policy of the current Actor network and obtains the actual set of offloading actions by adding Gaussian noise to the initial set of offloading actions through the following formula : ​​ ; In the formula: represents Gaussian noise with a mean of 0 and a variance of ; Then apply the actual offloading action set to the computing power network environment and record the current state set , the offloading action set , the immediate reward , the next state set and the value of the boolean variable . If the current state set is the termination state set , the value is false; otherwise the value is true, indicating that there is a next state set; The experience generated at each time step is The experience replay pool is stored in the form of , and update the current state set, let ; Assign priority to each experience according to the time series error TD-error, select the most important experience for the Actor and Critic network update for training, and the TD-error of each experience is calculated by the current Critic network Values ​​and Goals The difference between the values ​​is calculated as follows: ; Among them, ; Wherein: represents the temporal difference TD-error of the j-th experience; represents the target obtained based on the offloading action with added Gaussian noise value; and are the values of the next state set calculated by the target network, representing future rewards; min represents taking the minimum value to avoid overestimation of the value; is the discount factor, used to balance the current immediate reward and future rewards, representing the weight of future rewards; Calculate the sampling probability of each piece of experience according to the TD-error, as shown in the following formula: ; where: represents the sampling probability of the j-th experience; represents the hyperparameter of prioritized experience replay, controlling the influence of TD-error on the sampling probability; k represents the number of all experiences in the experience replay pool ; represents the sum of the -th power of the absolute value of the temporal difference error TD-error of all experiences in the experience replay pool after the power operation; After sampling B pieces of data from the experience replay pool according to the priority, optimize the loss function of the Critic network with the following loss function Update the parameters of the Critic network and , making the output by the Critic network closer to the target value : ; In the formula: represents the average over the batch size B, reducing the impact of individual data on the update; represents the importance sampling weight, used to correct the bias brought by priority sampling; If the current training step is an integer multiple of the delayed update step size Z, the parameters of the Actor network are updated using the following policy gradient method , select the optimal offloading action to maximize the value calculated by the Critic network ; In the formula: represents the policy gradient of the Actor network; represents the expectation; represents sampling from the experience replay pool ; represents the gradient with respect to the set of offloading actions; represents the derivative of the Critic network with respect to the set of offloading actions, which is used to guide the Actor network to learn better offloading actions; represents the gradient with respect to the parameters of the Actor network ; represents the gradient representation of the set of offloading actions with respect to the parameter ; After each round of training is completed, update the parameters of the target network based on the following soft update formula and , so that it gradually approaches the parameters of the current Actor network and Critic network , and , and achieve smooth adjustment of the target network parameters; ; When all training rounds are completed, a trained deep learning model is obtained.

8. An electronic terminal, characterized in that, It includes a processor and a memory connected to the processor. A computer program is stored in the memory. When the computer program is executed by the processor, the steps of the computing offloading method for a 6G computing power network according to any one of claims 1 to 7 are executed.

Citation Information

Patent Citations

  • Mobile edge computing unloading method and system based on Harris eagle algorithm

    CN116932086A

  • Satellite mobile edge computing unloading decision-making method based on deep reinforcement learning

    CN117579126A