A multi-target dynamic preference satellite-ground collaborative computing offloading method for a low earth orbit satellite constellation

By constructing a multi-objective dynamic preference satellite-ground collaborative computing offloading model through deep reinforcement learning, the problem of user preference changes in low-Earth orbit satellite constellations is solved, the computing offloading performance of low-Earth orbit satellite edge networks is improved, and the dynamic needs of users in terms of latency and energy consumption are met.

CN116566466BActive Publication Date: 2025-12-30SHANXI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310510405.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-08
Publication Date
2025-12-30
Estimated Expiration
2043-05-08

AI Technical Summary

Technical Problem

Existing methods for calculating offloading from low-Earth orbit satellite constellations fail to effectively consider the dynamic changes in users' preferences for different targets, resulting in an inability to meet users' dynamic needs in terms of latency and energy consumption, and making it difficult to adapt to dynamic and complex network environments.

Method used

A multi-objective dynamic preference satellite-ground collaborative computing offloading model is constructed using deep reinforcement learning. By minimizing the total cost, the optimal task execution order is designed. Combining the dynamic preference requirements of user latency and energy consumption, an optimization function with multiple users and multiple edge nodes is constructed, and the optimal offloading strategy is solved using deep reinforcement learning.

Benefits of technology

It improves the performance of computational offloading in low-Earth orbit satellite edge networks, meets users' dynamic preferences in terms of latency and energy consumption, enhances computational efficiency and flexibility, and adapts to dynamic and complex network environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116566466B_ABST
    Figure CN116566466B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of multi-target dynamic preference star-ground collaborative computing unloading methods for low-orbit satellite constellation, comprising the following contents: for dynamic preference task computing unloading problem in low-orbit satellite constellation, under the constraint conditions such as sufficient consideration to user dynamic preference demand, low-orbit satellite limited computing capacity and star-ground communication time, the dynamic preference task multi-objective star-ground collaborative computing unloading optimization function of multiple users and multiple edge nodes is built;According to the model created, the computing unloading problem is described as Markov decision process, and multi-objective reinforcement learning is used to solve unloading decision;Comprehensively consider multiple optimization goals and the preference demand of task time-varying, design composite reward function, obtain optimal unloading decision.Compared with traditional solving method, the present application can adapt to dynamic complex network environment, consider the different preference demand of multiple goals, and can generate optimal unloading decision under different requirements using trained model, improve solving efficiency and flexibility, can be widely used in low-orbit satellite constellation edge computing environment for computing unloading.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of computational offloading in mobile communication technology, and relates to edge computing networks for low-Earth orbit satellite constellations, particularly a multi-objective dynamic preference satellite-ground collaborative computational offloading method for low-Earth orbit satellite constellations. Background Technology

[0002] Elon Musk's SpaceX Starlink project is the world's most famous low-Earth orbit (LEO) satellite constellation plan. Once fully deployed, Starlink will provide high-bandwidth, low-latency network services globally, offering users worldwide communication, navigation, and remote sensing services. Clearly, LEO satellite constellations have enormous market potential and economic value in the foreseeable future. my country has also begun developing its own LEO satellite constellation and has made some progress. Space-ground collaborative computing is the fundamental mode for LEO satellite constellations to process data and execute tasks. How to rationally and efficiently transfer some computing tasks from ground equipment to LEO satellites in space-ground collaborative computing scenarios (i.e., computation offloading technology) is a hot research topic in this field.

[0003] Current research on computational offloading in low-Earth orbit (LEO) satellite constellations mainly includes traditional optimization methods, machine learning-based optimization methods, and reinforcement learning-based optimization methods. Traditional optimization methods primarily include convex optimization, heuristic algorithms, and game theory, but these methods are ill-suited to dynamic and complex network environments. When facing multi-user, multi-edge-node computational offloading decision-making problems, traditional algorithms lack sufficient search breadth and are prone to getting trapped in local optima; furthermore, they are time-consuming when processing large-scale data, making them unsuitable for the time-sensitive requirements of network data. Deep learning methods, on the other hand, require a large amount of sample data for training, and the generation of this data requires manual acquisition; the scale and quality of the data directly affect the learning performance. Since deep reinforcement learning can interact with the environment in real time and make optimal decisions through trial and error, it is more suitable for dynamic task requests and wireless channel communication environments. Therefore, computational offloading methods based on deep reinforcement learning (DRL) can be effectively utilized in multi-user, multi-edge-node LEO satellite constellations.

[0004] However, existing reinforcement learning methods for studying computational offloading in low-Earth orbit (LEO) satellite missions often neglect multi-objective problems. Most studies focus solely on minimizing latency or energy consumption, and some fail to consider the dynamic preferences of ground terminal equipment regarding mission computation latency and energy consumption when making comprehensive trade-offs. They rely solely on converting multi-objective weights into a single objective to build the model, and ignore the time constraints on user communication caused by the high-speed operation of LEO satellites. Since users have different preferences for different objectives at different times, it is difficult to determine appropriate weights, making it difficult for these methods to meet user needs. Summary of the Invention

[0005] The purpose of this invention is to overcome the shortcomings of the prior art. Considering the time-varying channel state and the randomness of the task, a multi-objective dynamic preference satellite-ground collaborative computational offloading method for low-Earth orbit satellite constellations is proposed. This method solves the problem that existing computational offloading methods for low-Earth orbit satellite constellations ignore the user's preference changes for different objectives, and effectively improves the computational offloading performance of low-Earth orbit satellite edge networks.

[0006] A multi-objective dynamic preference satellite-ground collaborative computational offloading method for low-Earth orbit satellite constellations is proposed. Based on the principle of minimizing total cost and according to a deep reinforcement learning strategy, the optimal task execution order is designed. The specific steps are as follows:

[0007] Step 1: Construct a dynamic preference task calculation model, communication model, and satellite-to-ground communication timing model for low-Earth orbit satellite constellations;

[0008] Step 2: Under the constraints of the user task's maximum reception latency, the maximum communication time between satellite and ground, the user's dynamic preference requirements for latency and energy consumption, and the limited computing power of low-orbit satellites, construct a dynamic preference task multi-objective satellite-ground cooperative computing offloading optimization function for multiple users and multiple edge nodes.

[0009] Step 3: Use deep reinforcement learning to solve the multi-objective optimization function of the low-orbit satellite-mobile edge computing system. The solution method is as follows: describe the computation offloading problem as a multi-objective Markov decision process (S, A, P, r, w), and transform it into the problem of solving the optimal computation offloading control strategy, where S represents the state space, A represents the action space, P represents the state transition probability, r is the reward, and w represents the preference space storing different preference values.

[0010] Step 4: Initialize the current network, target network, and experience pool size of the deep reinforcement learning model, and generate preference vectors using a random method to assign current user preference values ​​to the two objectives of latency and energy consumption;

[0011] Step 5: Select samples from the experience pool using random sampling to train the deep reinforcement learning algorithm;

[0012] Step 6: Obtain the system state and preferences for the current time slot, input the system state into the trained deep reinforcement learning algorithm, and use the trained deep reinforcement learning algorithm to obtain the task offloading decision for each time slot.

[0013] The task latency T of the local computing model in step 1 i L It can be represented as:

[0014]

[0015] Where S i This represents the number of CPU cycles required to complete task i on the terminal device. This indicates the local computing power of the terminal device.

[0016] The task energy consumption of the local computing model can be expressed as: It can be represented as:

[0017]

[0018] Where P i L T represents the computing power of the terminal device. i L This indicates the time required to calculate the task;

[0019] Task latency of the task unloading computation model It can be represented as:

[0020]

[0021] Where D i V represents the size (in bits) of the task data for user i. ij S is the uplink transmission rate for calculating the offloading of ground terminal equipment i and LEO satellite j via the satellite-to-ground link. i The number of CPU cycles required to complete task i The computing power allocated to terminal task i by LEO satellite j, where c represents the speed of light, and d ij The distance between ground terminal device i and LEO satellite j can be expressed as:

[0022]

[0023] Task energy consumption of the task unloading computation model It can be represented as:

[0024]

[0025] Where P i up The uplink transmission power P of task i generated by the ground terminal equipment. i L The computing power for local users.

[0026] Longest communication time in low-Earth orbit satellite coverage time model It can be represented as:

[0027]

[0028] Where V S For the LEO satellite's operating speed, Lij The arc length corresponding to the communication between ground terminal device i and LEO satellite j can be expressed as:

[0029] L ij =2(R+h)·γ ij

[0030] Where R is the Earth's radius, h is the distance between the ground terminal and the low-Earth orbit satellite, and γ ij The geocentric angle of low-orbit satellite j relative to the coverage area of ​​ground terminal equipment i can be expressed as:

[0031]

[0032] Where θ ij Let the elevation angle between ground terminal equipment i and low-Earth orbit satellite j be expressed as:

[0033]

[0034] Where d ij The distance between ground terminal i and low-orbit satellite j;

[0035] In step 2, the total latency L1(a) of all users' corresponding uninstallation decisions under different latency preference requirements i ,,b i ,k i This can be represented as:

[0036]

[0037] Where a i =0 indicates that task i is not unloaded and performs local computation, a i =1 indicates that task i is offloaded to a low-Earth orbit satellite server for computation; b i =M indicates that task i is offloaded to the Mth satellite edge node via the satellite-to-ground link, k i =K indicates that task i selects the Kth channel to transmit data, W i d This represents the latency preference value of ground terminal device i.

[0038] The total energy consumption L2(a) of all users' corresponding offloading decisions under different energy consumption preferences i ,,b i ,k i This can be represented as:

[0039]

[0040] Among them W i e This represents the energy consumption preference value of ground terminal device i.

[0041] The objective function is then a multi-objective optimization problem with two tasks:

[0042]

[0043] C1:

[0044] C2:

[0045] C3:

[0046] C4:

[0047] C5:

[0048] Constraints C1 and C2 respectively indicate that the completion time of terminal device task i must not exceed the maximum coverage time of satellite j under the offloading decision. And the maximum tolerance time T of task i i max C3 represents the computing resources allocated to all users by the low-Earth orbit satellite j, not exceeding its own maximum computing power F. j max C4 represents the latency and energy consumption preferences of ground terminal device i; C5 represents the offloading decision of ground terminal device i, which indicates whether to perform computation offloading, which low-Earth orbit satellite server to perform computation, and which channel to select for data transmission.

[0049] The aforementioned multi-objective optimization problem is mapped to a deep reinforcement learning problem. A multi-objective Markov decision process can be represented by a tuple (S, A, P, r, w), namely, the state space S, the action space A, the state transition probability P, the reward r, and the preference space w. In state s, the reinforcement learning algorithm selects action a to be executed, obtaining an optimal computational unloading control policy π* to maximize the long-term cumulative reward obtained during the device movement process.

[0050] The state space of a certain time slot is defined as follows: Where f i L (t) represents the local computing power of terminal user i at time slot t, D i I(t) represents the amount of task data generated by terminal user i in time slot t, and I(t) represents the priority sorting list of terminal devices in time slot t. denoted by k(t), which represents the computational power of low-Earth orbit satellite j in time slot t, and k(t) represents the channel transmission situation in time slot t.

[0051] The action space of a certain time slot is defined as A = {a i ,b i ,k i}, where ai =0 indicates that task i is not unloaded and performs local computation, a i =1 indicates that task i is offloaded to a low-Earth orbit satellite server for computation; b i =M indicates that task i is offloaded to the Mth satellite edge node via the satellite-to-ground link, k i =K indicates that task i selects the Kth channel to transmit data.

[0052] The composite reward function is set as follows:

[0053]

[0054] in r i d and r i e These represent user latency reward and energy consumption reward under different preferences, respectively. and Let z represent the latency and energy consumption preferences of terminal i, respectively, where z is a negative number.

[0055] The training method for the deep reinforcement learning model is as follows:

[0056] Step 1: Initialize the parameters of the main network and the target network in deep reinforcement learning, construct an experience pool of size N, set the number of iterations, and initialize the state to lay the foundation for the training process;

[0057] Step 2: Based on the system state and user preferences of the current time slot, input them into the main network to obtain the Q-value vector. Combine the current deep neural network parameters and use the inner product of the preference and the Q-vector as the selection criterion. Use the ε-greedy greedy strategy to decide the action a under the current state s, and calculate the immediate reward obtained by taking the decision action under the current state and the next state s′.

[0058] Step 3: Store the obtained system state, system actions, immediate rewards, preferences, and system state of the next time slot in the experience pool;

[0059] Step 4: Randomly select a portion of the experience buffer as experience samples for training, minimize the loss function according to gradient descent and update the main network parameters, and synchronize the main network parameters to the target network every 300 iterations.

[0060] Step 5: Train the system to maximize rewards and obtain the optimal unloading decision. The optimal unloading decision is obtained when the long-term cumulative reward reaches its optimum after the iterations are completed.

[0061] The advantages and positive effects of this invention are:

[0062] 1. This invention addresses the computational offloading problem for low-Earth orbit (LEO) satellite missions. Under the constraints of user dynamic preference requirements, limited computational capabilities of LEO satellites, and satellite-to-ground communication time, it constructs a multi-user, multi-edge-node dynamic preference mission multi-objective satellite-to-ground collaborative computational offloading optimization function.

[0063] 2. In this invention, the multi-objective optimization problem is modeled as a multi-objective Markov decision process. Unlike the traditional Markov decision process, the multi-objective Markov decision process extends the Q value into a vector form, where each element corresponds to an objective. Multiple objectives are optimized simultaneously, and the weights are dynamically adjusted to meet different user preferences.

[0064] 3. This invention employs a multi-objective reinforcement learning method to solve the edge computing offloading problem of low-Earth orbit satellites, seeking the optimal offloading strategy under the dynamic preferences of users, minimizing multiple objectives such as latency and energy consumption of task computation in the preference space, thereby meeting user needs and improving the computational efficiency of the satellite system.

[0065] 4. This invention comprehensively considers multiple optimization objectives and time-varying preference requirements of the task, designs a composite reward function, and obtains the optimal offloading decision. Simultaneously, the algorithm stores the current user's preference and experience sequences, and learns from the preference and experience pool to better learn the optimal strategy. Therefore, it can meet the user's constantly changing preferences, obtain the optimal solution that meets the user's needs, improve solution efficiency and flexibility, and can be widely used for computational offloading in low-Earth orbit satellite edge computing environments. Attached Figure Description

[0066] Figure 1 This is a schematic diagram of the process of the present invention.

[0067] Figure 2 This is a time delay comparison diagram between the present invention and other methods.

[0068] Figure 3 This is a comparison chart of energy consumption between the present invention and other methods.

[0069] Figure 4 This is a comparison chart of the rewards of this invention and other methods. Detailed Implementation

[0070] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0071] The purpose of this invention is to overcome the shortcomings of existing technologies. Considering the time-varying channel states and the randomness of tasks, it proposes a multi-objective dynamic preference-based satellite-ground collaborative computational offloading method for low-Earth orbit (LEO) satellite constellations. This method addresses the shortcomings of existing LEO satellite constellation computational offloading methods, which ignore user preferences for different objectives and cannot adapt to dynamic and complex network problems, thereby further improving the computational offloading performance of LEO satellite edge networks. Figure 1 As shown, the implementation steps are as follows:

[0072] Step 1: Construct a low-Earth orbit satellite constellation network system and its various models.

[0073] To achieve multi-objective reinforcement learning computation offloading for dynamic preference tasks in low-Earth orbit (LEO) satellite constellation networks, this step constructs a LEO-assisted MEC system. Each LEO satellite carries an MEC server, enabling users to offload tasks within its coverage area. The computation offloading problem is modeled as a multi-objective optimization problem, incorporating LEO satellite coverage time constraints, aiming to simultaneously minimize multiple objectives, including latency and energy consumption, of the LEO-MEC system.

[0074] This step models multiple objectives (latency and energy consumption) in the LEO-MEC environment. The specific method is as follows: This invention considers a multi-user, multi-access point MEC system and uses Orthogonal Frequency Division Multiple Access (OFDMA) wireless communication technology to construct a low-Earth orbit satellite edge computing offloading model. The system includes M LEO satellites carrying MEC servers, N ground mobile terminal devices, and K orthogonal subcarriers, which can be represented as M = {1, 2, ..., M}, N = {1, 2, ..., N}, and K = {1, 2, ..., K}, respectively.

[0075] Furthermore, assuming that each mobile terminal user has only one computational task to perform, and that the computational task cannot be divided, the task i generated by the terminal device can be represented by a triple W. i ={D i ,S i ,T i max} represent the data size (bits) of task i, the number of CPU cycles required to complete the task, and the maximum tolerable latency to complete the task, respectively.

[0076] The end user makes an intelligent migration decision based on the generated task information and the availability of computing and communication resources in the edge network. The offloading strategy of terminal device i in time slot t is A = [a i ,b i ,k i ], where a i b1 ∈ {0,1}, b1 ∈ {0,1,K,M}, k i ∈{1,K,K}. Where a i=0 indicates that task i is not unloaded and performs local computation, a i =1 indicates that task i is offloaded to a low-Earth orbit satellite server for computation; b i =M indicates that task i is offloaded to the Mth satellite edge node via the satellite-to-ground link, k i =K indicates that task i selects the Kth channel to transmit data.

[0077] The following sections will explain the communication model, local calculation model, offloading calculation model, and low-Earth orbit satellite coverage model for calculating offloading.

[0078] Low Earth Orbit (LEO) satellite communication model: Assuming each ground terminal device can only transmit mission data offload to one LEO satellite, multiple ground terminal devices share the same spectrum resources, and different terminal devices have different uplink and downlink rates when offloading mission calculations to different edge nodes, the uplink rate of mission i generated by the terminal device to edge node j via the satellite-to-ground link is as follows:

[0079]

[0080] Among them B ij The available spectrum bandwidth allocated by low-Earth orbit satellite edge node j to user i, g ij P represents the channel power gain from ground terminal i to low-Earth orbit satellite edge node j. i up Let σ be the transmission power of user task i in the uplink. 2 This represents the power of Gaussian white noise.

[0081] Furthermore, considering that the size of the calculated result is much smaller than the size of the input data, and that the downlink transmission rate of the LEO satellite is much greater than the uplink rate of the ground user, this invention ignores the downlink transmission delay caused by the LEO satellite transmitting the calculated result to the ground user.

[0082] Local computing model: In a local computing scheme, the completion time and energy consumption of a computing task depend only on the local computing power f. i L (CPU cycles / s) is related. Therefore, the task computation execution time T of the ground terminal device i is related to the CPU cycles / s. i L It can be represented as:

[0083]

[0084] Where S i This represents the number of CPU cycles required to complete task i on the terminal device. This indicates the local computing power of the terminal device.

[0085] Energy consumption cost of task execution calculated locally by ground terminal device i It can be represented as:

[0086]

[0087] Where P i L This indicates the computing power of the terminal device.

[0088] Task offloading computation model: In the offloading computation scheme, the computational task of ground terminal device i is offloaded to the satellite edge node via the satellite-to-ground link. Let D... i V represents the size (in bits) of the task data for user i. ij S is the uplink transmission rate for calculating the offloading of ground terminal equipment i and LEO satellite j via the satellite-to-ground link. i The number of CPU cycles required to complete task i The computing power (CPU cycles / s) allocated to terminal task i for LEO satellite j, d ij Let represent the distance between the ground terminal equipment and the low Earth orbit (LEO) satellite, and 'c' represent the speed of light. Due to the relatively long distance between the ground user and the LEO satellite, the ground user experiences propagation delays when communicating with the LEO satellite. Therefore, the computation execution time for task i to be offloaded from the ground terminal equipment to the LEO satellite j is [not specified]. It consists of three parts: transmission delay, computation delay, and propagation delay. The time required for task offloading is:

[0089]

[0090] Among them (D) i / V ij The uplink transmission delay for task i generated by the ground terminal equipment to send data to the LEO satellite. The computational latency of task i generated for ground terminal equipment on LEO satellite j, (d ij / c) Propagation delay between ground terminal equipment and LEO satellite. Where d ij Let i be the distance between ground terminal device i and LEO satellite j, and it is expressed as:

[0091]

[0092] In addition, the energy consumption of task i generated by the ground terminal equipment when unloading computation to LEO satellite j It consists of two parts: the energy consumption for task transmission and the energy consumption for maintaining the operation of local equipment during task computation and propagation, which can be expressed as:

[0093]

[0094] Where P i up The uplink transmission power P of task i generated by the ground terminal equipment. i L The computing power for local users.

[0095] Low Earth Orbit (LEO) Satellite Coverage Time Model: Considering the high-speed motion of LEO satellites, the communication time for ground terminal users during mission unloading is limited by the LEO satellite coverage time. Due to the high-speed rotation of LEO satellites, their positions dynamically change over time. Therefore, ground terminal users cannot communicate with LEO satellites at any time; interaction is only possible when the LEO satellites cover the user's location. θ ij Let i be the elevation angle between ground terminal equipment i and low-Earth orbit satellite j. The formula for the elevation angle is:

[0096]

[0097] Where h is the distance between the ground terminal and the low-Earth orbit satellite, R is the Earth's radius, and d ij γ is the distance between the ground terminal and the low-Earth orbit satellite. ij It is the geocentric angle of the low-orbit satellite j relative to the coverage area of ​​the ground terminal equipment i, and it is expressed as:

[0098]

[0099] set up Let be the longest communication time between ground user i and low-Earth orbit satellite j. Then the longest communication time is:

[0100]

[0101] Where V S For the LEO satellite's operating speed, L ij Let L be the arc length corresponding to the communication between ground terminal device i and LEO satellite j. The arc length is expressed as: L ij =2(R+h)·γ ij .

[0102] Step 2: Construct a dynamic preference task multi-target satellite-ground collaborative computing offloading optimization function based on the edge computing model.

[0103] In this invention, we consider latency and energy consumption as two core indicators for evaluating network performance in optimizing computation offloading decisions. While satisfying the limited computing power and coverage time constraints of low-Earth orbit satellites, we aim to minimize the total latency and energy consumption of all users' tasks. The objective function is a multi-objective optimization problem with two tasks:

[0104]

[0105] in This represents minimizing the total latency of all users' corresponding uninstallation decisions under different preference requirements. Minimize the total energy consumption of all users' corresponding offload decisions under different preference requirements, W i d and W i e These represent the preferred requirements of ground terminal device i for latency and energy consumption, respectively.

[0106] In addition, in minimizing system latency and energy consumption, we also need to follow the following constraints:

[0107] Constraint 1: The completion time of terminal device task i must not exceed the maximum coverage time of satellite j under the offloading decision.

[0108] Constraint 2: The completion time of task i on the terminal device must not exceed the maximum tolerance time T of task i under the offloading decision. i max ;

[0109] Constraint 3: The computing resources allocated to all users by low-Earth orbit satellite j shall not exceed its own maximum computing power F. j max ;

[0110] Constraint 4: Ground terminal equipment i's preferred requirements for latency and energy consumption;

[0111] Constraint 5: Represents the offloading decision for ground terminal device i, indicating whether to perform computational offloading, which low-Earth orbit satellite server to use for computation, and which channel to select for data transmission.

[0112] The constrained optimization problem described above can be expressed as:

[0113]

[0114] C1:

[0115] C2:

[0116] C3:

[0117] C4:

[0118] C5:

[0119] Constraints C1 and C2 respectively indicate that the completion time of terminal device task i must not exceed the maximum coverage time of satellite j under the offloading decision. And the maximum tolerance time T of task i i max C3 represents the computing resources allocated to all users by the low-Earth orbit satellite j, not exceeding its own maximum computing power F. j max C4 represents the latency and energy consumption preferences of ground terminal device i; C5 represents the offloading decision of ground terminal device i, which indicates whether to perform computation offloading, which low-Earth orbit satellite server to perform computation, and which channel to select for data transmission.

[0120] Step 3: Describe the computational unloading problem as a Markov decision process, and transform it into solving the optimal unloading strategy using deep reinforcement learning.

[0121] The specific implementation method of this step is as follows: For each unloading task solved using deep reinforcement learning, a task unloading model is constructed through a multi-objective Markov decision process. The multi-objective Markov decision process can be represented by tuples. Let S be the state space, A be the action space, P be the state transition probability, r be the reward, and w be the preference space.

[0122] In this invention, computational unloading is defined as a multi-objective problem, so Q is represented as a vector, with each element representing an objective.

[0123] Step 4: Initialize the neural network parameters and experience pool size. Generate two preference values ​​using a random method and assign the current user preference (weight) to the two objectives of latency and energy consumption.

[0124] The neural network in the deep reinforcement learning includes a main network and a target network, where θ is the parameter of the main network and θ′ is the parameter of the target network.

[0125] Step 5: In deep reinforcement learning, the agent begins to interact with the MEC environment. On one hand, the agent obtains the current state from the environment; on the other hand, the environment returns the current reward value and the next state based on the action selected by the agent, and updates the experience pool with the experience carrying the preferences. The composite reward function is set as follows:

[0126]

[0127] in r i d and r i e These represent user latency reward and energy consumption reward under different preferences, respectively. and Let z represent the terminal's latency and energy consumption preferences, respectively, where z is a negative number.

[0128] The agent selects actions using a greedy strategy based on the sum of the Q-vector output by the neural network and the inner product of preferences, as shown below:

[0129]

[0130] Step 6: Determine if the sample buffer has overflowed.

[0131] If it overflows, the quintuple will be...<s,a,r,w,s'> Store in sequence to the experience pool; otherwise, store the quintuple.<s,a,r,w,s'> Randomly stored in the experience pool to replace the sample.

[0132] Step 7: Randomly sample m samples from the sample pool and train them. The goal of training is to maximize the reward and obtain the optimal unloading decision. During the training process, the network input is the current state s and the preference, and the output is the Q vector.

[0133] Step 8: Calculate the objective function and loss function.

[0134] The objective function is:

[0135]

[0136] Where γ is the discount factor, arg Q Q(s) represents the supremum corresponding to multiple objective values. t+1 ,a t ,w′;θ t ′) is the state value function, representing the average cumulative reward obtained by performing action a under state s and preference w′.

[0137] The loss function is:

[0138] L(θ t ) = E s,a,w [(y t -Q(s t ,a t ,w;θ t )) 2 ]

[0139] In the above formula, Q represents the Q value obtained by the Q network, γ represents the reward discount factor, s′ is the next state output by the Q network, the Q network is updated using the loss function value, and the Q network parameters are synchronized to the target Q network every 300 iterations.

[0140] Step 9: Update the parameters θ of the main network using gradient descent.

[0141] The gradient descent (LD) method can be expressed as follows: That is, by taking the derivative along all parameter directions of the main network, we can obtain the direction of minimization of the objective function.

[0142] Step 10: When the preset period is reached, perform parameter copying, i.e., θ′←θ

[0143] The preset cycle is 300. After every 300 iterations, the values ​​of all parameters of the main network are copied to the corresponding target network as the parameter values ​​of the target network, thus achieving the purpose of periodic updating and replacement.

[0144] Step 11: If the optimal strategy is obtained, stop training. The optimal strategy, denoted by π*, refers to the action 'a' chosen to obtain the maximum cumulative reward, which can be expressed as the computational offloading and resource allocation strategy to obtain the maximum cumulative reward.

[0145] The above description is only a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiments. Any equivalent modifications or changes made by those skilled in the art based on the content disclosed in the present invention should be included within the scope of protection set forth in the claims.

[0146] The effectiveness of the invention is verified through simulation experiments below:

[0147] The simulation experiment was conducted in PyCharm, considering 5 ground terminal devices and 3 low-Earth orbit satellites carrying MEC servers at an altitude of 1000KM. The computing power of the server carried by each satellite follows a uniform distribution between [5,17] GHz. Each ground device has only one computing task that needs to be executed locally or unloaded. This task is randomly generated at each time interval. The task data size D∈[50,100]MB, the number of CPU cycles required to complete the task is set to 1000 cycles / bit, the elevation angle between the user terminal and the low-Earth orbit satellite is in the range of [10,30], the Earth's radius is 6371KM, and the BATCH SIZE is set to 64 and MEMORY CAPACITY to 5000.

[0148] Figure 2 The graph shows a comparison of latency between this invention and other methods under the same environment. The graph reveals that the algorithm of this invention optimizes multiple objectives and converges to a minimum value through multiple iterations. The higher latency due to all tasks being computed locally is attributed to the lower computing power of local devices, the longer local task buffer queue, and the fact that tasks are randomly generated by users, resulting in some fluctuation. In contrast, the original DDQN algorithm, because it did not consider the latency preferences of ground devices during the learning and training process and did not introduce rewards that take latency preferences into account, converged to a larger latency value.

[0149] Figure 3This is a comparison chart of the energy consumption of this invention and other methods under the same environment. The chart shows that the algorithm of this invention takes into account the conflict of multiple targets and can converge to the minimum energy consumption. The higher energy consumption of all local computations is due to the lower computing power of local devices, the longer local task buffer queue, and the fact that tasks are randomly generated by users, resulting in some fluctuation. In contrast, the original DDQN algorithm did not consider the energy consumption preferences of ground devices during the learning and training process, nor did it introduce rewards that take into account energy consumption preferences, thus its energy consumption converged to a larger value.

[0150] Figure 4 This is a comparison chart of the rewards of this invention and other methods under the same environment. The chart shows that the algorithm of this invention takes into account the conflict of multiple objectives during the learning and training process, sets a composite reward that meets the dynamic preference requirements, and can converge to a larger reward value. In contrast, the DDQN algorithm does not use a composite reward that meets the dynamic preference requirements, and therefore converges to a smaller reward value.

Claims

1. A multi-objective dynamic preference satellite-ground collaborative computational offloading method for low-Earth orbit satellite constellations, characterized in that: The method comprises the following steps: Step 1, constructing a multi-target dynamic preference task calculation model, a communication model and a satellite-ground communication time model for a low-orbit satellite constellation; Step 2, under the constraint conditions of a maximum user task acceptance time delay, a maximum satellite-ground communication time, dynamic preference requirements of a user for time delay and energy consumption and limited calculation capacity of a low-orbit satellite, constructing a dynamic preference task multi-target satellite-ground cooperative calculation and unloading optimization function for multiple users and multiple edge nodes; Step 3, a method of deep reinforcement learning is used to solve the multi-objective optimization function of the low-orbit satellite-mobile edge computing system, and the solving method is: the computing offloading problem is described as a multi-objective Markov decision process , which is converted into a problem of solving the optimal computing offloading control strategy; Step 4, initializing a current network and a target network of the deep reinforcement learning model and a size of an experience pool, and generating a preference vector by using a random method, so as to allocate current user preference values for the two targets of time delay and energy consumption; Step 5, training the deep reinforcement learning model by using a random sampling method to select samples from the experience pool; Step 6, obtaining a system state and a preference of a current time slot, inputting the system state into the trained deep reinforcement learning model, and obtaining a task unloading decision of each time slot by using the trained deep reinforcement learning model; The total time delay of all user corresponding unloading decisions under the dynamic time delay preference requirement is: , The total energy consumption of all user corresponding unloading decisions under the dynamic energy consumption preference requirement is: , wherein and respectively represent the preference value of the ground terminal device i for latency and energy consumption; The multi-target optimization function is a multi-target optimization problem with two tasks as follows: , , , , , , Among the constraints and This indicates that the completion time of task i of the terminal device must not exceed the maximum coverage time of satellite j under the offloading decision. and the maximum tolerance time for task i , This represents the local computation time for the task generated by ground terminal user i; This represents the computation time for a task generated by ground terminal user i to be offloaded to LEO satellite j; This represents the computing resources allocated to all users by low-Earth orbit satellite j, not exceeding its own maximum computing power. , The computing power (CPU cycles / s) allocated to LEO satellite j for terminal task i. middle and These represent the preferred latency and energy consumption requirements of ground terminal device i, respectively. This represents the offloading decision of ground terminal device i. Indicates whether to perform calculation and unloading. This indicates which low-Earth orbit satellite server the calculations will be performed on. Indicates which channel to select for data transmission; The calculation unloading problem is described as a multi-target Markov decision process, and is converted into a problem of solving an optimal calculation unloading control strategy, Multi-objective Markov decision processes can be represented by tuples. Representation, i.e., state space Action space State transition probability ,award and preference space In state s, the action a is selected based on the reinforcement learning algorithm to obtain an optimal unloading control strategy π*, which maximizes the long-term cumulative reward obtained during the device movement: A certain time slot state space is defined as wherein represents the local computing power of the end user i at time slot t, represents the amount of task data generated by the end user i at time slot t, represents the priority ranking list of the terminal device at time slot t, represents the computing power of the low-orbit satellite j at time slot t, represents the channel transmission situation at time slot t; A time slot action space is defined as wherein denotes that task i does not offload for local computation, denotes that task i offloads to a low earth orbit satellite server for computation; denotes that task i offloads to the Mth satellite edge node through a satellite-ground link, denotes that task i selects the Kth channel to transmit data The composite reward reward function is set as follows: , wherein , , and respectively represent the user latency reward and energy consumption reward under the preference, and respectively represent the latency and energy consumption preference of the terminal, z is a negative number; The training method of the deep reinforcement learning model is as follows: Step 1: parameter initialization is performed on a main network and a target network in the deep reinforcement learning, an experience pool with a size of N is constructed, an iteration number is set, and a state is initialized, thereby laying a foundation for a training process; Step 2: According to the system state and user preference of the current time slot, input it into the main network to get the Q value vector, combine the current deep neural network parameters and use the inner product of the preference and the Q vector as the selection basis, and use -greedy greedy strategy to decide the action a under the current state s, and calculate the immediate reward obtained by taking the decision action in the current state and the next state s´; Step 3: the obtained system state of a current time slot, a system action, an instant reward, a preference and a system state of a next time slot are stored in the experience pool; Step 4: a part of the experience pool is randomly selected as experience samples for training, a loss function is minimized according to gradient descent, and a main network parameter is updated, and the main network parameter is synchronized to a target network every 300 iterations; Step 5: an optimal unloading decision is obtained by training in order to maximize the reward; after the iteration is ended, a long-term cumulative reward reaches an optimum, and the optimal unloading decision is obtained.

Citation Information

Patent Citations

  • Edge computing task unloading method based on deep reinforcement learning in ultra-dense network

    CN115499441A

  • Satellite internet task unloading method and system and readable storage medium

    CN115499875A