A Satellite Network Resource Management Method Based on Deep Reinforcement Learning

By building LEO satellite mobility and SIoT network models, combined with deep reinforcement learning algorithms, optimizing task offloading and resource allocation, the problem of limited computing and energy resources in SIoT systems is solved, and the system overhead is reduced and the optimal allocation of resources is achieved.

CN119031394BActive Publication Date: 2025-07-11NANJING UNIV OF INFORMATION SCI & TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411521242.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-29
Publication Date
2025-07-11
Estimated Expiration
2044-10-29

AI Technical Summary

Technical Problem

In the existing satellite-assisted Internet of Things (SIoT) systems, the computing and energy resources of IoT devices are limited, resulting in high latency and energy consumption of sensing data processing, making it difficult to achieve load balancing between terminals, edges and cloud computing nodes.

Method used

The LEO satellite mobility model and SIoT network model are built, combined with the deep reinforcement learning algorithm, and the load balancing between terminals, edges and cloud computing nodes is achieved through task offloading and resource allocation algorithms. The model-assisted adaptive deep reinforcement learning (MADRL) algorithm is used to optimize task offload decisions and resource configuration.

Benefits of technology

It effectively reduces system overhead, improves task processing efficiency, significantly reduces system delay and energy consumption in various test environments, and realizes optimized resource allocation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119031394B_ABST
    Figure CN119031394B_ABST
Patent Text Reader

Abstract

The present invention provides a satellite network resource management method based on deep reinforcement learning, including: Step 1, constructing a LEO satellite mobility model; Step 2, constructing a SIoT network model with enhanced inter-satellite cooperation ISC; Step 3, respectively calculating the end-to-end delay and system energy consumption under three models of local computing, edge computing, and cloud computing; Step 4, establishing a task offloading and resource allocation algorithm and task offloading decision based on deep reinforcement learning to achieve load balancing among terminals, edge, and cloud computing nodes. The present invention proposes a model-assisted adaptive deep reinforcement learning algorithm for dynamic task arrival scenarios, which can realize the joint configuration of task offloading decisions, communication resources, and computing resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of information and communication technologies, and particularly relates to a satellite network resource management method based on deep reinforcement learning. Background Art

[0002] Existing wide-area Internet of Things (IoT) networks are mainly built on the basis of the fifth-generation mobile communication (5G) terrestrial network. Compared with terrestrial networks, satellite networks are rapidly becoming important communication infrastructures with advantages such as seamless coverage, flexibility, and reliability, driving the development of global connectivity and communication technologies. Therefore, satellite networks can be used as a supplement to terrestrial networks to provide services for IoT devices in remote suburbs or disaster areas, known as satellite-assisted Internet of Things (SIoT). Since IoT devices usually have limited power and few available resources for communication, computing, and caching, the sensing data generated by IoT devices usually needs to be forwarded to cloud or edge computing nodes for further processing. Mobile edge computing endows the edge of the network with computing power. By content caching and deploying edge computing servers near IoT devices, the processing delay can be effectively reduced. With the development of on-board processing technology, satellites with on-board processing units can also be regarded as an edge computing node to provide computing services for terrestrial users. In the current SIoT, terminals forward data to the terrestrial cloud data center through satellites for processing. Compared with satellite nodes, the cloud has higher computing power and sufficient energy supply. Summary of the Invention

[0003] Object of the Invention: The technical problem to be solved by the present invention is to provide a satellite network resource management method based on deep reinforcement learning for the deficiencies of the prior art, including the following steps:

[0004] Step 1, construct a LEO satellite mobility model;

[0005] Step 2, construct a SIoT network model with enhanced inter-satellite cooperation (ISC);

[0006] Step 3, calculate the end-to-end delay and system energy consumption under three models of local computing, edge computing, and cloud computing respectively;

[0007] Step 4, establish a task offloading and resource allocation algorithm and task offloading decision based on deep reinforcement learning to achieve load balancing among terminals, edge, and cloud computing nodes.

[0008] In Step 1, the LEO satellite mobility model includes:

[0009] LEO satellites are in orbit at an altitude above the ground H and fly at a speed uniformly. Let s be the angle between the LEO satellite and the positive horizontal direction at time slot n, and s be the geometric angle corresponding to the remaining coverage arc length from the LEO satellite to satellite user m at time slot n . Let R be the radius of the earth. When , the LEO satellite s can establish a communication link with device m;

[0010] According to geometric relationships, when , is expressed as:

[0011] (1),

[0012] When , is expressed as:

[0013] (2),

[0014] The linear distance between satellite s and satellite user m at time slot n is expressed as:

[0015] (3).

[0016] Step 2 includes: establishing a system for analysis according to the SIoT network model. The system contains M IoT devices, S LEO satellites and 1 cloud computing center. M = D + O, where O IoT devices are in the suburbs and D IoT devices are in the disaster area. The set of M IoT devices , the set of IoT devices in the disaster area , the set of IoT devices in the suburbs , and the LEO satellites are represented as the set . is the access satellite, is the cooperative satellite, and the set of time slots , where N represents the total number of time slots.

[0017] In step 2, let the length of a time slot be , and a period is divided into N time slots. It is assumed that the channel state remains unchanged within a time slot; the task n of satellite user m at time slot is modeled as . indicates that the task contains Bit data needs to be completed within time . The workload is , and the number of CPU cycles required for processing the task is . Use to represent the offloading decision of task .

[0018] In step 2, a satellite channel model is established, including communication loss, rainfall attenuation, and cloud attenuation. The communication loss is expressed as:

[0019] (4),

[0020] where is the communication distance, is the wavelength, is the carrier frequency, and c is the speed of light;

[0021] The rainfall attenuation is:

[0022] (5),

[0023] where is the rainfall rate exceeding 0.01% per year, is the effective path, and b are regression coefficients related to the raindrop size distribution, temperature, and frequency;

[0024] The attenuation of the cloud is:

[0025] (6),

[0026] where L is the total columnar content of liquid water in the cloud, is the specific attenuation coefficient of the cloud layer, and are the real and imaginary parts of the dielectric constant of water, respectively;

[0027] The total channel fading h during satellite communication is expressed as:

[0028] (7).

[0029] The present invention also provides an electronic device, including a processor and a memory. The memory stores program code. When the program code is executed by the processor, the processor executes the steps of the method.

[0030] The present invention proposes a dynamic SIoT network based on terminal-edge-cloud, and at the same time uses ISC (Inter-Satellite Collaboration) technology to improve the performance of the edge. In the constructed network, the tasks generated by IoT devices can be processed locally, offloaded to LEO (Low Earth Orbit) satellites for processing, or forwarded by LEO satellites to the cloud computing center for processing, that is, the tasks can be offloaded to the most suitable nodes to achieve load balancing among the terminal, edge, and cloud computing nodes. In addition, the present invention simultaneously considers the requirements of latency and energy consumption, takes the weighted sum of the latency and energy consumption of task processing as the overhead, and determines the weight between latency and energy consumption according to different task processing methods. The problem of jointly allocating task offloading decisions, CPU operating frequencies, and transmission powers in the SIoT network is modeled as a mixed integer nonlinear programming (MINLP) problem with resource constraints. To analyze and solve this problem, the present invention combines traditional optimization algorithms with deep reinforcement learning algorithms, and proposes a model-assisted adaptive deep reinforcement learning (MADRL) algorithm to minimize the system overhead.

[0031] The present invention has the following beneficial effects: By introducing a model-assisted adaptive deep reinforcement learning (MADRL) algorithm, the present invention effectively realizes the optimization of task offloading decisions and resource allocation. By introducing the mobility of satellites and ISC technology, the system overhead is significantly reduced. In addition, compared with traditional algorithms, the algorithms included in the present invention show lower system overhead in various test environments, indicating its significant advantages in practical applications. Brief Description of the Drawings

[0032] Figure 1 It is a geometric relationship diagram between LEO satellite s and satellite user m.

[0033] Figure 2 It is a schematic diagram of the system model in the first half of a cycle of the satellite.

[0034] Figure 3 It is a schematic diagram of the system model in the second half of a cycle of the satellite.

[0035] Figure 4 It is a flowchart of the MADRL algorithm.

[0036] Figure 5 It is the change of the algorithm reward value with the number of iterations at different learning rates.

[0037] Figure 6 The cumulative overhead of the algorithm over time for different learning rates.

[0038] Figure 7 The change of the algorithm's reward value with the number of iterations for different decay factors.

[0039] Figure 8 The change of the cumulative overhead of the algorithm proposed in the present invention over time for different exploration decay factors. Detailed implementation manners

[0040] The following further specifically describes the present invention in conjunction with the accompanying drawings and specific implementation manners, and the above and / or other advantages of the present invention will become clearer.

[0041] This embodiment provides a satellite network resource management method based on deep reinforcement learning, including:

[0042] Step 1, constructing a LEO satellite mobility model;

[0043] Step 2, constructing a SIoT network model with enhanced inter-satellite cooperation ISC;

[0044] Step 3, respectively calculating the end-to-end delay and system energy consumption under three models of local computing, edge computing, and cloud computing;

[0045] Step 4, establishing a task offloading and resource allocation algorithm and task offloading decision based on deep reinforcement learning to achieve load balancing among terminals, edge, and cloud computing nodes.

[0046] In step 1, the LEO satellite mobility model includes:

[0047] LEO satellite s The geometric relationship with satellite user m is as Figure 1 shown. Considering the mobility of LEO satellites, LEO satellite s flies at a constant speed H in an orbit at a height above the ground, is the angle between LEO satellite s and the positive horizontal direction at time slot n, is the geometric angle corresponding to the remaining coverage arc length from LEO satellite s to satellite user m at time slot n , R is the radius of the earth, and when , LEO satellite s can establish a communication link with device m;

[0048] According to the geometric relationship, when , is expressed as:

[0049] ​ (1),

[0050] When , It is expressed as:

[0051] (2),

[0052] The linear distance between satellite s and satellite user m in time slot n It is expressed as:

[0053] (3).

[0054] Step 2 includes: As Figure 2 shown, a system for analysis is established according to the SIoT network model. The system contains M IoT devices, S LEO satellites and 1 cloud computing center. M = D + O, where O IoT devices are in the suburbs and D IoT devices are in the disaster area. The set of M IoT devices , the set of IoT devices in the disaster area , the set of IoT devices in the suburbs , the LEO satellites are expressed as the set , is the access satellite, is the cooperative satellite, and the set of time slots , N represents the total number of time slots.

[0055] In step 2, let the length of a time slot be , a period is divided into N time slots, and it is set that the channel state remains unchanged within a time slot; the task n of satellite user m in time slot is modeled as , represents that the task contains bits of data and needs to be completed within time . The workload is , and the number of CPU cycles required to process the task . Use to represent the offloading decision of the task .

[0056] In step 2, for satellite communication, the channel fading is completely different from that of terrestrial communication. Here, this embodiment considers a satellite channel model closer to the real scenario, including communication loss, rainfall attenuation and cloud attenuation, where the communication loss is expressed as:

[0057] (4),

[0058] where is the communication distance (unit: km), is the wavelength, is the carrier frequency (unit: GHz), and c is the speed of light;

[0059] When the carrier frequency is above 10 GHz, rainfall is one of the main reasons for the attenuation of satellite communication signals. The rainfall attenuation is:

[0060] (5),

[0061] where is the rainfall rate exceeding 0.01% per year, is the effective path, and b are regression coefficients related to the raindrop size distribution, temperature, and frequency;

[0062] The attenuation of clouds is:

[0063] (6),

[0064] where L is the total columnar content of liquid water in the cloud, is the specific attenuation coefficient of the cloud layer, and are the real and imaginary parts of the dielectric constant of water, respectively;

[0065] The total channel fading h during satellite communication is expressed as:

[0066] (7).

[0067] Step 3 includes:

[0068] Step 3-1, constructing a local computing model:

[0069] Let the CPU operating frequency of satellite user m in time slot n be , and the task processing delay n of satellite user m in time slot is expressed as:

[0070] (8),

[0071] The task processing energy consumption n of satellite user m in time slot is expressed as:

[0072] (9),

[0073] where Represents the electrical coefficient;

[0074] The overhead of satellite user m in time slot n is expressed as: is expressed as:

[0075] (10),

[0076] where is the delay sensitivity coefficient in the local computing mode;

[0077] Step 3-2, construct an edge computing model:

[0078] Let the moment of the intermediate state of satellite motion be , as Figure 2 shown, when , the disaster area and the cloud computing center are within the service range of satellite , and the outer suburbs are within the service range of satellite , t represents the moment of the satellite motion state; as Figure 3 shown, when , the cloud computing center is within the service range of satellite , and the disaster area and the outer suburbs are within the service range of satellite ; when inter-satellite collaborative computing tasks are required, the task transmission link involves ISL (Inter-Satellite Link).

[0079] Let the CPU operating frequencies of satellites , in time slot n be and respectively, and the transmission powers be and respectively; the signal transmission rate is calculated by the Shannon formula. It is assumed that the Internet of Things devices have a sufficient number of orthogonal channels, and the channel allocation between multiple devices can be ignored. This assumption is based on the application of multi-frequency time division multiple access technology, which can dynamically allocate these orthogonal channels according to the current demand and traffic conditions. This enables the system to effectively manage the available resources and ensure that there are a sufficient number of channels for devices to transmit data.

[0080] The delay n of satellite user m in time slot is expressed as:

[0081] (11),

[0082] where indicates that the task transmission involves inter-satellite cooperation and requires an inter-satellite link; Indicates that the task transmission link does not involve inter-satellite cooperation and does not require an inter-satellite link; Indicates the first half of a cycle of the satellite, Indicates the second half of a cycle of the satellite; Indicates that satellite user m offloads a task to the satellite Signal transmission rate, Indicates that satellite user m offloads a task to the satellite Signal transmission rate, Indicates the signal transmission rate between the access satellite and the cooperative satellite, Indicates satellite user m and the satellite Distance between; Indicates satellite user m and the satellite Distance between; Indicates the distance between the access satellite and the cooperative satellite;

[0083] Energy consumption of satellite user m in time slot n Is expressed as:

[0084] (12),

[0085] Wherein, Indicates the transmission power of satellite user m in time slot n ;

[0086] Overhead of satellite user m in time slot n in the edge computing mode Is:

[0087] (13),

[0088] Wherein, Is the delay sensitivity coefficient in the edge computing mode;

[0089] Step 3-3, construct a cloud computing model:

[0090] Assume that the working frequency of a single-core CPU in the cloud computing center is , and the number of cores is ;

[0091] Delay of satellite user m in time slot n Is expressed as:

[0092] (14),

[0093] Where Indicates the satellite Offloads a task to the satellite Signal transmission rate, Indicates the satellite The signal transmission rate for offloading tasks to the cloud computing center represents the satellite and the satellite the distance between them represents the satellite the distance between it and the cloud computing center;

[0094] The energy consumption of satellite user m in time slot n is expressed as:

[0095] (15),

[0096] The total overhead of satellite user m in the cloud computing mode in time slot n is: For:

[0097] (16),

[0098] where is the delay sensitivity coefficient in the cloud computing mode;

[0099] Step 3 - 4, problem optimization: The end - to - end delay n of the tasks generated under different processing methods by satellite user m in time slot and the energy consumption are respectively:

[0100] (17),

[0101] (18),

[0102] where represents the task processing method. When is equal to 0, it means the task is processed locally. When is equal to 1, it means the task is offloaded to the access satellite for processing. When is equal to 2, 3, 4, 5, it means the task is offloaded to the cooperative satellite for processing. When is equal to 6, it means the task is offloaded to the cloud computing center for processing; A simple analysis shows that local processing has less energy consumption, but due to limited computing power, it may not meet the delay requirements for task processing. Processing tasks at the edge node helps reduce processing delay but increases satellite energy consumption. Processing tasks at the cloud center reduces computing delay but increases propagation delay. Therefore, the offloading decision and resource allocation for each time slot have a significant impact on the task completion delay and energy consumption. The present invention studies the problem of how to minimize the weighted sum of system delay and energy consumption under delay and energy consumption constraints within one period in a dynamic network scenario. Joint optimization of task offloading decisions, communication, and computing resource allocation is carried out. Equation (19) gives the target optimization problem P1.

[0103] The optimization objective is to minimize the weighted sum of system delay and energy consumption within one period. The optimization variables are the task offloading decision, the transmission power and CPU operating frequency of IoT devices, and the transmission power and CPU operating frequency of LEO satellites. Under different processing methods, the delay sensitivity coefficient is different. Because at the local node, the device needs to ensure battery life and is sensitive to energy consumption, the delay sensitivity coefficient should be set smaller. At the edge node, the low-earth orbit satellite can reduce the processing delay but increase the communication delay. At the same time, the computing power of the low-earth orbit satellite is limited, and it is sensitive to both delay and energy consumption. In the cloud node, the propagation delay increases significantly, and the computing power of the cloud center is much higher than that of the device and the LEO satellite. Therefore, it is sensitive to delay, and the delay sensitivity coefficient should be set relatively high.

[0104] (19),

[0105] where, represents the maximum tolerable delay of the task generated by the IoT device m in time slot n and represents the maximum energy consumption of the IoT device m . represents the maximum energy consumption of the satellite s . represents the maximum CPU operating frequency of the IoT device m . represents the maximum transmission power of the IoT device m . represents the maximum CPU operating frequency of the satellite s . represents the maximum transmission power of the satellite s ; Constraints C9 and C10 are continuity constraints to ensure that the state change of the system at adjacent times is not too drastic. represents a non-zero positive number used to measure the closeness of the function value at the next moment to the current moment function value, is the value of in constraint C9, is the value of in constraint C10; represents the delay sensitivity coefficient, that is, , and one of them; Constraint C1 indicates that the task processing methods include local computing, offloading to LEO satellites (including access satellites and cooperative satellites) for computing, and offloading to the cloud computing center for computing. Constraint C2 indicates that the end-to-end delay of task processing in any time slot should be less than the required delay of the task. Constraint C3 indicates that the energy consumption of satellite user m in any time slot should be less than its maximum energy consumption. Constraint C4 indicates that in any time slot, the satellites The energy consumption should be less than its maximum energy consumption. Constraint C5 means that the CPU operating frequency of satellite user m in any time slot should be less than its maximum operating frequency. Constraint C6 means that the transmission power of satellite user m in any time slot should be less than its maximum power. Constraint C7 means that the CPU s operating frequency of the satellite should be less than its maximum operating frequency. Constraint C8 means that the s transmission power of the satellite should be less than its maximum power. Constraints C9 and C10 are continuity constraints to ensure that the state change of the system at adjacent times is not too drastic.

[0106] Step 4 includes: In formula (19), the CPU operating frequency and transmission power are continuous variables, and the offloading decision is a discrete variable. The objective function is non-linear with respect to these variables. Therefore, formula (19) is a MINLP problem. It is difficult for traditional optimization methods to find the optimal solution. The present invention uses a MADRL algorithm to solve this problem. The first layer optimizes the CPU operating frequency and transmission power through model assistance and uses the binary search algorithm. The second layer uses the adaptive DRL algorithm to adapt to the dynamic network scenario and generates the Q network through self-learning. The task offloading and resource allocation algorithm based on deep reinforcement learning includes: optimizing the CPU operating frequency and transmission power through model assistance:

[0107] When the task adopts local computing, the optimization problem is transformed into:

[0108] (20),

[0109] The overhead function F1 in the local computing mode is expressed as

[0110] (21),

[0111] Extreme point is , and constraints C1 and C2 are simplified to , ; The upper bound and the lower bound of the feasible solution range of F1 are:

[0112] (22),

[0113] The optimal overhead of local computing is expressed as:

[0114] (23),

[0115] where represents the value when the overhead function takes the upper bound , represents the overhead function at the extreme point The value indicates the lower bound of the overhead function when taking the value;

[0116] When the task adopts edge computing, the optimization problem is transformed into:

[0117] (24),

[0118] When the task adopts cloud computing, the optimization problem is transformed into:

[0119] (25).

[0120] In step 4, the task offloading decision is determined using Deep Q-Network (DQN) and Double Deep Q-Network (DDQN). There are two neural networks in DQN, namely the online network and the target network. The online network is used to select the action a. When the input state is state, the output is the Q value of each action, and then the action a corresponding to the maximum Q value is selected. The target network is used to estimate the Q value to update the model parameters. The parameter update method is to overwrite the target network with the online network every X steps. Since the Q value output by the target network in DQN will produce overestimation, essentially it is because the action selected by the target network is incorrect, resulting in an inappropriate estimated Q value. DDQN solves this problem. In DDQN, the selection of the action is directly done by the online network. So the overall structure of DDQN is roughly the same as that of DQN, and the only difference lies in the calculation method of the y value during the learning process. In DDQN, the calculation of the target value y is related to both the target network and the online network. That is, first, the state is input into the online network to determine the action a. Then the state is input into the target network to output the Q values of all actions, and the Q value of the action a is directly used to calculate the target value y. Even if the Q value corresponding to the action a in the output of the target network is not the largest, it is still used to calculate the target value y. However, in the case of a small amount of data, DDQN usually does not perform as well as DQN. This is because when the amount of data is small, the parameters of these two networks are difficult to effectively train, resulting in a large deviation in the estimation of the target value, thus reducing the performance of DDQN. This exactly corresponds to the different situations in the disaster area and the far suburbs. The population density in the disaster area is high, the number of devices is large, and the amount of data generated is large, which is more suitable for processing with DDQN. The population density in the far suburbs is low, the amount of data is small, and better performance can be obtained using the DQN algorithm. Therefore, this embodiment proposes an adaptive DRL algorithm to train the Q networks in the far suburbs and the disaster area, and then obtain the task offloading decision. The adaptive deep reinforcement learning DRL algorithm includes the following steps:

[0121] Step a1, initialization and preparation:

[0122] Step a1-1, initialize the parameters of the online network Q and the target network Q_hat;

[0123] Step a1-2, set the training parameters, including the learning rate lr, the exploration rate epsilon, the discount factor γ, the target network update frequency target_update_freq, the exploration rate decay epsilon_decay, the size of the experience replay pool D, the minimum batch size batch_size, and the number of algorithm running episodes num_episode;

[0124] Step a2, perform the following training process:

[0125] Step a2-1, for each episode from 1 to num_episode, do the following:

[0126] Step a2-1-1, initialize the state state;

[0127] Step a2-1-2, for each step step, do the following:

[0128] Step a2-1-1-1, randomly select an action a with probability epsilon;

[0129] Step a2-1-1-2, otherwise, select the action a that maximizes Q(state, a); If the current step does not randomly select an action with probability epsilon (i.e., if it passes a probability-based "if" condition check and the condition is not met, i.e., no random action is selected), then the steps after "otherwise" will be executed. Q(state, a) represents the expected return of performing action a in state state;

[0130] Step a2-1-1-3, execute action a and observe the obtained reward r and the new state state';

[0131] Step a2-1-1-4, store the transition tuple (state, a, r, state') in the experience replay pool D;

[0132] Step a2-1-1-5, update the state state to state';

[0133] Step a2-2, if the size of the experience replay pool D is greater than or equal to batch_size, do the following:

[0134] Step a2-2-1, randomly sample a batch of transition tuples (state, a, r, state') from the experience replay pool D;

[0135] Step a2-2-2, for each extracted transition tuple, calculate the target value y: If state' is a terminal state, then y = r, otherwise, for the deep Q-network DQN, ; For the double deep Q-network DDQN , ; where max_a' represents the maximum value among all possible actions a'; Q_ hat(state', a') represents the expected return of executing action a' in state state'; argmax_a' represents the action a' that maximizes the Q value; Q(state', a') represents the expected return of executing action a' in state state';

[0136] Step a2-2-3, use the calculated target value y and the expected return of a certain predicted action in the given state to calculate the loss function;

[0137] Step a2-2-4, use the gradient descent method to update the parameters of the online network Q;

[0138] Step a2-3, every target_update_freq steps, copy the parameters of the online network Q to the target network Q_hat;

[0139] Step a2-4, after all episodes end, return the trained online network Q.

[0140] In Step 4, in order to achieve a comprehensive and effective solution, considering the optimization strategy comprehensively, combining model-assisted resource allocation with learning to optimize the offloading decision, a model-assisted adaptive deep reinforcement learning (Model-assisted Adaptive Deep Reinforcement Learning, MADRL) is proposed. The flowchart of the MADRL algorithm is as Figure 4 shown. The first layer optimizes the CPU operating frequency and transmission power through model assistance and the binary search algorithm, solves the optimal overhead of each task situation and processing method, and obtains the optimal overhead matrices cost_matrix_disaster and cost_matrix_remote for the disaster area and the suburban area. The second layer uses the adaptive DRL algorithm to train the Q-network and obtains the task offloading decision according to the actual task state.

[0141] Each element in the Q-network is described as follows: The deep Q-network DQN includes the following elements:

[0142] State space: In each time slot, the system observes the current state and obtains environmental information; the state state is represented by the task status of each user, and the task status of each user includes the task size , workload and maximum tolerable delay . ; The state space at time slot n is defined as , where represents the task size of device M in time slot n , represents the task workload of device M in time slot n, represents the maximum tolerable delay of the task generated by device M in time slot n;

[0143] Action space: After the online network obtains the state space , corresponding discrete offloading decisions will be generated, , ;

[0144] Reward function: After taking action a in state state, the environment enters the next state state' and returns a reward r, and the reward value r is defined as the reciprocal of the system overhead.

[0145] In the specific embodiments of the present invention, the performance of the proposed algorithm is evaluated through simulation analysis and the performance of the proposed algorithm is compared with that of the benchmark algorithm.

[0146] The main parameter settings used in the simulation include: for the number of devices, it is assumed that D = 300 and O = 5. At the same time, the number of episodes num_episode of the MADRL algorithm is 1000, the batch size batch_size is 32, the learning rate lr is 0.001, the discount factor γ is 0.9, the initial exploration probability epsilon is 1, the exploration probability decay rate epsilon_decay is 0.995, and the minimum exploration probability epsilon_min is 0.01. Based on the laser link deployment, the communication capacity of the ISL is set to 10 Gbps. The number of devices in the disaster area is 300, the number of devices in the outer suburbs is 5, the task size is [1e2, 1e3, 1e4, 1e5, 1e6] bit, the task workload is [1, 1.5] kcycle / bit, the maximum tolerable delay is [0.05, 0.1] s, the electrical coefficient is 10^(-28), and the channel bandwidth is 10 Mb;

[0147] The antenna gain is 20 dB, the noise temperature is 290°, the maximum power consumption of IoT satellite user m is 5 w, the maximum power consumption of LEO satellite s is 2000 w, the working frequency of a single-core CPU in the cloud computing center is 1.45 GHz, the number of cores in the cloud computing center is 256, the number of episodes num_episode is 1000, the batch size batch_size is 32, the learning rate lr is 0.001, the discount factor γ is 0.9, the initial exploration probability epsilon is 1, the exploration probability decay rate epsilon_decay is 0.995, and the minimum exploration probability epsilon_min is 0.01.

[0148] Through a series of simulation experiments, the effects of the learning rate and exploration rate on the convergence performance of the proposed algorithm were compared. Figure 5 For the change of the algorithm reward value with the number of iterations under different learning rates, when the learning rates are 0.1 and 0.01, the training effect is not good, and the reward value fluctuates violently with the increase of the number of iterations. When the learning rate is reduced to 0.001 and 0.0001, the convergence performance is better, and with the progress of training, the reward value fluctuates little. Figure 6 For the change of the cumulative overhead of the algorithm with time under different learning rates, when the learning rate is 0.001, the offloading decision made by the trained model minimizes the system overhead. Therefore, the learning rate of 0.001 is selected in the present invention. Figure 7 For the change of the algorithm reward value with the number of iterations under different decay factors, when the decay factor is too large, the exploration probability decreases too slowly, and a large amount of exploration is still carried out in the later stage of training, which will cause the model to be unable to make full use of the knowledge already learned and unable to converge to the optimal strategy. When the decay factor is too small, the exploration probability decreases too fast, and the existing strategy is overused in the early stage of training, which may cause the Q-network to be unable to fully explore the environment and fall into a local optimum. Therefore, the decay factor cannot be too large or too small. Figure 8 For the change of the cumulative overhead of the algorithm proposed in the present invention with time under different exploration decay factors, when the decay factors are 0.95 and 0.9995, the system overhead is greater than 0.995. Therefore, the decay factor of this algorithm is set to 0.995, which can find a good balance between exploration and exploitation, can fully explore in the initial stage of training, and can better utilize the learned strategy in the later stage, and can achieve the faster convergence of the reward value to a higher level.

[0149] The present invention provides a satellite network resource management method based on deep reinforcement learning. There are many methods and ways to specifically implement this technical solution. The above is only the preferred implementation manner of the present invention. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. Each component not clearly defined in this embodiment can be implemented by the prior art.

Claims

1. A satellite network resource management method based on deep reinforcement learning, characterized in that, It includes the following steps: Step 1, construct a LEO satellite mobility model; Step 2, construct a SIoT network model with enhanced inter-satellite cooperation (ISC); Step 3, calculate the end-to-end delay and system energy consumption under the three models of local computing, edge computing, and cloud computing respectively; Step 4, establish a task offloading and resource allocation algorithm and task offloading decision based on deep reinforcement learning to achieve load balancing among terminals, edge, and cloud computing nodes; In Step 1, the LEO satellite mobility model includes: LEO satellites fly uniformly at a speed of V in an orbit at a height H above the ground s and α m,s [n] is the angle between the LEO satellite s and the positive horizontal direction at time slot n, and γ m,s [n] is the geometric angle corresponding to the remaining coverage arc length from the LEO satellite s to the satellite user m at time slot n. R is the radius of the earth. When 0° < α m,s [n] < 180°, the LEO satellite s can establish a communication link with the device m; According to the geometric relationship, when 0° < α m,s [n] < 90°, γ m,s [n] is expressed as: When 90° < α m,s [n] < 180°, γ m,s [n] is expressed as: The linear distance D between satellite s and satellite user m at time slot n m,s is expressed as Step 2 includes: establishing a system for analysis according to the SIoT network model. The system includes M IoT devices, S LEO satellites and 1 cloud computing center, where M = D + O. Among them, O IoT devices are in the suburbs, and D IoT devices are in the disaster area. The set of M IoT devices The set of IoT devices in the disaster area The set of IoT devices in the suburbs The LEO satellites are represented as a set {S1, S2} are the access satellites, {S c1 , S c2 , S c3} are the cooperative satellites, and the set of time slots N represents the total number of time slots; In step 2, let the length of a time slot be τ, a period is divided into N time slots, and it is assumed that the channel state remains unchanged within a time slot; the task of satellite user m in time slot n is modeled as denotes the task contains bits of data and needs to be completed within time . The workload is the number of CPU cycles required to process the task Let denote the offloading decision of task ; In step 2, a satellite channel model is established, including communication loss, rainfall attenuation, and cloud attenuation, where the communication loss φ fs is expressed as: where d F is the communication distance, λ is the wavelength, and f C is the carrier frequency, and c is the speed of light; Rain attenuation φ rain is as follows: Among them, is the rainfall rate exceeding 0.01% per year, d eff is the effective path, and a and b are regression coefficients related to the raindrop size distribution, temperature, and frequency; Attenuation φ of cloud cloud is as follows: where L is the total columnar content of liquid water in the cloud, and k c is the specific attenuation coefficient of the cloud layer, and ε' and ε” are the real and imaginary parts of the dielectric constant of water, respectively; The total channel fading h during satellite communication is expressed as: h = φ fs φ rain φ cloud (7) Step 3 includes: Step 3-1, construct a local computing model: Let the CPU operating frequency of satellite user m at time slot n be The task processing delay of satellite user m at time slot n is expressed as: The task processing energy consumption of satellite user m in time slot n Expressed as: where ε x represents the electrical coefficient; Overhead of satellite user m in time slot n Expressed as: Among them, μ l is the latency sensitivity coefficient in the local computing mode; Step 3-2, construct an edge computing model: Let the time of the intermediate state of satellite motion be \(t\). mid , when \(t \lt t\) mid , the disaster area and the cloud computing center are within the service range of satellite \(s1\), and the suburban area is within the service range of satellite \(s2\). \(t\) represents the time of the satellite motion state; when \(t \gt t\) mid , the cloud computing center is within the service range of satellite \(s1\), and the disaster area and the suburban area are within the service range of satellite \(s2\). Let the CPU operating frequencies of satellites s1 and s2 in time slot n be and respectively, and their transmission powers be and respectively. The delay of satellite user m in time slot n is expressed as: Among them, ISL indicates that the task transmission involves inter-satellite cooperation and requires an inter-satellite link; No ISL indicates that the task transmission link does not involve inter-satellite cooperation and does not require an inter-satellite link; t < t mid indicates the first half of a satellite's cycle, t > t mid indicates the second half of a satellite's cycle; represents the signal transmission rate at which satellite user m offloads a task to satellite s1, represents the signal transmission rate at which satellite user m offloads a task to satellite s2, c s represents the signal transmission rate between the access satellite and the cooperative satellite, represents the distance between satellite user m and satellite s1, represents the distance between satellite user m and satellite s2, D S represents the distance between the access satellite and the cooperative satellite; The energy consumption of satellite user m in time slot n Expressed as: Among them, represents the transmission power of satellite user m in time slot n; Overhead of satellite user m in time slot n under the edge computing mode is as follows: where μ e is the latency sensitivity coefficient in the edge computing mode; Step 3-3, construct a cloud computing model: Let the working frequency of a single-core CPU in the cloud computing center be f c , and the number of cores be N ct ; The latency of satellite user m at time slot n It is expressed as: Among them represents the signal transmission rate at which satellite user m offloads tasks to satellite s1, represents the signal transmission rate at which satellite user m offloads tasks to satellite s2, represents the signal transmission rate at which satellite s1 offloads tasks to satellite s2, represents the signal transmission rate at which satellite s1 offloads tasks to the cloud computing center, represents the distance between satellite s1 and satellite s2, represents the distance between satellite s1 and the cloud computing center; The energy consumption of satellite user m at time slot n Expressed as: Total cost of satellite user m at time slot n under the cloud computing mode is as follows: Among them, μ c is the latency sensitivity coefficient in the cloud computing mode; Step 3-4, Problem Optimization: The end-to-end delay and energy consumption of the tasks generated by satellite user m under different processing methods in time slot n are respectively: Among them indicates the task processing method. When is equal to 0, it means the task is processed locally. When is equal to 1, it means the task is offloaded to the access satellite for processing. When is equal to 2, 3, 4, or 5, it means the task is offloaded to the cooperative satellite for processing. When is equal to 6, it means the task is offloaded to the cloud computing center for processing; The target optimization problem P1 is: Among them, represents the maximum tolerable delay for the IoT device m to generate a task in time slot n, represents the maximum energy consumption of the IoT device m, represents the maximum energy consumption of the satellite s, represents the maximum CPU operating frequency of the IoT device m, represents the maximum transmission power of the IoT device m, represents the maximum CPU operating frequency of the satellite s, represents the maximum transmission power of the satellite s; ε represents a non-zero positive number used to measure the closeness of the function value at the next moment to the function value at the current moment, ε f is the value of ε in constraint C9, ε p is the value of ε in constraint C10; μ represents the delay sensitivity coefficient; In Step 4, the task offloading and resource allocation algorithm based on deep reinforcement learning includes: optimizing the CPU working frequency and transmission power through model assistance: When the task adopts local computing, the optimization problem is transformed into: The overhead function F1 in the local computing mode is expressed as: Extreme point are constraints C1 and C2 are simplified to The upper bound f1 and lower bound f2 of the feasible solution range of F1 are: The optimal overhead of local computing is expressed as: where F1(f1) represents the value of the overhead function when taking the upper bound f1, represents the value of the overhead function at the extreme point and F1(f2) represents the value of the overhead function when taking the lower bound f2; When the task adopts edge computing, the optimization problem is transformed into: When the task adopts cloud computing, the optimization problem is transformed into: In Step 4, the task offloading decision is determined using a deep Q-network (DQN) and a double deep Q-network (DDQN). The deep Q-network DQN includes an online network Q and a target network Q_hat; Train the deep Q-network DQN in the suburbs and disaster areas using an adaptive deep reinforcement learning (DRL) algorithm to obtain the task offloading decision. The adaptive deep reinforcement learning DRL algorithm includes the following steps: Step a1, initialization and preparation: Step a1-1, initialize the parameters of the online network Q and the target network Q_hat; Step a1-2, set the training parameters, including the learning rate lr, exploration rate epsilon, discount factor γ, target network update frequency target_update_freq, exploration rate decay epsilon_decay, the size of the experience replay pool D, the minimum batch size batch_size, and the number of algorithm running episodes num_episode; Step a2, perform the following training process: Step a2-1, for each episode from 1 to num_episode, perform the following operations: Step a2-1-1, initialize the state state; Step a2-1-2, for each step step, perform the following operations: Step a2-1-1-1, randomly select an action a with probability epsilon; Step a2-1-1-2, otherwise, select the action a that maximizes Q(state,a); Q(state,a) represents the expected return of performing action a in state state; Step a2-1-1-3, execute action a and observe the obtained reward r and the new state state'; Step a2-1-1-4, store the transition tuple (state,a,r,state') in the experience replay pool D; Step a2-1-1-5, update the state state to state'; Step a2-2, if the size of the experience replay pool D is greater than or equal to batch_size, perform the following operations: Step a2-2-1, randomly sample a batch of transition tuples (state, a, r, state') from the experience replay pool D; Step a2-2-2, for each sampled transition tuple, calculate the target value y: if state’ is a terminal state, then y = r, otherwise, for the deep Q-network DQN, y = r + γmax_a'Q_hat(state', a'); for the double deep Q-network DDQN, y = r + γQ_hat(state', argmax_a'Q(state', a')); where max_a' represents the maximum value among all possible actions a'; Q_hat(state', a') represents the expected return for taking action a' in state state'; argmax_a' represents the action a' that maximizes the Q value; Q(state', a') represents the expected return for taking action a' in state state’; Step a2-2-3, calculate the loss function using the calculated target value y and the expected return of a certain action predicted in the given state; Step a2-2-4, update the parameters of the online network Q using the gradient descent method; Step a2-3, every target_update_freq steps, copy the parameters of the online network Q to the target network Q_hat; Step a2-4, after all episodes end, return the trained online network Q; In Step 4, the deep Q-network DQN includes the following elements: State space: In each time slot, the system observes the current state and obtains environmental information; the state state is represented by the task status of each user, and the state state includes the task size d m , the workload c m and the maximum tolerable delay The state space at time slot n is defined as where represents the task size of device M at time slot n, represents the task workload of device M at time slot n, represents the maximum tolerable delay for task generation by device M at time slot n; Action space: The online network obtains the state space g n and then generates the corresponding discrete offloading decision a n , Reward function: After taking action a in state state, the environment enters the next state and returns a reward r, and the reward value r is defined as the reciprocal of the system overhead.

2. An electronic device, characterized in that, It includes a processor and a memory, and the memory stores program code. When the program code is executed by the processor, the processor is caused to execute the steps of the method according to Claim 1.

Citation Information

Patent Citations

  • Calculation unloading method for distributed deep learning in satellite-ground collaborative network

    CN114153572A