Satellite-ground collaborative dynamic task division and resource allocation method based on deep reinforcement learning, electronic equipment and medium
By building a two-layer computing architecture of ground equipment and LEO satellites and a Markov decision-making process for deep reinforcement learning, the task division and resource allocation in the LEO satellite network are optimized, and the dynamic nature of task offloading and resource allocation imbalance are solved, and the calculation delay is minimized and resource utilization efficiency is improved.
Patent Information
- Application Number
- CN202510383116.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-01
AI Technical Summary
In LEO satellite networks, the existing technology is difficult to effectively solve the dynamic nature of mission offloading and resource allocation imbalance, resulting in high computing delays and low resource utilization efficiency.
Build a two-layer computing architecture for ground equipment and LEO satellites, establish computing and communication models through inter-satellite link technology, use deep reinforcement learning to optimize task division and resource allocation, and use the GDPG algorithm to generate task segmentation rate and resource allocation strategies.
It effectively reduces the total system latency, balances the computing load of the satellite network, improves resource utilization efficiency, and reduces dependence on remote clouds.
Smart Images

Figure CN120234150A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of edge computing, and more specifically, to a satellite-ground collaborative dynamic task partitioning and resource allocation method, an electronic device, and a medium based on deep reinforcement learning. Background Art
[0002] In traditional terrestrial networks, communication mainly relies on fixed infrastructures such as base stations and optical fibers. However, it is difficult for terrestrial networks to cover special environments such as remote areas, oceans, and the sky. In contrast, satellite networks, especially LEO satellite networks, provide a seamless coverage communication solution. In addition, LEO satellite networks can be quickly deployed and expanded, without relying on terrestrial infrastructures, and can still maintain efficient and stable communication services in the face of natural disasters or man-made damage. With the popularization of LEO satellite networks, extending mobile edge computing to LEO satellites has become a new option. By deploying edge nodes on satellites, the satellites are enabled with computing capabilities, and users can offload computing tasks to the satellites for execution, reducing the dependence on remote clouds, thereby saving valuable bandwidth resources and thus greatly reducing the computing latency and energy consumption.
[0003] In STIN, due to the high-speed movement and constantly changing orbital positions of LEO satellites, their channel conditions also change continuously with the positions of ground terminals. Such time-varying channel characteristics often affect the transmission stability. Therefore, most of the existing research on terrestrial MEC task offloading cannot be directly applied to satellite networks. In addition, the computing and storage capabilities of MEC servers carried by LEO satellites are limited compared with those of ground base stations. Under such limitations, it is necessary to accurately select appropriate task offloading decisions and resource allocations. Due to the randomness of task requirements and the differences in the number of users covered by satellites, some satellites may be in a load saturation state, while other satellites may be in a low-load state. This imbalance in resource utilization will reduce the efficiency of the entire satellite network.
[0004] Currently, most of the research on task offloading in satellite networks adopts binary offloading strategies, that is, completely offloading computing tasks to satellites or clouds for processing. This strategy does not fully utilize the advantages of collaborative computing of each layer architecture. Although some research adopts multi-layer computing strategies, it ignores the advantages of load balancing and resource utilization brought by the unique ISL cooperation technology in LEO satellites. Moreover, in the dynamic network of LEO satellites, task partitioning and resource allocation are often difficult to directly solve by convex optimization methods due to the time-varying nature of their channel states and satellite resources.
[0005] Therefore, it is necessary to develop a satellite-ground collaborative dynamic task partitioning and resource allocation method, an electronic device, and a medium based on deep reinforcement learning.
[0006] The information disclosed in the background section of the present invention is only intended to enhance the understanding of the general background of the present invention, and should not be regarded as an admission or any form of implication that this information constitutes the prior art known to those skilled in the art. Summary of the Invention
[0007] The present invention proposes a method, an electronic device, and a medium for satellite-ground collaborative dynamic task partitioning and resource allocation based on deep reinforcement learning, which can minimize the computing latency by optimizing task partitioning and resource allocation.
[0008] In a first aspect, an embodiment of the present disclosure provides a method for satellite-ground collaborative dynamic task partitioning and resource allocation based on deep reinforcement learning, including:
[0009] Construct a two-layer computing architecture for the collaboration between ground devices and LEO satellites;
[0010] Based on the two-layer computing architecture, establish a computing model and a communication model of the satellite-ground integrated network through inter-satellite link technology;
[0011] Establish an objective function, and construct the task partitioning and resource allocation problems in the satellite-ground integrated network as a Markov decision process;
[0012] Solve the task partitioning and resource allocation problems to generate a task segmentation rate and a resource allocation strategy in a continuous action space.
[0013] Preferably, constructing a two-layer computing architecture for the collaboration between ground devices and LEO satellites includes:
[0014] The computing task at time slot t is Q u,s (t) = {d u,s (t), c u,s , τ u,s (t)}, where d u,s (t) is the task size, c u,s is the number of CPU cycles required to compute one bit of data, and τ u,s (t) represents the maximum tolerance time of the task;
[0015] Let be the task allocation ratios for local processing, processing on the edge server integrated with the access satellite, and processing on the adjacent satellite of the access satellite, respectively, and then construct a two-layer computing architecture for the collaboration between ground devices and LEO satellites.
[0016] Preferably, establishing the computing model includes:
[0017] Construct a local computing model and a LEO access satellite edge computing model according to the splitting ratio of the task, the computing capabilities of the local device and the satellite. Define the amount of resources allocated by satellite s to adjacent satellites according to the ISL technology, share the inter-satellite resources, and construct a collaborative satellite computing model using ISL.
[0018] Preferably, establishing the communication model includes:
[0019] Let the spatial coordinates of the user at time slot t be O u,s (t)=[O u1,s (t), O u2,s (t), O u3,s (t)], where O u1,s (t), O u2,s (t), O u3,s (t) are the longitude, latitude and altitude of the user equipment respectively;
[0020] Let the position coordinates of the LEO satellite be O s (t)=[O s1 (t), O s2 (t), O s3 (t)], then the distance between the user and the satellite at this time slot is:
[0021]
[0022] Among them, R is the radius of the earth;
[0023] Let the user share the spectrum resources in the way of orthogonal frequency division multiple access, then the uplink transmission rate of the user to the satellite is:
[0024]
[0025] Among them, B is the uplink bandwidth, p u,s (t) is the transmission power of the ground device, N0 is the satellite noise spectral density, h u,s (t) is the channel gain between the device and the satellite, expressed as h u,s (t)=G u,s |g u,s (t)h u,s (t)D u,s (t) -a | 2 Among them, G u,s is the antenna gain of the device, g u,s (t) is the complex Gaussian variable of Rayleigh fading, h u,s (t) is the fading component, and a is the path exponent;
[0026] For user-to-satellite communication, assuming the access satellite is s, the user equipment sends the task to the access satellite, and then the access satellite further sends the task to the cooperative satellite. Therefore, the user uplink transmission delay is as follows:
[0027]
[0028] The propagation delay is:
[0029]
[0030] where c is the speed of light;
[0031] The communication delay from the user to the access satellite is the sum of the transmission delay and the propagation delay, which is expressed as:
[0032]
[0033] Preferably, the total time delay of the calculation model and the communication model is used as the objective function.
[0034] Preferably, constructing the task partitioning and resource allocation problem in the satellite-ground integrated network as a Markov decision process includes:
[0035] The delay is divided into two parts. The first part is The second part Then the total delay of the system is:
[0036]
[0037] The goal is to minimize all delays, jointly optimize the task splitting ratio and resource allocation parameters, and construct the task partitioning and resource allocation problem in the satellite-ground integrated network as a Markov decision process.
[0038] Preferably, the task partitioning and resource allocation problem is solved by a dynamic offloading algorithm that fuses a gated recurrent unit and deep deterministic policy gradient.
[0039] In a second aspect, an embodiment of the present disclosure also provides an electronic device, which includes:
[0040] A memory storing executable instructions;
[0041] A processor that runs the executable instructions in the memory to implement the satellite-ground collaborative dynamic task partitioning and resource allocation method based on deep reinforcement learning.
[0042] In a third aspect, an embodiment of the present disclosure further provides a computer-readable storage medium storing a computer program, which when executed by a processor, implements the method for satellite-ground collaborative dynamic task partitioning and resource allocation based on deep reinforcement learning.
[0043] The beneficial effects are as follows:
[0044] The present invention designs a two-layer computing offloading framework for satellite-ground collaboration, aiming to minimize task latency. By combining the computing capabilities of ground terminal devices and LEO satellites, it effectively shares the task computing load. By introducing an inter-satellite cooperation mechanism, it supports dynamic offloading and resource sharing of tasks between different satellites, balancing the computing load capacity of the satellite network. A GDPG reinforcement learning algorithm is designed to solve the problems of computing offloading and resource allocation in a dynamic environment. In this algorithm, the actor network adopts a GRU structure, which can capture the temporal dependence of task offloading and resource status through historical information. Compared with traditional fully connected networks, the GRU network not only significantly improves the decision-making ability of the model but also accelerates the convergence speed.
[0045] The methods and devices of the present invention have other characteristics and advantages, which will be obvious from the accompanying drawings incorporated herein and the subsequent detailed description, or will be described in detail in the accompanying drawings incorporated herein and the subsequent detailed description. These drawings and detailed description are used together to explain the specific principles of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] By describing the exemplary embodiments of the present invention in more detail in conjunction with the accompanying drawings, the above and other objects, features, and advantages of the present invention will become more obvious. Among them, in the exemplary embodiments of the present invention, the same reference numerals generally represent the same components.
[0047] Figure 1 A flowchart showing the steps of a method for satellite-ground collaborative dynamic task partitioning and resource allocation based on deep reinforcement learning according to an embodiment of the present invention.
[0048] Figure 2 A schematic diagram showing a LEO satellite edge network model according to an embodiment of the present invention.
[0049] Figure 3 A schematic diagram showing the structure of an offloading algorithm according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0050] The preferred embodiments of the present invention will be described in more detail below. Although the preferred embodiments of the present invention are described below, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein.
[0051] To facilitate the understanding of the solutions and effects of the embodiments of the present invention, four specific application examples are given below. Those skilled in the art should understand that these examples are only for facilitating the understanding of the present invention, and any specific details are not intended to limit the present invention in any way.
[0052] Example 1
[0053] Figure 1 The flowchart shows the steps of a satellite-ground collaborative dynamic task partitioning and resource allocation method based on deep reinforcement learning according to an embodiment of the present invention.
[0054] As Figure 1 shown, the satellite-ground collaborative dynamic task partitioning and resource allocation method based on deep reinforcement learning includes:
[0055] Step 101, constructing a two-layer computing architecture for the collaboration between ground devices and LEO satellites;
[0056] Step 102, based on the two-layer computing architecture, establishing a computing model and a communication model of the satellite-ground integrated network through inter-satellite link technology;
[0057] Step 103, establishing an objective function and constructing the task partitioning and resource allocation problems in the satellite-ground integrated network as a Markov decision process;
[0058] Step 104, solving the task partitioning and resource allocation problems to generate the task segmentation rate and resource allocation strategy in the continuous action space.
[0059] In one example, constructing a two-layer computing architecture for the collaboration between ground devices and LEO satellites includes:
[0060] The computing task at time slot t is Q u,s (t) = {d u,s (t), c u,s , τ u,s (t)}, where d u,s (t) is the task size, c u,s is the CPU cycles required to compute one bit of data, and τ u,s (t) represents the maximum tolerance time of the task;
[0061] Let be the task allocation ratios for local processing, processing on the edge server integrated with the access satellite, processing on the adjacent satellite of the access satellite, respectively, and then constructing a two-layer computing architecture for the collaboration between ground devices and LEO satellites.
[0062] In one example, establishing the computing model includes:
[0063] Construct a local computing model and a LEO access satellite edge computing model according to the splitting ratio of the task, the computing capabilities of local devices and satellites. Define the amount of resources allocated by satellite s to adjacent satellites according to the ISL technology, share inter-satellite resources, and construct a cooperative satellite computing model using ISL.
[0064] In one example, establishing a communication model includes:
[0065] Let the spatial coordinates of the user at time slot t be O u,s (t) = [O u1,s (t), O u2,s (t), O u3,s (t)], where O u1,s (t), O u2,s (t), O u3,s (t) are the longitude, latitude and altitude of the user equipment respectively;
[0066] Let the position coordinates of the LEO satellite be O s (t) = [O s1 (t), O s2 (t), O s3 (t)], then the distance between the user and the satellite at this time slot is:
[0067]
[0068] Among them, R is the radius of the earth;
[0069] Let the user share spectrum resources in the way of orthogonal frequency division multiple access, then the uplink transmission rate of the user to the satellite is:
[0070]
[0071] Among them, B is the uplink bandwidth, p u,s (t) is the transmission power of the ground device, N0 is the satellite noise spectral density, h u,s (t) is the channel gain between the device and the satellite, expressed as h u,s (t) = G u,s |g u,s (t)h u,s (t)D u,s (t) -a | 2 Among them, G u,s is the antenna gain of the device, g u,s (t) is a complex Gaussian variable of Rayleigh fading, h u,s (t) is the fading component, and a is the path exponent;
[0072] For user-to-satellite, assuming the access satellite is s, the user equipment sends the task to the access satellite, and then the access satellite further sends the task to the cooperative satellite. Therefore, the user uplink transmission delay is:
[0073]
[0074] The propagation delay is:
[0075]
[0076] where c is the speed of light;
[0077] The communication delay from the user to the access satellite is the sum of the transmission delay and the propagation delay, expressed as:
[0078]
[0079] In one example, the total time delay of the computing model and the communication model is used as the objective function.
[0080] In one example, constructing the task partitioning and resource allocation problem in the satellite-ground integrated network as a Markov decision process includes:
[0081] The delay is divided into two parts. The first part is The second part Then the total delay of the system is:
[0082]
[0083] The goal is to minimize all delays, jointly optimize the task splitting ratio and resource allocation parameters, and construct the task partitioning and resource allocation problem in the satellite-ground integrated network as a Markov decision process.
[0084] In one example, the task partitioning and resource allocation problem is solved by a dynamic offloading algorithm that fuses a gated recurrent unit and deep deterministic policy gradient.
[0085] Figure 2 The schematic diagram of the LEO satellite edge network model according to an embodiment of the present invention is shown.
[0086] Specifically, as Figure 2 shown, in the LEO satellite edge network model, a cooperative two-layer computing architecture is composed of ground users and LEO satellites. Each LEO satellite has integrated an edge server and can provide computing resources and storage space; assuming the satellite set is S is the number of LEO satellites; the user set under the coverage of each satellite is where U s is the number of user equipment.
[0087] Step 1: Construct a two - layer computing architecture for the cooperation between ground devices and LEO satellites through the task information in the satellite edge network.
[0088] Suppose that at time slot t, the terminal generates a batch of computing tasks Q u,s (t) = {d u,s (t), c u,s , τ u,s (t)}, where d u,s (t) is the task size, c u,s is the CPU cycles required to compute one bit of data, and τ u,s (t) represents the maximum tolerance time of the task. The user can choose to offload most of the tasks to the connectable LEO satellites for processing, or forward the tasks to adjacent satellites through LEO satellites for collaborative processing; assume that the offloaded tasks can be divided in any proportion. Let be the task allocation ratios for local processing, processing on the edge server integrated in the access satellite, processing on the adjacent satellite of the access satellite, and
[0089]
[0090] Step 2: Based on the two - layer computing architecture, establish the computing model and the communication model of the space - ground integrated network through the inter - satellite link technology (ISL).
[0091] Regarding the local computing model, for the tasks processed by the user on the local device, the delay is:
[0092]
[0093] where is the computing power of the user device, with the unit of cycles / s.
[0094] LEO access satellite edge computing model: When α u,s (t)≠1, for the computing tasks offloaded to the LEO satellite, they are either processed on the access satellite of the device or processed on the adjacent satellite through the ISL. Suppose the offloading satellite is s, and the satellite has multi - core computing ability and can process the computing tasks of multiple devices in parallel. Suppose the computing power allocated to the user by the satellite at time t is Then the processing delay on the satellite is:
[0095]
[0096] Cooperative satellite computing model using the ISL: When , it means that the device tasks are not only processed on satellite s but also by the adjacent satellite of satellite s. Let The partial computing resources allocated to satellite s by adjacent satellites, which may be the previous satellite s-1 or the next satellite s+1 of satellite s. The processing delay of the cooperative satellite is as follows:
[0097]
[0098] Among them is the computing resources allocated by the cooperative satellite to the users under satellite s, satisfying At the same time, let be the resource size allocated by satellite s to adjacent satellites. For satellite s, it is considered that the received computing resources and the allocated resources One of them is zero or both are zero. Because for satellite s, when the load of this satellite is too high, the best choice is to receive resources from adjacent satellites with lower load, and it should not allocate its own computing resources at this time to increase its own load. For the allocated resources, there is Define When the value of is negative, it means that the satellite allocates its own resources to adjacent satellite s-1, When it is positive, it allocates its own resources to adjacent satellite s+1. Therefore, the computing resources received by satellite s can be obtained as:
[0099]
[0100] For the communication model, the distance between the user and the MEC server changes rapidly with time, and this change causes dynamic fluctuations in the channel conditions. Assume that the spatial coordinates of the user at time slot t are O u,s (t) = [O u1,s (t), O u2,s (t), O u3,s (t)], where O u1,s (t), O u2,s (t), O u3,s (t) are the longitude, latitude and altitude of the user equipment respectively. Assume that the position coordinates of the LEO satellite are O s (t) = [O s1 (t), O s2 (t), O s3 (t)], then the distance between the user and the satellite in this time slot is:
[0101]
[0102] Among them,
[0103] R is the radius of the earth.
[0104] Assume that users share spectrum resources in an orthogonal frequency division multiple access manner. Then, the uplink transmission rate of users to the satellite is as follows:
[0105]
[0106] where B is the uplink bandwidth, p u,s (t) is the transmission power of the ground device, N0 is the satellite noise spectral density, and h u,s (t) is the channel gain between the device and the satellite, expressed as h u,s (t) = G u,s |g u,s (t)h u,s (t)D u,s (t) -a | 2 , where G u,s is the antenna gain of the device, g u,s (t) is a complex Gaussian variable of Rayleigh fading, h u,s (t) is the fading component, including shadow fading, rain fading, etc., and a is the path exponent.
[0107] For the user-to-satellite scenario, let the access satellite be s. The user device first sends the task to the access satellite, and then the access satellite further sends the task to the cooperative satellite. Therefore, the user uplink transmission delay is:
[0108]
[0109] Considering the distance between the user device and the satellite, the propagation delay cannot be ignored. The propagation delay is:
[0110]
[0111] where c is the speed of light.
[0112] Since the data volume of the task result is usually very small, the transmission delay of the task result can be ignored. Then, the communication delay from the user to the access satellite is the sum of the transmission delay and the propagation delay, expressed as:
[0113]
[0114] For some satellite-to-satellite inter-satellite cooperative offloading tasks, due to the progress of ISL technology, especially the development of low-earth orbit satellite constellations, the channel bandwidth of inter-satellite optical communication can reach the GHZ level, and the delay is greatly reduced compared with traditional communication methods. Therefore, the communication delay of inter-satellite cooperation is ignored.
[0115] Step 3: Take the total time delay of the calculation model and the communication model as the objective function, construct the task partitioning and resource allocation problem in the satellite-ground integrated network, and construct this problem as a Markov decision process.
[0116] The total latency of a task includes communication latency and computing latency. During the edge offloading process of LEO satellites, task dependencies have an important impact on the offloading strategy. Assuming that tasks have a hierarchical order, first, ground users need to complete some initial tasks. Before the LEO satellite can further execute a task, it needs to receive the execution results of the previous user equipment to continue with subsequent tasks.
[0117] The latency of the system is divided into two parts. The first part is The second part Then the total latency of the system is:
[0118]
[0119] The goal is to minimize the latency of all devices for the task splitting ratio and resource allocation parameters, which include the resources allocated by the satellite to the user The resources allocated for inter-satellite cooperation And the resource allocation variables for the user through adjacent satellites Jointly optimize these parameters and model the problem:
[0120]
[0121] s.t.
[0122]
[0123]
[0124] (12a) and (12b) are the constraints for the task splitting ratio and resource allocation decision variables. (12c) indicates that the total computing resources allocated by the satellite to the user equipment should not be greater than its own resource limit minus the resources allocated to adjacent satellites. (12e) is the resource constraint for the resources allocated to adjacent satellites. (12f) indicates that the computing resources allocated by adjacent satellites to each user should not be greater than the total resources received by the satellite. (12g) represents that the completion time of the task should be less than the tolerance latency of the task.
[0125] Construct the above problem into a Markov model:
[0126] State space: At time t, the state space of the system consists of various environmental information of the system, including the task information Q u,s (t) of the device, the channel gain h u,s (t) between the device and the satellite, and the latency Then the state of the system is defined as:
[0127]
[0128] Action space: At time t, the agent can make a choice of action based on the current state information, and the action includes the device task allocation ratio and satellite resource allocation information Define a u,s (t) as the action, then there is:
[0129] a(t) = {A u,s (t), f u,s (t)} (14)
[0130] Reward function: The reward function is the immediate return obtained after the agent takes an action. The goal is to minimize the delay. Therefore, the immediate reward of the user equipment is defined as the opposite of the delay, and there is:
[0131]
[0132] where ξ is the penalty for task timeout.
[0133] Step 4: Solve the task partitioning and resource allocation problem through a dynamic offloading algorithm that fuses a gated recurrent unit and deep deterministic policy gradient, and generate the task segmentation rate and resource allocation strategy in the continuous action space.
[0134] Figure 3 shows a schematic diagram of the offloading algorithm structure according to an embodiment of the present invention.
[0135] The GDPG algorithm proposed by the present invention is as Figure 3 shown. On the traditional DDPG network structure, the actor network adopts a GRU network, which has a hidden state h(t) that can capture medium-term and short-term dependencies. At the end of the GRU network, we connect its output result to a fully connected layer to generate the offloading decision a(t); the GRU includes two main structures, the reset gate and the update gate. The calculation method of the reset gate is:
[0136] g(t) = σ(W gs ·s(t) + W gh ·h(t - 1) + b g ) (16)
[0137] where W gs and W gh are the weight matrices of the input state and the hidden state, and b g is the bias term. The role of the update gate is to determine the proportion of the previous hidden state in the current hidden state, specifically:
[0138] z(t) = σ(W zs ·s(t) + W zh ·h(t - 1) + b z(17)
[0139] The current hidden state is the output of the GRU network, which is jointly determined by the state at the previous moment and the candidate hidden state. The calculation method is as follows:
[0140]
[0141] where ⊙ is the Hadamard product, and the candidate hidden state is:
[0142]
[0143] Update the critic network parameter θ with the extracted training samples Q , and the update method is as follows: First, calculate the mean squared error loss function L(θ Q ):
[0144]
[0145] where N is the number of samples required for one training, and y u (t) is the target value, which is:
[0146] y(t) = r(t) + γQ′(s(t + 1), μ′(s(t + 1)|θ μ′ )|θ Q′ ) (21)
[0147] where γ is the discount factor. Subsequently, use the gradient descent method to update the critic network parameter:
[0148]
[0149] where α Q is the learning rate. For the actor network, its gradient calculation method is the derivative of the value function with respect to the action function parameter, and then use the gradient ascent method to update the parameter specifically as:
[0150]
[0151] For the target network, use the soft update method to update the target network parameter, specifically as:
[0152] θ Q ″ ← ωθ Q + (1 - ω)θ Q ″(25)
[0153] θ μ ← ωθ μ + (1 - ω)θ μ
[0154] The pseudocode of the present invention is given in Algorithm 1: Task partitioning and resource allocation of the GDPG algorithm.
[0155]
[0156]
[0157] Example 2
[0158] The present disclosure provides an electronic device, which includes: a memory storing executable instructions; and a processor that runs the executable instructions in the memory to implement the above-mentioned satellite-ground collaborative dynamic task partitioning and resource allocation method based on deep reinforcement learning.
[0159] The electronic device according to an embodiment of the present disclosure includes a memory and a processor.
[0160] The memory is used to store non-transitory computer-readable instructions. Specifically, the memory may include one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc.
[0161] The processor may be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device to perform desired functions. In an embodiment of the present disclosure, the processor is used to run the computer-readable instructions stored in the memory.
[0162] Those skilled in the art should understand that, in order to solve the technical problem of how to obtain good user experience effects, this embodiment may also include well-known structures such as communication buses, interfaces, etc., and these well-known structures should also be included in the protection scope of the present disclosure.
[0163] For a detailed description of this embodiment, reference may be made to the corresponding descriptions in the foregoing embodiments, and details will not be repeated here.
[0164] Example 3
[0165] The embodiment of the present disclosure provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, it implements the above-mentioned satellite-ground collaborative dynamic task partitioning and resource allocation method based on deep reinforcement learning.
[0166] A computer-readable storage medium according to an embodiment of the present disclosure stores non-transitory computer-readable instructions. When the non-transitory computer-readable instructions are run by a processor, all or part of the steps of the methods of the various embodiments of the present disclosure described above are executed.
[0167] The above computer-readable storage medium includes, but is not limited to: optical storage media (such as CD-ROMs and DVDs), magneto-optical storage media (such as MOs), magnetic storage media (such as magnetic tapes or external hard drives), media with built-in rewritable non-volatile memories (such as memory cards), and media with built-in ROMs (such as ROM cartridges).
[0168] Those skilled in the art should understand that the purpose of the above description of the embodiments of the present invention is only to exemplarily illustrate the beneficial effects of the embodiments of the present invention, and is not intended to limit the embodiments of the present invention to any of the examples given.
[0169] The various embodiments of the present invention have been described above. The above description is exemplary and not exhaustive, and is also not limited to the disclosed embodiments. Many modifications and variations will be obvious to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments.
Claims
1. A method for dynamic task division and resource allocation for satellite-ground collaboration based on deep reinforcement learning, characterized in that: include: Build a two-layer computing architecture that coordinates ground equipment and LEO satellites; Based on the two-layer computing architecture, a computing model and a communication model of a satellite-ground fusion network are established through inter-satellite link technology; Establishing an objective function, and constructing the task division and resource allocation problem in the satellite-ground fusion network as a Markov decision process; Solve the problem of task division and resource allocation, and generate the task division rate and resource allocation strategy in the continuous action space.
2. The method for dynamic task division and resource allocation for satellite-ground collaboration based on deep reinforcement learning according to claim 1, wherein: The two-layer computing architecture for building collaboration between ground equipment and LEO satellites includes: The computational task in time slot t is Q u,s (t) = {d u,s (t),c u,s ,τ u,s (t)},d u,s (t) is the task size, c u,s The CPU cycles required to calculate one bit of data, τ u,s (t) represents the maximum tolerable time of the task; Let α u,s (t),β u,s (t), The task allocation ratios are local processing, processing on the edge server integrated with the access satellite, and processing on the adjacent satellite of the access satellite, thereby building a two-layer computing architecture for collaboration between ground equipment and LEO satellites.
3. The method for dynamic task division and resource allocation for satellite-ground collaboration based on deep reinforcement learning according to claim 1, wherein: Establishing the calculation model includes: According to the task segmentation ratio, the computing power of local equipment and satellites, a local computing model and a LEO access satellite edge computing model are constructed. Based on the ISL technology, the size of resources allocated by satellites to adjacent satellites is defined to share intersatellite resources and build a collaborative satellite computing model using ISL.
4. The method for dynamic task division and resource allocation for satellite-ground collaboration based on deep reinforcement learning according to claim 1, wherein: Establishing the communication model includes: Assume that the spatial coordinate of the user at time slot t is O u,s (t) = [O u1,s (t),O u2,s (t),O u3,s (t)], where O u1,s (t),O u2,s (t),O u3,s (t) are the longitude, latitude and altitude of the user equipment, respectively; Assume the position coordinates of the LEO satellite are O s (t) = [O s1 (t),O s2 (t),O s3 (t)], then the distance between the user and the satellite in this time slot is: in, R is the radius of the Earth; Assuming that users share spectrum resources using orthogonal frequency division multiple access, the uplink transmission rate of users to the satellite is: Where B is the uplink bandwidth, p u,s (t) is the transmission power of the ground equipment, N0 is the satellite noise spectrum density, h u,s (t) is the channel gain between the device and the satellite, expressed as h u,s (t) = G u,s |g u,s (t)h u,s (t)D u,s (t) -a | 2 , where G u,s is the antenna gain of the device, g u,s (t) is the complex Gaussian variable of Rayleigh fading, h u,s (t) is the fading component, a is the path index; For user-to-satellite transmission, let the access satellite be s. The user equipment sends the task to the access satellite, which then sends the task to the cooperative satellite. Therefore, the user uplink transmission delay is: The propagation delay is: Where c is the speed of light; The communication delay from the user to the access satellite is the sum of the transmission delay and the propagation delay, expressed as:
5. The method for dynamic task division and resource allocation for satellite-ground collaboration based on deep reinforcement learning according to claim 1, wherein: The total time delay of the calculation model and the communication model is used as the objective function.
6. The method for dynamic task division and resource allocation for satellite-ground collaboration based on deep reinforcement learning according to claim 5, wherein: The task division and resource allocation problem in the satellite-ground fusion network is constructed as a Markov decision process, which includes: The delay is divided into two parts. The first part is Part 2 The total delay of the system is: The goal is to minimize all delays, jointly optimize the task partitioning ratio and resource allocation parameters, and construct the task partitioning and resource allocation problems in the satellite-ground fusion network as a Markov decision process.
7. The method for dynamic task division and resource allocation for satellite-ground collaboration based on deep reinforcement learning according to claim 1, wherein: The task division and resource allocation problems are solved through a dynamic offloading algorithm that combines gated recurrent units with deep deterministic policy gradients.
8. An electronic device, characterized in that: The electronic device comprises: A memory storing executable instructions; A processor, wherein the processor runs the executable instructions in the memory to implement the satellite-ground collaborative dynamic task division and resource allocation method based on deep reinforcement learning as described in any one of claims 1-7.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the satellite-ground collaborative dynamic task division and resource allocation method based on deep reinforcement learning as described in any one of claims 1-7.