Space-ground cooperative task offloading method based on time series prediction deep reinforcement learning

By constructing a three-layer computing system and a duel-based dual-depth Q-network, combined with a battery-aware dynamic adaptive QoE function, the slow convergence speed and insufficient QoE optimization of the space-air-ground collaborative MEC task offloading method in highly dynamic environments are solved, achieving efficient task offloading decision-making and improved user experience quality.

CN122633265APending Publication Date: 2026-08-25GUANGZHOU UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610657189.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-13
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

Existing space-air-ground collaborative MEC task offloading methods suffer from slow convergence speed, lack of time series modeling capabilities, low exploration efficiency, and lack of personalized QoE optimization in highly dynamic environments, making it difficult to achieve full coverage and high-resilience services under millisecond-level response requirements.

Method used

A three-layer computing system is constructed by employing the LSTM-integrated Dual-Depth Q-Network (PA-D3QN) algorithm and combining it with a battery-aware dynamic adaptive QoE function. This system calculates satellite position and channel quality in real time, optimizes latency and energy consumption through battery-aware dynamic weight adjustment, and uses historical load data to predict future resource status, thereby achieving intelligent task offloading decisions.

Benefits of technology

It significantly improves task response speed and system throughput efficiency, reduces user device power consumption, and achieves a comprehensive improvement in latency, power consumption and user experience quality, while dynamically adjusting decision preferences according to user status.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122633265A_ABST
    Figure CN122633265A_ABST
Patent Text Reader

Abstract

The application aims to provide a space-ground cooperative task offloading method based on time series prediction deep reinforcement learning, comprising: acquiring satellite dynamic basic parameters; acquiring time delay and energy consumption quantization results of different offloading decisions for the user equipment; constructing a battery-aware dynamic adaptive QoE function, generating corresponding decision rewards by distinguishing task completion and timeout discard states; converting the offloading decision problem into a sequence decision optimization problem; extracting spatial features and time series dependence characteristics of the current state of the intelligent agent, combining a duel network architecture with a long short-term memory network to obtain an optimal action and the value of the optimal action; and outputting an optimal task offloading decision for the current system state. The present disclosure significantly reduces the energy consumption of the user equipment, fully embodying the intelligent feature of dynamically adjusting the decision preference of the algorithm according to the user state.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of energy management and service quality assurance technology for mobile computing devices, and in particular to a method for offloading space-ground collaborative tasks based on temporal prediction deep reinforcement learning. Background Technology

[0002] With the development of communication technologies, emerging applications such as Augmented Reality (AR), Virtual Reality (VR), and autonomous driving place stringent demands on computing resources and network latency. They must not only guarantee millisecond-level end-to-end latency and highly reliable task execution, but also balance the limited battery life of terminals and the dynamically changing user experience (QoE). However, relying solely on terrestrial mobile edge computing (MEC) is insufficient to achieve full coverage and highly resilient services: terrestrial base station coverage is limited, communication blind spots exist in remote areas, oceans, and airspace, and the system is susceptible to natural disasters or network congestion, resulting in insufficient system robustness. In contrast, Low Earth Orbit (LEO) satellites, with their wide-area coverage, high mobility, and strong resilience, provide crucial support for building an integrated space-ground network. By integrating space-based and terrestrial edge resources, three-dimensional synergy of computing power and service continuity can be achieved, providing ubiquitous computing power for highly mobile and widely distributed intelligent terminals. Therefore, intelligent task offloading mechanisms for space-ground collaborative environments have become a key path to overcome the bottleneck of ground computing power and achieve intelligent services across the entire domain. However, the high dynamism of LEO networks poses a severe challenge to the real-time and adaptive nature of offloading decisions. Traditional methods such as convex optimization and integer programming rely on static models and are difficult to solve online; while heuristic algorithms such as genetic algorithms and particle swarm optimization cannot meet the millisecond-level response requirements due to their high computational complexity and slow convergence. Deep reinforcement learning (DRL), with its powerful nonlinear fitting capabilities and online adaptive characteristics, can learn the optimal strategy through trial and error in unknown and non-stationary environments, naturally meeting the complexity and real-time requirements of space-ground collaborative offloading, and thus has become the most promising technological paradigm.

[0003] However, existing space-air-ground collaborative MEC task offloading methods still suffer from problems such as slow convergence speed, lack of time series modeling capability, low exploration efficiency, and lack of personalized QoE optimization in highly dynamic environments.

[0004] In view of this, this disclosure provides a method that combines the LSTM-integrated Dual Deep Q Network (PA-D3QN) algorithm with a battery-aware dynamic adaptive QoE function to optimize intelligent task offloading decisions in an integrated air-space-ground network, thereby achieving a comprehensive improvement in latency, energy consumption, and user experience quality. Summary of the Invention

[0005] The purpose of this disclosure is to provide a method for unloading space-ground collaborative tasks based on temporal prediction deep reinforcement learning, which aims to solve at least one technical problem in the prior art.

[0006] The technical solution disclosed herein is:

[0007] A method for offloading space-ground collaborative tasks based on temporal prediction deep reinforcement learning includes:

[0008] Construct a three-layer computing system comprising user equipment, ground edge servers, and low-Earth orbit satellites to obtain dynamic basic parameters of the satellites;

[0009] For any of the user equipment, a local computing queue, an edge transmission queue, and a satellite transmission queue are established; the edge transmission queue and the satellite transmission queue are calculated independently, and a first-in-first-out scheduling strategy is adopted. Then, latency and energy consumption calculation models are constructed under three paths: local task execution, edge offloading, and satellite offloading, respectively, to obtain the latency and energy consumption quantification results for different offloading decisions for the user equipment.

[0010] Based on the latency and energy consumption quantification results, a battery-aware dynamic adaptive QoE function is constructed. The optimization weights of latency and energy consumption are dynamically adjusted according to the real-time battery power of the user device, and corresponding decision rewards are generated by distinguishing between task completion and timeout discard states.

[0011] Based on the satellite dynamic basic parameters, the local computing queue, the edge transmission queue, the satellite transmission queue, and the user equipment battery level, the state space and the action space for local execution, edge offloading, and satellite offloading are obtained, and the offloading decision problem is transformed into a sequence decision optimization problem.

[0012] Treating any user device as an independent intelligent agent, a duel dual-deep Q network fused with LSTM is constructed to extract the spatial features and temporal dependencies of the agent's current state. The optimal action and its value are obtained by combining the duel network architecture with a long short-term memory network.

[0013] Dual Q-learning is used to evaluate the optimal action and its value, and finally outputs the optimal task offloading decision for the current system state.

[0014] The system comprises a three-layer computing system consisting of user equipment, ground edge servers, and low-Earth orbit satellites, and acquires dynamic basic parameters of the satellite, including one or more of the following: satellite position update, channel quality, satellite computing power, and dynamic transmission capacity.

[0015] The channel quality is smoothly mapped to the [0.1, 0.95] interval based on the received signal-to-noise ratio (SNR) using the Sigmoid function. The received SNR of the user equipment and satellite link is expressed as: ;

[0016] in, This refers to the satellite's launch power. and These are the transmit and receive antenna gains, respectively. Free space path loss; Indicates atmospheric loss; This represents the satellite elevation angle at time t; This represents the satellite slant range at time t;

[0017] Among them, elevation angle Based on spherical geometry calculations, it can be expressed as:

[0018] ;

[0019] The radius of the Earth; The geocentric angle between user equipment u and satellite s is expressed as: ;

[0020] Let 'a' be the semi-versus of the geocentric angle formula, expressed as:

[0021] ;

[0022] User equipment The position is ,satellite In the time slot The position is The difference between latitude and longitude is expressed as:

[0023] ;

[0024] ;

[0025] in, The slope distance is expressed as:

[0026] ;

[0027] in, For each LEO satellite At altitude;

[0028] The dynamic transmission capacity of the satellite link is expressed as: ;

[0029] in, This is the transmission efficiency factor; and These represent the minimum and maximum transmission rates supported by the link, respectively. Indicates the load adjustment factor;

[0030] The satellite In the time slot computing power Represented as:

[0031] ;

[0032] in, and These are the satellite's basic and maximum computing capabilities, respectively. It is the normalized satellite-to-ground channel quality factor.

[0033] The Methods for obtaining the load adjustment factor include:

[0034] The load adjustment factor is based on satellite. Current load status To dynamically adjust the effective rate allocated to each user, expressed as:

[0035] ;

[0036] in, Indicates in time slot For satellite The number of users served; Designated as a satellite The maximum number of concurrent users that can be served simultaneously; This represents the minimum load factor.

[0037] The method for obtaining the normalized satellite-to-ground channel quality factor includes:

[0038] ;

[0039] in, and These represent the signal-to-noise ratio thresholds corresponding to the worst and best acceptable channel quality, respectively.

[0040] The aforementioned battery-aware dynamic adaptive QoE function dynamically adjusts the optimization weights of latency and energy consumption based on the real-time battery level of the user device, and generates corresponding decision rewards by distinguishing between task completion and timeout / discard states, including:

[0041] Construct a QoE function Q(s,a) to distinguish between two cases: successful task completion and task timeout / discard.

[0042] ;

[0043] in, It is a fixed basic reward obtained upon successful completion of a task; It is the total cost incurred in performing the task; It is the penalty coefficient for task abandonment; It is about minimizing latency;

[0044] Constructing the cost function and make the cost function Task processing delay Total energy consumption generated by user equipment The weighted sum is composed of a weight that is related to the device's battery power. Related adaptive functions Dynamic adjustment;

[0045] By setting a non-linear deadline-sensitive reward function Incentivize agents to prioritize tasks that are about to time out.

[0046] The construction cost function and make the cost function Task processing delay Total energy consumption generated by user equipment The weighted sum is composed of a weight that is related to the device's battery power. Related adaptive functions Dynamic adjustment, including:

[0047] ;in, It is a weighting coefficient, and ;

[0048] in, It is a sensitivity parameter; It is the battery power threshold;

[0049] The method involves setting a non-linear deadline-sensitive reward function. Incentivize agents to prioritize tasks that are about to time out, including:

[0050] ;

[0051] in, It is the steepness parameter of the reward curve; It is the maximum tolerable time to complete the task; This is the threshold ratio of task completion delay to maximum tolerable delay.

[0052] The process involves treating any user device as an independent intelligent agent, constructing a duel-based dual-depth Q-network fused with LSTM, extracting the spatial features and temporal dependencies of the agent's current state, and utilizing a duel network architecture combined with a long short-term memory network to obtain the optimal action and its value. This includes:

[0053] The current task information, queue status, satellite coverage, and battery level information are used to form an instantaneous state vector; and the load levels of edge servers and satellites over the past L time slots are recorded through the historical load matrix H.

[0054] The instantaneous state branch is used to extract the spatial features of the current state through a two-layer fully connected network. At the same time, the historical load branch is used to capture the temporal dependence and changing trend of the load. The feature vectors extracted by the two branches are concatenated to form a comprehensive feature vector.

[0055] The action value function is decomposed into a state value function and an action advantage function through a duel network architecture. The outputs of the two branches are merged through a special aggregation layer to obtain the final Q value.

[0056] Specifically, the action value function is implemented through a duel network architecture. Decomposed into state value function and action advantage function The outputs of the two branches are merged through a special aggregation layer to obtain the final Q value, including:

[0057] ;

[0058] in, Assess the quality of the situation itself; Evaluate the advantage of choosing a particular action compared to the average action in this state; s represents the current environment state; s represents the action performed by the task. These are all the parameters of the state-value network; These are all the parameters of the action advantage network; It is the action space. Indicates the total number of actions.

[0059] The process of obtaining the optimal action and the value of the optimal action also includes: the weights of the fully connected layer. and bias The steps for introducing learnable parameterized noise are as follows:

[0060] The learnable parameterized noise is expressed as: ;in, As weight; For bias; It is the feature vector that is input into the noisy network after state preprocessing.

[0061] The weight and bias The acquisition process is as follows:

[0062] ;

[0063] ;

[0064] in, , and , ε is a trainable parameter; ε is a sampled noise variable.

[0065] The beneficial effects of this disclosure include at least the following:

[0066] The method described in this disclosure overcomes the limitations of terrestrial networks by utilizing satellite dynamic modeling. It integrates a satellite orbital dynamics model based on Kepler's laws to calculate LEO satellite positions, channel quality, and time-varying propagation delays in real time, effectively overcoming the coverage limitations of terrestrial MECs in remote, marine, and disaster-stricken scenarios. Simultaneously, by embedding an LSTM module into the Duel Dual-Depth Q-network, it uses historical load sequences to predict future resource states, proactively avoiding congested nodes, reducing average task processing latency, and significantly improving task response speed and system throughput efficiency. Furthermore, this method uses a dynamic weighting function based on real-time battery power to achieve smooth adjustment of latency-energy consumption weights, significantly reducing user equipment energy consumption and fully demonstrating the intelligent characteristics of the algorithm in dynamically adjusting decision preferences based on user status. Attached Figure Description

[0067] Figure 1 A three-tier architecture system model for space-ground collaboration

[0068] Figure 2 Here is a diagram of the PA-D3QN algorithm architecture;

[0069] Figure 3 This is a flowchart of the PA-D3QN algorithm. Detailed Implementation

[0070] The technical solution of this disclosure will be further described below with reference to the accompanying drawings.

[0071] In recent years, research on task offloading for integrated air-space-ground MEC networks has gradually incorporated the DRL framework, and has made initial progress in both single-objective optimization and multi-objective optimization.

[0072] In single-objective optimization, existing research mainly focuses on a single performance metric, with typical objectives including minimizing task execution latency or reducing terminal device power consumption. These methods typically model the offloading problem as a standard Markov decision process, employing basic deep reinforcement learning algorithms such as DQN and DDPG to encode environmental information such as channel quality, computational load, or queue length in the state space, and using a single scalar (e.g., average latency or total power consumption) as a reward signal to drive policy learning. While achieving lower latency or power consumption in static or slowly changing network environments, their fundamental limitation lies in simplifying the complex user experience to a single dimension, ignoring the multi-dimensional coupling characteristics inherent in QoE. In space-ground co-location scenarios, excessively pursuing low latency may lead to frequent offloading to high-power satellite links, accelerating terminal power consumption; while simply optimizing power consumption may cause task timeouts due to insufficient local computing power. Therefore, single-objective optimization struggles to simultaneously meet users' comprehensive needs for response speed, energy efficiency, and task reliability, failing to support the dynamic and personalized service experience in real-world applications.

[0073] In multi-objective optimization, existing research focuses on the collaborative optimization of multiple conflicting performance metrics such as task latency, terminal energy consumption, and task success rate. This is typically achieved by constructing composite reward functions or introducing multi-objective decision-making mechanisms. Common approaches include linearly weighting multiple objectives into a single scalar reward, or employing Pareto-optimal policy search frameworks to maintain solution diversity. Some works have also designed hierarchical reinforcement learning structures to adjust the priority of sub-objectives through high-level policies. While these methods alleviate the conflict between objectives to some extent, they still face significant challenges in the highly dynamic environment of integrated air-space-ground networks: weighted summation methods rely on preset and fixed weight coefficients, making it difficult to adapt to real-time changes in user battery status, task type, or network load; Pareto-based methods, while generating diverse policies, suffer from high computational complexity and slow convergence, failing to meet the timeliness requirements of millisecond-level offloading decisions; and hierarchical architectures often lack the ability to model the temporal dynamics of the environment, failing to predict future resource availability, resulting in multi-objective trade-offs based solely on the current instantaneous state, leading to short-sighted decision-making. More importantly, existing solutions generally fail to incorporate users' personalized needs into the QoE evaluation system, leaving multi-objective optimization still at the "system-centric" perspective, making it difficult to achieve truly user-centric differentiated service guarantees.

[0074] While the aforementioned methods alleviate the target conflict problem to some extent, when applied to space-ground integrated MEC scenarios, they require a large number of samples to converge in high-dimensional state and action spaces, resulting in slow convergence speeds. Furthermore, they lack temporal modeling capabilities, making it impossible to effectively utilize historical information from edge servers and satellite loads to predict future resource availability. Additionally, the standard ε-greedy exploration strategy is inefficient in complex environments. Moreover, the fixed-weight optimization objective does not consider the dynamic impact of user device battery status on decision preferences, failing to achieve truly user-centric experience quality optimization.

[0075] To address the above issues, given limited terrestrial network coverage, high-speed satellite movement, and limited user device battery power, this paper proposes a computational offloading scheme that balances task processing latency, device power consumption, task completion rate, and user experience quality, and adaptively makes optimal offloading decisions. This scheme constructs a three-layer heterogeneous computing architecture by integrating low Earth orbit satellites, and addresses dynamic user needs in remote areas and over wide regions. Furthermore, it utilizes an improved deep reinforcement learning algorithm to dynamically optimize task offloading decisions, thereby achieving the best user experience quality. Specific Implementation Example 1:

[0077] This disclosure provides an embodiment:

[0078] A method for offloading space-ground collaborative tasks based on temporal prediction deep reinforcement learning includes the following steps:

[0079] Step S1: Construct a three-tiered heterogeneous computing offloading architecture:

[0080] Establish containing A three-tier computing system consisting of one user equipment (UE), M ground edge servers (ES), and K low Earth orbit (LEO) satellites, such as Figure 1 As shown, an integrated satellite orbital dynamics model is used to calculate satellite position, coverage area, satellite-to-ground distance, and elevation angle in real time. Based on free-space path loss, atmospheric attenuation, and elevation angle factors, the satellite-to-ground channel quality, propagation delay, satellite computing power, and transmission capacity are dynamically calculated.

[0081] The satellite position update uses the Kepler orbital dynamics equations, taking into account the effects of Earth's rotation.

[0082] Meanwhile, the channel quality is smoothly mapped to the [0.1, 0.95] interval based on the received signal-to-noise ratio (SNR) using the Sigmoid function. The received signal-to-noise ratio (SNR) of the user equipment-satellite link is shown below:

[0083] in, For satellite launch power, and These represent the transmit and receive antenna gains, respectively. For free space path loss, This indicates atmospheric loss.

[0084] Satellite computing and transmission capabilities are dynamically adjusted based on channel quality and current load factor. The dynamic transmission capacity of the satellite link is shown below:

[0085]

[0086] in, For transmission efficiency factor, and These represent the minimum and maximum transmission rates supported by the link, respectively, simulating the physical performance limitations of the communication module. This represents the load adjustment factor, and its specific calculation is shown below:

[0087]

[0088] Load adjustment factor based on satellite Current load status To dynamically adjust the effective rate allocated to each user. Indicates in time slot For satellite The number of users of the service. Designated as a satellite The maximum number of concurrent users that can be served at the same time. The minimum load factor is set to ensure that the link remains available even under heavy load conditions.

[0089] In order to dynamically model the computing power of satellites, the satellites In the time slot computing power The definition is as follows:

[0090]

[0091] in, and These are the satellite's basic structure and maximum computing power, respectively. This is the normalized satellite-to-ground channel quality factor. A higher value indicates better channel conditions. The specific calculation formula is as follows:

[0092]

[0093] in, and These represent the signal-to-noise ratio (SNR) thresholds corresponding to the worst and best acceptable channel quality, respectively. When the actual SNR is lower than... When the channel quality factor is 0, it indicates that the link is unavailable; when it is higher than 0, the channel quality factor is 0. When the factor is 1, it indicates that the channel quality is at its best.

[0094] This step integrates satellite orbital dynamics models to accurately model the dynamic changes in satellite position, coverage status, and channel quality. It addresses the shortcomings of existing technologies that treat satellites as static nodes, solves the problem of limited coverage in terrestrial MEC networks, and provides wide-area computing services for remote areas, oceans, disaster areas, and other regions.

[0095] Step S2: Establish the queue model and latency-energy consumption model:

[0096] A local computing queue, an edge transmission queue, and a satellite transmission queue are established for each user device; a separate computing queue is maintained for each user for each edge server and satellite. The queues employ a first-in, first-out (FIFO) scheduling strategy, and a maximum tolerable latency for each task is set. Tasks that time out will be discarded.

[0097] The system employs a first-in, first-out (FIFO) scheduling strategy, establishing independent queues for each computation and transmission process within the three-layer architecture.

[0098] Local compute queue: Each user device maintains a local compute queue. The queue length is determined by the amount of tasks remaining in the previous time slot, the amount of task data arriving in the current time slot, the amount of data processed locally, and the amount of data discarded due to timeout. The local processing rate depends on the device's CPU frequency, time slot duration, and task computation density.

[0099] Edge computing queues: Each edge server maintains an independent computing queue for each user device it is associated with. Queue updates take into account the amount of data arriving from user devices, the amount of data processed by the server, and the amount of data discarded due to timeout. The server's processing capacity is evenly distributed among connected users.

[0100] Satellite computing queue Each satellite maintains a computing queue for each user equipment, with update rules similar to those for edge computing queues. The satellite's computing power is dynamic, derived from the basic computing power and maximum computing power through linear interpolation using a normalized satellite-to-ground channel quality factor, reflecting the impact of channel conditions on the allocation of satellite computing resources.

[0101] Edge transmission queue When a task is unloaded to the edge server, it first enters the edge transmission queue on the user device side. The queue update is determined by the amount of unloaded task data arriving, the amount of data actually transmitted to the edge server (determined by the smaller of the channel transmission capacity and the available data in the queue), and the amount of data discarded due to timeout.

[0102] Satellite transmission queue Tasks offloaded to satellites enter the satellite transmission queue on the user equipment side. Their status update rules are similar to those of the edge transmission queue, and the transmission rate is determined by the dynamic channel conditions from the user equipment to the low-Earth orbit satellite.

[0103] All queues track task waiting time in real time. When the cumulative delay of a task in the queue exceeds its deadline, the task will be discarded by the system to ensure the effective use of system resources.

[0104] The latency model described above measures the total latency experienced by a task from its generation to its final completion in units of time slots, defined as the time from task generation (time slot t) to its complete processing (time slot t). The total number of time slots actually experienced. The accumulation process of delay varies depending on the offloading decision:

[0105] Local offload latency: Tasks are placed in a local computing queue and wait for service. Processing begins when a task reaches the head of the queue. The processing rate is determined by the device's local computing power until the task is completed.

[0106] Edge offloading latency: Task execution consists of two sequential phases: transmission and computation. The task first enters the edge transmission queue and is transmitted to the edge server via a wireless channel. After transmission, it enters the edge server's computation queue for processing until the task is completed.

[0107] Satellite offload delay: Mission execution consists of three sequential phases: signal propagation, transmission, and computation. First, there is a fixed two-way signal propagation delay. Then, the signal enters the satellite transmission queue. After transmission, the signal enters the satellite computation queue for processing until the mission is completed.

[0108] The energy consumption model described herein focuses on the energy consumption of user equipment. Total energy consumption is the sum of energy consumed by the device in each time slot throughout the lifecycle of a task. The power state of the equipment is determined by the position of the task in that time slot.

[0109] Local offload power consumption: Device power consumption occurs only during actual computation. Computational power is related to CPU frequency and is determined by the effective capacitance factor and the cube of the frequency. No additional standby power consumption is included while tasks are queuing in the local queue.

[0110] Edge offloading power consumption: Device power consumption consists of two parts: data transmission power consumption and standby power consumption. During task transmission, the device operates at a constant transmission power; after transmission is completed, when the task is queued in the server's computing queue or being processed, the device is in standby mode, consuming standby power.

[0111] Satellite offloading energy consumption: The energy consumption structure is similar to that of edge offloading, but the transmission power is dynamically changing, and its value depends on the satellite-to-ground distance. The power during this stage is shown in the following formula:

[0112]

[0113] in and These are the minimum and maximum transmission power, respectively. and These are the preset minimum and maximum reference distances, respectively. During the waiting period for the results, in addition to the server's computation time, there is also the signal propagation time; during this period, the devices are in standby mode.

[0114] Step S3: Design a battery-aware dynamic adaptive QoE function:

[0115] Define a comprehensive QoE function Q(s,a) to distinguish between two cases: successful task completion and task timeout / discard. The specific formula is as follows:

[0116]

[0117] in, It is a fixed basic reward that can be obtained upon successful completion of the task. It is the total cost incurred in performing the task. This is the penalty coefficient for task abandonment. For successfully completed tasks, the total reward consists of a base reward and an additional deadline-sensitive reward.

[0118] Comprehensive cost function It is caused by task processing delay Total energy consumption generated by user equipment The weighted sum is composed of a factor related to the device's battery level. Related adaptive functions Dynamic adjustment, as shown below:

[0119]

[0120] To ensure the numerical comparability of the two components, this model directly uses the original value of the time delay (unit: number of time slots), and obtains the original energy consumption after normalization and scaling. .

[0121] This is the core of the dynamic QoE model, mapping the user's battery state to a weighted coefficient. This embodiment uses a hysteresis-based sigmoid function to smoothly adjust the weights, aiming to build a highly stable decision-making system that prioritizes performance when the battery is fully charged and ensures that minor fluctuations in battery level do not cause drastic changes in decision-making near a threshold. See below:

[0122]

[0123] in, It is a sensitivity parameter used to adjust the steepness of the weight curve. It is the battery power threshold; when the battery power is low... Higher than hour, The value of is close to 1, making the cost function more focused on minimizing the time delay. When the battery level is low hour, The value of is close to 0, making the cost function more focused on minimizing energy consumption. In the formula A hysteresis effect is partially introduced, the main purpose of which is to increase the stability of the decision and prevent frequent and unnecessary fluctuations in the weights when the battery level fluctuates slightly around the threshold.

[0124] To incentivize the agent to prioritize tasks that are about to time out, the total reward includes a non-linear deadline-sensitive reward function. As shown below:

[0125]

[0126] It is the steepness parameter of the reward curve. This is the maximum tolerable time for task completion. When the ratio of task completion delay to the maximum tolerable delay exceeds... At that point, the reward will decrease sharply. This function allows the additional reward to be obtained as quickly as the task is completed.

[0127] This step involves designing a battery-aware, dynamically adaptive QoE function that smoothly adjusts the latency-energy consumption weights based on the real-time battery level of the user's device, resulting in higher task completion rates, lower average latency, better energy management, and a better user experience.

[0128] Step S4: Construct a Markov Decision Process (MDP):

[0129] 1. State Space: In Markov decision-making, the user equipment... In the time slot The state space is defined as The vector representation is shown in the following formula:

[0130]

[0131] This indicates the current data size of the task that has arrived. This is a local computation queue. and For transmission queues. and For calculating the queue. This refers to the coverage status of each satellite for user u. This represents the battery percentage of the user's device. To predict future load on edge servers and satellites, a historical load information matrix is ​​also included in the disclosure. Record the past The load variation of each time slot is represented by the following vector:

[0132]

[0133] in, and These represent edge servers. and satellite In the time slot The load level, i.e., how many user devices are served by the edge server and satellite respectively in this time slot.

[0134] Motion space: Motion space This defines the set of all possible offload decisions for each user equipment (UE) at time t. Actions Defined as: .in, =0 indicates that the task is processed locally. This indicates that the data will be offloaded to edge server e. The corresponding unload is to satellite S.

[0135] Step S5: Construct a Duel Dual-D3QN network fused with LSTM (PA-D3QN):

[0136] To address the complexity and dynamism of computational offloading in integrated space-air-ground MEC networks, this embodiment employs a predictive adaptive dueling dual-depth Q-network-based collaborative offloading algorithm (PA-D3QN). The PA-D3QN algorithm architecture diagram and flowchart are shown below. Figure 2 , Figure 3As shown, the algorithm employs a distributed multi-agent architecture, where each user device acts as an independent agent, learning the optimal offloading strategy through deep reinforcement learning. The core innovation of the algorithm lies in combining a duel network architecture with a long short-term memory network, enabling the agent to accurately assess the value of the current state and effectively utilize historical information to predict future network states.

[0137] The overall algorithm architecture comprises four main components: a state-aware module, a history learning module, a decision network module, and an adaptive reward module. The state-aware module collects current environmental information, including task characteristics, queue status, and battery level. The history learning module processes historical load information through an LSTM network to provide predictive guidance for decision-making. The decision network module employs a duel architecture to learn state value and action advantage separately. The adaptive reward module dynamically adjusts the optimization objective based on the device's battery status.

[0138] The network input layer receives two types of input data: the instantaneous state vector s contains information such as the current task information, queue status, satellite coverage, and battery level; and the historical load matrix H records the load levels of edge servers and satellites over the past L time slots, providing a data foundation for time series prediction.

[0139] Feature Extraction Layer: Feature extraction employs a dual-branch structure. The immediate state branch extracts the spatial features of the current state through two fully connected layers; the historical load branch inputs the historical load matrix H into the LSTM layer, utilizing its gating mechanism to capture the temporal dependence and changing trends of the load. Finally, the feature vectors extracted from the two branches are concatenated to form a comprehensive feature vector that integrates immediate information and historical patterns.

[0140] Duel Architecture: Duel network architecture incorporates action value functions Decomposed into state value function and action advantage function The outputs of the two branches are merged through a special aggregation layer to obtain the final Q value as shown below:

[0141]

[0142] in, Assess the quality of the situation itself. The decomposition evaluates the advantage of choosing a particular action in a given state compared to the average action. This decomposition allows the network to learn state values ​​independently without requiring a complete evaluation for each action, thus accelerating convergence and improving sample utilization efficiency, making it particularly suitable for scenarios with large action spaces.

[0143] Weights in the fully connected layer and bias Learnable parameterized noise is introduced, in the following form:

[0144]

[0145] In noisy networks, their weights and bias It is no longer a deterministic value, but consists of a mean term and a noisy standard deviation term, as shown in the formulas:

[0146]

[0147]

[0148] Here, μ and σ are trainable parameters, and ε is a sampled noise variable. Compared with the traditional ε-greedy strategy, this mechanism shifts the exploration from the action space to the parameter space, allowing the exploration level to be automatically adjusted as training progresses, and the randomness of exploration varies in different states, thus achieving intelligent adaptive exploration based on state uncertainty.

[0149] This step utilizes a duel-style dual-depth Q-network algorithm that combines LSTM and Noisy Net to predict future resource availability using historical load information, thereby enabling intelligent real-time offloading decisions.

[0150] Step S6: Implement dual Q learning and experience replay training:

[0151] The system employs a dual-network architecture: an evaluation network and a target network. The evaluation network is responsible for action selection and parameter updates, while the target network provides a stable benchmark for value assessment. During training, the agent processes the experience tuples from each time slot. The data is stored in a playback buffer of size C, and temporal correlation is broken through random sampling.

[0152] Dual Q-learning mechanism: using an evaluation network to select the optimal action for the next state. However, the value of the action is evaluated using the target network. This decouples action selection and value assessment, effectively avoiding the systematic overestimation of Q-values ​​in standard Q-learning. The network is optimized by minimizing the mean square loss of the TD error, and the target network employs a soft update strategy. Slowly track and evaluate the network to ensure training stability. Specific Implementation Example 2:

[0154] This disclosure also provides an embodiment:

[0155] An electronic device includes: a storage medium and a processing unit; wherein the storage medium is used to store a computer program, and the processing unit exchanges data with the storage medium, and is used to execute the computer program through the processing unit during the unloading of a space-ground collaborative mission, performing the steps of the method as described in Specific Embodiment 1.

[0156] A computer-readable storage medium storing a computer program; when the computer program is run, it performs the steps of the method as described in Specific Embodiment 1.

[0157] In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wireless, wireline, optical fiber, RF, etc., or any suitable combination thereof.

[0158] The above disclosures only cover a few specific implementation scenarios. However, this disclosure is not limited to these, and any variations that can be conceived by those skilled in the art should fall within the protection scope of this disclosure. The serial numbers in this disclosure are for descriptive purposes only and do not represent the superiority or inferiority of the implementation scenarios.

Claims

1. A method for unloading space-ground collaborative tasks based on temporal prediction deep reinforcement learning, characterized in that, include: Construct a three-layer computing system comprising user equipment, ground edge servers, and low-Earth orbit satellites to obtain the satellite's dynamic basic parameters; For any of the user equipment, a local computing queue, an edge transmission queue, and a satellite transmission queue are established; the edge transmission queue and the satellite transmission queue are calculated independently, and a first-in-first-out scheduling strategy is adopted. Then, latency and energy consumption calculation models are constructed under three paths: local task execution, edge offloading, and satellite offloading, respectively, to obtain the latency and energy consumption quantification results for different offloading decisions for the user equipment. Based on the latency and energy consumption quantification results, a battery-aware dynamic adaptive QoE function is constructed. The optimization weights of latency and energy consumption are dynamically adjusted according to the real-time battery power of the user device, and corresponding decision rewards are generated by distinguishing between task completion and timeout discard states. Based on the satellite dynamic basic parameters, the local computing queue, the edge transmission queue, the satellite transmission queue, and the user equipment battery level, the state space and the action space for local execution, edge offloading, and satellite offloading are obtained, and the offloading decision problem is transformed into a sequence decision optimization problem. Treating any user device as an independent intelligent agent, a duel dual-deep Q network fused with LSTM is constructed to extract the spatial features and temporal dependencies of the agent's current state. The optimal action and its value are obtained by combining the duel network architecture with a long short-term memory network. Dual Q-learning is used to evaluate the optimal action and its value, and finally outputs the optimal task offloading decision for the current system state.

2. The method for unloading space-ground collaborative tasks based on temporal prediction deep reinforcement learning according to claim 1, characterized in that, The system comprises a three-layer computing system consisting of user equipment, ground edge servers, and low-Earth orbit satellites, and acquires dynamic basic parameters of the satellite, including one or more of the following: satellite position update, channel quality, satellite computing power, and dynamic transmission capacity. The channel quality is smoothly mapped to the [0.1, 0.95] interval based on the received signal-to-noise ratio (SNR) using the Sigmoid function. The received SNR of the user equipment and satellite link is expressed as: ; in, This refers to the satellite's launch power. and These are the transmit and receive antenna gains, respectively. Free space path loss; Indicates atmospheric loss; This represents the satellite elevation angle at time t; This represents the satellite slant range at time t; B represents the power spectral density of Gaussian white noise; B represents the system channel bandwidth. The dynamic transmission capacity of the satellite link is expressed as: ; in, This is the transmission efficiency factor; and These represent the minimum and maximum transmission rates supported by the link, respectively. Indicates the load adjustment factor; The satellite in the time slot computing power Represented as: ; in, and These are the satellite's basic and maximum computing capabilities, respectively. It is the normalized satellite-to-ground channel quality factor.

3. The method for unloading space-ground collaborative tasks based on temporal prediction deep reinforcement learning according to claim 2, characterized in that, The Methods for obtaining the load adjustment factor include: The load adjustment factor is based on satellite. Current load status To dynamically adjust the effective rate allocated to each user, expressed as: ; in, Indicates in time slot For satellite The number of users served; Designated as a satellite The maximum number of concurrent users that can be served simultaneously; This represents the minimum load factor.

4. The space-ground collaborative task offloading strategy according to claim 2, characterized in that, The method for obtaining the normalized satellite-to-ground channel quality factor includes: ; in, and These represent the signal-to-noise ratio thresholds corresponding to the worst and best acceptable channel quality, respectively.

5. The method for unloading space-ground collaborative tasks based on temporal prediction deep reinforcement learning according to claim 1, characterized in that, The aforementioned battery-aware dynamic adaptive QoE function dynamically adjusts the optimization weights of latency and energy consumption based on the real-time battery level of the user device, and generates corresponding decision rewards by distinguishing between task completion and timeout / discard states, including: Construct a QoE function Q(s,a) to distinguish between two cases: successful task completion and task timeout / discard. ; in, It is a fixed basic reward obtained upon successful completion of a task; It is the total cost incurred in performing the task; It is the penalty coefficient for task abandonment; It is about minimizing latency; Constructing the cost function and make the cost function Task processing delay Total energy consumption generated by user equipment The weighted sum is composed of a weight that is related to the device's battery power. Related adaptive functions Dynamic adjustment; By setting a non-linear deadline-sensitive reward function Incentivize agents to prioritize tasks that are about to time out.

6. The method for unloading space-ground collaborative tasks based on temporal prediction deep reinforcement learning according to claim 5, characterized in that, The construction cost function and make the cost function Task processing delay Total energy consumption generated by user equipment The weighted sum is composed of a weight that is related to the device's battery power. Related adaptive functions Dynamic adjustment, including: ;in, It is a weighting coefficient, and ; in, It is a sensitivity parameter; It is the battery power threshold; The method involves setting a non-linear deadline-sensitive reward function. Incentivize agents to prioritize tasks that are about to time out, including: ; in, It is the steepness parameter of the reward curve; It is the maximum tolerable time to complete the task; This is the threshold ratio of task completion delay to maximum tolerable delay.

7. The method for unloading space-ground collaborative tasks based on temporal prediction deep reinforcement learning according to claim 1, characterized in that, The process involves treating any user device as an independent intelligent agent, constructing a duel dual-depth Q-network fused with LSTM, extracting the spatial features and temporal dependencies of the agent's current state, and utilizing a duel network architecture combined with a long short-term memory network to obtain the optimal action and its value. This includes: The current task information, queue status, satellite coverage, and battery level information are used to form an instantaneous state vector; and the load levels of edge servers and satellites over the past L time slots are recorded through the historical load matrix H. The instantaneous state branch is used to extract the spatial features of the current state through a two-layer fully connected network. At the same time, the historical load branch is used to capture the temporal dependence and changing trend of the load. The feature vectors extracted by the two branches are concatenated to form a comprehensive feature vector. The action value function is decomposed into a state value function and an action advantage function through a duel network architecture. The outputs of the two branches are merged through a special aggregation layer to obtain the final Q value.

8. The method for unloading space-ground collaborative tasks based on temporal prediction deep reinforcement learning according to claim 7, characterized in that: The action value function is implemented through a duel network architecture. Decomposed into state value function and action advantage function The outputs of the two branches are merged through a special aggregation layer to obtain the final Q value, including: ; in, Used to assess whether the state itself is good or bad; Used to evaluate the advantage of choosing a particular action compared to the average action in this state; Indicates the current environment state; s represents the action performed by the task; This represents all parameters of the state-value network; This represents all parameters of the action advantage network; Represents the action space, Indicates the total number of actions.

9. The method for unloading space-ground collaborative tasks based on temporal prediction deep reinforcement learning according to claim 7, characterized in that, The process of obtaining the optimal action and the value of the optimal action also includes: the weights of the fully connected layer. and bias The steps for introducing learnable parameterized noise are as follows: The learnable parameterized noise is expressed as: ;in, As weight; For bias; This is the feature vector input to the noisy network after state preprocessing.

10. The method for unloading space-ground collaborative tasks based on temporal prediction deep reinforcement learning according to claim 9, characterized in that, The weight and bias The acquisition process is as follows: ; ; in, , and , ε is a trainable parameter; ε is a sampled noise variable.