Task offloading method and device based on space-air-ground integrated network

By dynamically adjusting the transmission power of IoT devices and the flight status of drones using the HASDRL algorithm, the resource utilization of the integrated air-space-ground network is optimized. This solves the problem that fixed transmission power and time slot duration cannot adapt to dynamic changes, and improves system performance and energy management efficiency.

CN119211875BActive Publication Date: 2025-11-04SOUTH CHINA NORMAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411118087.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-15
Publication Date
2025-11-04
Estimated Expiration
2044-08-15

AI Technical Summary

Technical Problem

In existing integrated air-space-ground network offloading and resource allocation strategies, the transmission power of IoT devices and the duration of time slots allocated to drones are fixed, which cannot adapt to dynamically changing network environments, resulting in reduced system performance and efficiency.

Method used

By using the Hybrid Action Space Deep Reinforcement Learning (HASDRL) algorithm, combined with the location, distance, and battery capacity information of IoT devices, the transmission power and drone flight status are dynamically adjusted to optimize communication quality and resource utilization, thereby achieving energy management for IoT devices and drones.

Benefits of technology

It improves the efficiency of task offloading, reduces energy consumption of IoT devices and drones, and adapts to dynamically changing system environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119211875B_ABST
    Figure CN119211875B_ABST
Patent Text Reader

Abstract

The present application relates to the field of task offloading, and particularly relates to a task offloading method and device based on space-air-ground integrated network, computer equipment and storage medium, which fully considers the performance gain brought by the real-time changes of the transmission power of the Internet of Things device and the time slot duration allocated to the Internet of Things device by the unmanned aerial vehicle to the space-air-ground integrated network, and combines the different transmission power required by the Internet of Things devices at different positions and different distances to optimize the communication quality and resource utilization, so that the Internet of Things device can adjust the power according to the actual demand, the unmanned aerial vehicle can dynamically adjust the flight condition to adapt to the dynamically changing system environment, improve the performance of the space-air-ground integrated network through efficient resource management, reduce the energy consumption of the Internet of Things device and the unmanned aerial vehicle, and improve the efficiency of task offloading.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of task offloading, and particularly relates to a task offloading method and device based on a space-air-ground integrated network, a computer device and a storage medium. BACKGROUND

[0002] With the rapid development and popularization of 5G technology, more and more computing-intensive and latency-sensitive applications are running on Internet of Things devices. However, relying solely on ground wireless networks may not be able to meet the computing needs of Internet of Things devices, especially in remote areas lacking communication facilities. Therefore, space-air-ground integrated networks are considered to be a key framework for providing computing offloading services for Internet of Things devices. In this integrated network architecture, how to make optimal offloading decisions while efficiently allocating communication network resources to reduce system latency and energy consumption is an important research direction in the relevant field.

[0003] In the existing research on space-air-ground integrated network offloading and resource allocation strategies, it is still a major challenge to balance the delay and energy consumption caused by computing offloading. Excessive reliance on offloading will increase communication delay, while too much local computing will increase the energy consumption of Internet of Things devices. In addition, many researches on space-air-ground integrated network resource allocation methods usually assume that the transmission power of Internet of Things devices and the time slot duration allocated to Internet of Things devices by unmanned aerial vehicles are fixed values. However, fixed transmission power cannot adapt to dynamically changing network environments, resulting in Internet of Things devices being unable to adjust power according to actual needs, thereby reducing the overall performance and efficiency of the system. SUMMARY

[0004] Therefore, the present application provides a task offloading method and device based on a space-air-ground integrated network, a computer device and a storage medium, which fully considers the performance gain brought by the real-time changes of the transmission power of Internet of Things devices and the time slot duration allocated to Internet of Things devices by unmanned aerial vehicles in the space-air-ground integrated network, and combines the different transmission powers required by Internet of Things devices at different locations and different distances to optimize communication quality and resource utilization. In this way, Internet of Things devices can adjust power according to actual needs, unmanned aerial vehicles can dynamically adjust flight conditions to adapt to dynamically changing system environments, and the performance of the space-air-ground integrated network can be improved through efficient resource management, reducing the energy consumption of Internet of Things devices and unmanned aerial vehicles, and improving the efficiency of task offloading.

[0005] In a first aspect, an embodiment of the present application provides a task offloading method based on a space-air-ground integrated network, the space-air-ground integrated network comprising a plurality of Internet of Things devices, unmanned aerial vehicles and a plurality of satellites; the method comprising the following steps:

[0006] obtain state space information of a current time slot of a space-air-ground integrated network and a preset task offloading model;

[0007] input the state space information of the current time slot into the task offloading model, obtain action space information of the current time slot, control the space-air-ground integrated network to perform task offloading according to the action space information of the current time slot, and obtain state space information of a next time slot;

[0008] perform reward calculation according to the state space information and the action space information of the current time slot, obtain reward information of the current time slot, combine the state space information, the action space information, the reward information of the current time slot, and the state space information of the next time slot, and construct a training information combination of the current time slot;

[0009] repeat the construction of the training information combination, obtain training information combinations of a plurality of time slots, update model parameters of the task offloading model according to the training information combinations of the plurality of time slots, obtain a target task offloading model, and configure the target task offloading model in a plurality of Internet of Things devices and unmanned aerial vehicles to perform task offloading execution operations.

[0010] In a second aspect, an embodiment of the present application provides a task offloading device based on a space-air-ground integrated network, including:

[0011] a data obtaining module configured to obtain state space information of a current time slot of a space-air-ground integrated network and a preset task offloading model;

[0012] a task offloading module configured to input the state space information of the current time slot into the task offloading model, obtain action space information of the current time slot, control the space-air-ground integrated network to perform task offloading according to the action space information of the current time slot, and obtain state space information of a next time slot;

[0013] a training information combination construction module configured to perform reward calculation according to the state space information and the action space information of the current time slot, obtain reward information of the current time slot, combine the state space information, the action space information, the reward information of the current time slot, and the state space information of the next time slot, and construct a training information combination of the current time slot;

[0014] a model parameter updating module configured to repeatedly execute the task offloading module and the training information combination construction module, obtain training information combinations of a plurality of time slots, update model parameters of the task offloading model according to the training information combinations of the plurality of time slots, obtain a target task offloading model, and configure the target task offloading model in a plurality of Internet of Things devices and unmanned aerial vehicles to perform task offloading execution operations.

[0015] In a third aspect, an embodiment of the present application provides a computer device, comprising a processor, a memory, and a computer program stored in the memory and capable of running on the processor; when the computer program is executed by the processor, the steps of the task offloading method based on the space-air-ground integrated network according to the first aspect are implemented.

[0016] In a fourth aspect, an embodiment of the present application provides a storage medium, which stores a computer program, and when the computer program is executed by a processor, the steps of the task offloading method based on the space-air-ground integrated network according to the first aspect are implemented.

[0017] In the embodiments of the present application, the task offloading method, device, computer device and storage medium based on the space-air-ground integrated network fully consider the performance gain brought by the real-time changes of the transmission power of the Internet of Things device and the time slot duration allocated to the Internet of Things device by the unmanned aerial vehicle to the space-air-ground integrated network, and combine the different transmission power required by the Internet of Things devices at different positions and different distances to optimize the communication quality and resource utilization, so that the Internet of Things device can adjust the power according to the actual demand, the unmanned aerial vehicle can dynamically adjust the flight condition to adapt to the dynamically changing system environment, improve the performance of the space-air-ground integrated network through efficient resource management, reduce the energy consumption of the Internet of Things device and the unmanned aerial vehicle, and improve the efficiency of task offloading.

[0018] In order to better understand and implement, the present application is described in detail below with reference to the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS

[0019] Figure 1 The application environment of the task offloading method based on the space-air-ground integrated network provided by an embodiment of the present application;

[0020] Figure 2 The flowchart of the task offloading method based on the space-air-ground integrated network provided by an embodiment of the present application;

[0021] Figure 3 The flowchart of S2 in the task offloading method based on the space-air-ground integrated network provided by an embodiment of the present application;

[0022] Figure 4 The flowchart of S3 in the task offloading method based on the space-air-ground integrated network provided by an embodiment of the present application;

[0023] Figure 5 The flowchart of S4 in the task offloading method based on the space-air-ground integrated network provided by an embodiment of the present application;

[0024] Figure 6This is a flowchart illustrating step S42 of a task offloading method based on an integrated air-space-ground network provided in one embodiment of this application.

[0025] Figure 7 This is a flowchart illustrating step S421 of a task offloading method based on an integrated air-space-ground network provided in one embodiment of this application.

[0026] Figure 8 A schematic diagram of the structure of a task offloading device based on an integrated air-space-ground network provided in one embodiment of this application;

[0027] Figure 9 This is a schematic diagram of the structure of a computer device provided in one embodiment of this application. Detailed Implementation

[0028] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0029] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0030] It should be understood that although the terms first, second, third, etc., may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."

[0031] The execution subject of the task offloading method based on the space-air-ground integrated network is an offloading device (hereinafter referred to as the offloading device) of the task offloading method based on the space-air-ground integrated network. The offloading device can be realized by software and / or hardware, and the task offloading method based on the space-air-ground integrated network can be realized by software and / or hardware. The offloading device can be composed of two or more physical entities, or one physical entity. The hardware pointed to by the offloading device is essentially a computer device, for example, the offloading device can be a computer, a mobile phone, a tablet computer or an interactive tablet computer. In an optional embodiment, the offloading device can be a server or a server cluster composed of multiple computer devices.

[0032] Referring to Figure 1 , Figure 1 The application environment of the task offloading method based on the space-air-ground integrated network provided by an embodiment of the present application is shown in the figure. The space-air-ground integrated network includes a plurality of Internet of Things devices 6, a plurality of unmanned aerial vehicles 2 and a plurality of satellites 1. In the area 8 lacking communication infrastructure, the Internet of Things devices 6 can establish communication with the unmanned aerial vehicles 2 through Wi-Fi, and the unmanned aerial vehicles 2 can act as edge devices or wireless access points to provide edge computing services for the Internet of Things devices 6 in the area. In addition, the satellites 1, specifically low-orbit satellites, can provide cloud computing services for the Internet of Things devices 6 in their coverage areas through 5G / 4G / Wi-Fi communication.

[0033] Specifically, the users move slowly in the area 8, and each user carries an Internet of Things device. Since the computing power of these Internet of Things devices is poor and the energy storage is limited, the Internet of Things devices need to select to offload part of the generated computing tasks to the unmanned aerial vehicles or the satellites. The unmanned aerial vehicles 2 will hover at each sensing access point 5 in each time slot, and randomly schedule an Internet of Things device 6 located in the coverage range 7 thereof to establish communication with the device through a Wi-Fi channel 4. The scheduled Internet of Things device 6 needs to adjust its transmission power and the task ratio offloaded to the unmanned aerial vehicles and the low-orbit satellites according to the weighted sum of the system delay and energy consumption, wherein the Internet of Things device 6 establishes communication with the satellite 1 through a 5G\4G\Wi-Fi channel 3. After receiving and processing the computing tasks offloaded by the Internet of Things device 6, the unmanned aerial vehicle 2 adjusts the flight angle and speed, and flies to the next sensing access point 5 to continue to provide computing offloading services for other Internet of Things devices 6.

[0034] Referring to Figure 2 , Figure 2 The flowchart of the task offloading method based on the space-air-ground integrated network provided by an embodiment of the present application is shown in the figure. The method includes the following steps:

[0035] S1: Obtain the state space information of the current time slot of the space-air-ground integrated network and a preset task offloading model.

[0036] In the embodiment, the offloading device is connected with the Internet of Things device, the unmanned aerial vehicle and the satellites to obtain state space information of a current time slot of the space-air-ground integrated network, wherein the state space information reflects environmental information of the current time slot, including Internet of Things device position information, unmanned aerial vehicle position information, distance information between the unmanned aerial vehicle and the Internet of Things device, distance information between the satellites and the Internet of Things device, generated task data volume information, residual task data volume information and battery capacity information.

[0037] Specifically, the Internet of Things device position information includes position parameters of the Internet of Things devices; the distance information between the unmanned aerial vehicle and the Internet of Things device includes distance parameters between the unmanned aerial vehicle and the Internet of Things devices; the distance information between the satellites and the Internet of Things device includes distance parameters between the satellites and the Internet of Things devices; the generated task data volume information is used to indicate the size of the task data volume generated by the Internet of Things devices in the current time slot, including generated task data volumes of the Internet of Things devices; the residual task data volume information is used to indicate the size of the task data volume to be processed by the allocation device in the current time slot; and the battery capacity information is used to indicate the residual battery capacity of the Internet of Things devices and the unmanned aerial vehicle in each time slot, including residual battery capacities of the Internet of Things devices and the unmanned aerial vehicle.

[0038] The offloading device obtains a preset task offloading model, wherein the task offloading model includes a discrete action network and a continuous action network; and the action space information includes discrete action space information and continuous action space information.

[0039] S2: inputting the state space information of the current time slot into the task offloading model to obtain action space information of the current time slot, and controlling the space-air-ground integrated network to perform task offloading according to the action space information of the current time slot to obtain state space information of a next time slot.

[0040] In the embodiment, the offloading device inputs the state space information of the current time slot into the task offloading model, adopts a hybrid action space deep reinforcement learning algorithm (HASDRL), which has a significant advantage in solving complex optimization problems compared with traditional algorithms, can effectively process the integrated space-ground-network system with a large-scale state and action space, so as to solve the approximate optimal solution of each parameter of the action space and obtain the action space information of the current time slot, wherein the action space information includes discrete action space information and continuous action space information, the discrete action space information includes Internet of Things device scheduling information, the Internet of Things device scheduling information is used to indicate a certain Internet of Things device called in a time slot; the continuous action space information includes Internet of Things device transmit power information, unmanned aerial vehicle flight information and task offloading ratio information; the Internet of Things device transmit power information includes the transmit power of a plurality of Internet of Things devices; the unmanned aerial vehicle flight information includes unmanned aerial vehicle flight speed, unmanned aerial vehicle flight angle and unmanned aerial vehicle flight time; and the task offloading ratio information includes the unmanned aerial vehicle task offloading ratio and the satellite task offloading ratio corresponding to a plurality of Internet of Things devices.

[0041] The offloading device controls the integrated space-ground-network to perform task offloading according to the action space information of the current time slot, and obtains the state space information of the next time slot.

[0042] Referring to Figure 3 , Figure 3 The flowchart of S2 in the task offloading method based on the integrated space-ground-network provided by an embodiment of the present application includes steps S21-S22, and specifically as follows:

[0043] S21: inputting the Internet of Things device position information and the unmanned aerial vehicle position information in the state space information into an evaluation network in the discrete action generation network, adopting a greedy algorithm to obtain the discrete action space information.

[0044] In the embodiment, the offloading device inputs the Internet of Things device position information and the unmanned aerial vehicle position information in the state space information into the evaluation network in the discrete action generation network, adopts a greedy algorithm to obtain the discrete action space information, wherein the discrete action space information includes Internet of Things device scheduling information, and the discrete action space information is as follows:

[0045]

[0046] In the formula, a is the discrete action space information, Z() is an evaluation network function, s i is the state space information of the i th time slot, a i is the discrete action space information of the i th time slot, θ z is the network parameter of the evaluation network.

[0047] S22: inputting the distance information between the UAV and the IoT device, the distance information between the satellite and the IoT device, the generated task data amount information, the residual task data amount information and the battery capacity information in the state space information into a policy network in the continuous action generation network to make an action decision, and obtaining the continuous action space information.

[0048] The distance between the UAV and each IoT device is an important basis for the UAV to decide to adjust the flight speed, flight angle and flight time of the UAV. In order to reduce the transmission delay and energy consumption of the system, the UAV will schedule the IoT device closest to itself, and adjust the flight speed, flight angle and flight time of the UAV according to the position parameters of the IoT device to obtain the UAV flight information, wherein the flight time is the time for the UAV to fly to the corresponding perception access point.

[0049] The residual task data amount information in the state space information is an important basis for judging whether the task is completed. When the residual task size is 0, it means that the UAV or the satellite has processed all the computing tasks generated by the IoT device.

[0050] In this embodiment, the offloading device inputs the distance information between the UAV and the IoT device, the distance information between the satellite and the IoT device, the generated task data amount information, the residual task data amount information and the battery capacity information in the state space information into a policy network in the continuous action generation network to make an action decision, and obtains the continuous action space information, wherein the continuous action space information includes IoT device transmit power information, UAV flight information and task offloading ratio information.

[0051] Specifically, the offloading device constructs the corresponding constraint condition by using the residual task data amount information and the battery capacity information, and calculates the channel gain of the UAV and the satellite to the IoT device by using the IoT device position information, the UAV position information and the distance information between the satellite and the IoT device in the state space information, and then obtains the data transmission rate of the UAV and the satellite to the IoT device. The transmit power of the IoT device will affect the data transmission rate, and the size of the data transmission rate will affect the transmission delay of the system. By optimizing the transmit power of the IoT device, the transmission delay of the system can be indirectly reduced, and the IoT device transmit power information can be obtained.

[0052] The generated task data amount information in the state space information is an important basis for optimizing the respective task offloading ratios of the unmanned aerial vehicle and the satellite. The offloading device utilizes the generated task data amount of the Internet of Things device in the generated task data amount information; obtains the task data amount sizes of the Internet of Things device offloaded to the unmanned aerial vehicle and the satellite, thereby calculating the time delay and energy consumption generated by offloading the task to the unmanned aerial vehicle and the satellite, and taking the calculated time delay and energy consumption as a basis, optimizing the task offloading ratio, obtaining the unmanned aerial vehicle task offloading ratio and the satellite task offloading ratio corresponding to the Internet of Things device, and obtaining the task offloading ratio information.

[0053] S3: performing reward calculation according to the state space information and the action space information of the current time slot, obtaining reward information of the current time slot; combining the state space information, the action space information, the reward information of the current time slot, and the state space information of the next time slot, and constructing a training information combination of the current time slot.

[0054] In the embodiment, the offloading device performs reward calculation according to the state space information and the action space information of the current time slot, and obtains reward information of the current time slot.

[0055] The offloading device combines the state space information, the action space information, the reward information of the current time slot, and the state space information of the next time slot, and constructs a training information combination of the current time slot. In an optional embodiment, in order to avoid the algorithm falling into a local optimal solution in the training process, and make the unmanned aerial vehicle and the Internet of Things device better explore the non-stationary environment, the offloading device adds Gaussian noise to all continuous actions in the action space information.

[0056] Please refer to Figure 4 , Figure 4 The flowchart of S3 in the task offloading method based on the space-air-ground integrated network provided by an embodiment of the present application is shown in FIG. 3, which includes the following step S31.

[0057] S31: obtaining reward information of the current time slot according to the state space information, the action space information of the current time slot, and a preset reward calculation algorithm.

[0058] The reward calculation algorithm is as follows:

[0059]

[0060] In the formula, r iis the reward information of the ith time slot, min is the minimum function, max is the maximum function, n = {n(i)}, N is the number of Internet of Things devices, n represents the nth Internet of Things device, n(i) is the Internet of Things device called in the ith time slot, U = {0(i), v(i)}, U is the unmanned aerial vehicle flight information, 0(i) is the unmanned aerial vehicle flight angle in the ith time slot, v(i) is the unmanned aerial vehicle flight speed in the ith time slot, R is the task offloading ratio information, is the unmanned aerial vehicle task offloading ratio corresponding to the nth Internet of Things device in the ith time slot, is the satellite task offloading ratio corresponding to the nth Internet of Things device in the ith time slot, P = {P n (i)}, P is the Internet of Things device transmit power information, P n (i) is the transmit power of the nth Internet of Things device in the ith time slot, T = {T fly (i)}, T fly (i) is the unmanned aerial vehicle flight time in the ith time slot, 5 is a weight parameter, indicating the proportion of delay and energy consumption, is the task delay of the nth Internet of Things device left for local processing, is the transmission delay and processing delay of the task of the nth Internet of Things device offloaded to the unmanned aerial vehicle, is the transmission delay and processing delay of the task of the nth Internet of Things device offloaded to the low-orbit satellite, is the energy consumption of the task of the nth Internet of Things device left for local processing, is the transmission energy consumption and processing energy consumption of the task of the nth Internet of Things device offloaded to the unmanned aerial vehicle, is the transmission energy consumption and processing energy consumption of the task of the nth Internet of Things device offloaded to the low-orbit satellite.

[0061] Since the time slot duration allocated to the Internet of Things device by the unmanned aerial vehicle is composed of two parts, namely where t n(i) the time slot duration allocated to the nth IoT device by the UAV, in the embodiment, the offloading device calculates the reward information of the current time slot by jointly optimizing the user scheduling, the UAV flight trajectory, the task offloading ratio, the IoT device transmission power and the flight time of the UAV in each time slot according to the state space information, the action space information and the preset reward calculation algorithm of the current time slot, so as to minimize the weighted sum of the delay and the energy consumption, fully considers the performance gain of the space-air-ground integrated network caused by the real-time changes of the IoT device transmission power and the time slot duration allocated to the IoT device by the UAV, and combines the different transmission powers required by the IoT devices at different positions and different distances to optimize the communication quality and resource utilization, so that the IoT devices can adjust the power according to the actual demand, and the efficiency of task offloading is improved.

[0062] S4: repeatedly constructing the training information combination to obtain training information combinations of a plurality of time slots, updating the model parameters of the task offloading model according to the training information combinations of the plurality of time slots, obtaining a target task offloading model, and configuring the target task offloading model in the plurality of IoT devices and the UAV to perform the task offloading execution operation.

[0063] In the embodiment, the offloading device repeatedly constructs the training information combination to obtain training information combinations of a plurality of time slots, updates the model parameters of the task offloading model according to the training information combinations of the plurality of time slots, and obtains a target task offloading model.

[0064] In order to improve the stability and convergence of the model training process, specifically, the offloading device adopts an experience replay mechanism, repeatedly performs steps S2-S3, constructs a plurality of initial training information combinations of time slots, and stores them in a preset database. The offloading device obtains a plurality of initial training information combinations of time slots by randomly sampling a minimum batch of initial training information combinations, updates the model parameters of the task offloading model according to the training information combinations of the plurality of time slots, and obtains a target task offloading model, so as to execute in the environment and obtain the maximum long-term return.

[0065] The offloading device configures the target task offloading model in a plurality of Internet of Things devices and unmanned aerial vehicles. The Internet of Things devices and unmanned aerial vehicles serve as agents and can obtain corresponding action space information based on corresponding information in state space information of a current time slot to perform task offloading execution operations. The performance gain brought by real-time changes in Internet of Things device transmission power and time slot duration allocated to the Internet of Things devices by the unmanned aerial vehicles to the space-air-ground integrated network is fully considered. Different transmission powers are required for Internet of Things devices at different positions and distances to optimize communication quality and resource utilization. The Internet of Things devices can adjust the power according to actual needs. The unmanned aerial vehicles dynamically adjust the flight time to the next sensing access point to adapt to the dynamically changing system environment. Efficient resource management improves the performance of the space-air-ground integrated network, reduces the energy consumption of the Internet of Things devices and unmanned aerial vehicles, and improves the efficiency of task offloading.

[0066] The discrete action network includes an evaluation network and a target network. Please refer to Figure 5 , Figure 5 The flowchart of S4 in the task offloading method based on the space-air-ground integrated network provided by an embodiment of the present application includes steps S411-S413, which are as follows:

[0067] S411: Obtain a discrete evaluation value parameter of a current time slot according to state space information, discrete action space information, and an evaluation network of the current time slot in the training information combination; and obtain discrete action space information of a next time slot according to state space information of the next time slot in the training information combination.

[0068] In this embodiment, the offloading device obtains a discrete evaluation value parameter of a current time slot according to state space information, discrete action space information, and an evaluation network of the current time slot in the training information combination. The discrete evaluation value parameter is Z(s i ,a i |θ Z ).

[0069] The offloading device obtains discrete action space information of a next time slot according to state space information of the next time slot in the training information combination. For specific embodiments, please refer to steps S21-S22, which are not described here.

[0070] S412: Input reward information of the current time slot, state space information of the next time slot, and discrete action space information of the next time slot in the training information combination into the target network, and obtain a discrete target value parameter of the current time slot according to a preset discrete target value calculation algorithm.

[0071] The discrete target value calculation algorithm is as follows:

[0072]

[0073] wherein z ′ i is the discrete target value parameter of the i th time slot, r i is the reward information of the i th time slot, γ 1 is a first discount factor, and z ′ is a target network function, and z () is an evaluation network function, is a maximum function, and is used for selecting an action, s i+1 is state space information of an i+1 th time slot, a i+1 is discrete action space information of the i+1 th time slot, θ Z is a network parameter of the evaluation network, θ Z′ is a network parameter of the target network.

[0074] In this embodiment, the offloading device inputs the reward information of the current time slot, the state space information of the next time slot, and the discrete action space information in the training information set to the target network, and obtains the discrete target value parameter of the current time slot according to a preset discrete target value calculation algorithm.

[0075] S413: According to the discrete evaluation value parameter of the current time slot and the discrete target value parameter, an expected loss value is obtained according to a preset expected loss algorithm, the parameters of the evaluation network are updated according to the expected loss value, an updated evaluation network is obtained, the parameters of the updated evaluation network are updated to the target network by using a soft update strategy, and an updated target network is obtained.

[0076] The expected loss algorithm is as follows:

[0077] L1=E[(z ′ i -z i ) 2 ]

[0078] wherein L1 is an expected loss value, E() is an expected function, z i is the discrete evaluation value parameter of the i th time slot.

[0079] In this embodiment, according to the discrete evaluation value parameter of the current time slot and the discrete target value parameter, an expected loss value is obtained according to a preset expected loss algorithm, the parameters of the evaluation network are updated according to the expected loss value, an updated evaluation network is obtained, the parameters of the updated evaluation network are updated to the target network by using a soft update strategy, and an updated target network is obtained.

[0080] The continuous action network comprises a value network, the value network comprises a predicted value network and a target value network, wherein the predicted value network is used to evaluate action value of a current policy in real time, and the target value network is used to calculate loss of the predicted value network; please refer to Figure 6 , Figure 6 A flowchart of S4 in the task offloading method based on the space-air-ground integrated network is provided in an embodiment of the present application, comprising steps S421-S423, and specifically as follows:

[0081] S421: inputting state space information and continuous action space information of a current time slot in the training information combination into the predicted value network, obtaining predicted value parameters of the current time slot according to a preset predicted value calculation algorithm, and updating parameters of the policy network by using a policy gradient training method according to the predicted value parameters of the current time slot, the state space information and the continuous action space information, to obtain an updated policy network.

[0082] In the embodiment, the offloading device inputs state space information and continuous action space information of a current time slot in the training information combination into the predicted value network, and obtains predicted value parameters of the current time slot according to a preset predicted value calculation algorithm, wherein the predicted value calculation algorithm is as follows:

[0083] y i = Q(s i ,a′ i | θ Q )

[0084] In the formula, y i are predicted value parameters of the i th time slot, Q() is a predicted value network function, s i is state space information of the i th time slot, a′ i is continuous action space information of the i th time slot, and θ Q is a network parameter of the predicted value network.

[0085] The offloading device updates parameters of the policy network by using a policy gradient training method according to the predicted value parameters of the current time slot, the state space information and the continuous action space information, to obtain an updated policy network.

[0086] Please refer to Figure 7 , Figure 7 A flowchart of S421 in the task offloading method based on the space-air-ground integrated network is provided in an embodiment of the present application, comprising step S4211, and specifically as follows:

[0087] S4211: obtaining a policy gradient loss value according to the predicted value parameter of the current time slot, the state space information, the continuous action space information, and a preset policy gradient loss algorithm, updating the parameters of the policy network according to the policy gradient loss value, and obtaining an updated policy network.

[0088] The policy gradient loss algorithm is:

[0089]

[0090] In the formula, is a policy gradient loss value, E s~D is a policy gradient function, and μ is a network parameter of the policy network, is a gradient of the continuous action space information, is a gradient of the network parameter of the policy network.

[0091] In the embodiment, the offloading device obtains a policy gradient loss value according to the predicted value parameter of the current time slot, the state space information, the continuous action space information, and a preset policy gradient loss algorithm, updates the parameters of the policy network according to the policy gradient loss value, and obtains an updated policy network.

[0092] S422: obtaining continuous action space information of a next time slot according to state space information of the next time slot in the training information combination and the updated policy network, inputting the state space information of the next time slot, the continuous action space information, and reward information of the current time slot into the target value network, and obtaining a continuous target value parameter of the current time slot according to a preset continuous target value calculation algorithm.

[0093] The continuous target value calculation algorithm is:

[0094] y ′ i = r i + γ1Q ′ (s i+1 , μ ′ (s i+1 | θ μ′ )| θ Q′ )

[0095] In the formula, y ′ i is a continuous target value parameter of the i th time slot, r i is reward information of the i th time slot, γ2 is a second discount factor, Q ′ () is a target value network function, s i+1is the state space information of the i+1th time slot, μ ′ is the action generation function of the updated policy network, θ μ′ is the network parameter of the updated policy network, θ Q′ is the network parameter of the target value network;

[0096] In this embodiment, the offloading device obtains the continuous action space information of the next time slot according to the state space information of the next time slot in the training information combination and the updated policy network, inputs the state space information of the next time slot, the continuous action space information, and the reward information of the current time slot into the target value network, obtains the continuous target value parameter of the current time slot according to a preset continuous target value calculation algorithm, adjusts the policy by updating the parameter of the policy network to generate the optimal action under the current state, thereby maximizing the long-term cumulative return and ensuring the stability of the policy.

[0097] S423: According to the target value parameter and the predicted value parameter of the current time slot, a mean square error loss value is obtained according to a preset mean square error loss algorithm, the parameter of the predicted value network is updated according to the mean square error loss value to obtain an updated predicted value network, and the parameter of the updated predicted value network is updated to the target value network by using a soft update strategy to obtain an updated target value network.

[0098] In this embodiment, the offloading device obtains the continuous action space information of the next time slot according to the state space information of the next time slot in the training information combination and the updated policy network, inputs the state space information of the next time slot, the continuous action space information, and the reward information of the current time slot into the target value network, obtains the continuous target value parameter of the current time slot according to a preset continuous target value calculation algorithm, adjusts the policy by updating the parameter of the policy network to generate the optimal action under the current state, thereby maximizing the long-term cumulative return and ensuring the stability of the policy.

[0099]

[0100] In the formula, L2 is the mean square error loss value, and I is the total number of the training information combination.

[0101] The offloading device updates the parameter of the updated predicted value network to the target value network by using a soft update strategy to obtain an updated target value network. Through the above steps, the predicted value network continuously improves its estimation of the Q value, thereby providing more accurate feedback for the predicted value network and helping the predicted value network optimize its policy.

[0102] Please refer to Figure 8 , Figure 8A structural schematic diagram of a task offloading device based on a space-air-ground integrated network is provided for an embodiment of the present application. The device can realize all or part of the task offloading device based on the space-air-ground integrated network through software, hardware, or a combination of both. The device 8 includes:

[0103] A data obtaining module 81 is configured to obtain state space information of a current time slot of the space-air-ground integrated network and a preset task offloading model.

[0104] A task offloading module 82 is configured to input the state space information of the current time slot into the task offloading model, obtain action space information of the current time slot, control the space-air-ground integrated network to perform task offloading according to the action space information of the current time slot, and obtain state space information of a next time slot.

[0105] A training information combination construction module 83 is configured to perform reward calculation according to the state space information and the action space information of the current time slot, obtain reward information of the current time slot, and combine the state space information, the action space information, and the reward information of the current time slot with the state space information of the next time slot to construct a training information combination of the current time slot.

[0106] A model parameter updating module 84 is configured to repeatedly execute the task offloading module and the training information combination construction module, obtain training information combinations of a plurality of time slots, update model parameters of the task offloading model according to the training information combinations of the plurality of time slots, obtain a target task offloading model, and configure the target task offloading model in a plurality of Internet of Things devices and unmanned aerial vehicles to perform task offloading execution operations.

[0107] In the embodiment of the present application, the state space information of the current time slot of the space-air-ground integrated network and a preset task offloading model are obtained by the data obtaining module; the state space information of the current time slot is input into the task offloading model by the task offloading module to obtain the action space information of the current time slot, and the space-air-ground integrated network is controlled to perform task offloading according to the action space information of the current time slot to obtain the state space information of the next time slot; the reward information of the current time slot is obtained by calculating the reward according to the state space information and the action space information of the current time slot by the training information combination construction module; the state space information, the action space information, the reward information of the current time slot and the state space information of the next time slot are combined to construct the training information combination of the current time slot; the task offloading module and the training information combination construction module are repeatedly executed by the model parameter updating module to obtain the training information combinations of several time slots, the model parameters of the task offloading model are updated according to the training information combinations of several time slots to obtain a target task offloading model, and the target task offloading model is configured in the several Internet of Things devices and unmanned aerial vehicles to perform task offloading operation. The real-time changes of the transmission power of the Internet of Things devices and the time slot duration allocated by the unmanned aerial vehicles to the Internet of Things devices are fully considered to bring performance gain to the space-air-ground integrated network, and different transmission powers are combined to optimize the communication quality and resource utilization of the Internet of Things devices at different positions and different distances, so that the Internet of Things devices can adjust the power according to the actual demand, the unmanned aerial vehicles can dynamically adjust the flight condition to adapt to the dynamically changing system environment, the performance of the space-air-ground integrated network is improved through efficient resource management, the energy consumption of the Internet of Things devices and the unmanned aerial vehicles is reduced, and the efficiency of task offloading is improved.

[0108] Please refer to Figure 9 , Figure 9 The structural schematic diagram of the computer device provided in an embodiment of the present application, the computer device 9 comprises a processor 91, a memory 92, and a computer program 93 stored in the memory 92 and executable on the processor 91; the computer device can store a plurality of instructions, the instructions are suitable for being loaded and executed by the processor 91 to perform the method steps shown in the above Figures 1 to 7 , and the specific execution process can be referred to the specific description shown in the above Figures 1 to 7 , which will not be described here.

[0109] The processor 91 can include one or more processing cores. The processor 91 connects various parts within the server by various interfaces and lines, performs various functions and processes data of the space-ground-earth integrated network based task offloading apparatus 8 by running or executing instructions, programs, code sets or instruction sets stored in the memory 92, and calling data in the memory 92, and can be implemented in at least one of a hardware form of a digital signal processing (Digital Signal Processing, DSP), a field-programmable gate array (Field-Programmable Gate Array, FPGA), and a programable logic array (Programble Logic Array, PLA). The processor 91 can be integrated with one or a combination of a central processing unit 91 (Central Processing Unit, CPU), a graphics processor 91 (Graphics Processing Unit, GPU), and a modem, etc. Among them, the CPU is mainly used to process operating systems, user interfaces, and application programs, etc.; the GPU is used to render and draw the content to be displayed on the touch display screen; and the modem is used to process wireless communication. It can be understood that the above-mentioned modem can also not be integrated into the processor 91, but can be realized by a separate chip.

[0110] The memory 92 can include a random access memory 92 (Random Access Memory, RAM) and a read-only memory 92 (Read-Only Memory). Optionally, the memory 92 includes a non-transitory computer-readable storage medium. The memory 92 can be used to store instructions, programs, codes, code sets or instruction sets. The memory 92 can include a program storage area and a data storage area, wherein the program storage area can store instructions for implementing an operating system, instructions for at least one function (such as touch instructions, etc.), instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area can store data involved in the above-mentioned various method embodiments, etc. The memory 92 can also be at least one storage device located away from the aforementioned processor 91.

[0111] The embodiments of the present application also provide a storage medium, which can store a plurality of instructions, the instructions being suitable for being loaded and executed by a processor to perform the above-mentioned Figures 1 to 7 The method steps shown in the specific description, which will not be described here. Figures 1 to 7

[0112] ​Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is exemplified, and in actual application, the above-mentioned functions can be completed by different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be realized in the form of hardware or software. In addition, the specific name of each functional unit and module is only for the convenience of mutual distinction, and does not limit the protection scope of the present application. The specific working process of the unit and module in the above system can refer to the corresponding process in the foregoing method embodiment, which will not be described here.

[0113] In the above embodiments, the description of each embodiment has its own emphasis, and the parts not described or recorded in detail in a certain embodiment can be referred to the related description of other embodiments.

[0114] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized in electronic hardware or a combination of computer software and electronic hardware. Whether the functions are executed in hardware or software depends on the specific application and design constraints of the algorithm. A person skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0115] In the embodiments provided by the present application, it should be understood that the disclosed apparatus / terminal device and method can be implemented by other ways. For example, the above-mentioned apparatus / terminal device embodiments are only schematic, and the division of the modules or units is only a logical function division, and there can be another division way in actual implementation, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual coupling or direct coupling or communication connection can be indirect coupling or communication connection through some interface, device or unit, and can be electrical, mechanical or other forms.

[0116] The units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed on multiple network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.

[0117] In addition, each function unit in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software function unit.

[0118] The integrated module / unit, if realized in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, all or part of the processes in the above-mentioned embodiment methods can also be completed by a computer program instructing related hardware, and the computer program can be stored in a computer readable storage medium. When the computer program is executed by a processor, the steps of each method embodiment described above can be implemented. The computer program includes computer program code, which can be in the form of source code, object code, an executable file, or some intermediate form.

[0119] The present application is not limited to the above-described embodiments, and various modifications or changes can be made to the present application without departing from the spirit and scope of the present application. Therefore, it is intended that the present application encompass all such modifications and changes and fall within the scope of the appended claims and their equivalents.

Claims

1. A task offloading method based on a space-air-ground integrated network, the space-air-ground integrated network comprising a plurality of Internet of Things devices, unmanned aerial vehicles, and a plurality of satellites; the method comprising: The method comprises the following steps: Obtaining state space information of a current time slot of a space-air-ground integrated network and a preset task offloading model; Inputting the state space information of the current time slot into the task offloading model to obtain action space information of the current time slot, and controlling the space-air-ground integrated network to perform task offloading according to the action space information of the current time slot to obtain state space information of a next time slot; According to the state space information and the action space information of the current time slot, reward calculation is performed to obtain reward information of the current time slot; the state space information, the action space information, the reward information and the state space information of the next time slot are combined to construct a training information combination of the current time slot; Repeat the construction of the training information combination to obtain training information combinations of several time slots, update the model parameters of the task offloading model according to the training information combinations of the several time slots to obtain a target task offloading model, and configure the target task offloading model in the several Internet of Things devices and unmanned aerial vehicles to perform task offloading execution operations.

2. The task offloading method based on the space-air-ground integrated network according to claim 1, wherein: The state space information comprises Internet of Things device position information, unmanned aerial vehicle position information, distance information between the unmanned aerial vehicle and the Internet of Things device, distance information between the satellite and the Internet of Things device, generated task data amount information, residual task data amount information and battery capacity information; The task offloading model comprises a discrete action network and a continuous action network; the action space information comprises discrete action space information and continuous action space information; The inputting of the state space information of the current time slot into the task offloading model to obtain the action space information of the current time slot comprises the following steps: The Internet of Things device position information and the unmanned aerial vehicle position information in the state space information are inputted into an evaluation network in the discrete action network to obtain the discrete action space information by using a greedy algorithm, wherein the discrete action space information comprises Internet of Things device scheduling information; The distance information between the unmanned aerial vehicle and the Internet of Things device, the distance information between the satellite and the Internet of Things device, the generated task data amount information, the residual task data amount information and the battery capacity information in the state space information are inputted into a policy network in the continuous action generation network to make an action decision to obtain the continuous action space information, wherein the continuous action space information comprises Internet of Things device transmission power information, unmanned aerial vehicle flight information and task offloading ratio information. 3.The task offloading method based on the space-air-ground integrated network according to claim 2, characterized in that: The Internet of Things device position information includes position parameters of a plurality of Internet of Things devices; the distance information between the satellites and the Internet of Things devices includes distance parameters between a plurality of satellites and a plurality of Internet of Things devices; the generated task data volume information is used to indicate the size of the task data volume generated by the Internet of Things devices in the current time slot, and includes generated task data volumes of a plurality of Internet of Things devices; the residual task data volume information is used to indicate the size of the task data volume that needs to be processed by the offloading device in the current time slot; and the battery capacity information is used to indicate the residual battery capacity of the Internet of Things devices and the unmanned aerial vehicle in each time slot, and includes residual battery capacities of a plurality of Internet of Things devices and the unmanned aerial vehicle. The Internet of Things device scheduling information is used to indicate a certain Internet of Things device called in a time slot; the Internet of Things device transmit power information includes transmit powers of a plurality of Internet of Things devices; the unmanned aerial vehicle flight information includes an unmanned aerial vehicle flight speed, an unmanned aerial vehicle flight angle, and an unmanned aerial vehicle flight time; and the task offloading ratio information includes unmanned aerial vehicle task offloading ratios and satellite task offloading ratios corresponding to a plurality of Internet of Things devices. 4.The task offloading method based on the space-air-ground integrated network according to claim 3, characterized in that, The reward information of the current time slot is obtained by performing reward calculation according to the state space information and the action space information of the current time slot, and includes the following steps: The reward information of the current time slot is obtained by performing reward calculation according to the state space information, the action space information, and a preset reward calculation algorithm of the current time slot, and the reward calculation algorithm is as follows: wherein r i is the reward information of the ith time slot, min is the minimum function, max is the maximum function, n = {n(i)}, N is the number of IoT devices, n represents the nth IoT device, n(i) is the IoT device called in the ith time slot, U = {θ(i), v(i)}, U is the UAV flight information, θ(i) is the UAV flight angle in the ith time slot, v(i) is the UAV flight speed in the ith time slot, R is the task offloading ratio information, is the UAV task offloading ratio corresponding to the nth IoT device in the ith time slot, is the satellite task offloading ratio corresponding to the nth IoT device in the ith time slot, P = {P n (i)}, P is the IoT device transmit power information, P n (i) is the transmit power of the nth IoT device in the ith time slot, T = {T fly (i)}, T fly (i) is the UAV flight time in the ith time slot, δ is a weight parameter, is the task delay of the nth IoT device left for local processing, is the transmission delay and processing delay of the task of the nth IoT device offloaded to the UAV, is the transmission delay and processing delay of the task of the nth IoT device offloaded to the low earth orbit satellite, is the energy consumption of the task of the nth IoT device left for local processing, is the transmission energy consumption and processing energy consumption of the task of the nth IoT device offloaded to the UAV, is the transmission energy consumption and processing energy consumption of the task of the nth IoT device offloaded to the low earth orbit satellite.

5. The task offloading method based on the space-air-ground integrated network according to any one of claims 2 to 4, characterized in that: The discrete action network includes a target network. The model parameters of the task offloading model are updated according to the training information combination of a plurality of time slots to obtain a target task offloading model, and the following steps are included: The discrete evaluation value parameters of the current time slot are obtained according to the state space information, the discrete action space information, and the evaluation network of the current time slot in the training information combination; and the discrete action space information of the next time slot is obtained according to the state space information of the next time slot in the training information combination. The reward information of the current time slot, the state space information of the next time slot, and the discrete action space information of the next time slot in the training information combination are input into the target network, and the discrete target value parameters of the current time slot are obtained according to a preset discrete target value calculation algorithm, and the discrete target value calculation algorithm is as follows: wherein z ′ i is a discrete target value parameter for the i-th time slot, r i is a reward information for the i-th time slot, γ1is a first discount factor, Z ′ is a target network function, Z() is an evaluation network function, is a maximum function, s i+1 is state space information for the i+1-th time slot, a i+1 is discrete action space information for the i+1-th time slot, θ Z is a network parameter of the evaluation network, θ Z′ is a network parameter of the target network; The expected loss value is obtained according to the discrete evaluation value parameters and the discrete target value parameters of the current time slot according to a preset expected loss algorithm, the parameters of the evaluation network are updated according to the expected loss value to obtain an updated evaluation network, the parameters of the updated evaluation network are updated to the target network by using a soft update strategy to obtain an updated target network, and the expected loss algorithm is as follows: L1 = E[(z ′ i -z i ) 2 ] where L1 is the expected loss value, E() is the expectation function, z i is the discrete evaluation value parameter for the i-th time slot.

6. The task offloading method based on the space-air-ground integrated network according to any one of claims 2 to 4, characterized in that: The continuous action network comprises a value network, the value network comprises a prediction value network and a target value network, wherein the prediction value network is used for evaluating action value of a current policy in real time, and the target value network is used for calculating loss of the prediction value network; The model parameter of the task offloading model is updated according to the training information combination of a plurality of time slots, and a target task offloading model is obtained, comprising the steps of: The state space information and the continuous action space information of the current time slot in the training information combination are input into the prediction value network, the prediction value parameter of the current time slot is obtained according to a preset prediction value calculation algorithm, the parameter of the policy network is updated by using a policy gradient training method according to the prediction value parameter of the current time slot, the state space information and the continuous action space information, and an updated policy network is obtained, wherein the prediction value calculation algorithm is: y i = Q(s i ,a′ i | θ Q ) In the formula, y i is the predicted value parameter of the i th time slot, Q() is the predicted value network function, s i is the state space information of the i th time slot, a′ i is the continuous action space information of the i th time slot, θ Q is the network parameter of the predicted value network; The state space information of the next time slot in the training information combination and the updated policy network are used to obtain the continuous action space information of the next time slot, the state space information of the next time slot, the continuous action space information and the reward information of the current time slot are input into the target value network, the continuous target value parameter of the current time slot is obtained according to a preset continuous target value calculation algorithm, wherein the continuous target value calculation algorithm is: y ′ i = r i + γ1Q ′ (s i+1 , μ ′ (s i+1 | θ μ′ ) | θ Q′ ) where y ′ i is the continuous target value parameter of the ith time slot, r i is the reward information of the ith time slot, γ2 is a second discount factor, Q ′ () is a target value network function, s i+1 is the state space information of the i+1th time slot, μ ′ () is an action generation function of the updated policy network, θ μ′ is the network parameter of the updated policy network, θ Q′ is the network parameter of the target value network; The mean square error loss value is obtained according to the target value parameter and the prediction value parameter of the current time slot according to a preset mean square error loss algorithm, the parameter of the prediction value network is updated according to the mean square error loss value, and an updated prediction value network is obtained; the parameter of the updated prediction value network is updated to the target value network by using a soft update strategy, and an updated target value network is obtained, wherein the mean square error loss algorithm is: In the formula, L2 is the mean square error loss value, and I is the total number of the training information combination. 7.The task offloading method based on the space-air-ground integrated network according to claim 6, characterized in that, The parameter of the policy network is updated by using a policy gradient training method according to the prediction value parameter, the state space information and the continuous action space information of the current time slot, and an updated policy network is obtained, comprising the steps of: The policy gradient loss value is obtained according to the prediction value parameter, the state space information, the continuous action space information of the current time slot and a preset policy gradient loss algorithm, the parameter of the policy network is updated according to the policy gradient loss value, and an updated policy network is obtained, wherein the policy gradient loss algorithm is: wherein is the policy gradient loss value, E s~D is the policy gradient function, μ() is the action generation function of the policy network, θ μ is the network parameter of the policy network, is the gradient of the information of the continuous action space, is the gradient of the network parameter of the policy network. 8.A task offloading apparatus based on a space-air-ground integrated network, the space-air-ground integrated network comprising a plurality of Internet of Things devices, unmanned aerial vehicles and a plurality of satellites, characterized in that, Comprise: The data obtaining module is used for obtaining state space information of a current time slot of the space-air-ground integrated network and a preset task offloading model; The task offloading module is used for inputting the state space information of the current time slot into the task offloading model to obtain action space information of the current time slot, and controlling the space-air-ground integrated network to perform task offloading according to the action space information of the current time slot to obtain state space information of a next time slot. The training information combination construction module is configured to calculate a reward according to the state space information and the action space information of the current time slot to obtain reward information of the current time slot; and combine the state space information, the action space information, the reward information of the current time slot, and the state space information of the next time slot to construct a training information combination of the current time slot. The model parameter updating module is configured to repeatedly execute the task offloading module and the training information combination construction module to obtain training information combinations of a plurality of time slots, update model parameters of the task offloading model according to the training information combinations of the plurality of time slots to obtain a target task offloading model, and configure the target task offloading model in the plurality of Internet of Things devices and the unmanned aerial vehicle to perform a task offloading operation.

9. A computer device, comprising: The computer program is executed by the processor to implement the steps of the task offloading method based on the space-air-ground integrated network according to any one of claims 1 to 7. The computer program is executed by the processor to implement the steps of the task offloading method based on the space-air-ground integrated network according to any one of claims 1 to 7.

10. A storage medium characterized by: ​

Citation Information

Patent Citations

  • Computing task unloading method and device in space-air-ground network and electronic equipment

    CN114884957A

  • Unmanned aerial vehicle emergency network task unloading method and system based on deep reinforcement learning

    CN117835324A