Collaborative optimization method and system for computing unloading reliability of space-air-ground integrated network, processing equipment and storage medium

By adopting the D-MAPPO algorithm and task offloading strategy optimization in the integrated air-space-ground network, the computing needs of high-density IoT terminals in remote areas are solved, the joint optimization of latency and energy consumption is achieved, and the network reliability and computing resource utilization efficiency are improved.

CN120769306APending Publication Date: 2025-10-10QINGHAI NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510922527.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-04
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

In an integrated air-space-ground network, the high-density IoT terminal computing needs in remote areas are difficult to be effectively supported by traditional edge computing systems, especially in extreme environments such as disaster areas, oceans, and deserts. Spectrum resources are limited and node load capacity is limited, resulting in increased computing task failure rates and increased service delays.

Method used

The D-MAPPO algorithm is used to divide computing tasks into urgent tasks and ordinary tasks based on their urgency, and corresponding offloading strategies are designed. Through the collaborative optimization of ground sensors, drones, satellites and cloud servers, communication, satellite coverage time and computing models are established, and the offloading strategy is optimized to improve network reliability and efficiency.

Benefits of technology

It achieves comprehensive support for the computing needs of high-density IoT terminals in remote areas, optimizes latency, energy consumption and offloading success rate, and improves the network's service quality and computing resource utilization efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120769306A_ABST
    Figure CN120769306A_ABST
Patent Text Reader

Abstract

The invention relates to an air-space-ground integrated network computing unloading reliability collaborative optimization method and system, processing equipment and a storage medium, and the method comprises the steps: dividing a computing task into an emergency task and a common task based on the emergency degree of the computing task; setting an unloading strategy corresponding to the emergency task and the common task; adopting a D-MAPPO algorithm, based on a set unloading strategy, generating an unloading strategy of a plurality of unmanned aerial vehicles in the air-space-ground integrated network, and decomposing the unloading strategy of the plurality of unmanned aerial vehicles into an unloading strategy of each unmanned aerial vehicle in an environment; according to the method, a D-MAPPO algorithm is adopted, the generated unloading strategy is optimized based on a communication model, a satellite coverage time model and a calculation model which are established in advance and a preset reliability mechanism, the space-air-ground integrated network works according to the optimized unloading strategy, and the method can be widely applied to the field of the space-air-ground integrated network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of integrated air-space-ground-integrated networks, and in particular to a method, system, processing device and storage medium for collaborative optimization of computing offloading reliability in an integrated air-space-ground-integrated network. Background Art

[0002] With the continued advancement of research on sixth-generation (6G) wireless communication systems, aerial access networks (Aerial Access Networks) and space-air-ground integrated networks (SAGIN) have become a key area of ​​focus for both industry and academia. In recent years, numerous studies have explored the potential of SAGIN in areas such as the Internet of Things (IoT), cognitive communications, and edge computing, further expanding its application scenarios in remote areas. By integrating satellites in space, unmanned aerial vehicles (UAVs) in the air, ground base stations, and data centers, SAGIN builds a cross-domain collaborative integrated network architecture, providing key support for achieving seamless global connectivity. This is particularly true in remote areas such as deserts, oceans, and sparsely populated areas, where traditional communication networks struggle to cover them due to the scarcity and high construction costs of ground base stations. However, data collected from environmental monitoring in remote areas is of great value to the global information system. With its wide-area coverage, high flexibility, and reliability, SAGIN is an ideal solution for real-time data collection and transmission in the Internet of Remote Things (IoRT), and is considered a key solution for addressing communication connectivity issues in remote areas.

[0003] Companies like SpaceX and OneWeb are driving the large-scale deployment of low-Earth orbit (LEO) constellations, redefining network architecture to achieve low latency, high capacity, and global services. This presents new opportunities for realizing global networks. Compared to medium-Earth orbit (MEO) and geostationary orbit (GEO) satellites, low-Earth orbit (LEO) satellites are closer to the Earth and better suited to supporting latency-sensitive communications on a global scale. LEO satellites, with their high capacity and wide coverage, are currently experiencing rapid growth. The low-latency communication capabilities of LEO satellites make them ideal for IoT applications, and the rise of large-scale LEO constellations is fundamentally changing network architecture to achieve low latency, high capacity, and global coverage. This lays a solid foundation for realizing the vision of a connected world, providing reliable connectivity to remote areas and narrowing the digital divide. Mobile Edge Computing (MEC), a revolutionary concept that enhances the low-latency and high-bandwidth capabilities of communication systems, has also been incorporated into the SAGIN architecture to enable the local deployment of computing resources to meet user demands for real-time processing and high-speed transmission.

[0004] Although SAGIN has shown great development potential, it still faces many challenges in practical applications. Due to the highly dynamic and resource-constrained nature of the satellite communication environment, the network often encounters interruptions and data loss problems. There is an urgent need to build a highly reliable and stable system to ensure communication quality. At the same time, latency and energy consumption issues remain key areas for SAGIN optimization. On the one hand, there are significant differences in user needs in different application scenarios in remote areas. For example, scenarios such as emergency rescue and telemedicine are extremely sensitive to communication delays, while applications such as environmental monitoring and resource exploration are more concerned with energy efficiency and coverage. How to optimize overall latency and energy consumption performance while meeting the needs of heterogeneous users has become an important topic of current research.

[0005] In an edge computing environment, terminal devices typically have limited computing power, and the number and resources of edge servers are similarly limited. Once the number of offload requests exceeds the threshold that the edge node can handle, the quality of service will inevitably decline, manifested as an increase in the failure rate of computing tasks and increased service latency. To address this issue, scholars have proposed a variety of resource allocation and task scheduling optimization strategies, attempting to alleviate the edge overload problem by building a reasonable offloading decision model. In addition, to expand the service scope and capacity of edge computing, researchers have introduced an auxiliary computing platform based on unmanned aerial vehicles (UAVs) to build a multi-level offloading system for air-ground collaboration. However, due to limited spectrum resources and limited node load capacity, especially in extreme environments such as disaster areas, oceans, and deserts, the coverage and computing resources of edge computing still face significant challenges, making it difficult to fully support the computing needs of high-density IoT terminals in remote areas. Summary of the Invention

[0006] In response to the above problems, the purpose of the present invention is to provide a method, system, processing equipment and storage medium for collaborative optimization of reliability of integrated air-space-ground network computing offloading, which can fully support the computing needs of high-density Internet of Things terminals in remote areas.

[0007] To achieve the above objectives, the present invention adopts the following technical solutions: In a first aspect, a method for collaborative optimization of reliability of space-ground integrated network computing offloading is provided, comprising:

[0008] Based on the urgency of computing tasks, computing tasks are divided into urgent tasks and ordinary tasks;

[0009] Set up offloading strategies for urgent tasks and common tasks;

[0010] The D-MAPPO algorithm is used to generate the offloading strategies of multiple UAVs in the air-ground integrated network based on the set offloading strategies, and the offloading strategies of multiple UAVs are decomposed into the offloading strategies of each UAV in the environment.

[0011] The D-MAPPO algorithm is used to optimize the generated offloading strategy based on the pre-established communication model, satellite coverage time model and calculation model and the pre-set reliability mechanism, so that the integrated air-space-ground network can work according to the optimized offloading strategy.

[0012] Furthermore, the offloading strategy for the urgent task is:

[0013] When the ground sensors of the integrated space-air-ground network generate an emergency task, the emergency task is directly offloaded to the satellite through the connection channel between the ground sensor and the satellite. The satellite processes the emergency task and transmits the calculation results directly back to the ground sensor.

[0014] The offloading strategy for the common task is:

[0015] The ground sensors of the integrated air-space-ground network offload common tasks to the drones through the connection channel between the ground sensors and the drones. The drones offload common tasks to satellites and cloud servers for processing according to the determined offloading ratio, and the remaining common tasks remain on the drones for processing.

[0016] Furthermore, the probability distribution function of the D-MAPPO algorithm is the Dirichlet probability distribution:

[0017]

[0018] in, represents the probability density function of the Dirichlet probability distribution; x J Indicates the uninstall ratio, satisfying α J represents the parameters of the Dirichlet probability distribution; B(α) represents the normalization term of the Beta function;

[0019] The advantage estimation function of the D-MAPPO algorithm for:

[0020]

[0021] Among them, δ t is the TD error, which represents the estimated difference between the reward at the current moment and the value of the next state; γ represents the discount factor; λ represents the GAE smoothing parameter;

[0022] The joint strategy of the D-MAPPO algorithm updates the objective function L clip (θ) is:

[0023]

[0024] Among them, r t(θ) represents the probability ratio; clip represents the clipping operation; ∈ represents the hyperparameter;

[0025] The value function of the D-MAPPO algorithm is:

[0026]

[0027] Wherein, L VF (φ) represents a loss function of the value function, used to measure the error between the estimated value of the current value function V φ and the actual cumulative discounted return R t ; o t represents the environment observation value at the current time; R t represents the cumulative discounted return.

[0028] Further, the D-MAPPO algorithm is used to optimize the generated offloading strategy based on the pre-established communication model, satellite coverage time model and calculation model and pre-set reliability mechanism, so that the space-ground integrated network works according to the optimized offloading strategy, including:

[0029] The pre-established satellite coverage time model and the pre-set reliability mechanism are used to calculate the coverage time and the maximum uploadable task amount of each satellite, and the emergency tasks are offloaded to the satellites in proportion according to the maximum uploadable task amount of the satellites;

[0030] The pre-established communication model between the unmanned aerial vehicle and the ground sensor and the satellite is used to calculate the transmission delay;

[0031] The pre-established satellite calculation model is used to process the emergency tasks and calculate the transmission delay, transmission energy consumption, calculation delay and calculation energy consumption of the emergency tasks;

[0032] Each unmanned aerial vehicle calculates the remaining maximum uploadable task amount of the satellite according to the pre-set reliability mechanism, then calculates the offloading strategy to offload the ordinary tasks in turn, and processes the ordinary tasks according to the pre-established communication model between the unmanned aerial vehicle and the ground sensor, the communication model between the unmanned aerial vehicle and the cloud server, the communication model between the unmanned aerial vehicle and the ground sensor and the satellite, and the calculation model of the unmanned aerial vehicle, the satellite and the cloud server, to calculate the transmission delay, transmission energy consumption, calculation delay and calculation energy consumption of the ordinary tasks;

[0033] When all the calculation tasks are processed, the energy consumption, delay and offloading success rate of the entire system model are obtained, and the reward score of the current offloading strategy is calculated according to the reward function and transmitted to the D-MAPPO algorithm;

[0034] The D-MAPPO algorithm calculates the probability ratio of each agent selecting each action under the current offloading strategy;

[0035] Clipping of probability ratios;

[0036] The clipped probability ratio is weighted by the advantage estimation function to construct the joint strategy update objective function;

[0037] Update the gradient of the objective function according to the joint strategy and update the offloading strategy;

[0038] After each update, the new uninstallation policy is used as the old uninstallation policy to prepare for the next sampling and update.

[0039] Furthermore, the communication model between the UAV and the ground sensor is:

[0040]

[0041] Among them, R i Indicates the maximum transmission rate of the ground sensor; R u Indicates the maximum transmission rate of the drone; B i Indicates the bandwidth of the ground sensor; B u represents the bandwidth of the drone; P i and P u Represent the transmission power of ground sensors and UAVs respectively; σ 2 represents the Gaussian noise power;

[0042] The communication model between the drone and the cloud server is:

[0043]

[0044] Among them, R c Indicates the maximum transmission rate of the cloud server; B c Indicates the bandwidth of the cloud server; P c Indicates the transmit power of the cloud server;

[0045] The communication model between the UAV, ground sensors and satellite is:

[0046]

[0047] in, They represent the maximum transmission rates of UAV, ground sensor and satellite under line-of-sight communication respectively; P s Indicates the available bandwidth of the satellite; G0 indicates the fixed antenna gain; P s represents the satellite's transmit power; N0 represents the spectral density of the additive white Gaussian noise.

[0048] Furthermore, the satellite coverage time model is:

[0049]

[0050] where T S denotes the coverage time of the satellite; V S denotes the flight speed of the satellite; L S denotes the coverage arc length of the satellite.

[0051] Further, the computing model of the UAV is:

[0052]

[0053] where, denotes the transmission time of the UAV u collecting the computing task; b>1 denotes the transmission overhead coefficient; N u denotes the set of common tasks collected by the UAV; denotes the data volume of the xth common task, R i,u denotes the maximum transmission rate of the ground sensor i to the UAV u; denotes the transmission energy consumption of the UAV u collecting the computing task ; P i,u denotes the transmission power of the ground sensor i to the UAV u; μ u denotes the proportion of the UAV u processing the task locally, 0≤μ u <1; denotes the total local computing delay of the UAV u; denotes the task complexity of the xth common task; f u denotes the computing resource of the UAV u; E denotes the total local computing energy consumption of the UAV u; τ denotes the energy coefficient; denotes the retransmission time of the UAV u retransmitting the result data to the ground sensor after the computation is completed; denotes the total data volume of the result data of the UAV u; denotes the energy consumption of the UAV u retransmitting the data;

[0054] The computing model of the cloud server is:

[0055]

[0056]

[0057] where, and denote the transmission time and the transmission energy consumption of the cloud server receiving the UAV task, respectively; R u,c denotes the maximum transmission rate of the uth UAV to the cloud server; P u,c denotes the transmission power of the uth UAV to the cloud server; and Represent the computing delay and computing energy consumption of the cloud server respectively; f C Represents the computing resources of the cloud server; and They represent the return delay of the calculation results of the computing task to the cloud server of the drone. and backhaul energy consumption; Indicates the total amount of server result data;

[0058] The calculation model of the satellite is:

[0059]

[0060] in, and They represent the transmission time and energy consumption of satellite s receiving the emergency mission and the unloading mission of the UAV respectively; represents the transmission time of the urgent mission in satellite s; represents the transmission time of ordinary missions in satellite s; M ur Indicates the total number of urgent tasks; represents the workload of the xth urgent task; R i,s represents the maximum transmission rate from ground sensor i to satellite s; R u,s represents the maximum transmission rate of UAV u to satellite s; P i,s represents the transmission power from ground sensor i to satellite s; P u,s represents the transmission power from UAV u to satellite s; and They represent the computational delay and computational energy consumption of satellite s in processing emergency tasks and ordinary tasks respectively; represents the task complexity of the xth urgent task; f s represents the computing resources of satellite s; and They represent the return delay and return energy consumption of satellite s when satellite s directly transmits the calculation results of the computing task to the ground sensor.

[0061] Secondly, a system for collaborative optimization of reliability of space-ground integrated network computing offloading is provided, including:

[0062] A task division module is used to divide computing tasks into urgent tasks and common tasks based on the urgency of the computing tasks;

[0063] The uninstallation strategy setting module is used to set the uninstallation strategies corresponding to emergency tasks and ordinary tasks;

[0064] An offloading strategy generation module is used to generate offloading strategies for multiple drones in an air-ground integrated network based on the set offloading strategies using the D-MAPPO algorithm, and decompose the offloading strategies of multiple drones into the offloading strategies of each drone in the environment;

[0065] The offloading strategy optimization module is used to optimize the generated offloading strategy using the D-MAPPO algorithm based on the pre-established communication model, satellite coverage time model and calculation model and the pre-set reliability mechanism, so that the integrated air-space-ground network can operate according to the optimized offloading strategy.

[0066] According to a third aspect, a processing device is provided, comprising computer program instructions, wherein the computer program instructions, when executed by the processing device, are used to implement the steps corresponding to the above-mentioned method for collaborative optimization of reliability of integrated air-space-ground network computing offloading.

[0067] In a fourth aspect, a computer-readable storage medium is provided, on which computer program instructions are stored, wherein the computer program instructions, when executed by a processor, are used to implement the steps corresponding to the above-mentioned method for collaborative optimization of the reliability of integrated air-space-ground network computing offloading.

[0068] The present invention has the following advantages due to the adoption of the above technical solution:

[0069] This paper proposes an integrated air-ground network framework suitable for remote areas, combining ground sensors, drones, satellites, and cloud servers. It models the network's communications, satellite coverage time, and costs. It also proposes a reliability mechanism for offloading tasks from ground sensors and drones to satellites.

[0070] 2. Based on the different characteristics of emergency tasks and ordinary tasks, this paper proposes a computational offloading problem that jointly optimizes network energy consumption and latency. Parameters such as bandwidth, computing resources, and the number of offloaded tasks that can be received by cloud servers and satellites are used as constraints. The optimization problem is modeled as an MDP, and the D-MAPPO algorithm is proposed to learn the optimal task offloading strategy.

[0071] 3. After a large number of simulation experiments, the present invention proves that the D-MAPPO algorithm has faster and more stable convergence, and the optimization effect of this algorithm in terms of delay, energy consumption and unloading success rate is better than Beta-MAPPO, PPO, local, offloading and random algorithms.

[0072] In summary, the present invention can be widely applied in the field of air-space-ground integrated network. BRIEF DESCRIPTION OF THE DRAWINGS

[0073] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiment below. The accompanying drawings are for illustration purposes only and are not to be considered as limiting the present invention. Throughout the drawings, the same reference numerals are used to denote the same components. In the drawings:

[0074] Figure 1 This is a flow chart of a method provided by one embodiment of the present invention;

[0075] Figure 2 This is a schematic diagram of an air-ground integrated network framework provided by an embodiment of the present invention;

[0076] Figure 3 1 is a schematic diagram of a satellite coverage model provided by an embodiment of the present invention;

[0077] Figure 4 Schematic diagram of the convergence reward of the D-MAPPO algorithm provided by one embodiment of the present invention under different learning rates, gamma and n_epochs;

[0078] Figure 5 This is a schematic diagram of the convergence of reward values ​​of different algorithms provided by an embodiment of the present invention;

[0079] Figure 6 This is a schematic diagram comparing the optimization of common task delays by different algorithms provided by an embodiment of the present invention;

[0080] Figure 7 This figure shows a comparison of the minimum latency of common tasks under different algorithm optimizations provided by an embodiment of the present invention, and a schematic diagram of the minimum latency changes of other algorithms based on the local algorithm.

[0081] Figure 8 This is a schematic diagram comparing the optimization of energy consumption of common tasks by different algorithms provided by an embodiment of the present invention;

[0082] Figure 9 This figure shows a comparison of the minimum energy consumption of common tasks under different algorithm optimizations provided by an embodiment of the present invention, and a schematic diagram of the changes in the minimum energy consumption of other algorithms based on the local algorithm.

[0083] Figure 10 This is a schematic diagram comparing the optimization of the success rate of common task offloading by different algorithms provided by an embodiment of the present invention;

[0084] Figure 11 This is a schematic diagram comparing the success rates of different task offloading when the rewards of different algorithms are the highest, provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0085] Exemplary embodiments of the present application will be described more fully hereinafter with reference to the accompanying drawings. While exemplary embodiments of the present application are shown in the drawings, it is to be understood that the present application can be embodied in many forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and fully convey the scope of the present application to those skilled in the art.

[0086] It is to be understood that the terminology used herein is for the purpose of describing particular example embodiments only and is not intended to be limiting. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. The terms "comprises", "comprising", "includes", "including" and the like are to be construed to be inclusive (i.e., to include both instances of open ended terms and instances of terms limiting to a specific number) unless otherwise indicated as otherwise limited by context. The methods described herein can be implemented as a method, an apparatus, a system, a computer program product, or any combination thereof.

[0087] Although the terms first, second, third, etc. can be used herein to describe various elements, components, regions, layers and / or sections, these elements, components, regions, layers and / or sections should not be limited by these terms. These terms can be only used to distinguish one element, component, region, layer or section from another region, layer or section. Terms such as "first", "second", and other numerical terms when used herein do not imply a sequence or order unless clearly indicated by the context. Thus, a first element, component, region, layer or section discussed below could be termed a second element, component, region, layer or section without departing from the teachings of the example embodiments.

[0088] Currently, scholars have proposed various resource allocation and task scheduling optimization strategies to try to alleviate the edge overload problem by building a reasonable offloading decision model. In addition, in order to expand the service range and capacity of edge computing, researchers introduce an auxiliary computing platform based on unmanned aerial vehicles (UAV), and build a multi-level offloading system of air-ground cooperation. However, due to the limited spectrum resources and limited node load capacity, especially in extreme environments such as disaster areas, oceans, and deserts, the coverage and computing resources of edge computing still face great challenges, and it is difficult to fully support the computing needs of high-density Internet of Things terminals in remote areas. The embodiment of the present invention provides a space-air-ground integrated network computing offloading reliability collaborative optimization method, which includes: based on the urgency of the computing task, the computing task is divided into urgent task and ordinary task; set the offloading strategy corresponding to the urgent task and the ordinary task; using the D-MAPPO algorithm, based on the set offloading strategy, the offloading strategy of multiple unmanned aerial vehicles in the space-air-ground integrated network is generated, and the offloading strategy of multiple unmanned aerial vehicles in the environment is decomposed into the offloading strategy of each unmanned aerial vehicle; using the D-MAPPO algorithm, based on the pre-established communication model, satellite coverage time model and computing model and the pre-set reliability mechanism, the generated offloading strategy is optimized, so that the space-air-ground integrated network works according to the optimized offloading strategy. The present invention unifies the cloud server and the satellite as part of the MEC, and gives the unmanned aerial vehicle (UAV) the role of decision maker, aiming to improve the service quality (QoS) of the entire network by selecting the optimal offloading strategy. Further, the present invention not only considers traditional network resources such as bandwidth and computing capacity as variable conditions, but also introduces the limiting factor of the number of offloading tasks that the cloud server and the satellite can receive, to more comprehensively analyze their influence on offloading decision and network performance. In addition, to meet the differentiated needs of different types of tasks, the present invention divides the tasks into urgent tasks and ordinary tasks, and designs two different offloading strategies to achieve multi-level service quality guarantee. Finally, since the traditional deep reinforcement learning (DRL) algorithm can adapt to dynamic environment optimization, but is mostly based on discrete action space, limiting the optimization efficiency, the present invention proposes a new multi-agent reinforcement learning algorithm based on Dirichlet distribution and multi-agent proximal policy optimization (MAPPO), named D-MAPPO algorithm. This algorithm can more efficiently deal with optimization problems in dynamic environments by introducing a continuous action space.

[0089] Embodiment 1

[0090] As Figure 1 shown, the embodiment provides a space-air-ground integrated network computing offloading reliability collaborative optimization method, which includes the following steps:

[0091] 1) Based on the urgency of the computing task, the computing task is divided into urgent task and ordinary task, specifically:

[0092] Specifically, emergency tasks are generated when monitoring computing tasks occur in emergencies, such as fires and earthquakes. Emergency tasks are a relatively small percentage of all tasks and are prioritized. Therefore, they are prioritized to minimize computing task delays and ensure their complete completion.

[0093] Common tasks are tasks generated by ground sensors, except for urgent tasks. This means that the majority of computing tasks generated by all ground equipment are common tasks. Therefore, optimizing common tasks is the focus of optimizing computational offloading for the entire system. Because common tasks are latency-tolerant, energy consumption and latency should be jointly optimized.

[0094] Specifically, if Figure 2 The figure shows the framework diagram of the integrated space-air-ground network, where I represents the ground sensor set, i∈I; U represents the UAV set, u∈U; S represents the satellite set, s∈S; C represents the cloud server, and there is only one cloud server; M represents the offload task set, m∈M, where M n Indicates ordinary tasks, M ur Indicates an urgent task.

[0095] 2) Set the corresponding offloading strategies for urgent tasks and ordinary tasks.

[0096] Specifically, the emergency task offloading strategy is as follows: When ground sensors in the integrated space-ground network generate an emergency task, the task is directly offloaded to the satellite via the connection channel between the ground sensor and the satellite. The satellite processes the emergency task and transmits the calculation results directly back to the ground sensor. The number of emergency tasks offloaded to the satellite is evenly distributed based on the service time of each satellite. In other words, the longer the satellite's service time, the more emergency tasks it receives.

[0097] The offloading strategy for common tasks is as follows: Ground sensors in the integrated air-ground-ground network offload common tasks to drones via the connection channel between the ground sensors and the drones. Based on the determined offloading ratios, the drones offload common tasks to satellites and cloud servers for processing, while the remaining common tasks remain on the drone itself for processing. The offloading ratio of common tasks to satellites and cloud servers for each drone is optimized using the following algorithm.

[0098] Specifically, each drone can perform two-way data transmission with ground sensors and cloud servers. Note that since a backhaul link has already been established between the satellite and the ground sensors, the satellite does not need to transmit the calculation results back to the drone. Therefore, only one-way data transmission is performed between the drone and the satellite.

[0099] 3) Based on the air-ground-integrated network framework of ground sensors, drones, satellites and cloud servers, establish the communication model, satellite coverage time model and calculation model in the air-ground-integrated network.

[0100] 3.1) Establishing a communication model in an integrated air-space-ground network:

[0101] 3.1.1) Establish a communication model between the UAV and ground sensors.

[0102] Specifically, the air-to-ground communication channel depends on the altitude, elevation angle, and propagation environment. The average path loss of the air-to-ground channel can be defined as:

[0103]

[0104] Among them, P Los represents the probability of line of sight (Loss) between the ground sensor and the UAV; h represents the flight altitude of the UAV; r represents the horizontal distance between the UAV and the ground sensor; η Los ,η NLos They represent the additional losses of the Los link and the non-Los link on the basis of the free space path loss, and (η Los ,η NLos ) value is (0.1,2.1); f c represents the carrier frequency; c represents the speed of light.

[0105] Assuming that the communication link between the ground sensor and the UAV operates in the C-band spectrum, the maximum transmission rate of the ground sensor is R i and the maximum transmission rate R of the drone u for:

[0106]

[0107] Among them, B i Indicates the bandwidth of the ground sensor; B u represents the bandwidth of the drone; P i and P u Represent the transmission power of ground sensors and UAVs respectively; σ 2 represents the Gaussian noise power.

[0108] 3.1.2) Establish a communication model between the drone and the cloud server.

[0109] Specifically, the cloud server is also a device deployed on the ground, so the communication is similar to that between the drone and the ground sensor. The difference is that the cloud server also needs to transmit the result data back to the drone, and the drone transmits the result data back to the ground device. Therefore, the maximum transmission rate of the cloud server is R cfor:

[0110]

[0111] Among them, B c Indicates the bandwidth of the cloud server; P c Indicates the transmit power of the cloud server.

[0112] 3.1.3) Build a communication model between UAVs and ground sensors and satellites.

[0113] Specifically, the communication link between the UAV and the satellite is mainly based on a clear line-of-sight (LoS) link, supplemented by a small amount of non-line-of-sight (NLoS) links. The communication channel between the UAV and the satellite is modeled as a Rician channel. Therefore, the channel gains of LoS and NLoS are integrated, so the channel coefficient between the UAV and the satellite is Among them, F represents the Rician factor, α represents the distance attenuation factor, ξ LoS and ξ NLoS They represent the LoS and NLoS channel gains between the satellite and the communication equipment respectively. Finally, the maximum transmission rate of the drone, ground sensor and satellite under line-of-sight communication can be obtained. for:

[0114]

[0115] Among them, B s represents the available bandwidth of the satellite; G0 represents the fixed antenna gain; P s represents the satellite's transmit power; N0 represents the spectral density of the additive white Gaussian noise (AWGN).

[0116] 3.2) Establish a satellite coverage time model in the integrated air-space-ground network.

[0117] Specifically, given the dynamic characteristics of low-Earth orbit satellites, ground sensors and drones can only communicate with satellites when they are within their coverage area. Therefore, the coverage time of these satellites must be modeled as a basic reference for making intelligent task offloading decisions. Figure 3 As shown, the geometric relationship between low-Earth orbit satellites and user equipment is shown, through which the low-Earth orbit coverage time model can be obtained. It is worth noting that the satellites are deployed in space 300km above the ground, while the drones are deployed in the sky within 100m from the ground. The height of the drone is very small compared to the height of the satellite, so in the satellite coverage model of the present invention, the height of the drone is negligible. The drone and the ground sensor use the same satellite coverage model. The present invention refers to drones and ground sensors as user equipment. The specific process of this step is:

[0118] 3.2.1) Calculate the elevation angle θ between the user equipment and the satellite communication G :

[0119]

[0120] Among them, d E represents the radius of the Earth; d o Indicates the orbital height of the satellite; d GS represents the distance between the user equipment and the satellite; θ c represents the coverage angle of the low-Earth orbit satellite, and:

[0121]

[0122] 3.2.2) Based on the coverage angle θ of the low-Earth orbit satellite c , calculate the satellite coverage arc length L S :

[0123] L S =2(d E +d o )θ c (7)

[0124] 3.2.3) According to the satellite coverage arc length L S , calculate the satellite coverage time T S :

[0125]

[0126] Among them, V S Indicates the satellite's flight speed.

[0127] 3.3) Establish computing models for drones, satellites, and cloud servers.

[0128] Specifically, in the working environment of the present invention, the devices that can process computing tasks include drones, satellites, and cloud servers. Ground sensors do not process tasks. For a computing task m given by a drone sensor, assuming in, represents the amount of data for computing task m, ρ m represents the complexity of computational task m, i.e., the number of CPU cycles required to process the computational task. Ordinary tasks require drones to collect data, which is then processed separately based on the drone's offload decision. Emergency tasks, on the other hand, are directly offloaded from ground sensors to satellites for processing, and the processing results are directly transmitted back to the ground sensors by the satellites. The specific process of this step is as follows:

[0129] 3.3.1) Establish the calculation model of the UAV (including formula (9) to formula (14)), including the delay model and energy consumption model of the UAV.

[0130] Specifically, the computational tasks generated by ground sensors collected by UAVs can be processed locally or offloaded. The transmission time of the computational tasks collected by UAV u is:

[0131]

[0132] Where b>1 represents the transmission overhead coefficient; N u represents the set of common tasks collected by the drone; Indicates the data volume of the xth common task, R i,u represents the maximum transmission rate from ground sensor i to UAV u.

[0133] UAV u collection computing tasks The transmission energy consumption is:

[0134]

[0135] Among them, P i,u represents the transmission power from ground sensor i to UAV u.

[0136] Let μ u is the proportion of UAV u processing tasks locally, 0≤μ u ≤1, then the total local computing delay of UAV u for:

[0137]

[0138] in, represents the task complexity of the xth common task; f u Represents the computing resources of UAV u.

[0139] The total local computing energy consumption e of UAV u is:

[0140]

[0141] Where τ represents the energy coefficient, which depends on the CPU structure of UAV u. When the calculation is completed, UAV u transmits the result data back to the ground sensor. The transmission time is:

[0142]

[0143] in, Indicates the total amount of drone u result data.

[0144] Energy consumption of UAV data transmission for:

[0145]

[0146] 3.3.2) Establish a calculation model for the cloud server (including formulas (15) to (20)).

[0147] Specifically, the computing resources of the cloud server are also limited. The number of tasks that the cloud server can process is recorded as ψ. When ψ is greater than zero, the cloud server receives the computing tasks unloaded by the drone for processing. The cloud server has to accept computing tasks from multiple drones. Let μ (u,C) is the offloading ratio of drone u to the cloud server, 0≤μ (u,C) ≤1, the cloud server receives the transmission time of the drone mission and transmission energy consumption for:

[0148]

[0149] Among them, R u,c represents the maximum transmission rate from the u-th UAV to the cloud server; P u,c represents the transmission power of the u-th UAV to the cloud server.

[0150] Computational latency of cloud servers and calculate energy consumption for:

[0151]

[0152] Among them, f C Represents the computing resources of a cloud server.

[0153] The cloud server needs to transmit the calculation results of the computing task back to the drone, and the return delay of the cloud server and backhaul energy consumption for:

[0154]

[0155] in, Indicates the total amount of server result data.

[0156] 3.3.3) Establish a satellite calculation model (including formulas (21) to (26)).

[0157] Specifically, the satellite not only has to process the computational tasks offloaded by the drone locally, but also has to process the urgent tasks uploaded by the ground sensors. Let μ (u,s) is the ratio of unloading ordinary tasks from UAV u to satellite s, 0≤μ (u,s) ≤1, it is worth noting Transmission time for satellites to receive emergency missions and unload missions from drones and transmission energy consumption for:

[0158]

[0159] in, represents the transmission time of the urgent mission in satellite s; represents the transmission time of ordinary missions in satellite s; M ur Indicates the total number of urgent tasks; Indicates the workload of the xth urgent task; R i,s represents the maximum transmission rate from ground sensor i to satellite s; R u,s represents the maximum transmission rate of UAV u to satellite s; P i,s represents the transmission power from ground sensor i to satellite s; P u,s represents the transmission power from UAV u to satellite s.

[0160] Satellites handle computing delays for urgent and routine tasks and calculate energy consumption for:

[0161]

[0162] in, represents the task complexity of the xth urgent task; f s Represents the computing resources of satellite s.

[0163] Satellite s directly transmits the calculation results of the computing task to the ground sensor, and the return delay of satellite s and backhaul energy consumption for:

[0164]

[0165] 4) Establish a reliable mechanism for ground sensors and drones to offload computing tasks to satellites.

[0166] Specifically, the task is considered to be successfully offloaded only when every bit of each computing task is successfully uploaded. However, the satellite coverage time is limited, so when offloading computing tasks to the satellite, the computing task may fail to be offloaded. If the offloading of an emergency task fails, the consequences will be disastrous. Therefore, the reliability mechanism for offloading computing tasks to the satellite set by the present invention includes:

[0167] ① Limit the minimum value of the maximum transmission rate of the device. The minimum value of the maximum transmission rate can be calculated based on the satellite coverage time model.

[0168] ② When each device offloads tasks to the satellite, it first calculates the maximum number of computing tasks that can be offloaded, and emergency tasks are offloaded to a limited extent.

[0169] ③The mission transmission time should be less than the satellite coverage time, that is:

[0170]

[0171] in, represents the total mission volume transmitted to the satellite, and R represents the maximum transmission rate of the user equipment.

[0172] ④ Restrictions on the maximum amount of data transmitted by user equipment:

[0173]

[0174] 5) Based on the Dirichlet probability distribution, determine the D-MAPPO algorithm.

[0175] Specifically, the present invention considers the delay and energy consumption issues of common and urgent tasks in SAGIN. Urgent tasks focus on whether the computing tasks can be completed and whether the computing tasks can be completed quickly; common tasks focus on whether the computing tasks can be completed with lower energy consumption and delay. According to the task offloading mechanism, urgent tasks are directly offloaded to the satellite for processing, ignoring energy consumption and delay, and only considering the completion status. Regardless of whether it is an urgent task or a common task, as long as it is successfully offloaded to the offloading target, it can be calculated as completed. Therefore, it is only necessary to consider the offloading success rate of all offloaded tasks. The offloading success rate δ of the offloaded tasks can be expressed as:

[0176]

[0177] Among them, N su and N up They are respectively expressed as the total number of successfully offloaded tasks and the total number of offloaded tasks. In summary, the optimization problem can be derived, that is, while ensuring the completion rate of emergency task offloading, improving the offloading completion rate of ordinary tasks, reducing the energy consumption and delay of ordinary tasks. The problem model is shown in P1:

[0178]

[0179] Where E represents the total energy consumption of the entire network system; T represents the total time of SAGIN operation; δ n represents the offloading success rate of common tasks; B represents bandwidth; B min Indicates the minimum bandwidth; B max represents the maximum bandwidth; f represents computing resources; f min Indicates the minimum computing resources; f max represents the maximum computing resources; δ ur Indicates the success rate of offloading urgent tasks; Indicates the workload of common tasks; Minimum task quantity of common task Maximum task quantity of common task n Task complexity of common task Minimum task complexity of common task Maximum task complexity of common task.a indicates that the bandwidth resources of all devices in the network are uncertain, but are within the limit range;b indicates that the maximum task data quantity of the unmanned aerial vehicle to the satellite should meet the reliability constraint;c indicates that the computing resources of all computing devices in the network are also uncertain, but are within the limit range;d-f indicates that the offloading rates in the offloading strategy should be between 0 and 1, and g indicates that the sum of all offloading rates is 1;h indicates that the offloading success rate of the emergency task is 100%;i indicates that the number of tasks offloaded by all unmanned aerial vehicles to the cloud server is less than the maximum number of tasks that can be processed by the cloud server;j-k indicates that the task quantity and task complexity of the common task are between the minimum value and the maximum value.

[0180] To solve the above-mentioned computing offloading optimization problem, the present application designs a reinforcement learning algorithm based on Dirichlet probability distribution and MAPPO, namely D-MAPPO algorithm, which has good adaptability to continuous and dynamic environment and can better adapt to the SAGIN environment. Therefore, the specific process of this step is:

[0181] 5.1) Establish an MDP model, including a state set, an action set, state transition and a reward function.

[0182] Specifically, Markov Decision Process (MDP) is a mathematical framework for modeling sequential decision-making problems, which is widely used in Reinforcement Learning (RL). MDP describes the interaction between the agent and the environment through state, action, state transition and reward function, so as to learn the optimal decision strategy. In the computing offloading problem of space-air-ground integrated network (SAGIN), the computing offloading is modeled as MDP to optimize the offloading strategy of computing tasks, improve the computing efficiency and reduce the energy consumption and delay. The total time of SAGIN operation is set as T, and T is divided into t time slots.

[0183] Specifically, the state set (State Space) defines the possible states of the environment, which describes the current situation of the system and provides the basis for the decision of the agent. In the computing offloading problem, the state set should be able to fully express the availability of computing resources and the execution state of tasks, so that the agent can select the appropriate offloading strategy accordingly. In the present application, each unmanned aerial vehicle is regarded as an agent, and the state s t is defined as:

[0184]

[0185] where f C represents the computing resource of the cloud server; B C represents the bandwidth resource of the cloud server; ψ represents the number of tasks that the cloud server can handle; f s represents the computing resource of the satellite s; B s represents the bandwidth resource of the satellite s; θ c,s represents the coverage angle of the satellite s, s∈S, note that there are multiple satellites in the space-air-ground integrated network of the present application, state s t records the states of all satellites; f u represents the computing resource of the UAV u; B u represents the bandwidth resource of the UAV u; N n represents the number of ordinary tasks; N ur represents the number of emergency tasks; ρ represents the task computing complexity; represents the amount of a single task.

[0186] Specifically, Action Space: The action space defines the operations that the agent can perform in each state. In the computing offloading problem, the action determines the allocation ratio of the computing task on different computing nodes, so the present application takes the offloading ratio μ u ,μ (u,s) ,…,μ (u,s) ,μ (u,C) as the action set, i.e. the action a t of each agent at time slot t a u ,μ (u,s) ,…,μ (u,S) ,μ (u,C) ), where the offloading ratios in the action set a t satisfy the constraint conditions d-g in the problem P1.

[0187] Specifically, State Transition: State transition describes the process of the system changing from one state to another after performing an action. In the computing offloading problem, state transition depends on the influence of the offloading decision on the resource and task state. The state transition of the present application is jointly influenced by the following factors: ① the change of the available computing resources of each device after the execution of the computing task; ② the arrival of new computing tasks after the completion of the computing task. The state transition can be described by a probability model , which reflects the influence of the offloading decision on the future state of the system.

[0188] Specifically, the reward function measures the pros and cons of the offloading decision, which is the basis for the reinforcement learning agent to optimize the strategy. The goal of the present application is to improve the offloading completion rate of ordinary tasks and reduce the energy consumption and delay of ordinary tasks while ensuring the offloading completion rate of emergency tasks, therefore, the reward function r(s t ,a t ) is:

[0189]

[0190] wherein k represents the weight coefficient between energy consumption and delay, which can be adjusted according to different user needs; i represents the penalty coefficient, and each time a task is not offloaded, it is penalized, and the total penalty amount will decrease as the task offloading success rate increases; (1-δ n )N n represents the number of ordinary tasks that are not offloaded; (1-δ ur )N ur represents the number of emergency tasks that are not offloaded.

[0191] 5.2) Based on the established MDP model, the D-MAPPO algorithm is obtained:

[0192] 5.2.1) Determine the Dirichlet probability distribution as the probability distribution function.

[0193] Specifically, in the MAPPO algorithm, the policy network (actor network) is responsible for generating the action of the agent. Since the calculation of offloading tasks involves continuous action space, the policy network needs to use a probability distribution function to generate an offloading allocation scheme that meets the constraint conditions. In the action space a t , all offloading proportions are limited to the interval [0, 1], and their sum must be equal to 1. However, traditional probability distribution functions usually cannot meet this constraint condition, and if a distribution that does not meet the constraint is directly used, additional regularization or output format adjustment needs to be applied during the training process, or the action restriction condition needs to be added in the environment, which will increase the computational overhead and reduce the convergence efficiency and optimization effect. To avoid introducing additional complex constraint processing during the policy training process, the present application selects the Dirichlet probability distribution. The Dirichlet probability distribution meets the constraint requirements of the offloading proportion, which is a multi-dimensional probability distribution defined between [0, 1], and the sum of all offloading proportions must be equal to 1, which is also consistent with the properties of the Dirichlet probability distribution, so the Dirichlet probability distribution is suitable for the calculation of the offloading environment of the present application. The probability density function of the Dirichlet probability distribution is:

[0194]

[0195] wherein, denotes the probability density function of Dirichlet probability distribution; x J denotes the offloading ratio, satisfying α J denotes the parameter of Dirichlet probability distribution; B(α) denotes the normalization term of Beta function.

[0196] 5.2.2) Based on the established MDP model, the advantage estimation function, the policy update target function and the value function of the MAPPO algorithm are determined.

[0197] Specifically, Multi-Agent Proximal Policy Optimization (MAPPO) is a multi-agent reinforcement learning algorithm based on the Proximal Policy Optimization (PPO) framework. By extending the trust region constraint mechanism of PPO to the multi-agent scenario, it solves the instability problem of traditional methods in non-steady-state environments and policy collaborative update. The algorithm adopts the Centralized Training Decentralized Execution (CTDE) architecture, in which agents share a global value function during the training phase to accelerate learning, and rely on independent policy networks for distributed decision-making during the execution phase. Experiments show that MAPPO improves the win rate by more than 23% in complex collaborative tasks such as StarCraft Multi-Agent Challenge (SMAC). The advantages of the MAPPO algorithm mainly lie in the following aspects: ① Strong stability: By inheriting the policy update clipping mechanism of PPO, MAPPO explicitly constrains the update amplitude of the multi-agent joint policy, avoiding the gradient explosion problem caused by collaborative optimization of policy networks; ② High sample efficiency: The shared value function design allows agents to use global state information for value estimation during the training phase, significantly alleviating the multi-agent credit assignment problem; ③ Adaptability to complex tasks: For partially observable environments (POE), MAPPO integrates a temporal modeling module (such as a Gated Recurrent Unit) into the policy network, allowing agents to infer the environment state based on local observation history.

[0198] Specifically, the key steps of the MAPPO algorithm include calculating the advantage estimation function (Advantage Estimation), policy update (Policy Update), and value function update (Value Function Update).

[0199] Specifically, in reinforcement learning, the advantage estimation function is used to measure whether a certain action is better than the average action, and it is a key quantity in the MAPPO algorithm, which helps to reduce variance and improve learning efficiency. The MAPPO algorithm uses the Generalized Advantage Estimation (GAE) The calculation formula is:

[0200]

[0201] Among them, δ t TD error (Temporal Difference Error), which represents the estimated difference between the reward at the current moment and the value of the next state, δ t =r t +γV φ (o t+1 )-V φ (o t ), γ represents the discount factor, which is used to measure the importance of future rewards; λ represents the GAE smoothing parameter, which controls the bias-variance trade-off in the advantage estimation function.

[0202] Specifically, in the MAPPO algorithm, the policy update uses the PPO clipping loss (Clipped PPO Loss) to avoid excessive policy updates and ensure training stability. For a system containing multiple agents, the joint policy update objective function is defined as:

[0203]

[0204] Among them, r t (θ) represents the probability ratio. π θ Indicates the current strategy; Indicates the previous old strategy; o t Represents the current environmental observation value; clip represents the clipping operation (Clipping), if r t If (θ) changes too much, the part exceeding [1-∈, 1+∈] will be directly truncated to avoid excessive policy updates; ∈ represents a hyperparameter used to control the clipping range, and its value is usually [0.1, 0.2].

[0205] Specifically, in reinforcement learning, the value function estimates the long-term benefits of the current state. In the MAPPO algorithm, the value function V φ Optimization by mean square error (MSE Loss):

[0206]

[0207] Among them, L VF (φ) represents the loss function of the value function, which is used to measure the current value function V φ The estimated value and actual cumulative discounted return R t The error between the target value R trepresents the cumulative discounted return, R t =r t +γr t+1 +γ 2 r t+2 +….

[0208] 6) Using the D-MAPPO algorithm, based on the neural network and Dirichlet probability distribution, the offloading strategies of multiple drones are generated based on the set offloading strategies, and the offloading strategies are transmitted to the environment. In the environment, the offloading strategies of multiple drones are decomposed into the offloading strategies of each drone for use by the drones.

[0209] 7) Using the D-MAPPO algorithm, based on the established communication model, satellite coverage time model, calculation model, and established reliability mechanism, the generated offloading strategy is optimized so that the air-ground integrated network operates according to the optimized offloading strategy. Specifically:

[0210] 7.1) Using the satellite coverage time model and reliability mechanism, calculate the coverage time and maximum uploadable task volume of each satellite. Based on the maximum uploadable task volume of the satellite, offload urgent tasks to the satellite in proportion.

[0211] 7.2) Use the communication model between the UAV and ground sensors and the satellite to calculate the transmission delay.

[0212] 7.3) Use the satellite computing model to process emergency tasks and calculate the transmission delay, transmission energy consumption, computing delay and computing energy consumption of emergency tasks.

[0213] 7.4) Each UAV calculates the maximum remaining uploadable tasks to the satellite based on the established reliability mechanism, then calculates an offloading strategy to sequentially offload common tasks. Based on the communication models between the UAV and ground sensors, the communication model between the UAV and the cloud server, the communication model between the UAV and ground sensors and the satellite, and the computational model between the UAV, satellite, and cloud server, these common tasks are processed and their transmission delay, transmission energy consumption, computational delay, and computational energy consumption are calculated.

[0214] 7.5) After all computing tasks are processed, the energy consumption, delay, and offloading success rate of the entire system model are obtained. Based on the reward function, the reward score of this round of offloading strategy is calculated and transmitted to the D-MAPPO algorithm.

[0215] 7.6) The D-MAPPO algorithm calculates the probability ratio of each agent to choose each action under the current unloading strategy, that is, the probability ratio of the new strategy to the old strategy.

[0216] 7.7) To ensure the stability of the update, the probability ratio is clipped and limited to a fixed range.

[0217] 7.8) The clipped probability ratios are weighted by the advantage estimation function to construct the joint strategy update objective function.

[0218] 7.9) The optimizer adjusts the parameters of the policy network according to the gradient of the joint policy update objective function, thereby completing the update of the offloading policy.

[0219] 7.10) After each update, the new uninstall policy is used as the old uninstall policy to prepare for the next sampling and update.

[0220] Specifically, during the value network update, the discounted reward and the estimated value of the state-value function are used to calculate the advantage estimate function for each time step. The mean squared error between the value network's output and the target reward is used as the loss function. The optimizer updates the value network's parameters based on the gradient of this loss function, enabling the value network to more accurately estimate the state value. This process is typically repeated multiple times after each sampling cycle to fully utilize the sampled data and improve the accuracy of the value function estimate. This process is repeated until the algorithm terminates.

[0221] The beneficial effects of the D-MAPPO algorithm proposed in the collaborative optimization method for reliability of space-ground integrated network computing offloading are described in detail below through specific embodiments:

[0222] I) Parameter settings

[0223] In the simulation, Python and PyTorch were used to train the D-MAPPO algorithm proposed in this paper. Unless otherwise specified, the parameters are set as follows: the transmission power P of the ground user i =0.1W, the UAV’s transmission power P u , the transmission power P of the cloud server c and satellite P s Bandwidth resources are variable resources, and the bandwidth of satellite and cloud servers is B s and B c All are set at around 1GHz, and the bandwidth of the drone is B u At around 10MHz, the bandwidth B of the ground sensor i Around 1MHz. The minimum elevation angle θ for ground equipment and drones to communicate with satellites G =40, the satellite's moving speed V s =7.8km / s. The data size of each computing task is randomly generated, ranging from 0.9MB to 1.1MB. The computing CPU cycles required for each computing task are 450Gcycles to 550Gcycles, and the energy coefficient τ = 10-25 The total number of common tasks is about 236, and the total number of emergency tasks is about 36, accounting for 13% of the total tasks. The computing power of each UAV is [450, 550] MHz, the computing power of the low earth orbit satellite and the cloud server is [0.9, 1.1] GHz. The number of UAVs U is set to 5, the number of low earth orbit satellites S is set to 3, and the number of cloud servers C is set to 1.

[0224] Based on the above parameter settings, the D-MAPPO algorithm of the application and the other five benchmark algorithms are used to optimize the offloading strategy under the same environmental parameter settings. A total of six algorithms are used for comparative experiments. The following is an introduction to the other five algorithms:

[0225] Beta-MAPPO: MAPPO algorithm using Beta distribution to model the action space. The action space needs to be trained to meet the requirements of the environment in the paper. It is also suitable for multi-agent continuous action space offloading decision problems.

[0226] PPO: a single-agent reinforcement learning algorithm. Since the total action set of the five UAVs in the experiment is 20-dimensional, Dirichlet distribution cannot be used, and Beta distribution cannot meet the performance requirements of training, so a discrete action space is used.

[0227] Local: Each UAV will collect the computing tasks and perform local computation without making other offloading decisions.

[0228] Offloading: All tasks are offloaded to the cloud server and satellite for processing.

[0229] Random: The action selection is completely random and does not consider the environmental state or policy optimization, often used as a lower limit baseline for algorithm performance comparison.

[0230] II) Convergence of D-MAPPO algorithm

[0231] To verify the convergence of the algorithm and select the optimal hyperparameters, the learning rate (lr), discount factor (gamma), and policy update rounds (n_epochs) under different hyperparameters are compared. The comparison results are shown in Figure 4 .

[0232] In Figure 4(a) compares the average reward of the D-MAPPO algorithm over the number of training episodes for three learning rate settings: 0.003, 0.0003, and 0.00003. The figure shows that with a learning rate of 0.0003 (red line), the algorithm exhibits the fastest convergence rate, converges within approximately 200 episodes, and ultimately stabilizes at the highest average reward, demonstrating good learning efficiency and performance. Despite a relatively slow convergence rate with a learning rate of 0.003 (blue line), the algorithm still achieves high performance after 400 episodes and ultimately exhibits good stability. With a learning rate of 0.00003 (green line), the algorithm exhibits significant oscillation during training, with an average reward significantly lower than the previous two settings, and fails to converge effectively. This suggests that the excessively small learning rate limits the policy updates, affecting learning effectiveness.

[0233] exist Figure 4 (b) shows the training performance of the D-MAPPO algorithm with discount factors of 0.8, 0.95, and 0.99, respectively. When the discount factor is 0.95 (red line), the average reward converges quickly and reaches a high stable value, indicating that the policy learning effect is better when considering longer-term rewards. When the discount factor is 0.99 (blue line), the optimization effect is the lowest before 700 rounds. After 700 rounds, the optimization effect improves, but still falls short of the optimization effect with a discount factor of 0.95. When the discount factor is 0.8 (green line), although the initial convergence is faster, the performance declines in the later stages of training, and the final reward is lower than the previous two, indicating that overly short-term reward estimates are not conducive to long-term policy optimization.

[0234] exist Figure 4 (c) compares the algorithm performance for 5, 7, and 10 policy update rounds. When n_epochs equals 5 and 7 (red and green lines), the average return converges quickly and relatively smoothly, indicating that a moderate number of policy updates can ensure sufficient learning while avoiding overfitting. The figure shows that when n_epochs equals 5, the optimization effect is better, the curve is smoother, and the convergence is faster. When n_epochs equals 10 (blue line), although the performance is good in the initial stage, the return decreases in the later stages of training, indicating that excessive policy updates may cause the policy to overfit the current sample, thereby reducing generalization performance.

[0235] After comparing the three hyperparameters, it was found that the D-MAPPO algorithm achieved better optimization results when lr = 0.0003, gamma = 0.95, and n_epochs = 5. Therefore, the above three hyperparameters were used in the subsequent comparison with other algorithms.

[0236] III) Performance comparison of D-MAPPO algorithms

[0237] After selecting the appropriate hyperparameters, to verify the optimization effect of D-MAPPO algorithm, the optimization effect of D-MAPPO algorithm and the other five algorithms is compared. When optimizing, the tasks that have not been offloaded are in the timeout state, and the timeout tasks are given fixed delay and energy consumption values.

[0238] i) Reward value comparison between different algorithms

[0239] As shown in Figure 5 , the training performance of the six algorithms in the same environment: D-MAPPO (red line), Beta-MAPPO (green line), PPO (blue line), local (black line), offloading (orange line), and random (sky blue line). The horizontal axis is the number of training rounds (episodes), and the vertical axis is the average reward. D-MAPPO performs best, with average reward rapidly rising in a very short time and reaching a stable state in about 200 rounds. The final converged reward value is -73, much higher than other algorithms, showing strong convergence speed and strategy quality. It shows significant advantages in dynamic variable action space and multi-agent coordination. Beta-MAPPO has very low reward in the initial stage, even lower than the random strategy, showing significant instability. But it rebounds after about 150 rounds and continues to rise, eventually stabilizing at -105, better than PPO and local, but still lagging behind D-MAPPO. This "lateness" is due to the unstable convergence of Beta distribution strategy in action sampling, although the long-term performance is acceptable, but the training efficiency is not as good as D-MAPPO. PPO has a relatively smooth convergence process, and eventually converges to around -90. Although it is better than local, offloading, random, and Beta-MAPPO strategies, it is significantly lower than D-MAPPO. This shows that single-agent strategies are difficult to capture the synergy in multi-agent environments and have limited adaptability. Local performance is basically stable at -126, as all normal tasks can be completed locally without offloading, so there is no penalty. Offloading will offload all tasks, and the tasks that can be received by the cloud server and satellite are limited, so many tasks cannot be successfully offloaded, resulting in a lot of accumulated penalties, making the reward value very low, around -500. The reward value of random fluctuates greatly, because the offloading strategy in random state has too much randomness, so the performance difference is also very large.

[0240] ii) Delay optimization comparison

[0241] As shown in Figure 6As shown, the performance of local, offloading (full offloading), and random algorithms in ordinary task delay optimization is compared. The results show that the D-MAPPO algorithm of the present application exhibits a significant delay reduction trend at the beginning of training, and converges to a minimum delay value (<0.6s) after about 200 rounds, which is much better than other algorithms. Although the initial delay of Beta-MAPPO is extremely high, it quickly stabilizes between 0.8s and 0.9s, which is not as good as D-MAPPO and PPO. In contrast, PPO converges slowly and the final delay is about 0.95s, which is second only to D-MAPPO. The local and random strategies stabilize at about 1.0s and 1.1s, respectively, showing the limitations of lacking dynamic scheduling ability; the offloading strategy always maintains a high delay (>3s) and has no obvious optimization trend. In summary, the D-MAPPO algorithm of the present application exhibits excellent learning ability and decision-making efficiency in delay optimization, verifying its superiority in complex multi-agent offloading scenarios.

[0242] As Figure 7 shown, the performance of the D-MAPPO algorithm of the present application and the Beta-MAPPO, PPO, local, offloading, and random algorithms in ordinary task minimum delay optimization, and the minimum delay change rate relative to the local algorithm is compared in the form of a columnar line graph. Analysis shows that the D-MAPPO algorithm of the present application performs best in minimum delay, with a value of 0.52 seconds, while the offloading algorithm has the highest minimum delay of 2.3 seconds. In terms of minimum delay change rate, the offloading algorithm has the highest change rate relative to the local algorithm, about 1.25, while the D-MAPPO algorithm has the lowest change rate relative to the local algorithm, -0.47. This indicates that the D-MAPPO algorithm of the present application has a significant advantage in delay optimization, while the high delay of the offloading algorithm is related to its specific offloading strategy, resulting in higher delay. Other algorithms are between the two in terms of minimum delay and delay change rate, with an average performance. This result provides an important reference for the selection and optimization of computational offloading algorithms, especially in delay-sensitive application scenarios, the D-MAPPO algorithm of the present application has more application potential.

[0243] iii) Energy consumption optimization comparison

[0244] As Figure 8As shown in the figure, the performance of several algorithms in optimizing energy consumption for common tasks is compared. The energy consumption of the D-MAPPO algorithm of the present invention drops rapidly in the early stage of training, stabilizes after about 100 episodes, and finally stabilizes at about 0.07KJ, with the best optimization effect. The energy consumption of the Beta-MAPPO algorithm drops slightly slower, and stabilizes at about 0.072KJ after about 200 episodes. The energy consumption of the PPO algorithm gradually drops from 0.7KJ to about 0.2KJ, and fluctuates greatly, and the optimization effect is unstable. The energy consumption of the local algorithm is always maintained at 0.025KJ because wireless data transmission requires a large transmission power, especially long-distance transmission, such as communication with satellites, and the local algorithm does not need long-distance communication with satellites and cloud servers, so the energy consumption under the local algorithm is the lowest, and the performance is stable but lacks flexibility. The energy consumption of the offloading algorithm is as high as 2.7KJ and remains unchanged, with the worst performance. The energy consumption of the random algorithm fluctuates greatly, but it is also always stable at 0.5~0.6KJ. In general, the D-MAPPO and Beta-MAPPO algorithms have obvious advantages in energy consumption optimization and can effectively reduce energy consumption, and the optimization performance of the D-MAPPO algorithm of the present invention is better than that of Beta-MAPPO.

[0245] like Figure 9 The figure shows a bar chart comparing the lowest energy consumption of common tasks and the rate of change of lowest energy consumption relative to the local algorithm. The horizontal axis is the algorithm name, the left vertical axis represents the average energy consumption (KJ), and the right vertical axis represents the rate of change of lowest energy consumption. The D-MAPPO algorithm has an energy consumption of 0.072KJ and an energy consumption change rate of nearly 1.68, performing the best. The Beta-MAPPO and PPO algorithms have slightly higher energy consumption and a higher rate of change than D-MAPPO, at 1.87 and 4.14, respectively. The local algorithm consumes about 0.025KJ because it does not have a task offloading function.

[0246] The offloading algorithm performed the worst, consuming a whopping 1.83 kJ and exhibiting a 71.8% variability. The random algorithm consumed 0.48 kJ and exhibited a variability of 18.27%. In summary, the D-MAPPO algorithm had the lowest energy consumption and was stable, the offloading algorithm had the highest energy consumption, the random algorithm had intermediate energy consumption and variability, the Beta-MAPPO and PPO algorithms performed second best, and the local algorithm, while energy-efficient, lacked flexibility.

[0247] iv) Comparison of task offloading success rate optimization

[0248] The comparison of the D-MAPPO algorithm of the present invention and the other five algorithms in optimizing the success rate of common task offloading is shown in the figure. Figure 10As shown. The D-MAPPO algorithm of the present invention quickly improves the offloading success rate in the early stage of training, reaching 97% after about 100 episodes and continuing to rise slowly, and finally tends to stabilize at 100%, showing good convergence and high efficiency. The Beta-MAPPO algorithm has a low success rate in the early stage of training, which increases significantly after about 200 episodes, and finally stabilizes at 98%-99%. The optimization effect is second only to D-MAPPO. The offloading success rate of the PPO algorithm slowly rises from 85%-90% to 93%-95%, which is at a medium level. The local offloading success rate remains at 100%, representing an extreme strategy for local processing. The success rate of the random algorithm fluctuates between 90%-93%, and the performance is unstable. Due to the limited number of cloud servers and satellite receiving tasks, the offloading success rate of the offloading algorithm is very poor. Overall, the D-MAPPO algorithm performs best in optimizing the task offloading success rate, followed by Beta-MAPPO, while the PPO and random algorithms are relatively inferior, and the local and all-off algorithms lack flexibility. This shows that the D-MAPPO and Beta-MAPPO algorithms based on reinforcement learning have significant advantages in optimizing the success rate of task offloading and can effectively improve performance.

[0249] The success rates of unloading urgent tasks, ordinary tasks and all tasks under different algorithms when the reward is the highest are as follows: Figure 10 As shown. Figure 11 It can be seen that the D-MAPPO algorithm of the present invention shows a 100% offloading success rate in all task types, indicating that it can efficiently complete offloading under different task types. The success rate of Beta-MAPPO and local algorithms in emergency tasks and ordinary tasks is also 100%. The performance of the PPO algorithm is slightly worse, but still maintains a high level. In contrast, the success rate of the offloading algorithm in all task types is significantly lower, about 0.7 to 0.75, while the success rate of the random algorithm is about 0.85, which is relatively consistent but lower than the algorithm based on reinforcement learning. This shows that the D-MAPPO algorithm of the present invention has significant advantages in processing different types of tasks and can effectively improve the offloading success rate.

[0250] Example 2

[0251] This embodiment provides an air-ground-space integrated network computing offloading reliability collaborative optimization system, including:

[0252] A task division module is used to divide computing tasks into urgent tasks and common tasks based on the urgency of the computing tasks;

[0253] The uninstallation strategy setting module is used to set the uninstallation strategies corresponding to emergency tasks and ordinary tasks;

[0254] The unloading strategy generation module is configured to generate the unloading strategy of the multiple unmanned aerial vehicles in the space-ground integration network based on the set unloading strategy by using the D-MAPPO algorithm, and decompose the unloading strategy of the multiple unmanned aerial vehicles into the unloading strategy of each unmanned aerial vehicle in the environment.

[0255] The unloading strategy optimization module is configured to optimize the generated unloading strategy based on the pre-established communication model, satellite coverage time model and calculation model and the pre-set reliability mechanism by using the D-MAPPO algorithm, so that the space-ground integration network works according to the optimized unloading strategy.

[0256] The system provided in the embodiment is used to execute the above-mentioned method embodiments, and the specific process and detailed content are referred to the above-mentioned embodiments, which will not be described here.

[0257] Embodiment 3

[0258] The embodiment provides a processing device corresponding to the space-ground integration network computing and unloading reliability collaborative optimization method provided in the embodiment 1. The processing device can be applied to the processing device of a client, such as a mobile phone, a notebook computer, a tablet computer, a desktop computer and the like, to execute the method of the embodiment 1.

[0259] The processing device includes a processor, a memory, a communication interface and a bus. The processor, the memory and the communication interface are connected through the bus to complete the communication among each other. The memory stores a computer program that can run on the processing device. When the processing device runs the computer program, the space-ground integration network computing and unloading reliability collaborative optimization method provided in the embodiment 1 is executed.

[0260] In some implementations, the memory can be a high-speed random access memory (RAM) and can also include a non-volatile memory, such as at least one disk memory.

[0261] In other implementations, the processor can be a central processing unit (CPU), a digital signal processor (DSP) and various types of general-purpose processors, which are not limited here.

[0262] In addition, the logic instructions in the memory described above can be implemented in the form of a software function unit and sold or used as an independent product, and can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or part of the technical solutions can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the method described in various embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

[0263] Those skilled in the art can understand that the structure of the computing device described above is only part of the structure related to the present application scheme, and does not constitute a limitation on the computing device to which the present application scheme is applied. The specific computing device can include more or fewer components, or combine certain components, or have a different component arrangement.

[0264] Embodiment 4

[0265] The embodiment provides a computer program product corresponding to the space-air-ground integrated network computing offloading reliability cooperative optimization method provided in the embodiment 1. The computer program product can include a computer readable storage medium, and the computer readable storage medium has loaded the computer readable program instructions for executing the space-air-ground integrated network computing offloading reliability cooperative optimization method provided in the embodiment 1.

[0266] The computer readable storage medium can be a tangible device that maintains and stores instructions for use by an instruction execution device. The computer readable storage medium can be, for example but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any combination thereof.

[0267] The computer readable storage medium provided in the above embodiment has similar implementation principles and technical effects to the above method embodiments, and details are not described here.

[0268] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart Figure 1 one or more functions specified in the flowchart or multiple flows and / or blocks. Figure 1 one or more functions specified in the flowchart or multiple flows and / or blocks.

[0269] These computer program instructions can also be stored in a computer readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer readable memory produce an article of manufacture including instructions which implement the function specified in the flowchart Figure 1 one or more functions specified in the flowchart or multiple flows and / or blocks. Figure 1 one or more functions specified in the flowchart or multiple flows and / or blocks.

[0270] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart Figure 1 one or more functions specified in the flowchart or multiple flows and / or blocks. Figure 1 one or more functions specified in the flowchart or multiple flows and / or blocks.

[0271] The above-described embodiments are merely intended to illustrate the present application, and the structure, connection manner, and manufacturing process of each component can be changed. Any equivalent changes and improvements made on the basis of the technical solutions of the present application shall not be excluded from the protection scope of the present application.

Claims

1. A collaborative optimization method for reliability of space-ground integrated network computing offloading, characterized in that: include: Based on the urgency of computing tasks, computing tasks are divided into urgent tasks and ordinary tasks; Set up offloading strategies for urgent tasks and common tasks; The D-MAPPO algorithm is used to generate the offloading strategies of multiple UAVs in the air-ground integrated network based on the set offloading strategies, and the offloading strategies of multiple UAVs are decomposed into the offloading strategies of each UAV in the environment. The D-MAPPO algorithm is used to optimize the generated offloading strategy based on the pre-established communication model, satellite coverage time model and calculation model and the pre-set reliability mechanism, so that the integrated air-space-ground network can work according to the optimized offloading strategy.

2. The method for collaborative optimization of space-ground integrated network computing offloading reliability according to claim 1, characterized in that: The offloading strategy for the urgent task is: When the ground sensors of the integrated space-air-ground network generate an emergency task, the emergency task is directly offloaded to the satellite through the connection channel between the ground sensor and the satellite. The satellite processes the emergency task and transmits the calculation results directly back to the ground sensor. The offloading strategy for the common task is: The ground sensors of the integrated air-space-ground network offload common tasks to the drones through the connection channel between the ground sensors and the drones. The drones offload common tasks to satellites and cloud servers for processing according to the determined offloading ratio, and the remaining common tasks remain on the drones for processing.

3. The method for collaborative optimization of space-ground integrated network computing offloading reliability according to claim 1, characterized in that: The probability distribution function of the D-MAPPO algorithm is the Dirichlet probability distribution: in, represents the probability density function of the Dirichlet probability distribution; x J Indicates the uninstall ratio, satisfying α J represents the parameters of the Dirichlet probability distribution; B(α) represents the normalization term of the Beta function; The advantage estimation function of the D-MAPPO algorithm for: Among them, δ t is the TD error, which represents the estimated difference between the reward at the current moment and the value of the next state; γ represents the discount factor; λ represents the GAE smoothing parameter; The joint strategy of the D-MAPPO algorithm updates the objective function L clip (θ) is: Among them, r t (θ) represents the probability ratio; clip represents the clipping operation; ∈ represents the hyperparameter; The value function of the D-MAPPO algorithm is: Among them, L VF (φ) represents the loss function of the value function, which is used to measure the current value function V φ The estimated value and actual cumulative discounted return R t The error between t Represents the environmental observation value at the current moment; R t represents the cumulative discounted return.

4. The method for collaborative optimization of space-ground integrated network computing offloading reliability according to claim 1, characterized in that: The D-MAPPO algorithm is used to optimize the generated offloading strategy based on a pre-established communication model, a satellite coverage time model, a calculation model, and a pre-set reliability mechanism, so that the air-ground integrated network operates according to the optimized offloading strategy, including: Using a pre-established satellite coverage time model and a pre-set reliability mechanism, the coverage time and the maximum uploadable workload of each satellite are calculated. Based on the maximum uploadable workload of the satellite, urgent tasks are offloaded to the satellite in proportion. Using pre-established communication models between drones and ground sensors and satellites, the transmission delay is calculated; Using pre-established satellite computing models, emergency tasks are processed and the transmission delay, transmission energy consumption, computing delay and computing energy consumption of emergency tasks are calculated; Each UAV calculates the maximum remaining uploadable tasks to the satellite based on a pre-set reliability mechanism, then calculates an offloading strategy to offload common tasks in sequence. Based on pre-established communication models between the UAV and ground sensors, between the UAV and cloud servers, between the UAV and ground sensors and satellites, and computational models between the UAV, satellite, and cloud servers, the UAV processes common tasks and calculates their transmission delay, transmission energy consumption, computational delay, and computational energy consumption. After all computing tasks are processed, the energy consumption, delay, and offloading success rate of the entire system model are obtained. Based on the reward function, the reward score of this round of offloading strategy is calculated and transmitted to the D-MAPPO algorithm. The D-MAPPO algorithm calculates the probability ratio of each agent to choose each action under the current unloading strategy; Clipping of probability ratios; The clipped probability ratio is weighted by the advantage estimation function to construct the joint strategy update objective function; Update the gradient of the objective function according to the joint strategy and update the offloading strategy; After each update, the new uninstallation policy is used as the old uninstallation policy to prepare for the next sampling and update.

5. The method for collaborative optimization of reliability of space-ground integrated network computing offloading according to claim 1, characterized in that: The communication model between the UAV and the ground sensor is: Among them, R i Indicates the maximum transmission rate of the ground sensor; R u Indicates the maximum transmission rate of the drone; B i Indicates the bandwidth of the ground sensor; B u represents the bandwidth of the drone; P i and P u Represent the transmission power of ground sensors and UAVs respectively; σ 2 represents the Gaussian noise power; The communication model between the drone and the cloud server is: Among them, R c Indicates the maximum transmission rate of the cloud server; B c Indicates the bandwidth of the cloud server; P c Indicates the transmit power of the cloud server; The communication model between the UAV, ground sensors and satellite is: in, Respectively represent the maximum transmission rate of UAV, ground sensor and satellite under line-of-sight communication; B s represents the available bandwidth of the satellite; G0 represents the fixed antenna gain; P s represents the satellite's transmit power; N0 represents the spectral density of the additive white Gaussian noise.

6. The method for collaborative optimization of space-ground integrated network computing offloading reliability according to claim 1, characterized in that: The satellite coverage time model is: Among them, T S Indicates the satellite coverage time; V S Indicates the flight speed of the satellite; L S Indicates the coverage arc length of the satellite.

7. The method for collaborative optimization of space-ground integrated network computing offloading reliability according to claim 1, characterized in that: The calculation model of the UAV is: in, represents the transmission time of the UAV u collecting computing tasks; b>1 represents the transmission overhead coefficient; N u represents the set of common tasks collected by the drone; Indicates the data volume of the xth common task, R i,u represents the maximum transmission rate from ground sensor i to UAV u; Indicates the drone u collects the computing task Transmission energy consumption; P i,u represents the transmission power from ground sensor i to UAV u; μ u Represents the proportion of UAV u processing tasks locally, 0≤μ u ≤1; represents the total local computation delay of UAV u; represents the task complexity of the xth common task; f u represents the computing resources of UAV u; E represents the total local computing energy consumption of UAV u; τ represents the energy coefficient; Indicates the time it takes for the UAV u to transmit the result data back to the ground sensor after the calculation is completed; Indicates the total amount of drone u result data; represents the energy consumption of the data transmitted by UAV u; The computing model of the cloud server is: in, and They represent the transmission time and energy consumption of the cloud server receiving the UAV mission; R u,c represents the maximum transmission rate from the u-th UAV to the cloud server; P u,c represents the transmission power of the u-th UAV to the cloud server; and Respectively represent the computing delay and computing energy consumption of the cloud server; f C Represents the computing resources of the cloud server; and They represent the return delay of the calculation results of the computing task to the cloud server of the drone. and backhaul energy consumption; Indicates the total amount of server result data; The calculation model of the satellite is: in, and They represent the transmission time and energy consumption of satellite s receiving the emergency mission and the unloading mission of the UAV respectively; represents the transmission time of the urgent mission in satellite s; represents the transmission time of ordinary missions in satellite s; M ur Indicates the total number of urgent tasks; Indicates the workload of the xth urgent task; R i,s represents the maximum transmission rate from ground sensor i to satellite s; R u,s represents the maximum transmission rate of UAV u to satellite s; P i,s represents the transmission power from ground sensor i to satellite s; P u,s represents the transmission power from UAV u to satellite s; and They represent the computational delay and computational energy consumption of satellite s in processing emergency tasks and ordinary tasks respectively; represents the task complexity of the xth urgent task; f s represents the computing resources of satellite s; and They represent the return delay and return energy consumption of satellite s when satellite s directly transmits the calculation results of the computing task to the ground sensor.

8. A space-ground integrated network computing offloading reliability collaborative optimization system, characterized by: include: A task division module is used to divide computing tasks into urgent tasks and common tasks based on the urgency of the computing tasks; The uninstallation strategy setting module is used to set the uninstallation strategies corresponding to emergency tasks and ordinary tasks; An offloading strategy generation module is used to generate offloading strategies for multiple drones in an air-ground integrated network based on the set offloading strategies using the D-MAPPO algorithm, and decompose the offloading strategies of multiple drones into the offloading strategies of each drone in the environment; The offloading strategy optimization module is used to optimize the generated offloading strategy using the D-MAPPO algorithm based on the pre-established communication model, satellite coverage time model and calculation model and the pre-set reliability mechanism, so that the integrated air-space-ground network can operate according to the optimized offloading strategy.

9. A processing device, characterized in that: It includes computer program instructions, wherein when the computer program instructions are executed by a processing device, they are used to implement the steps corresponding to the collaborative optimization method for reliability of integrated air-space-ground network computing offloading according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer program instructions, wherein the computer program instructions, when executed by the processor, are used to implement the steps corresponding to the collaborative optimization method for reliability of integrated air-space-ground network computing offloading according to any one of claims 1 to 7.