Satellite-assisted multi-unmanned aerial vehicle calculation unloading method and device and program product

By acquiring a sample set of interaction states between UAVs and satellite mobile edge computing networks, and using a near-end policy optimization algorithm to update the UAV computation offloading strategy, the problem of global optimization and dynamic changes in cross-domain resource collaboration of the SatMEC network is solved. This enables low-energy, low-latency computation offloading of UAVs and improves the overall network performance.

CN121984571APending Publication Date: 2026-05-05DONGGUAN UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
DONGGUAN UNIV OF TECH
Filing Date
2026-02-14
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing SatMEC networks struggle to achieve global optimization in cross-domain resource collaboration. Traditional resource allocation schemes lack the flexibility to cope with dynamic network changes, leading to increased equipment energy consumption and computational latency, especially in areas with insufficient coverage such as drone operation airspace and high-altitude logistics channels.

Method used

By acquiring a sample set of interaction states between the UAV and the satellite mobile edge computing network, the computation offloading strategy of the UAV is updated using a near-end policy optimization algorithm until the network reaches a stable state, achieving optimal computation offloading. The total loss function is optimized by combining gradient pruning and backpropagation to ensure the stability and accuracy of the strategy.

Benefits of technology

It achieves low power consumption and low computational latency for UAVs in highly dynamic network environments, reduces conflicts between UAVs, improves beam load balancing and computing resource utilization, and ensures the overall stability of network performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121984571A_ABST
    Figure CN121984571A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computing resource allocation, and discloses a satellite-assisted multi-unmanned aerial vehicle computing unloading method and device and a program product, and the method comprises the steps: obtaining a state sample set generated by interaction between each unmanned aerial vehicle and a satellite mobile edge computing network and stored in an experience playback buffer area; judging whether the number of state samples in each state sample set is greater than or equal to a preset batch size; and when the number of the state samples in the state sample set is greater than or equal to a preset batch size, updating the initial calculation unloading strategy of each unmanned aerial vehicle by using a near-end strategy optimization algorithm until the satellite mobile edge computing network is in a stable state, and obtaining an optimal calculation unloading strategy of the unmanned aerial vehicle, according to the invention, low energy consumption and low calculation delay of the unmanned aerial vehicles are ensured in a high-dynamic network environment, and conflicts among the unmanned aerial vehicles when the unmanned aerial vehicles acquire calculation resources are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computing resource allocation technology, specifically to a satellite-assisted multi-UAV computing offloading method, apparatus, and program product. Background Technology

[0002] Satellite Mobile Edge Computing (SatMEC) networks, through the collaboration of low-Earth orbit (LEO) satellites and ground devices, provide wide-area coverage and low-latency services for scenarios such as the Internet of Things (IoT) and real-time communication. However, with the increase in IoT devices, the demand for dynamic computing resources places a heavy burden on the resources allocated to ground base stations. Furthermore, the coverage of ground base stations is limited, making it difficult to cover low-altitude areas (such as drone operating airspace and high-altitude logistics corridors). Given the rapid development of drones in recent years, low-altitude areas will be a key development focus. Against this backdrop, uneven distribution of computing resources leads to increased device energy consumption.

[0003] Currently, research on resource allocation in SatMEC networks is progressing in multiple directions, aiming to ensure efficient cross-domain collaboration and rational resource allocation in SatMEC networks, including:

[0004] 1. Emerging Technology Integration: With continuous technological advancements, emerging technologies such as artificial intelligence and edge computing are gradually being integrated into the field of satellite network security. Artificial intelligence and machine learning algorithms excel in SatMEC network resource allocation. Through reinforcement learning or evolutionary game theory, they can address the high dynamism of the network, or combine reinforcement learning with predictive scheduling algorithms to deploy computing resources in high-demand areas in advance. The introduction of edge computing technology enables data to be processed at edge nodes close to data sources or users, reducing data transmission latency and improving network response speed.

[0005] 2. Optimization Based on Digital Twins: Digital twins, by constructing a virtual mirror synchronized in real time with the physical network, provide a virtual environment for resource allocation, making them a powerful tool for global optimization and algorithm verification. Resource allocation strategies are first tested, evaluated, and optimized within the digital twin before the optimal strategy is deployed to the physical network for execution. This solves the problem of physical networks being unable to conduct high-risk or large-scale algorithm experiments.

[0006] However, existing technologies have shortcomings: Firstly, cross-domain resource coordination is insufficient. The SatMEC network involves the coordination of heterogeneous resources such as low-Earth orbit satellites, ground base stations, and drones, but existing methods struggle to achieve global optimization. Resource allocation in multi-conflict domain scenarios needs to consider the dynamic nature of overall resources and data transmission rates, and traditional game theory models have limited ability to coordinate optimization across multiple conflict domains. Secondly, traditional resource allocation schemes lack flexibility in responding to dynamic network changes. Service requests in satellite networks are random and fluctuating, while existing static resource allocation strategies cannot be flexibly adjusted according to real-time network conditions, making it difficult to improve resource utilization efficiency. Furthermore, in complex satellite network environments, existing methods lack sufficient balance and robustness, and the stability of game equilibrium is limited by network topology changes and external interference. For example, traditional resource allocation methods such as Q-learning and optimization algorithms are infeasible in high-dimensional state and action spaces, or their decision-making speed is slow. Summary of the Invention

[0007] This invention provides a satellite-assisted multi-UAV computational offloading method, apparatus, and program product to solve the problem of highly dynamic cross-domain network resource allocation optimization in existing SatMEC network resources.

[0008] In a first aspect, the present invention provides a satellite-assisted multi-UAV computational offloading method, the method comprising: Obtain the state sample set generated by the interaction between each UAV and the satellite mobile edge computing network stored in the experience replay buffer; determine whether the number of state samples in each state sample set is greater than or equal to the preset batch size; when the number of state samples in the state sample set is greater than or equal to the preset batch size, update the initial computation offloading strategy of each UAV using the near-end policy optimization algorithm until the satellite mobile edge computing network is in a stable state, and obtain the optimal computation offloading strategy of the UAV.

[0009] The satellite-assisted multi-UAV computational offloading method provided by this invention acquires the state sample set generated by the interaction between each UAV and the satellite mobile edge computing network, stored in the experience replay buffer. This provides independent interaction data that aligns with the dynamic characteristics of the satellite mobile edge computing network for training the near-end policy optimization algorithm of each UAV, thus adapting to the distributed decision-making needs of multiple UAVs and helping to ensure the authenticity and relevance of single-UAV policy optimization. Furthermore, by determining whether the number of state samples in each state sample set is greater than or equal to the preset batch size, policy update deviations caused by insufficient sample size for a single UAV are avoided, ensuring sample support for the training of the near-end policy optimization algorithm of each UAV and improving the accuracy and stability of independent policy optimization for each UAV. Furthermore, by updating the initial computational offloading strategy of each UAV using a near-end policy optimization algorithm when the number of state samples in the state sample set is greater than or equal to a preset batch size, until the satellite mobile edge computing network is in a stable state, the optimal computational offloading strategy of the UAV is obtained. This achieves distributed autonomous optimization of computational offloading strategies for multiple UAVs, adapting to the highly dynamic cross-domain characteristics of the satellite mobile edge computing network. This enables joint optimization of the entire network with low energy consumption and low latency, maximizing beam load balancing and computing resource utilization, while ensuring the overall performance stability of the satellite mobile edge computing network. Therefore, by implementing this invention, low energy consumption and low computational latency of UAVs are guaranteed in a highly dynamic network environment, and conflicts between UAVs when acquiring computing resources are reduced.

[0010] In one optional implementation, when the number of state samples in the state sample set is greater than or equal to a preset batch size, the initial computational offloading strategy for each UAV is updated using a near-end policy optimization algorithm until the satellite mobile edge computing network is in a stable state, thus obtaining the optimal computational offloading strategy for the UAV, including: Using the initial computational offloading strategy of each UAV as the old strategy, a new real-time computational offloading strategy is obtained for each UAV. The initial network weight parameters of the real-time computational offloading strategy are consistent with those of the initial computational offloading strategy. Based on the state sample set, initial computational offloading strategy, and real-time computational offloading strategy of each UAV, a total loss function for each UAV is constructed through a near-end strategy optimization algorithm. The total loss function of each UAV is optimized using gradient pruning and backpropagation algorithms, and the initial network weight parameters of the real-time computational offloading strategy are iteratively updated. The updated real-time computational offloading strategy is used as the new initial computational offloading strategy, and the steps of constructing the total loss function for each UAV are returned. This process is repeated until the satellite mobile edge computing network is in a stable state, thus obtaining the optimal computational offloading strategy for each UAV.

[0011] The satellite-assisted multi-UAV computational offloading method provided by this invention uses the initial policy as the old policy and obtains a real-time computational offloading policy with consistent initial network weights, ensuring the continuity of the old and new policies, avoiding training fluctuations caused by policy mutations, and improving the stability of policy updates. Furthermore, based on the sample set and the old and new policies, a total loss function is constructed through a near-end policy optimization algorithm, which can integrate policy performance differences and evaluation errors, providing a clear objective for policy optimization and ensuring that the update direction aligns with optimal performance requirements. Furthermore, gradient pruning and backpropagation are used to optimize the total loss function and update the real-time policy network weights, achieving iterative policy optimization, limiting the update magnitude to avoid gradient explosion, and improving training stability and sample utilization efficiency. Furthermore, the updated policy is used as the new initial policy, and the utility value calculation step is iterated repeatedly until the system stabilizes, promoting continuous policy optimization, adapting to dynamic network changes, and ultimately achieving a dynamic balance between beam load and device task offloading, maximizing beam utilization.

[0012] In one optional implementation, based on the state sample set, initial computational offloading strategy, and real-time computational offloading strategy of each UAV, a total loss function for each UAV is constructed after processing by a near-end policy optimization algorithm, including: Each state sample set is input into the value network of the near-end policy optimization algorithm for computation, resulting in multiple advantage values. Each advantage value represents the degree of superiority or inferiority of the current action state of each UAV relative to the average value of the computational offloading policy. The probability policy ratio for the same action is calculated for each initial computational offloading policy and each new real-time computational offloading policy. Based on the multiple advantage values, a preset pruning function is used to restrict each probability policy ratio within a preset interval, and a target pruning function for each UAV is constructed. The value loss function of the critic network in the near-end policy optimization algorithm is obtained. Based on the target pruning function and value loss function for each UAV, a total loss function for each UAV is constructed.

[0013] The satellite-assisted multi-UAV computational offloading method provided by this invention quantifies the relative merits of each UAV's individual actions by inputting a sample set into a value network to calculate advantage values. This provides precise guidance for the independent policy update direction of each UAV, optimizes the training stability of single-UAV algorithms, and enhances the rationality of individual policy selection for multiple UAVs. Furthermore, by calculating the probability-policy ratio of the old and new policies on the same actions, the differences in action selection preferences between the old and new policies can be clarified, helping to provide a basis for controlling the subsequent policy update magnitude. Furthermore, by using a pruning function to limit the probability-policy ratio and constructing a target pruning function, training oscillations caused by excessively large policy update magnitudes for each UAV are avoided, enhancing the stability of individual policy updates for multiple UAVs, improving the sample utilization efficiency of each UAV, and ensuring the consistency of overall network policy optimization. Furthermore, by obtaining the value loss function of the critic network, the deviation between the predicted and actual policy value values ​​of each UAV is quantified, optimizing the evaluation of individual policy values ​​for multiple UAVs, thereby helping to improve the accuracy of independent policy decisions for each UAV. Furthermore, by constructing a total loss function based on the target pruning function and the value loss function, the joint optimization of policy loss and evaluation loss is achieved, ensuring that policy updates are both efficient and stable.

[0014] In one alternative implementation, each set of state samples is input into the value network of the near-end policy optimization algorithm for computation, resulting in multiple advantage values, including: Based on the location of each UAV and the target beam selected by each UAV, the policy utility value of the initial computation offloading strategy of each UAV is calculated using a weighted utility function; based on the load value of the target beam selected by each UAV, the beam load reward value of each target beam is calculated; based on the distance between each UAV and each target beam, multiple flight distance penalty values ​​are calculated; based on the latency of each UAV's current computation offloading strategy and the latency of each UAV's local computation strategy, multiple low-latency reward values ​​are calculated; based on multiple policy utility values, multiple beam load reward values, multiple flight distance penalty values, and multiple low-latency reward values, multiple instant reward values ​​of the UAV are determined; based on multiple state sample sets and multiple instant reward values, multiple advantage values ​​are obtained through value network calculation in the near-end policy optimization algorithm.

[0015] The satellite-assisted multi-UAV computational offloading method provided by this invention calculates the policy utility value of each UAV's initial computational offloading strategy using a weighted utility function based on the location and target beam selected by each UAV. This quantifies the comprehensive performance of each UAV's execution strategy in real time, providing a core foundation for multi-UAV individual reward calculation and enabling joint independent evaluation of the energy consumption, latency, and resource costs of each UAV's strategy. Furthermore, by calculating the beam load reward value for each target beam, each UAV can be guided to independently prioritize low-load beams, achieving a globally balanced distribution of satellite beam load, reducing inter-beam interference, and improving the overall utilization efficiency of satellite beams. Furthermore, by calculating multiple flight distance penalty values, each UAV can avoid blindly pursuing low latency, resulting in excessive flight distance, effectively controlling the flight energy consumption of multiple UAVs, reducing energy redundancy among UAVs, and lowering the overall network energy consumption. Furthermore, by calculating multiple low-latency reward values, each UAV can be incentivized to independently select low-latency offloading strategies, promoting the real-time performance of multi-UAV individual task processing, thereby reducing the overall network computational latency. Furthermore, by determining multiple instantaneous reward values ​​for each drone, a multi-dimensional reward evaluation system is constructed for each drone, enabling precise feedback on the quality of each drone's actions and providing a reliable independent evaluation benchmark for calculating the advantage value of each drone. Moreover, converting the absolute reward of each drone into a relative advantage value eliminates the interference of average strategy performance on individual drones, thus providing more accurate directional guidance for the independent strategy updates of each drone.

[0016] In one optional implementation, based on the location of each UAV and the target beam selected by each UAV, the policy utility value of the initial computational offloading strategy for each UAV is calculated using a weighted utility function, including: Obtain the computational resource cost generated by the satellite mobile edge computing network; calculate the total latency and total energy consumption of each UAV executing the initial computational offloading strategy based on the decision type of each initial computational offloading strategy and the positional relationship between each UAV and each target beam; calculate the strategy utility value of each UAV using a weighted utility function based on the computational resource cost, the total latency and total energy consumption of each UAV.

[0017] The satellite-assisted multi-UAV computational offloading method provided by this invention acquires the computational resource cost generated by the satellite mobile edge computing network and incorporates satellite resource consumption into the strategy evaluation system of each UAV. This achieves a comprehensive global consideration of integrated space-ground network resources and improves the global rationality of resource allocation under distributed decision-making by multiple UAVs. Furthermore, based on the decision type, location, and target beam position relationship, the total latency and total energy consumption are calculated, accurately quantifying the core performance indicators for each UAV under different decision types. This provides crucial data support for the utility evaluation of individual UAVs and can adapt to different scenario requirements for local computation and beam offloading by each UAV. Furthermore, based on computational resource cost, total latency, and total energy consumption, a weighted utility function is used to calculate the strategy utility value, achieving a multi-dimensional weighted comprehensive evaluation of each UAV. This accurately reflects the actual execution effect of each UAV's strategy and provides a core basis for subsequent calculation of rewards and advantage values ​​for individual UAVs.

[0018] In one alternative implementation, the method further includes: Obtain the initial upper limit value of the load of a single beam and the real-time load value of each beam in the satellite mobile edge computing network; determine whether the real-time load value of all beams in the satellite mobile edge computing network is equal to the initial upper limit value; when the real-time load value of all beams is equal to the initial upper limit value, determine the optimal computing offload strategy as the local computing offload strategy.

[0019] The satellite-assisted multi-UAV computational offloading method provided by this invention obtains the initial upper limit of the load of a single beam and the real-time load value of each beam in the satellite mobile edge computing network. This provides a unified benchmark and real-time data for global beam load judgment, facilitating accurate global perception of satellite beam load status and adapting to the dynamic resource change characteristics of the satellite mobile edge computing network. Furthermore, by determining whether the real-time load value of all beams in the satellite mobile edge computing network is equal to the initial upper limit, the global saturation state of beam resources in the network can be accurately identified, providing a clear and unified basis for multi-UAV collective strategy switching and avoiding a decline in overall network service quality due to beam overload. Moreover, when the load of all beams is equal to the initial upper limit, the optimal computational offloading strategy is determined to be the local computational offloading strategy. This avoids invalid access to saturated beams by multiple UAVs, reduces inter-beam interference and global data transmission congestion, and thus controls the overall offloading energy consumption and latency of multiple UAVs, ensuring the normal and orderly execution of UAV tasks across the network.

[0020] In a second aspect, the present invention provides a satellite-assisted multi-UAV computational offloading device, the device comprising: The acquisition module is used to acquire the state sample set generated by the interaction between multiple UAVs and the satellite mobile edge computing network stored in the experience replay buffer; the judgment module is used to determine whether the number of state samples in the state sample set is greater than or equal to the preset batch size; the update module is used to update the initial computation offloading strategy of each UAV using the near-end policy optimization algorithm when the number of state samples in the state sample set is greater than or equal to the preset batch size, until the satellite mobile edge computing network is in a stable state, and obtain the optimal computation offloading strategy of the UAV.

[0021] Thirdly, the present invention provides an electronic device comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the computer instructions to perform the satellite-assisted multi-UAV computational offloading method described in the first aspect or any corresponding embodiment thereof.

[0022] Fourthly, the present invention provides a computer-readable storage medium storing computer instructions for causing a computer to execute the satellite-assisted multi-UAV computational offloading method of the first aspect or any corresponding embodiment thereof.

[0023] Fifthly, the present invention provides a computer program product, including computer instructions for causing a computer to execute the satellite-assisted multi-UAV computational offloading method of the first aspect or any corresponding embodiment thereof. Attached Figure Description

[0024] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0025] Figure 1 This is a schematic diagram of an application scenario according to an embodiment of the present invention; Figure 2 This is a schematic flowchart of a satellite-assisted multi-UAV computational unloading method according to an embodiment of the present invention; Figure 3 This is a schematic diagram of the dynamic strategy selection process of the satellite mobile edge computing service network according to an embodiment of the present invention; Figure 4 This is a schematic diagram of a strategy selection process based on PPO according to an embodiment of the present invention; Figure 5 This is a comparison chart of average latency according to an embodiment of the present invention; Figure 6 This is a comparison chart of average energy consumption according to embodiments of the present invention; Figure 7 This is a beam load comparison diagram according to an embodiment of the present invention; Figure 8 This is a structural block diagram of a satellite-assisted multi-UAV computational unloading device according to an embodiment of the present invention; Figure 9 This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present invention. Detailed Implementation

[0026] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0027] It is understood that before using the technical solutions disclosed in the various embodiments of the present invention, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in the present invention and their authorization should be obtained in accordance with relevant laws and regulations through appropriate means.

[0028] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0029] As an optional application scenario of this invention, the specific application environment architecture or specific hardware architecture on which the satellite-assisted multi-UAV computational offloading method depends is described herein. For example... Figure 1 As shown, the architecture system may include at least one terminal device and at least one server. Figure 1 The system is illustrated in the example, which includes a computer 101, a mobile terminal 102, and a server 103, and the terminal devices such as the computer 101 and the mobile terminal 102 are connected to the server 103 through a network 110.

[0030] Specifically, the terminal device can be a smartphone, tablet, laptop, PDA, desktop computer, game console, smart TV, smart wearable device, in-vehicle terminal, VR (Virtual Reality) device, AR (Augmented Reality) device, etc. Server 103 can be a standalone physical server, a server cluster, a distributed system, or a cloud server providing cloud services. Network 110 can be a wired or wireless network, examples of which include, but are not limited to, the Internet, corporate intranet, local area network, wide area network, mobile communication network, and combinations thereof.

[0031] Current allocation of computing resources for the SatMEC network mainly focuses on the following aspects, but there are still shortcomings: 1. Beam tracking-based computational resource allocation strategy.

[0032] Existing solutions allocate resources by directly covering IoT devices with beams from low-Earth orbit (LEO) satellites or ground base stations. For example, LEO satellites utilize point beams to cover devices within their coverage area. While this approach can improve computing resources within the beam range, it suffers from the following problems: (1) Large inter-beam interference: For a region, the computing resources of a single beam may be insufficient and the data transmission rate may be low. Multiple adjacent beams may cause inter-beam interference, reducing beam utilization efficiency.

[0033] (2) High energy consumption: Ground base stations or low-orbit satellites calculate and allocate energy in a unified manner. For situations with large and dynamic demand, this leads to increased energy consumption of base stations and low-orbit satellites.

[0034] 2. Traditional game theory optimization methods.

[0035] Existing research attempts to model resource allocation problems using game theory, but most of them are based on the following assumptions: (1) Insufficient cross-domain resource coordination mechanism: The SatMEC network involves the coordination of heterogeneous resources such as low-orbit satellites, ground base stations and drones, but existing game theory models often fail to achieve global optimization; (2) Imbalance between theory and actual needs: Game theory models usually assume that participants are completely rational and information is symmetrical, but in the SatMEC network, factors such as limited resources and environmental interference may cause decisions to deviate from expectations.

[0036] This invention provides a satellite-assisted multi-UAV computation offloading method, which ensures low power consumption and low computation latency of UAVs in a highly dynamic network environment, and reduces conflicts between UAVs when acquiring computing resources.

[0037] According to an embodiment of the present invention, a satellite-assisted multi-UAV computational offloading method is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0038] This embodiment provides a satellite-assisted multi-UAV computational offloading method, which can be used on the aforementioned mobile terminals, such as mobile phones and tablets. Figure 2 This is a flowchart of a satellite-assisted multi-UAV computational offloading method according to an embodiment of the present invention, such as... Figure 2 As shown, the process includes the following steps: Step S201: Obtain the state sample set generated by each UAV's interaction with the satellite mobile edge computing network, stored in the experience replay buffer.

[0039] In one optional embodiment, the Satellite Mobile Edge Computing (SatMEC) network represents a space-ground integrated collaborative computing network. Its core is to deploy edge computing nodes on low-Earth orbit (LEO) satellites, placing computing resources within the satellite beam coverage area close to terminal devices such as drones, replacing traditional ground base stations, and providing computing services for low-altitude operations and wide-area coverage scenarios.

[0040] Furthermore, the Satellite Mobile Edge Computing (SatMEC) network consists of low-Earth orbit satellites (providing edge computing resources and beam communication coverage) and drones (mission requesters), eliminating the need for ground base stations. Moreover, the satellites cover specific areas with point beams, receiving computation offloading requests from drones and completing task computations at the satellite edge nodes, reducing data transmission distance and latency. Furthermore, the satellite beam coverage area dynamically changes with satellite orbital motion, and the drone's location, mission requirements, and beam load all exhibit randomness and fluctuation.

[0041] In one alternative embodiment, a set of state samples containing complete decision-related information generated by each UAV during its independent interaction with the satellite mobile edge computing network can be extracted from the experience replay buffer built for the Proximal Policy Optimization (PPO) algorithm. Each UAV has an independent set of state samples, which adapts to the architecture requirements of distributed autonomous decision-making for multiple UAVs.

[0042] Among them, the Proximal Policy Optimization (PPO) algorithm is a reinforcement learning algorithm that introduces a policy update pruning mechanism to ensure that each parameter update is within a safe range, thus avoiding a sharp drop in model performance due to excessive update magnitude.

[0043] In one optional embodiment, the drone has completed multiple rounds of interaction with the SatMEC network environment as an intelligent agent. In each round of counting and unloading decision, each drone generates raw interaction data including its own state, actions, rewards, and next state.

[0044] Furthermore, the aforementioned states generated by each drone ,action ,award Next state The original interaction data of the quadruples will be stored in real time and independently in the experience replay buffer. The buffer allocates an independent data storage area for each drone to ensure that the state sample set of each drone is not confused during subsequent extraction, which is suitable for the needs of distributed training of multiple drones.

[0045] Furthermore, for the independent data stored in the experience replay buffer, batch extraction is performed on a per-drone basis to form a state sample set for each drone, and each sample in the state sample set contains complete decision-related information about the interaction between the drone and the SatMEC network.

[0046] For example, drones interact with the environment, and each drone gets a turn. The state of the drone at that time The state is represented by the following relation (1): (1) In the formula: Indicates the current round; Indicates a round drones The initial coordinates; Indicates a round drones The utility; Indicates a round drones Decision-making choices; This indicates all drones in the current round. A collection of states at time.

[0047] Furthermore, PPO, through interaction with the environment, uses its own policy network to collect samples of the state, actions, rewards, and next state of individual drones. This indicates that a large number of samples can be acquired by allowing the agent to run multiple rounds in the environment, forming a corresponding set of state samples. Among these...? Step S202: Determine whether the number of state samples in each state sample set is greater than or equal to the preset batch size.

[0048] In an optional embodiment, the preset batch size can be set by combining the number of beams of the SatMEC network, the size of the UAV mission, and the hyperparameters of the PPO algorithm (such as learning rate and pruning factor).

[0049] Furthermore, using the preset batch size as a unified judgment criterion, the number of samples in the independent state sample set corresponding to each UAV in the experience playback buffer is counted, and it is determined whether the actual number of samples in each sample set reaches or exceeds the preset batch size.

[0050] Furthermore, if the actual number of state samples of a certain UAV is greater than or equal to the preset batch size, it is determined that the policy update condition is met, the UAV is marked, and the subsequent PPO algorithm policy update step for the UAV is triggered.

[0051] Furthermore, if the actual number of state samples of a certain UAV is less than the preset batch size, it is determined that the policy update condition is not met. At this time, no policy update operation is performed on the UAV. Then, the experience replay buffer continues to collect new samples generated by the interaction between the UAV and the network until the number of samples reaches the preset batch size.

[0052] Step S203: When the number of state samples in the state sample set is greater than or equal to the preset batch size, the initial computational offloading strategy of each UAV is updated using the near-end strategy optimization algorithm until the satellite mobile edge computing network is in a stable state, and the optimal computational offloading strategy of the UAV is obtained.

[0053] In one optional embodiment, the initial computation offloading strategy refers to the task processing scheme selected by the UAV based on its initial position after system initialization, which includes only two types of decisions: one is local computation, that is, without relying on satellite beams, the task is completed directly through its own processor; the other is beam offloading, that is, flying to the selected target satellite beam and completing the task computation through the satellite edge node.

[0054] In one optional embodiment, the satellite mobile edge computing network being in a stable state means that the UAV's decision-making strategy is no longer dynamically adjusted, and the beam load of the satellite mobile edge computing network and the average utility of the system are dynamically balanced. At this time, the total latency, total energy consumption, and satellite resource utilization of the UAV in performing the mission are all in an optimal and stable state.

[0055] In one optional embodiment, the optimal computation offloading strategy represents the task processing scheme adapted to the current network environment after multiple rounds of iterative optimization by the PPO algorithm. It can achieve multi-objective joint optimization with low latency, low energy consumption, and low satellite resource cost, while ensuring satellite beam load balance. It is the optimal decision choice (local computation or specific beam offloading) under the stable state of the satellite mobile edge computing network.

[0056] In one optional embodiment, provided that the number of state samples of a single UAV in the experience replay buffer meets the batch training requirements of the PPO algorithm, the initial computational offloading strategy of each UAV is taken as the optimization starting point. The near-end policy optimization (PPO) algorithm is executed independently for each UAV to perform iterative policy updates. This update process continues until the satellite mobile edge computing network reaches a globally stable equilibrium state and then stops. At this point, the computational offloading strategy of each UAV after iterative convergence is the optimal computational offloading strategy that adapts to the dynamic characteristics of the network and its own operational needs.

[0057] The satellite-assisted multi-UAV computation offloading method provided in this embodiment acquires the state sample set generated by the interaction between each UAV and the satellite mobile edge computing network, stored in the experience replay buffer. This provides independent interaction data that aligns with the dynamic characteristics of the satellite mobile edge computing network for training the near-end policy optimization algorithm of each UAV, thus adapting to the distributed decision-making needs of multiple UAVs and helping to ensure the authenticity and relevance of single-UAV policy optimization. Furthermore, by determining whether the number of state samples in each state sample set is greater than or equal to the preset batch size, policy update deviations caused by insufficient sample size for a single UAV are avoided, ensuring sample support for the training of the near-end policy optimization algorithm of each UAV and improving the accuracy and stability of independent policy optimization for each UAV. Furthermore, by updating the initial computational offloading strategy of each UAV using a near-end policy optimization algorithm when the number of state samples in the state sample set is greater than or equal to a preset batch size, until the satellite mobile edge computing network is in a stable state, the optimal computational offloading strategy of the UAV is obtained. This achieves distributed autonomous optimization of computational offloading strategies for multiple UAVs, adapting to the highly dynamic cross-domain characteristics of the satellite mobile edge computing network. This enables joint optimization of the entire network with low energy consumption and low latency, maximizing beam load balancing and computing resource utilization, while ensuring the overall performance stability of the satellite mobile edge computing network. Therefore, by implementing this invention, low energy consumption and low computational latency of UAVs are guaranteed in a highly dynamic network environment, and conflicts between UAVs when acquiring computing resources are reduced.

[0058] In some optional implementations, step S203 above includes: Step S2031: Using the initial computational offloading strategy of each UAV as the old strategy, obtain the new real-time computational offloading strategy for each UAV.

[0059] In one optional embodiment, the initial network weight parameters for real-time calculation of the offloading strategy are consistent with the initial network weight parameters for calculating the offloading strategy.

[0060] For example, the policy update employs a fixed dual-network architecture weight copying mechanism: for two network instances, the initial policy network and the old policy network, before each policy update, the weight parameter tensor of the policy network is copied to the old policy network, keeping the network architecture unchanged; the policy gradient is calculated based on the parameter synchronization mechanism (i.e., saving the old policy network parameters), and the policy network parameters are optimized and updated; the weight copying and optimization update are repeatedly executed to avoid frequent creation and destruction of network objects. That is, in this embodiment, the initial network weight parameters for real-time calculation of the unloading policy are consistent with the network weight parameters for the initial calculation of the unloading policy through the above magnitude mechanism.

[0061] In an optional embodiment, the current initial compute offloading strategy to be optimized is defined as the old strategy, while a new real-time compute offloading strategy is initialized. The initial weight parameters of its policy network are exactly the same as those of the old policy.

[0062] Furthermore, using the initial policy as the old policy ensures the continuity between the old and new policies, avoiding training fluctuations caused by policy mutations. Moreover, setting the initial weights of the new policy to be consistent with the old policy allows for fine-tuning and optimization based on the old policy, thereby improving the stability and convergence speed of policy updates and adapting to highly dynamic network environments.

[0063] Step S2032: Based on the state sample set, initial calculation unloading strategy, and real-time calculation unloading strategy of each UAV, the total loss function of each UAV is constructed after processing by the near-end policy optimization algorithm.

[0064] In one optional embodiment, the total loss function represents the core objective function used to optimize the policy in the PPO algorithm. It integrates the policy loss (output of the objective pruning function), the value loss (calculated by the commentator network), and the entropy regularization term. By minimizing this function, stable iterative optimization of the policy can be achieved.

[0065] Specifically, step S2032 includes: Step a1: Input each state sample set into the value network of the near-end policy optimization algorithm for calculation to obtain multiple advantage values.

[0066] In one optional embodiment, each advantage value represents the degree to which the current operational state of each UAV is superior or inferior to the average value of the computed offloading strategy. Furthermore, the average value of the computed offloading strategy is the average performance level of all possible computed offloading strategies (local computed or beam-based offloading) in the same network environment. It is calculated by statistically analyzing the strategy utility values ​​of all UAVs in the network via satellite and serves as a benchmark for measuring the superiority or inferiority of a single strategy.

[0067] In one alternative embodiment, the value network is one of the core components of the PPO algorithm, used to estimate the long-term expected return (state value) of the UAV in a specific state, providing a basis for advantage value calculation and assisting in judging the superiority or inferiority of the current action relative to the average strategy.

[0068] In one alternative embodiment, the state sample set is input into the value network of the PPO algorithm to obtain the value estimates of the current state and the next state, and then combined with the immediate reward, the advantage value can be calculated.

[0069] For example, setting the initial advantage value Then, the first state sample set Input a value network, output a value estimate for the current state. ; the second state sample set Input a value network, output the value estimate of the next state. .

[0070] Furthermore, the dominance value can be calculated using the time-difference formula. The following relation (2) is shown: (2) In the formula: and Indicates hyperparameters; This represents the discount factor; a value close to 1 indicates consideration of long-term returns, while a value close to 0 indicates consideration of current returns. The parameter is used to balance variance and bias. When it is close to 1, the bias is low but the variance is large. When it is close to 0, the variance is small and the bias is high. This indicates the end of the cycle; 1 indicates the end of the cycle, and 0 indicates continuation. and All observation tensors are obtained by forward propagating them through a policy network; This represents the immediate reward at the current time step.

[0071] In some alternative implementations, step a1 above includes: Step a11: Based on the location of each UAV and the target beam selected by each UAV, calculate the policy utility value of the initial computational offloading strategy for each UAV using a weighted utility function.

[0072] In one optional embodiment, the weighted utility function represents a multi-objective evaluation function that comprehensively considers the total latency, total energy consumption, and satellite computing resource costs of the UAV's mission. It balances the proportions of these three factors in the strategy evaluation through preset weighting factors, using quantitative values ​​to reflect the quality of the strategy. Furthermore, a larger value indicates a better strategy.

[0073] In an optional embodiment, the strategy utility value represents the quantitative result calculated by the weighted utility function, which can comprehensively reflect the overall performance of the initial computation offloading strategy in terms of low latency, low energy consumption, and low satellite resource cost.

[0074] In one optional embodiment, based on the initial position of the UAV and the selected target beam, the overall performance quantification value of the initial computation offloading strategy, i.e. the strategy utility value, can be calculated by comprehensively considering the total delay, total energy consumption and satellite computing resource cost of the task execution through a preset weighted utility function.

[0075] Step a11 further includes: Step a111: Obtain the computing resource cost generated by the satellite mobile edge computing network.

[0076] In one optional embodiment, the computational resource cost represents the specific cost incurred when a low-orbit satellite provides edge computing resources to a current drone in a satellite mobile edge computing network. It is a key indicator for measuring the cost of satellite resource occupancy and is used to reflect the scarcity and cost of using satellite resources.

[0077] In one alternative embodiment, the satellite, as a mobile edge computing node, has limited computing resources (such as processor computing power and storage resources). Therefore, providing computing services to drones will consume the satellite's own resources and generate operating costs.

[0078] Furthermore, by acquiring the computing resource cost generated by the satellite mobile edge computing network in this embodiment, the cost of satellite resource occupation can be incorporated into the strategy evaluation system. This avoids excessive consumption of satellite resources due to focusing only on the performance of the UAV end, and realizes global resource optimization of the space-ground integrated network.

[0079] Step a112: Based on the decision type of each initial computational unloading strategy and the positional relationship between each UAV and each target beam, calculate the total latency and total energy consumption of each UAV executing the initial computational unloading strategy.

[0080] In one optional embodiment, the total latency represents the complete time consumed by the UAV in executing the initial computational offloading strategy, which can be divided into two categories according to different decision types: local computation only includes local computation time; beam offloading includes flight time, data transmission latency and satellite edge computation time.

[0081] In one optional embodiment, total energy consumption represents the total energy consumed by the UAV in executing the initial computational offloading strategy, which can be divided into two categories according to different decision types: local computation includes local computation energy consumption and hovering energy consumption; beam offloading includes flight energy consumption and hovering energy consumption.

[0082] In one optional embodiment, the decision type of the initial computation offloading strategy (local computation / beam offloading) directly determines the difference in the task execution process, while the positional relationship (distance) between the UAV's initial position and the target beam affects flight time, flight energy consumption, and data transmission quality, thereby determining the total latency and total energy consumption. Therefore, by distinguishing the decision type and combining it with positional relationship calculations, the performance differences between different strategies can be accurately quantified.

[0083] For example, the decision type is first determined based on the initial selection of the drone, as shown in the following relation (3): (3) In the formula: Indicates drone The decision.

[0084] Furthermore, the total latency of the UAV executing the initial computational unloading strategy can be calculated according to the decision type and scenario, as shown in the following relationship (4): (4) In the formula: Indicates the latency of local computation (unit: seconds); Indicates drone Task size (unit: bits); Indicates drone Locally calculated frequency (unit: Hz); Indicates drone In the round Time flies towards the beam Total delay of unloading task (unit: seconds); Indicates drone Flying towards the beam Required flight time (in seconds); Indicates drone via beam Transmission delay of the transmission task (unit: seconds); The time for calculating the edge of a low-Earth orbit satellite (in seconds): (5) In the formula: Indicates drone Fly to the beam Straight-line distance (unit: m); Indicates drone Flight speed (unit: m / s); Indicates drone The amount of task data, This represents computational density, used to divide the task cycle. Convert to bit units ; Indicates a round Time satellite as beam Allocated computing resources (unit: Hz); Indicates a round drones In beam The data transmission rate (unit: bit / s) in the data is shown in the following equation (6): (6) In the formula: Indicates beam Channel bandwidth (unit: Hz); Indicates beam Select the number of drones to unload the task; Indicates beam The instantaneous signal-to-noise ratio is shown in the following equation (7): (7) In the formula: Indicates the drone's transmission power; Indicates beam Channel gain of the coverage area; Indicates the noise power spectral density; Indicates beam The channel bandwidth.

[0085] Furthermore, the total energy consumption of the drone executing the initial computational offloading strategy can be calculated based on the decision type and scenario. The following relation (8) is shown: (8) In the formula: Indicates drone Locally calculated energy consumption (unit: J); Indicates drone Hovering energy consumption (unit: J); Indicates drone Flight to beam Energy consumption (unit: J): (9) In the formula: The energy efficiency coefficient of a drone processor (unit: J / (bit)) Hz 2 )); Indicates the local calculation frequency; This indicates the rotor power of the drone (unit: W). The induced power of the drone (in W) is the energy required to generate enough lift to offset the weight of the drone. This indicates the total flight power of the drone (unit: W). This indicates the drone's flight speed (unit: m / s). Furthermore, , , The following relationships (10) and (11) are shown: (10) (11) In the formula: Air density (unit: kg / m³) 3 ); Indicates the blade drag coefficient; Represents the area of ​​the rotor disk (unit: m²) 2 ); Indicates the rotor angular velocity (unit: rad / s); Indicates the rotor radius (unit: m); Indicates the mass of the drone (unit: kg); This represents the acceleration due to gravity (unit: m / s²). Indicates the rotor induction speed (unit: m / s); Indicates the rotor blade tip velocity (unit: m / s); Indicates the drag coefficient; The reference area of ​​the drone body (unit: m) 2 ).

[0086] Step a113: Based on the computational resource cost, the total latency and total energy consumption of each UAV, calculate the policy utility value of each UAV using a weighted utility function.

[0087] In one optional embodiment, low latency, low energy consumption, and low satellite resource cost are the three optimization objectives of the computation offloading strategy, and these three objectives are mutually trade-offs (e.g., reducing latency may increase energy consumption). Furthermore, the weighted utility function balances the importance of the three objectives through weighting factors, which can transform the multi-objective optimization problem into an optimization problem of a single quantitative index, thereby achieving a comprehensive and objective evaluation of the initial computation offloading strategy.

[0088] For example, the computational resource cost, total latency, and total energy consumption are substituted into a preset weighted utility function. The proportions of the three indicators are adjusted by weighting factors. After weighted summation and inversion, the first strategy utility value is obtained, thus completing the quantitative evaluation of the overall performance of the initial computational offloading strategy, as shown in the following relationship (12): (12) In the formula: Indicates drone In the round The first strategy utility value at that time; , , The weighting factors represent the proportions of execution time, total energy consumption, and satellite resource costs in the overall utility, respectively. Reflecting the importance of total delay in utility assessment, Reflecting the importance of total energy consumption in utility assessment, This reflects the importance of satellite resource costs in utility assessment; Indicates drone The total delay is obtained through the above relationship (4); Indicates drone The total energy consumption is obtained through the above relationship (8); This indicates the cost of calculating resources.

[0089] Step a12: Calculate the beam load bonus value for each target beam based on the load value of the target beam selected by each UAV.

[0090] In one optional embodiment, the load of satellite beams has an upper limit threshold. If a single beam is selected by a large number of drones, it will cause problems such as beam overload, decreased data transmission rate, and a surge in computing latency; while the resources of low-load beams will be idle. Therefore, by constructing a reward function associated with the real-time load value of the beams, high positive rewards are given to low-load beams, and low or even negative rewards are given to high-load / overloaded beams. This allows drones to spontaneously favor low-load beams in their strategy selection, ultimately achieving dynamic load balancing of all satellite beams, maximizing the overall beam resource utilization efficiency, and avoiding problems such as inter-beam interference and uneven resource allocation.

[0091] In an optional embodiment, the real-time load value of the target beam selected by each UAV can be quantified into a corresponding beam load bonus value. The following relation (13) is shown: (13) In the formula: Indicates target beam Real-time load value; Indicates the load factor, used to characterize the degree of beam load. The larger the value, the higher the load on the beam and the less resources are left. Indicates the preset maximum load threshold of the beam; Indicates a low load reward coefficient; This indicates a high load penalty coefficient.

[0092] Among them, when the beam is unloaded At that time, the reward is calculated using the exponential decay formula, and the load rate is... The smaller, The higher the value, the higher the beam load reward value, which creates a strong positive incentive for the UAV to select the low load beam.

[0093] Furthermore, when the beam is overloaded... At this time, a linear formula is used to calculate the reward, and the load rate is... , If the value is negative, the beam load reward is negative, which creates a reverse constraint on the selection of the overcarrier beam by the UAV, guiding the UAV to avoid the beam.

[0094] Step a13: Calculate multiple flight distance penalty values ​​based on the distance between each UAV and each target beam.

[0095] In one alternative embodiment, the farther the UAV flies towards the satellite beam, the more non-linearly the flight time and energy consumption increase. If only low beam load is pursued while ignoring flight costs, the overall strategy utility will decrease. Therefore, by constructing a penalty function linked to flight distance, the distance cost is quantified into a specific penalty value, allowing the UAV to consider both beam load and flight distance in strategy selection, thereby avoiding blindly choosing long-distance, low-load beams.

[0096] In an optional embodiment, a one-to-one distance penalty quantization relationship is established between each UAV and each of its selectable satellite target beams, and the corresponding flight distance penalty value for each UAV when selecting different beams is calculated. The following relation (14) is shown: (14) In the formula: Indicates that the drone flies to the beam Straight-line distance (unit: m); This indicates the maximum distance a drone can fly in a single mission. It can be set based on the drone's endurance and operational scenarios to prevent drones from blindly pursuing low latency, which would increase flight distance and cause a surge in drone flight energy consumption. This represents the distance penalty coefficient.

[0097] Furthermore, the penalty value is a numerical result, and the greater the distance, the greater the penalty value, which will directly offset the total reward value and reduce the strategic attractiveness of the long-distance beam.

[0098] Step a14: Calculate multiple low-latency reward values ​​based on the latency of the current computational offloading strategy of each drone and the latency of the local computational strategy of each drone.

[0099] In one optional embodiment, one of the core objectives of satellite-assisted unloading is to reduce computational latency. If the latency of the beam unloading strategy selected by the UAV is lower than the local computational latency, a positive reward should be given to strengthen the strategy. If the unloading latency is higher, the reward value will be reduced, which can reflect the latency disadvantage of the strategy and also conform to the training logic of the PPO algorithm to positively incentivize superior strategies and weaken inferior strategies.

[0100] In an optional embodiment, for each UAV, the total unloading latency when selecting different target beams for unloading is calculated and compared with the UAV's own local computation latency, thereby quantifying the latency optimization ratio into a low-latency reward value. The following relation (15) is shown: (15) In the formula: This indicates the local computation latency, i.e., the latency of the local computation strategy; Indicates drone Select target beam The total uninstallation time during uninstallation, i.e. the time required to calculate the current uninstallation strategy; This represents the delay penalty coefficient.

[0101] Furthermore, the low-latency reward value is a numerical result; the lower the offloading latency compared to the local computation latency, the larger the low-latency reward value.

[0102] Step a15: Determine multiple instantaneous reward values ​​for the UAV based on multiple policy utility values, multiple beam load reward values, multiple flight distance penalty values, and multiple low latency reward values.

[0103] In one optional embodiment, the merits of the UAV computational offloading strategy need to be comprehensively evaluated from four core dimensions: overall strategy utility, beam load, flight distance, and latency optimization. Therefore, by integrating the quantitative indicators of the four dimensions into a single instant reward value through linear fusion, the PPO algorithm can judge the merits of the strategy through this single indicator. At the same time, reasonable quantitative weights are set for each dimension to ensure that the reward can accurately reflect the actual execution effect of the strategy.

[0104] In an optional embodiment, the instantaneous reward value is calculated using the following formula (16) by combining the obtained multiple policy utility values, multiple beam load reward values, multiple flight distance penalty values, and multiple low latency reward values. : (16) In the formula: Indicates drone The utility value of the selected current computational unloading strategy; Indicates the reward amplification factor; Indicates the basic reward; This indicates the basic penalty.

[0105] Furthermore, when the latency of the unloading strategy is greater than the latency of local computation, the reward is directly set to -100, so that the drone avoids selecting the high-load beam through negative rewards.

[0106] Step a16: Based on multiple state sample sets and multiple immediate reward values, multiple advantage values ​​are obtained through value network calculation in the near-end policy optimization algorithm.

[0107] In one optional embodiment, the immediate reward value only reflects the absolute gain of a certain step of the UAV's strategy, while the advantage value needs to quantify the relative superiority or inferiority of a certain action (strategy selection) relative to the average performance of the current strategy. Therefore, in this embodiment, the UAV's state sample set and immediate reward value are processed through the PPO value network, and combined with the temporal difference (TD) method, the absolute reward is transformed into a relative advantage value. This not only reduces the variance of the reward value, but also provides a more accurate gradient update signal for the policy network, ensuring the stability of the PPO algorithm training.

[0108] In an alternative embodiment, an initial advantage value is set. Then, the current state Input a value network, output a value estimate for the current state. ; the next state Input a value network, output the value estimate of the next state. .

[0109] Furthermore, the corresponding advantage value can be calculated using the above relationship (2).

[0110] Step a2: Calculate the probability policy ratio of each initial computation unloading policy and each new real-time computation unloading policy for the same action.

[0111] In an optional embodiment, the probability-policy ratio represents the ratio of the probability of the old policy (initial calculation unloading policy) and the new policy (real-time calculation unloading policy) selecting the same action under the same conditions, and is used to measure the difference in action selection preferences between the old and new policies.

[0112] In an optional embodiment, the probability difference between the old and new strategies on the same action directly reflects the update direction and magnitude of the strategy. Therefore, for each state in the first state sample set, the probability of the old strategy (initial calculation unloading strategy) and the new strategy (real-time calculation unloading strategy) selecting the same action (local calculation or beam unloading) is calculated respectively, and the ratio of the two is the probability strategy ratio.

[0113] For example, extracting the state sample set corresponding actions This refers to the actions performed under the old strategy.

[0114] Furthermore, through the old strategy Calculation in state Select action probability At the same time, through new strategies Calculation in the same state Select the same action below probability Furthermore, the probability strategy ratio is calculated using the following relationship (17): (17) In the formula: This represents the weight parameters of the policy network.

[0115] Furthermore, if If the probability is 0, it means that the current strategy is more inclined to choose that action; otherwise, the current strategy will reduce the probability of choosing the highest-probability action.

[0116] Step a3: Based on multiple advantage values, a preset clipping function is used to limit the ratio of each probability strategy within a preset range and construct a target clipping function for each UAV.

[0117] In one optional embodiment, the preset pruning function is a function used in the PPO algorithm to limit the policy update magnitude. It can constrain the probability-policy ratio within a preset range, thereby avoiding training fluctuations caused by excessive differences between the old and new policies and ensuring the stability of policy updates. The target pruning function is the core policy loss function of the PPO algorithm. It balances the efficiency and stability of policy optimization by taking the smaller value between the product of the probability-policy ratio and the advantage value and the product of the pruned probability-policy ratio and the advantage value.

[0118] In an optional embodiment, a clipping factor hyperparameter is set. By using a preset pruning function, the probability strategy ratio is limited to a certain range. Within this framework, combined with the advantage value, a target clipping function can be constructed.

[0119] For example, the probability policy ratio is processed through a pruning function. Furthermore, if The cropped value is ;like The cropped value is Otherwise, the original ratio remains unchanged.

[0120] Furthermore, combining the advantage value Construct the target clipping function as shown in the following relation (18): (18) In the formula: This represents the target pruning function (policy loss).

[0121] Step a4: Obtain the value loss function of the commentator network in the near-end policy optimization algorithm.

[0122] In one optional embodiment, the Critic Network is the network component responsible for state value evaluation in the PPO algorithm. It optimizes the accuracy of value estimation by calculating the value loss function, reduces the evaluation variance of the advantage value, and assists the policy network in achieving more accurate optimization.

[0123] In one alternative embodiment, the value loss function is the optimization objective function of the commentator network, which improves the accuracy of state value estimation by minimizing the mean square error between the predicted value and the target value.

[0124] In one optional embodiment, the core role of the commentator network is to accurately estimate state value. The value loss function guides the commentator network to optimize by measuring the difference between the predicted value and the target value, thereby reducing the variance of value estimation. Therefore, in this embodiment, to make policy evaluation more accurate and assist policy optimization, a value loss function of the commentator network is introduced. By combining TD objectives with immediate rewards and bootstrapping estimation, variance is reduced, as expressed by the following relationship (19): (19) in: (20) In the formula: This indicates the network of critics' opinions on the current state. The projected value; The target value estimate for the next state is calculated through the policy network of PPO, and the remaining parameters are consistent with the advantage value formula, i.e., the above relationship (2).

[0125] Step a5: Based on the target pruning function and value loss function of each UAV, construct the total loss function for each UAV.

[0126] In one optional embodiment, the total loss function is formed by integrating the policy loss corresponding to the target pruning function, the value loss corresponding to the value loss function, and the entropy regularization term according to preset weights.

[0127] The entropy regularization term is used to ensure the exploratory nature of the new strategy and prevent the strategy from converging to a local optimum too early.

[0128] Furthermore, by combining these three elements, the joint goals of strategy optimization, value assessment optimization, and exploratory safeguards are achieved, ensuring that strategy updates are efficient and stable.

[0129] Step S2033: Optimize the total loss function of each UAV using gradient pruning and backpropagation algorithms, and iteratively update the initial network weight parameters of the real-time calculation unloading strategy.

[0130] In one alternative embodiment, gradient clipping is a technique to limit the magnitude of gradient updates in a neural network and avoid gradient explosion; the backpropagation algorithm is the core algorithm that updates the weight parameters in reverse by calculating the gradient of the total loss function with respect to the network weights. The combination of the two can achieve stable iterative optimization of the strategy.

[0131] In one optional embodiment, the gradient of the total loss function with respect to the network weights of the real-time computational unloading strategy (new strategy) is calculated using the backpropagation algorithm. After limiting the gradient magnitude through gradient pruning, the weight parameters are iteratively updated along the gradient descent direction, thereby gradually improving the overall performance of the new strategy.

[0132] Step S2034: The updated real-time computation offloading strategy is used as the new initial computation offloading strategy. The step of constructing the total loss function for each UAV is returned. The process is iterated repeatedly until the satellite mobile edge computing network is in a stable state, and the optimal computation offloading strategy for each UAV is obtained.

[0133] In an optional embodiment, the real-time computation offloading strategy after weight update is redefined as a new initial computation offloading strategy, and the process returns to step S2032. Steps S2032 to S2034 are repeated until the satellite mobile edge computing network reaches a stable state. Finally, the optimal computation offloading strategy is output, at which point the entire satellite mobile edge computing network will maximize its service revenue.

[0134] In some optional implementations, the above method further includes: Step b1: Obtain the initial upper limit of the load of a single beam and the real-time load value of each beam in the satellite mobile edge computing network.

[0135] Step b2: Determine whether the real-time load values ​​of all beams in the satellite mobile edge computing network are equal to the initial upper limit value.

[0136] Step b3: When the real-time load values ​​of all beams are equal to the initial upper limit value, the optimal computation offloading strategy is determined to be the local computation offloading strategy.

[0137] In one alternative embodiment, the computing resources (computing power, bandwidth) of a satellite beam are limited. Excessive drone access can lead to overload, causing problems such as increased transmission latency and data packet loss. Therefore, by setting an initial upper limit for the load of a single beam, the resource carrying capacity boundary of the beam can be clearly defined, avoiding service quality issues caused by resource depletion.

[0138] For example, a satellite can preset an initial upper limit value for the load of a single beam (unit: number, i.e., the maximum number of drones that can be connected to a single beam) based on the computing power (such as processor frequency) and channel bandwidth of its own edge computing nodes, combined with the average resource consumption of drone missions.

[0139] Furthermore, the satellite can broadcast the initial upper limit of the load for a single beam to all drones in the network, thereby ensuring that all drones obtain a uniform load constraint standard.

[0140] Furthermore, the drone acquires the current number of drones connected to the target beam in real time (i.e., the current real-time load value), compares it with the acquired initial upper limit value, and determines whether the target beam is in an overload state.

[0141] Furthermore, when the load on all satellite beams reaches its initial limit, the satellite network's computing resources are completely exhausted, leaving no resources available for the drone to offload the mission. At this point, attempting to switch beams or optimize the offloading strategy is pointless; failure to trigger a special strategy in a timely manner could lead to the drone mission failing or a sharp decline in service quality. Therefore, by assessing this extreme scenario, a fallback plan can be quickly triggered to ensure the feasibility of mission execution.

[0142] Furthermore, the local computing strategy does not rely on satellite resources and can complete the task independently. It is the only feasible fallback solution in this extreme scenario, which can ensure the normal execution of UAV missions and improve system robustness.

[0143] For example, the drone can acquire the current load of all satellite beams in the satellite mobile edge computing network in real time, and compare it with the acquired initial upper limit value to determine whether all beams have reached the load limit.

[0144] Furthermore, when it is determined that the load of all satellite beams has reached the initial upper limit, the PPO algorithm iteration update process is skipped directly, and the local computation offloading strategy is determined as the optimal computation offloading strategy for the UAV, thereby ensuring that the mission can still be executed smoothly without satellite resource support.

[0145] In one example, a method for allocating computational resources based on satellite-assisted multi-UAV offloading is provided, such as... Figure 3 As shown, it includes the following steps: Step 1: In the satellite mobile edge computing network, the system is first initialized by setting the upper limit of the load of a single beam and randomly initializing the positions of the beams and drones.

[0146] Step 2: Each UAV calculates the utility of this strategy based on its own position and the selected beam. The utility calculation takes into account two key factors: the time and energy required to complete the task. The utility is represented by the above relationship (10).

[0147] Furthermore, time reflects the data transmission rate and task computation time, while energy consumption reflects the UAV's flight energy consumption, hovering energy consumption, and local computing energy consumption. Low-Earth orbit (LEO) satellites broadcast information from all beams to the UAV so that it can make subsequent strategy adjustments. LEO satellites receive information on the number of users connected to the beams and calculate the average utility based on this information. The average utility provides users with a reference benchmark for the overall system performance.

[0148] Step 3: Adjust the strategy using the PPO algorithm. Each UAV compares its utility (calculated from the above relation (10)) with its own local computation utility. If the utility of the strategy is lower than the utility of its own local computation, it means that the performance of the currently selected satellite beam is not optimal. The strategy's utility being lower than the local computation utility is optimal for itself. If the beam load reaches the upper limit, then local computation is the optimal strategy for itself. Therefore, the UAV selects different satellite beams or local computation to explore better service options. Conversely, if the node's utility is higher than its own local computation utility, the current selection remains unchanged. In addition, if the load of all beams reaches the maximum value, then local computation is selected. The specific process is as follows: (1) The drone interacts with the environment, and each drone gets a turn. The state of the drone at that time The state is represented by the above relation (1).

[0149] Furthermore, PPO, through interaction with the environment, uses its own policy network to collect samples of the state, actions, rewards, and next state of individual drones. This means that a large number of samples can be obtained by having the agent run through the environment for multiple rounds.

[0150] (2) Using the collected samples and the PPO value network, the advantage function can be calculated. The advantage function can quantify the superiority or inferiority of a specific action relative to the average performance of the policy, thereby guiding the policy update direction and optimizing training stability. First, initialize the system to make the system more stable. , The calculation is shown in the above relation (2).

[0151] (3) In order to measure the probability difference between the new strategy and the old strategy on the same action, so as to prune the objective function in the future, PPO introduced the probability strategy ratio, as shown in the above relation (17).

[0152] Furthermore, to stabilize training and limit the update magnitude of the policy, PPO introduces a pruning objective function. As shown in the above relation (18).

[0153] (4) In order to make the evaluation of the strategy more accurate and to assist in the optimization of the strategy, the value loss function of the commentator network is introduced. By combining the TD objective with the immediate reward and bootstrap estimation, the variance is reduced, as expressed by the above relation (19).

[0154] Step Four: The above steps will be repeated until the system reaches a stable state. During this process, the user's policy selection will be dynamically adjusted according to the real-time network environment, eventually reaching an equilibrium point. The policy at this point can be considered the optimal service selection policy. This dynamic adjustment process ensures that the system can adapt to the dynamic characteristics of the satellite network and maximize the efficiency of computing resource utilization.

[0155] Through the above process, users can switch strategies and make optimal decisions in the satellite mobile edge network. When the system reaches an equilibrium state, the entire satellite mobile edge computing network will maximize its service revenue.

[0156] In some implementations, the method of satellite-assisted multi-UAV unloading missions, combined with Figure 4 You can follow these steps: (1) Set the upper limit of the satellite beam load based on the actual scenario. The setting of this upper limit lays the foundation for subsequent service selection, which will affect the initial allocation of network resources and the initial service experience of users. In addition, the position of the beam can be set to fixed each time, which can significantly reduce training time.

[0157] (2) Evaluating the effectiveness of the strategy using a utility model. In this invention, the time and energy consumption of the UAV in executing the strategy are used to construct a utility model through weighted parameters, as shown in equation (12). The UAV can choose from strategies such as local computation ( ) and flying towards the One beam offload task ( If the drone chooses the offloading strategy, then The total time consists of the UAV flight time, data transmission time, and satellite computation time; if the UAV chooses local computation, the total time consists only of local computation time. Correspondingly, if the UAV chooses an offload strategy, the total energy consumption consists of flight energy consumption and hovering energy consumption during data transmission; if the UAV chooses local computation, the energy consumption consists only of hovering energy consumption. After this, the satellite broadcasts beamload information, providing a basis for subsequent strategy selection.

[0158] (3) Determine whether the strategy needs to be adjusted based on utility: If the utility of the strategy selected by the UAV is less than its local computational utility, that is: If so, it can choose to fly to different satellite beams to offload tasks or perform local calculations; conversely, if If the current satellite beam selection remains unchanged, then the local calculation strategy is selected if the load of each beam reaches its maximum value. The above steps constitute one cycle of the loop, combined with the following procedure: Initialize the policy network (policy) and the old policy network (old_policy). for episode = 1 to max_episodes: # Dynamically adjust hyperparameters Adjust the learning rate lr = lr Attenuation coefficient Adjust the PPO update epochs (ppo_epochs) based on the training phase. Adjust the discount factor gamma and the clipping coefficient clip_epsilon Adjusting the entropy coefficient entropy_coeff For each intelligent agent (drone): Sampling actions based on the current strategy Storage transferred to buffer # Experience replay buffer update check if buffer.size >= batch_size: # Strategy Optimization Phase Randomly sample a batch of data from the buffer. Calculate the estimate of the advantage function using the TD error. Calculate return estimates # Multi-round PPO optimization for k = 1 to ppo_epochs: Update network parameters for the old policy: old_policy ← policy For each mini-batch of data, do: # Strategy Loss Calculation (PPO Core) Calculate the probability of the new strategy Calculate the probability of the old policy Calculate the probability ratio: ratio = new_probs / old_probs # Pruning Objective Function surr1 = ratio Advantages surr2 = clip(ratio, 1-ε, 1+ε) advantages (used to limit the update range) actor_loss = -min(surr1, surr2).mean() Calculate the value function loss (critic_loss) Calculate the entropy regularization term entropy_loss Calculate the total loss function total_loss = actor_loss + vf_coeff critic_loss + entropy_coeff entropy_loss Backpropagation and gradient clipping end end end end This process is repeated until the system reaches a steady state. Once the system reaches a steady state, the user's policy choice no longer changes; this policy is the optimal service selection policy. Ultimately, based on the steady state obtained through the dynamic adjustment process, each user in the system can receive optimal service, achieving optimized network resource utilization.

[0159] The computational resource allocation method based on satellite-assisted multi-UAV offloading provided in this example has the following advantages: 1. By introducing a time and energy consumption assessment model based on UAV mission execution, this invention can quantify the strategic utility of UAVs in real time and provide comprehensive decision evaluation. Combined with a dynamic service selection strategy, the system can adjust decisions in real time according to changes in the network environment, optimize the trade-off between total energy consumption and computation time, and prioritize low-load satellite beam offloading tasks for UAVs, thereby maximizing the average utility of the system (e.g., Figures 5 to 7 As shown), ensuring stable operation in dynamic satellite networks. Among them, Figure 5 This is a comparison chart of average latency. Average Dleay represents average latency; Drone Count represents the number of drones; Random represents the random policy; Local represents the local computation policy; Uninstall represents the fixed uninstallation policy; DQN represents the deep Q-network algorithm. Figure 6 This is a graph comparing average energy consumption; Energy Consumption represents energy consumption. Figure 7The diagram shows a comparison of beam load. "Comparison of UAV strategy selection based on three algorithms" represents a comparison of UAV strategy selection based on three algorithms; "Strategy selection of UAV" represents the strategy selection of UAVs; "Local Computing" represents local computation; "Beam1", "Beam2", "Beam3", and "Beam4" represent beams 1, 2, 3, and 4, respectively; and "Number of UAVs" represents the number of UAVs.

[0160] 2. The distributed service selection algorithm in this example exhibits rapid convergence, reaching a stable equilibrium state within a short time. It utilizes the PPO framework to provide a stable policy update mechanism, limiting the policy update magnitude to avoid training fluctuations and significantly improving the overall system performance. The architecture is essentially centralized training with distributed autonomous decision-making, eliminating the need for real-time coordination with other drones; a drone failure does not affect the decisions of other drones.

[0161] This embodiment also provides a satellite-assisted multi-UAV computational offloading device, which is used to implement the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that performs a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0162] This embodiment provides a satellite-assisted multi-UAV computational offloading device, such as... Figure 8 As shown, the device includes: The acquisition module 801 is used to acquire a set of state samples generated by the interaction between multiple UAVs and satellite mobile edge computing networks, stored in the experience playback buffer.

[0163] The judgment module 802 is used to determine whether the number of state samples in the state sample set is greater than or equal to the preset batch size.

[0164] The update module 803 is used to update the initial computational offloading strategy of each UAV using a near-end strategy optimization algorithm when the number of state samples in the state sample set is greater than or equal to the preset batch size, until the satellite mobile edge computing network is in a stable state, and obtain the optimal computational offloading strategy of the UAV.

[0165] The satellite-assisted multi-UAV computational offloading device provided in this embodiment of the invention can execute the satellite-assisted multi-UAV computational offloading method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects for executing the method. Further functional descriptions of the above modules are the same as in the corresponding embodiments described above, and will not be repeated here.

[0166] Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention.

[0167] The following is a detailed reference. Figure 9 This diagram illustrates a suitable structural schematic for implementing an electronic device according to embodiments of the present invention. The electronic device may include a processor (e.g., a central processing unit, graphics processor, etc.) 901, which can perform various appropriate actions and processes based on a program stored in read-only memory (ROM) 902 or a program loaded from memory 909 into random access memory (RAM) 903. RAM 903 also stores various programs and data required for the operation of the electronic device. The processor 901, ROM 902, and RAM 903 are interconnected via bus 904. An input / output (I / O) interface 905 is also connected to bus 904.

[0168] Typically, the following devices can be connected to I / O interface 905: input devices 906 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 907 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; memory devices 909 including, for example, magnetic tapes, hard disks, etc.; and communication devices 909. Communication device 909 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 9 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown, and more or fewer devices may be implemented or have instead.

[0169] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 909, or installed from a memory 909, or installed from a ROM 902. When the computer program is executed by the processor 901, it performs the functions defined in the satellite-assisted multi-UAV computational offloading method of the embodiments of the present invention.

[0170] Figure 9The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of the present invention.

[0171] This invention also provides a computer-readable storage medium. The methods described above according to embodiments of the invention can be implemented in hardware or firmware, or implemented as computer code that can be recorded on a storage medium, or implemented as computer code downloaded via a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that the computer, processor, microprocessor controller, or programmable hardware includes storage components capable of storing or receiving software or computer code. When the software or computer code is accessed and executed by the computer, processor, or hardware, the satellite-assisted multi-UAV computational offloading method shown in the above embodiments is implemented.

[0172] A portion of this invention can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to the invention through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.

[0173] Although embodiments of the invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the invention, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. A satellite-assisted multi-UAV computational unloading method, characterized in that, The method includes: Obtain the state sample set generated by each UAV's interaction with the satellite mobile edge computing network, stored in the experience replay buffer; Determine whether the number of state samples in each state sample set is greater than or equal to the preset batch size; When the number of state samples in the state sample set is greater than or equal to the preset batch size, the initial computational offloading strategy of each UAV is updated using the near-end strategy optimization algorithm until the satellite mobile edge computing network is in a stable state, thus obtaining the optimal computational offloading strategy of the UAV.

2. The method according to claim 1, characterized in that, When the number of state samples in the state sample set is greater than or equal to the preset batch size, the initial computational offloading strategy of each UAV is updated using a near-end policy optimization algorithm until the satellite mobile edge computing network is in a stable state, thus obtaining the optimal computational offloading strategy for the UAV, including: Using the initial computational offloading strategy of each UAV as the old strategy, a new real-time computational offloading strategy is obtained for each UAV, wherein the initial network weight parameters of the real-time computational offloading strategy are consistent with the network weight parameters of the initial computational offloading strategy. Based on the state sample set, the initial computational offloading strategy, and the real-time computational offloading strategy of each UAV, the total loss function of each UAV is constructed after processing by the near-end strategy optimization algorithm. The total loss function of each UAV is optimized using gradient pruning and backpropagation algorithms, and the initial network weight parameters of the real-time calculation unloading strategy are iteratively updated. The updated real-time computation offloading strategy is used as the new initial computation offloading strategy. The step of constructing the total loss function for each UAV is returned and iterated repeatedly until the satellite mobile edge computing network is in a stable state, thus obtaining the optimal computation offloading strategy for each UAV.

3. The method according to claim 2, characterized in that, Based on the state sample set, the initial computational offloading strategy, and the real-time computational offloading strategy of each UAV, and after processing by the near-end policy optimization algorithm, a total loss function for each UAV is constructed, including: Each state sample set is input into the value network in the near-end policy optimization algorithm for calculation to obtain multiple advantage values. Each advantage value is the degree of superiority or inferiority of the current action state of each UAV relative to the average value of the calculated unloading policy. Calculate the probability strategy ratio of each initial computation offloading strategy and each new real-time computation offloading strategy for the same action; Based on multiple advantage values, a preset pruning function is used to limit the probability strategy ratio of each of the above-mentioned probabilities within a preset range and to construct a target pruning function for each of the above-mentioned UAVs. Obtain the value loss function of the commentator network in the near-end policy optimization algorithm; Based on the target pruning function and the value loss function of each UAV, the total loss function of each UAV is constructed.

4. The method according to claim 3, characterized in that, Each state sample set is input into the value network of the near-end policy optimization algorithm for calculation, resulting in multiple advantage values, including: Based on the location of each UAV and the target beam selected by each UAV, the strategy utility value of the initial computational offloading strategy of each UAV is calculated using a weighted utility function; Based on the load value of the target beam selected by each of the UAVs, calculate the beam load bonus value for each target beam; Based on the distance between each of the UAVs and each of the target beams, multiple flight distance penalty values ​​are calculated; Based on the latency of the current computational offloading strategy of each UAV and the latency of the local computational strategy of each UAV, calculate multiple low-latency reward values; Based on multiple policy utility values, multiple beam load reward values, multiple flight distance penalty values, and multiple low latency reward values, multiple instant reward values ​​for the UAV are determined. Based on the multiple state sample sets and the multiple instant reward values, the multiple advantage values ​​are obtained through value network calculation in the near-end policy optimization algorithm.

5. The method according to claim 4, characterized in that, Based on the location of each UAV and the target beam selected by each UAV, the strategy utility value of the initial computational offloading strategy for each UAV is calculated using a weighted utility function, including: Obtain the computing resource cost generated by the satellite mobile edge computing network; Based on the decision type of each initial computational offloading strategy, the position of each UAV and the positional relationship between each target beam, calculate the total latency and total energy consumption of each UAV executing the initial computational offloading strategy; Based on the computational resource cost, the total latency of each UAV, and the total energy consumption, the policy utility value of each UAV is calculated using the weighted utility function.

6. The method according to claim 1, characterized in that, The method further includes: Obtain the initial upper limit of the load of a single beam and the real-time load value of each beam in the satellite mobile edge computing network; Determine whether the real-time load values ​​of all beams in the satellite mobile edge computing network are equal to the initial upper limit value; When the real-time load values ​​of all beams are equal to the initial upper limit value, the optimal computation offloading strategy is determined to be the local computation offloading strategy.

7. A satellite-assisted multi-UAV computational unloading device, characterized in that, The device includes: The acquisition module is used to acquire a set of state samples generated by the interaction between multiple UAVs and satellite mobile edge computing networks, stored in the experience replay buffer. The judgment module is used to determine whether the number of state samples in the state sample set is greater than or equal to the preset batch size; The update module is used to update the initial computational offloading strategy of each UAV using a near-end strategy optimization algorithm when the number of state samples in the state sample set is greater than or equal to a preset batch size, until the satellite mobile edge computing network is in a stable state, thereby obtaining the optimal computational offloading strategy of the UAV.

8. An electronic device, characterized in that, include: A memory and a processor are communicatively connected, the memory stores computer instructions, and the processor executes the computer instructions to perform the satellite-assisted multi-UAV computational offloading method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing the computer to execute the satellite-assisted multi-UAV computational offloading method according to any one of claims 1 to 6.

10. A computer program product, characterized in that, Includes computer instructions for causing a computer to execute the satellite-assisted multi-UAV computational offloading method according to any one of claims 1 to 6.