Joint optimization method and device for offloading task of internet of things equipment to satellite

By employing a joint optimization method combining multi-agent deep reinforcement learning and convex optimization models, the data processing challenges faced by IoT devices in areas lacking ground infrastructure are addressed. This approach enables efficient task offloading and resource allocation, reduces energy consumption and latency, and enhances system flexibility and resource utilization.

CN120301491BActive Publication Date: 2026-05-01BEIJING UNIV OF POSTS & TELECOMM +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING UNIV OF POSTS & TELECOMM
Filing Date
2025-04-08
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

In existing technologies, IoT devices struggle to effectively process massive amounts of data in areas lacking adequate terrestrial communication and computing infrastructure. Single-satellite offloading solutions are limited by coverage area, while multi-satellite offloading solutions lead to increased device energy consumption and high system costs.

Method used

A three-layer offloading system for IoT devices, gateways, and satellites is constructed by employing a multi-agent deep reinforcement learning model and a convex optimization model. By jointly optimizing offloading decisions, network access selection, and computing resource allocation, transmission power is dynamically adjusted to optimize the task offloading path.

Benefits of technology

It enables efficient task offloading and resource allocation in complex environments, reduces system energy consumption and transmission latency, improves equipment efficiency and resource utilization, and adapts to the flexibility of multiple devices and network changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120301491B_ABST
    Figure CN120301491B_ABST
Patent Text Reader

Abstract

The application provides a joint optimization method and device for offloading tasks of Internet of Things devices to satellites, and relates to the technical field of satellite computing, comprising: collecting current environment data through multiple Internet of Things devices; obtaining observation information of an i-th Internet of Things device at a time slot t; inputting the observation information of the i-th Internet of Things device at the time slot t into a policy network to generate an offloading decision scheme and a network access selection scheme; under the premise of the offloading decision scheme and the network access selection scheme, using a convex optimization model to determine at least one of the following: a computing resource amount allocated to the i-th Internet of Things device by an s-th satellite, a transmission power of the i-th Internet of Things device at the time slot t, and a transmission power of a g-th gateway at the time slot t. The joint optimization method improves the resource utilization of the system while ensuring the energy efficiency and service quality of the Internet of Things devices by comprehensively considering multiple factors such as task offloading, resource allocation and power control.
Need to check novelty before this filing date? Find Prior Art

Description

A joint optimization method and apparatus for offloading tasks from IoT devices to satellites Technical Field

[0001] This invention relates to the field of satellite computing technology, and more specifically to a joint optimization method and apparatus for offloading tasks from Internet of Things (IoT) devices to satellites. Background Technology

[0002] With the rapid development of Internet of Things (IoT) technology, an increasing number of sensor devices are being used in fields such as data acquisition, environmental monitoring, and disaster early warning. In approximately 80% of the world's land areas, due to a lack of adequate terrestrial communication and computing infrastructure, IoT devices in these regions face the challenge of effectively processing massive amounts of data, especially in resource-constrained environments such as remote areas and disaster zones. Utilizing low Earth orbit (LEO) satellites as communication and computing infrastructure, offloading data from IoT devices to satellites, and performing data processing and task offloading on the satellites, has become an effective solution.

[0003] However, most related satellite mission offloading schemes rely on a single satellite or multiple satellites for mission offloading. In single-satellite-assisted offloading, IoT devices rely on only one satellite for data offloading, thus eliminating the satellite selection issue. However, this scheme cannot effectively handle complex coverage problems and is easily limited by the coverage capabilities of a single satellite. In multi-satellite-assisted offloading, although multiple satellites can be used to provide more services, most studies adopt the approach of IoT devices directly offloading their missions to the satellites. This architecture faces the problem of excessive device transmission power, leading to increased device energy consumption and higher system design costs. Summary of the Invention

[0004] This invention provides a joint optimization method and apparatus for offloading tasks from IoT devices to satellites, aiming to solve the problems existing in the background art.

[0005] To solve the above-mentioned technical problems, the present invention is implemented as follows:

[0006] In a first aspect, embodiments of the present invention provide a joint optimization method for offloading tasks from IoT devices to satellites, applied to a three-layer offloading system comprising multiple IoT devices, multiple gateways, and multiple satellites. In this three-layer offloading system, a single IoT device is simultaneously covered by multiple gateways, and a single gateway is simultaneously covered by multiple satellites. The method includes:

[0007] The current environmental data is collected by multiple IoT devices, and the current environmental data includes at least one of the following: images, temperature, and wind speed.

[0008] Obtain the observation information of the i-th IoT device in time slot t;

[0009] The observation information of the i-th IoT device in time slot t is input into the policy network deployed locally by the i-th IoT device to generate an offloading decision scheme and a network access selection scheme. The offloading decision scheme indicates whether the i-th IoT device offloads the computing task to the satellite through the gateway in time slot t, and the network access selection scheme indicates whether the g-th gateway accesses the s-th satellite through the q-th sub-frequency band.

[0010] Given the offloading decision scheme and the network access selection scheme, at least one of the following is determined using a convex optimization model: the amount of computing resources allocated by the s-th satellite to the i-th IoT device, the transmit power of the i-th IoT device in time slot t, and the transmit power of the g-th gateway in time slot t.

[0011] Optionally, the policy network locally deployed by the i-th IoT device is obtained by training a multi-agent deep reinforcement learning model;

[0012] The reward function of the multi-agent deep reinforcement learning model includes a device utility reward and a task timeout penalty. The device utility reward is determined based on the task latency and energy consumption of the i-th IoT device in local computing mode and in satellite computing mode. The task timeout penalty is determined based on the maximum tolerable latency of the task and the actual completion latency of the task.

[0013] The observation space of the multi-agent deep reinforcement learning model is defined by the observation information of the i-th IoT device in time slot t. The observation information includes: the amount of task data, computation density and maximum tolerable latency of the i-th IoT device in time slot t, the transmit power of the i-th IoT device in time slot t, the transmit power of the g-th gateway in time slot t, the computational resources allocated to the i-th IoT device by the s-th satellite, the channel gain from the g-th gateway to the s-th satellite in the q-th sub-frequency band, and the actual task completion latency of the i-th IoT device in time slot t-1.

[0014] The action space of the multi-agent deep reinforcement learning model is defined by a first binary decision variable, which determines whether the i-th IoT device unloads its computing task to the satellite through the gateway in time slot t, and a second binary decision variable, which determines whether the g-th gateway accesses the s-th satellite through the q-th sub-frequency band.

[0015] Optionally, the task latency of the i-th IoT device in local computing mode and the task latency in satellite computing mode are determined by the following steps:

[0016] Based on the first binary decision variable, determine whether the i-th IoT device chooses to offload the computing task to the satellite through the gateway in time slot t;

[0017] Without selecting to offload computing tasks to the satellite via the gateway, the task latency of the i-th IoT device in local computing mode is calculated based on the local computing capability of the i-th IoT device and the task data volume of the un-offloaded computing tasks, and the task latency of the i-th IoT device in local computing mode is used as the actual task completion latency.

[0018] If the task is to be offloaded through a gateway, the target gateway to which the i-th IoT device is connected and the target satellite corresponding to the target gateway are determined based on the second binary decision variable.

[0019] Based on the transmission rate from the i-th IoT device to the target gateway, the communication delay from the i-th IoT device to the target gateway is determined, and based on the transmission rate from the target gateway to the target satellite, the total communication delay from the target gateway to the target satellite is determined.

[0020] Based on the communication latency from the i-th IoT device to the target gateway, the total communication latency from the target gateway to the target satellite, and the computation processing latency of the computing tasks offloaded by the target satellite to the i-th IoT device, the task latency of the i-th IoT device in satellite computing mode is obtained, and the task latency of the i-th IoT device in satellite computing mode is taken as the actual task completion latency.

[0021] Optionally, the communication delay from the i-th IoT device to the target gateway is determined based on the transmission rate from the i-th IoT device to the target gateway, including:

[0022] Based on the amount of task data of the computing task offloaded by the i-th IoT device and the transmission rate of the target gateway, the transmission delay from the i-th IoT device to the target gateway is determined;

[0023] Based on the transmission rate from the target gateway to the target satellite, the total communication delay from the target gateway to the target satellite is determined, including:

[0024] Based on the amount of task data of the computing tasks offloaded by the target gateway and the transmission rate of the target satellite, the transmission delay from the target gateway to the target satellite is determined.

[0025] Based on the signal transmission distance and the speed of light between the target gateway and the target satellite, the propagation delay from the target gateway to the target satellite is determined;

[0026] The total communication delay from the target gateway to the target satellite is determined by summing the transmission delay from the target gateway to the target satellite and the propagation delay from the target gateway to the target satellite.

[0027] Optionally, calculating the task latency of the i-th IoT device in local computing mode based on the local computing power of the i-th IoT device and the amount of task data of the unloaded computing tasks includes:

[0028] The task data volume, computation density, and local computing power of the computation task that is not unloaded in time slot t of the i-th IoT device are obtained, wherein the computation density represents the number of computation cycles required per bit of task data.

[0029] The total computational requirements of the task are determined based on the amount of task data and computational density.

[0030] Dividing the total computing requirements by the local computing power yields the task latency of the i-th IoT device in local computing mode.

[0031] Optionally, the task power consumption of the i-th IoT device in local computing mode and the task power consumption in satellite computing mode are determined by the following steps:

[0032] Based on the first binary decision variable, determine whether the i-th IoT device chooses to offload the computing task to the satellite through the gateway in time slot t;

[0033] Without selecting to offload computing tasks to the satellite via a gateway, the task energy consumption of the i-th IoT device in local computing mode is calculated based on the local computing capability of the i-th IoT device, the task latency in local computing mode, and the energy consumption coefficient.

[0034] When choosing to offload computing tasks to a satellite via a gateway, the task energy consumption of the i-th IoT device in satellite computing mode is calculated based on the transmit power of the i-th IoT device in time slot t and the communication delay from the i-th IoT device to the target gateway.

[0035] Alternatively, the utility maximization problem can be constructed by following these steps:

[0036] With the goal of maximizing the average utility of all IoT devices in a single time slot, and under the conditions of satisfying the total satellite computing resource constraints and the maximum transmission power constraints of each IoT device and gateway, the utility maximization problem is established. The utility maximization problem includes the following variables: the first binary decision variable, the second binary decision variable, the transmission power allocation variable, and the computing resource allocation variable.

[0037] Optionally, the convex optimization model is configured to independently solve multiple resource allocation subproblems of the utility maximization problem, the resource allocation subproblems including a first subproblem, a second subproblem, and a third subproblem;

[0038] The method further includes:

[0039] The utility maximization problem is decomposed into: a first sub-problem for optimizing the transmission power from the IoT device to the target gateway, a second sub-problem for optimizing the transmission power from the target gateway to the target satellite, and a third sub-problem for optimizing the allocation of computing resources to the target satellite.

[0040] Optionally, under the premise of the offloading decision scheme and the network access selection scheme, at least one of the following is determined using a convex optimization model: the computing resources allocated by the s-th satellite to the i-th IoT device, the transmit power of the i-th IoT device in time slot t, and the transmit power of the g-th gateway in time slot t, including:

[0041] For the first sub-problem, based on the total amount of task data of all tasks offloaded to the s-th satellite and the satellite resource capacity of the s-th satellite, a convex optimization model is constructed. The computing resources of the s-th satellite are allocated to each IoT device offloaded to the s-th satellite in proportion, and the computing resources allocated by the s-th satellite to the i-th IoT device are determined.

[0042] The second subproblem is modeled as a linear programming problem, and the interior point method is used to solve for the transmission power of the i-th IoT device in time slot t under the constraint of the maximum transmission power of the i-th IoT device.

[0043] The third subproblem is modeled as a convex optimization problem, and the transmission power of the g-th gateway in time slot t is solved using convex optimization tools under the constraint of the maximum transmission power of the g-th gateway.

[0044] Solve the first, second, and third subproblems independently to obtain a closed solution to the first subproblem, and update the solutions to the second and third subproblems iteratively.

[0045] Under the condition of satisfying the preset iteration, the optimal solutions to the second subproblem and the third subproblem are obtained;

[0046] Based on the closed solution of the first subproblem, and the optimal solutions of the second and third subproblems, the computing resources allocated by the s-th satellite to the i-th IoT device, the transmission power of the i-th IoT device in time slot t, and the transmission power of the g-th gateway in time slot t are obtained.

[0047] Secondly, embodiments of the present invention provide a joint optimization device for offloading tasks from IoT devices to satellites, applied to a three-layer offloading system comprising multiple IoT devices, multiple gateways, and multiple satellites. In this three-layer offloading system, a single IoT device is simultaneously covered by multiple gateways, and a single gateway is simultaneously covered by multiple satellites. The device includes:

[0048] The data acquisition module is used to collect current environmental data through multiple IoT devices. The current environmental data includes at least one of the following: image, temperature, and wind speed.

[0049] The acquisition module is used to acquire the observation information of the i-th IoT device in time slot t;

[0050] The input module is used to input the observation information of the i-th IoT device in time slot t into the policy network deployed locally by the i-th IoT device, and generate an offloading decision scheme and a network access selection scheme. The offloading decision scheme represents whether the i-th IoT device offloads the computing task to the satellite through the gateway in time slot t, and the network access selection scheme represents whether the g-th gateway accesses the s-th satellite through the q-th sub-frequency band.

[0051] The convex optimization module is used to determine at least one of the following using a convex optimization model, under the premise of the offloading decision scheme and the network access selection scheme: the amount of computing resources allocated by the s-th satellite to the i-th IoT device, the transmit power of the i-th IoT device in time slot t, and the transmit power of the g-th gateway in time slot t.

[0052] Optionally, the policy network locally deployed by the i-th IoT device is obtained through training a multi-agent deep reinforcement learning model; the device further includes a multi-agent deep reinforcement learning model definition module, wherein the reward function of the multi-agent deep reinforcement learning model includes a device utility reward and a task timeout penalty, the device utility reward being determined based on the task latency and energy consumption of the i-th IoT device in local computing mode, and the task latency and energy consumption in satellite computing mode; the task timeout penalty being determined based on the maximum tolerable latency of the task and the actual completion latency of the task;

[0053] The observation space of the multi-agent deep reinforcement learning model is defined by the observation information of the i-th IoT device in time slot t. The observation information includes: the amount of task data, computation density and maximum tolerable latency of the i-th IoT device in time slot t, the transmit power of the i-th IoT device in time slot t, the transmit power of the g-th gateway in time slot t, the computational resources allocated to the i-th IoT device by the s-th satellite, the channel gain from the g-th gateway to the s-th satellite in the q-th sub-frequency band, and the actual task completion latency of the i-th IoT device in time slot t-1.

[0054] The action space of the multi-agent deep reinforcement learning model is defined by a first binary decision variable, which determines whether the i-th IoT device unloads its computing task to the satellite through the gateway in time slot t, and a second binary decision variable, which determines whether the g-th gateway accesses the s-th satellite through the q-th sub-frequency band.

[0055] Optionally, the device further includes:

[0056] The first judgment module is used to determine, based on the first binary decision variable, whether the i-th IoT device chooses to offload the computing task to the satellite through the gateway in time slot t;

[0057] The first computing module is used to calculate the task latency of the i-th IoT device in local computing mode based on the local computing capability of the i-th IoT device and the task data volume of the un-unloaded computing task when the computing task is not selected to be unloaded to the satellite through the gateway, and to use the task latency of the i-th IoT device in local computing mode as the actual completion latency of the task.

[0058] The second calculation module is used to determine the target gateway to which the i-th IoT device is connected and the target satellite corresponding to the target gateway, based on the second binary decision variable, when the task is selected to be unloaded through the gateway.

[0059] The third calculation module is used to determine the communication delay from the i-th IoT device to the target gateway based on the transmission rate from the i-th IoT device to the target gateway, and to determine the total communication delay from the target gateway to the target satellite based on the transmission rate from the target gateway to the target satellite.

[0060] The fourth calculation module is used to obtain the task delay of the i-th IoT device in satellite computing mode based on the communication delay from the i-th IoT device to the target gateway, the total communication delay from the target gateway to the target satellite, and the calculation processing delay of the computing task offloaded by the target satellite to the i-th IoT device. The task delay of the i-th IoT device in satellite computing mode is then used as the actual completion delay of the task.

[0061] Optionally, the third computing module includes:

[0062] The first calculation submodule is used to determine the transmission delay from the i-th IoT device to the target gateway based on the amount of task data of the calculation task offloaded by the i-th IoT device and the transmission rate of the target gateway.

[0063] The third calculation module also includes:

[0064] The first calculation submodule is used to determine the transmission delay from the target gateway to the target satellite based on the amount of task data of the calculation tasks offloaded by the target gateway and the transmission rate of the target satellite.

[0065] The second calculation submodule is used to determine the propagation delay from the target gateway to the target satellite based on the signal transmission distance and the speed of light between the target gateway and the target satellite;

[0066] The third calculation submodule is used to determine the total communication delay from the target gateway to the target satellite by summing the transmission delay from the target gateway to the target satellite and the propagation delay from the target gateway to the target satellite.

[0067] Optionally, the first computing module includes:

[0068] The fourth calculation submodule is used to obtain the task data volume, calculation density and local computing power of the calculation task of the i-th IoT device that is not unloaded in time slot t, wherein the calculation density represents the number of calculation cycles required per bit of task data.

[0069] The fifth calculation submodule is used to determine the total computational requirements of the task based on the amount of task data and the computational density.

[0070] The sixth calculation submodule is used to divide the total calculation requirement by the local computing capability to obtain the task latency of the i-th IoT device in local computing mode.

[0071] Optionally, the device further includes:

[0072] The second judgment module is used to determine, based on the first binary decision variable, whether the i-th IoT device chooses to offload the computing task to the satellite through the gateway in time slot t;

[0073] The fifth computing module is used to calculate the task energy consumption of the i-th IoT device in local computing mode, based on the local computing capability of the i-th IoT device, the task latency in local computing mode, and the energy consumption coefficient, when the computing task is not selected to be offloaded to the satellite through the gateway.

[0074] The sixth calculation module is used to calculate the task energy consumption of the i-th IoT device in satellite computing mode, based on the transmit power of the i-th IoT device in time slot t and the communication delay from the i-th IoT device to the target gateway, when the computing task is selected to be offloaded to the satellite through the gateway.

[0075] Optionally, the device further includes:

[0076] A module is established to address the utility maximization problem, which aims to maximize the average utility of all IoT devices in a single time slot, while satisfying the constraints of total satellite computing resources and the maximum transmission power constraints of each IoT device and gateway. The utility maximization problem includes the following variables: a first binary decision variable, a second binary decision variable, a transmission power allocation variable, and a computing resource allocation variable.

[0077] Optionally, the convex optimization model is configured to independently solve multiple resource allocation subproblems of the utility maximization problem, the resource allocation subproblems including a first subproblem, a second subproblem, and a third subproblem; the apparatus further includes:

[0078] The decomposition module is used to decompose the utility maximization problem into: a first sub-problem for optimizing the transmission power from the IoT device to the target gateway, a second sub-problem for optimizing the transmission power from the target gateway to the target satellite, and a third sub-problem for optimizing the allocation of computing resources of the target satellite.

[0079] Optionally, the convex optimization module includes:

[0080] The first convex optimization submodule is used to construct a convex optimization model for the first subproblem based on the total amount of task data of all tasks offloaded to the s-th satellite and the satellite resource capacity of the s-th satellite. The model allocates the computing resources of the s-th satellite to each IoT device offloaded to the s-th satellite in proportion, and determines the computing resources allocated by the s-th satellite to the i-th IoT device.

[0081] The second convex optimization submodule is used to model the second subproblem as a linear programming problem and use the interior point method to solve for the transmission power of the i-th IoT device in time slot t under the constraint of the maximum transmission power of the i-th IoT device.

[0082] The third convex optimization submodule is used to model the third subproblem as a convex optimization problem and use convex optimization tools to solve for the transmit power of the g-th gateway in time slot t under the constraint of the maximum transmit power of the g-th gateway.

[0083] The fourth convex optimization submodule is used to independently solve the first subproblem, the second subproblem, and the third subproblem, obtain the closed solution of the first subproblem, and update the solutions of the second subproblem and the third subproblem through iteration;

[0084] The fifth convex optimization submodule is used to obtain the optimal solutions to the second subproblem and the third subproblem under the condition of satisfying the preset iteration conditions;

[0085] The sixth convex optimization submodule is used to obtain the computing resources allocated by the s-th satellite to the i-th IoT device, the transmission power of the i-th IoT device in time slot t, and the transmission power of the g-th gateway in time slot t based on the closed solution of the first subproblem and the optimal solutions of the second and third subproblems.

[0086] The technical solutions provided by the embodiments of the present invention bring at least the following beneficial effects:

[0087] First, this invention effectively solves the task offloading and network access selection problems in complex environments by jointly optimizing a three-layer offloading system involving multiple IoT devices, gateways, and satellites. Through deep reinforcement learning algorithms, it dynamically makes offloading and gateway access selection decisions based on device observation information, avoiding the limitations of traditional static decision-making. It can also adjust strategies according to the real-time environment, ensuring efficient task offloading and rational allocation of computing resources. Second, by optimizing offloading decisions, transmission power, and computing resource allocation, this invention significantly reduces system energy consumption and transmission latency, ensuring minimal energy consumption for IoT devices and gateways while improving task processing response speed. This not only improves device efficiency but also ensures efficient data processing with low energy consumption. Third, this invention adapts to complex environments with multiple IoT devices, gateways, and satellites, exhibiting strong flexibility and adaptability. Regardless of changes in the number of devices, gateways, or satellites or network conditions, it can adjust strategies based on real-time information to ensure stable operation, maximize the use of existing resources, and avoid resource waste or insufficiency. In summary, the joint optimization method of the present invention, by comprehensively considering multiple factors such as task offloading, resource allocation, and power control, not only ensures the energy efficiency and service quality of IoT devices, but also improves the resource utilization of the system. Attached Figure Description

[0088] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0089] Figure 1 is a schematic diagram of the overall concept of a joint optimization method for offloading tasks from an Internet of Things device to a satellite according to an embodiment of the present invention;

[0090] Figure 2 is a schematic diagram of the steps of a joint optimization method for offloading tasks from an Internet of Things (IoT) device to a satellite according to an embodiment of the present invention;

[0091] Figure 3 is a structural block diagram of a joint optimization device for offloading tasks from an Internet of Things (IoT) device to a satellite, according to an embodiment of the present invention. Detailed Implementation

[0092] Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without inventive effort are within the scope of protection of this invention. In the description of the embodiments of this invention, unless otherwise stated, " / " means "or," for example, A / B can mean A or B; "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone. In this invention, "at least one" refers to one or more, and "more" refers to two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of a single item or a plurality of items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.

[0093] Approximately 80% of the world's landmass lacks adequate terrestrial communication and computing infrastructure, leading to difficulties in natural disaster monitoring and severe economic losses. Data acquisition and artificial intelligence (AI) analysis based on IoT devices are feasible solutions, but the computing power, energy constraints, and high deployment costs of IoT devices in remote areas are significant challenges. LEO satellites, due to their wide coverage and low latency, are ideal for mission offloading. However, in related technologies, the coverage area of ​​a single satellite is limited, making it unable to handle the concurrent mission offloading needs of multiple devices. Furthermore, many technologies employ a mode where IoT devices directly offload missions to satellites, requiring high-power transmission from sensor devices, significantly increasing energy consumption and hardware costs. This makes it difficult to minimize system energy consumption and latency in multi-coverage scenarios, thus hindering the maximization of system utility.

[0094] To address this issue, this invention establishes a three-layer communication and computing architecture involving multiple devices, multiple gateways, and multiple satellites. It jointly optimizes task offloading paths, access relationships, transmission power, and computing resources to maximize overall system utility and reduce energy consumption and task latency. Figure 1 is a schematic diagram of the overall concept of a joint optimization method for offloading tasks from IoT devices to satellites, provided by an embodiment of this invention. Referring to Figure 1, this invention aims to maximize system utility and comprehensively considers the following factors: task completion latency, device power consumption, and fairness of computing resource allocation, etc., to construct a joint optimization problem, namely, the utility maximization problem (the main problem P shown in Figure 1, used to maximize utility).

[0095] To address the primary problem (utility maximization), a multi-agent deep reinforcement learning decision-making framework is used in the outer model. Each IoT device is modeled as an agent, independently executing decision-making tasks. Agents generate actions based on local observations, deciding whether to offload tasks and selecting the appropriate access gateway and target satellite. All agents share a centralized value network deployed on the satellite, achieving global collaborative optimization through experience replay and centralized training.

[0096] The mathematical model of the main problem can be constructed as follows:

[0097]

[0098] In the formula, For offloading decisions, it indicates whether IoT devices offload tasks to the gateway; For gateway-satellite access decisions, it characterizes which satellite and sub-frequency band the gateway selects for transmission; Assign transmission power to IoT devices and gateways, representing the transmission power from IoT devices to gateways and from gateways to satellites; Satellite computing resource allocation represents the computing resources that the satellite allocates to each device. It is the utility function of IoT device i in time slot t.

[0099] For the main problem, the following three sub-problems are solved sequentially in the inner model:

[0100] The first subproblem P1: Given the current access and offloading scheme, this invention transforms the computational resource allocation problem into a decomposable convex optimization problem, and solves how much computational resources to allocate to the device's tasks in each time slot in order to minimize offloading latency and energy consumption.

[0101] The second subproblem P2: For the link from the IoT device to the gateway, construct a linear programming problem to minimize the device energy consumption during the task offloading process.

[0102] The third subproblem P3: Further optimize the transmission energy consumption during the process of the gateway uploading data to the satellite, and establish a convex optimization problem.

[0103] By promoting the nested optimization of outer-layer intelligent agent strategy generation and inner-layer three-type resource allocation, this invention effectively improves the timeliness and energy efficiency of task processing, and demonstrates superior performance and scalability in scenarios involving multiple devices and large-scale satellite access.

[0104] This invention is applied to a three-layer offloading system for multiple IoT devices, multiple gateways, and multiple satellites. In the three-layer offloading system, a single IoT device is simultaneously covered by multiple gateways, and a single gateway is simultaneously covered by multiple satellites.

[0105] This invention applies to a three-layer communication and computing offloading architecture, which consists of an Internet of Things (IoT) layer, a relay gateway layer, and a low Earth orbit (LEO) satellite layer. The three layers are connected by a hierarchical design and wireless links to form the entire path for task offloading.

[0106] The IoT device layer consists of numerous sensing or edge computing IoT devices. These devices are deployed in geographically dispersed, remote areas with limited communication infrastructure, and are characterized by energy sensitivity, weak computing power, and limited processing capabilities. IoT devices can perform local processing or offload tasks to upper-layer nodes.

[0107] At the relay gateway layer, the gateway is deployed within the communication radius of IoT devices, possessing strong wireless communication capabilities. It is responsible for receiving task data from multiple devices and further uploading the data to the satellite. The gateway uses low-power wide-area communication technologies (such as LoRa) to establish connections with IoT devices and uses Ka-band or other high-frequency channels to communicate with the satellite.

[0108] At the low Earth orbit (LEO) satellite level, satellites act as high-altitude relay nodes, providing extensive coverage, simultaneously serving multiple gateways, and possessing some edge computing capabilities. Satellites also have computing servers that handle computational tasks from IoT devices and interconnect with ground stations to perform further processing or storage of these tasks.

[0109] For the aforementioned three-layer communication and computation offloading architecture, this embodiment pre-constructs a network model.

[0110] For the IoT device layer, the set of IoT devices is represented as: It includes multiple IoT devices deployed in remote areas, where each IoT device has limited computing power and local computing capabilities. It adjusts dynamically over time, and the maximum transmission power is Transmission power Allocation to be optimized; equipment It can be covered by multiple gateways, and the set of covered gateways is: .

[0111] For the gateway layer, the gateway set is represented as Each gateway receives device data via LoRa technology and communicates with satellites via the Ka band; gateway It can be covered by multiple satellites, and the set of covered satellites is: The gateway's maximum transmission power is Transmission power Allocation needs to be optimized.

[0112] For the satellite layer, the satellite set is represented as A single satellite can cover multiple gateways, forming a dynamic coverage relationship; satellite Provide computing resources in time slot t Assigned to device Its computing power is Multiple satellites achieve load balancing and resource scheduling through inter-satellite links.

[0113] This invention employs a time-slot mechanism to schedule the aforementioned network model.

[0114] Specifically, gateway candidate satellite set Determined by satellite orbital position and coverage time window. The time axis is divided into... Each time slot Within each time slot, an IoT device generates a computing task, and the task model is defined as follows: ;

[0115] in, The amount of data for the task (in bits).

[0116] Calculate the number of cycles required per bit (cycles / bit).

[0117] Maximum tolerable latency for the task (seconds).

[0118] Figure 2 is a schematic diagram of the steps of a joint optimization method for offloading tasks from an IoT device to a satellite according to an embodiment of the present invention. As shown in Figure 2, based on the above modeling, the method includes:

[0119] Step S101: Collect current environmental data through multiple IoT devices. The current environmental data includes at least one of the following: image, temperature, and wind speed.

[0120] In many remote areas, such as mountains, deserts, and forests, inadequate ground communication infrastructure makes real-time environmental data monitoring difficult. Traditional monitoring methods may be limited by geographical environment, cost, and technology. However, by deploying IoT devices, automated and real-time environmental monitoring can be achieved. IoT devices collect current environmental data through equipped sensors or other monitoring modules. Specifically, cameras or imaging sensors can collect images, suitable for visual analysis scenarios in environmental monitoring; temperature sensors can collect ground temperature data, air temperature data, etc.; and wind speed sensors can collect wind speed data in different regions to help assess the availability of wind energy resources. When collecting environmental data, IoT devices face computational demands, primarily involving processing, analyzing, and making decisions based on the collected data. For example, in some application scenarios, image data processing often involves computational tasks such as target recognition, image segmentation, and anomaly detection.

[0121] Step S102: Obtain the observation information of the i-th IoT device in time slot t.

[0122] In time slot t, the observation information of the i-th IoT device It includes the following key parameters to describe the current device status, network environment, and historical performance:

[0123]

[0124] The specific parameters are defined in Table 1.

[0125] Table 1

[0126]

[0127] Observational information serves as input to multi-agent deep reinforcement learning, directly influencing the output of the policy network, including offloading decisions, gateway selection, satellite access, and power adjustment.

[0128] Step S103: Input the observation information of the i-th IoT device in time slot t into the policy network deployed locally by the i-th IoT device to generate an offloading decision scheme and a network access selection scheme. The offloading decision scheme indicates whether the i-th IoT device offloads the computing task to the satellite through the gateway in time slot t, and the network access selection scheme indicates whether the g-th gateway accesses the s-th satellite through the q-th sub-frequency band.

[0129] As previously mentioned, the policy network of the i-th IoT device receives observation information in time slot t, including the data volume, computational density, maximum tolerable latency of the current task, transmit power of the IoT device and gateway, computational resources allocated by the satellite, channel gain, and historical task completion latency. The policy network uses deep reinforcement learning algorithms to extract features and make decision inferences from the input information, weighing the energy consumption and latency differences between local computation and satellite offloading, and dynamically generating the optimal offloading and access strategies.

[0130] The policy network outputs a first binary decision variable to indicate whether the i-th IoT device should offload the computing task through the gateway in time slot t. If the decision is to offload, a target gateway is selected from multiple candidate gateways covering the device for data transmission; if the decision is to process locally, the task is executed by the device itself without interaction with the gateway or satellite.

[0131] If the offloading decision is to transmit via a gateway, the policy network further outputs a second binary decision variable to determine the access relationship between the target gateway and the target satellite. Specifically, the target gateway selects one or more available sub-frequency bands to access the target satellite within its coverage area based on channel quality and satellite load status to ensure efficient and reliable data transmission. Each sub-frequency band is allocated to only one gateway on the same satellite to avoid signal interference between multiple gateways.

[0132] The offloading decision and access selection scheme are generated through the output layer of the policy network, and then the i-th IoT device sends the decision result to the centralized evaluation network at the satellite layer. The decisions of all IoT devices are aggregated at the satellite end to form a global state and joint actions. The parameters of the policy network are optimized through experience replay and centralized training to improve the coordination of subsequent decisions and system effectiveness.

[0133] Through the above steps, the policy network can dynamically generate optimal offloading paths and network access schemes that take into account both energy consumption and latency based on real-time observation information, providing a basis for decision-making in subsequent resource allocation.

[0134] Step S104: Under the premise of the offloading decision scheme and the network access selection scheme, use a convex optimization model to determine at least one of the following: the amount of computing resources allocated by the s-th satellite to the i-th IoT device, the transmission power of the i-th IoT device in time slot t, and the transmission power of the g-th gateway in time slot t.

[0135] In this step, the first optimization objective is to rationally allocate satellite computing resources to minimize task processing latency and improve energy efficiency, given the decision to offload equipment to the satellite.

[0136] Within each time slot, the satellite dynamically allocates computing resources to each device based on the number of currently connected devices and task requirements. Resource allocation must satisfy two constraints: first, the upper limit of the total computing power of a single satellite; and second, the fairness requirement among different devices. Optionally, by constructing a convex optimization problem and employing numerical methods such as gradient descent or interior-point methods, the computing resource quota for each device on the satellite can be determined. The resource allocation result is proportional to the computational density and local processing capacity of the device's tasks, ensuring that high-load tasks receive more resources.

[0137] The final optimization result of the first optimization objective is to dynamically adjust resource allocation based on the computational requirements of each offloading task, prioritizing tasks with near-term latency constraints. The allocation results are distributed to gateways and IoT devices via the satellite computing server to ensure timely task initiation.

[0138] The second optimization objective is to optimize the equipment's transmission power to reduce energy consumption while ensuring communication quality.

[0139] Power allocation is modeled as a linear programming problem, with the objective function being to minimize the total energy consumption of the devices, and the constraint being the maximum transmit power of the IoT devices. Optionally, based on the LoRa communication rate model (related to the spreading factor and bandwidth), the minimum power required to meet the data transmission rate is calculated. The interior-point solver in the Python SciPy library is used to quickly obtain the optimal transmit power value for each device. The power allocation results dynamically adapt to changes in channel conditions, avoiding excessive power consumption.

[0140] The final optimization result of the second optimization objective is: IoT devices adjust their transmission power based on the gateway signal quality, reducing power when channel conditions are good and increasing power when conditions are poor, but not exceeding the hardware limit. The power allocation result is fed back to the IoT devices through the gateway to ensure a balance between data transmission success rate and energy efficiency.

[0141] The third optimization objective is to optimize the gateway transmit power to reduce uplink power consumption while meeting the communication quality requirements of satellite access.

[0142] A convex optimization model is constructed, with the objective function comprehensively considering transmission rate, channel gain, and interference level. Constraints include the gateway's maximum transmit power and the satellite's receive signal-to-noise ratio requirements. Optionally, the CVXPY toolkit is used for modeling and solving, and the optimal power value is determined using the Lagrange multiplier method or iterative optimization algorithm. Orthogonal frequency band allocation avoids interference between multiple gateways, ensuring that the same satellite sub-band serves only a single gateway.

[0143] The optimal result of the third optimization objective is: the gateway dynamically selects available frequency bands (Ka-band sub-bands) based on satellite coverage and monitors satellite link status in real time, adjusting power to maintain a stable connection. The power allocation result is linked to the satellite load status, prioritizing the transmission of critical tasks during high load periods.

[0144] First, this invention effectively solves the task offloading and network access selection problems in complex environments by jointly optimizing a three-layer offloading system involving multiple IoT devices, gateways, and satellites. Through deep reinforcement learning algorithms, it dynamically makes offloading and gateway access selection decisions based on device observation information, avoiding the limitations of traditional static decision-making. It can also adjust strategies according to the real-time environment, ensuring efficient task offloading and rational allocation of computing resources. Second, by optimizing offloading decisions, transmission power, and computing resource allocation, this invention significantly reduces system energy consumption and transmission latency, ensuring minimal energy consumption for IoT devices and gateways while improving task processing response speed. This not only improves device efficiency but also ensures efficient data processing with low energy consumption. Third, this invention adapts to complex environments with multiple IoT devices, gateways, and satellites, exhibiting strong flexibility and adaptability. Regardless of changes in the number of devices, gateways, or satellites or network conditions, it can adjust strategies based on real-time information to ensure stable operation, maximize the use of existing resources, and avoid resource waste or insufficiency. In summary, the joint optimization method of the present invention, by comprehensively considering multiple factors such as task offloading, resource allocation, and power control, not only ensures the energy efficiency and service quality of IoT devices, but also improves the resource utilization of the system.

[0145] In one alternative implementation, the policy network locally deployed by the i-th IoT device is obtained by training a multi-agent deep reinforcement learning model.

[0146] Each IoT device acts as an independent intelligent agent, equipped with a local policy network μ, employing a deep neural network structure. All agents share a centralized evaluation network deployed on a satellite for collaborative optimization of the global policy.

[0147] During the training phase, the agent feeds back local observation information (part of the observation information), actions, and rewards to the satellite to construct a global state and joint action space. It stores samples through an experience replay buffer and updates the network parameters of the policy network and the centralized evaluation network using offline data. The policy network parameters θ are updated using the policy gradient method, while the target network is updated using soft updates (τ is the smoothing coefficient).

[0148] After training, each IoT device independently generates offloading and access decisions based on its local policy network, without the need for real-time communication.

[0149] The reward function of the multi-agent deep reinforcement learning model includes a device utility reward and a task timeout penalty. The device utility reward is determined based on the task latency and energy consumption of the i-th IoT device in local computing mode and in satellite computing mode. The task timeout penalty is determined based on the maximum tolerable latency of the task and the actual completion latency of the task.

[0150] The reward function involved in training the policy network combines device utility rewards and task timeout penalties, and is designed as follows:

[0151] The reward function is determined based on the performance difference between local computing and satellite computing. The following formula is the utility function that measures the performance difference between local computing and satellite computing for the i-th IoT device:

[0152]

[0153] in: Let represent the latency and energy consumption of the i-th IoT device in local computing mode, respectively.

[0154] Latency and energy consumption of the i-th IoT device in satellite computing mode.

[0155] The weighting coefficients for latency and energy consumption are respectively, satisfying... ,and ∈[0,1].

[0156] The penalty mechanism of the reward function is as follows: a linear penalty is applied to tasks that exceed the maximum tolerable latency, and the penalty term is...

[0157]

[0158] Where δ is the penalty coefficient. When the penalty is triggered, .

[0159] Therefore, the total reward function is defined as:

[0160] The observation space is constructed based on partial observation information, which has been explained in detail above and will not be repeated here.

[0161] The action space is jointly defined by two levels of binary decision variables, that is, it is determined by the first binary decision variable and the second binary decision variable.

[0162] The first binary decision variable is represented as follows: , that is, It takes a value between the binary values ​​1 and 0.

[0163] When the first binary decision variable takes the value 1 (indicating the existence of one and only one target gateway g corresponding to IoT device i that can be accessed by IoT device i), that is:

[0164]

[0165] At this point, the first binary decision variable represents selecting a single target gateway (the g-th gateway) for task offloading within the set of gateways covering the i-th IoT device.

[0166] When the first binary decision variable takes the value of 0, that is... At this point, the first binary decision variable is represented by local computation.

[0167] The second-level decision is the choice between gateway and satellite access, and the second binary decision variable is represented as follows:

[0168] Each sub-band of the same satellite Only one gateway is allowed to access, that is:

[0169]

[0170] Based on the first binary decision variable and the second binary decision variable, the joint action vector representing the final decision action is as follows:

[0171]

[0172] exist In this case, it means that IoT device i uses a sub-frequency band through a certain gateway. Access satellite Unload data.

[0173] exist In this case, it indicates that the IoT device i is performing the task locally.

[0174] It is understandable that this embodiment involves the idea of ​​multiple summations, which can be understood as performing traversal summations on the device set, gateway set, satellite set, or sub-band set to determine an optimal task execution method in order to achieve comprehensive optimization of the overall system utility.

[0175] The observation space of the multi-agent deep reinforcement learning model is defined by the observation information of the i-th IoT device in time slot t. The observation information includes: the amount of task data, computation density and maximum tolerable latency of the i-th IoT device in time slot t, the transmit power of the i-th IoT device in time slot t, the transmit power of the g-th gateway in time slot t, the computational resources allocated to the i-th IoT device by the s-th satellite, the channel gain from the g-th gateway to the s-th satellite in the q-th sub-frequency band, and the actual task completion latency of the i-th IoT device in time slot t-1.

[0176] Specifically, the task data volume is the size (in bits) of the data collected by the i-th IoT device in time slot t.

[0177] Computational density represents the computational resources (such as CPU cycles) required per bit of a task, while maximum tolerable latency is the maximum acceptable task completion latency for IoT device i. Computational density and maximum tolerable latency together determine device i's requirements and tolerance for task processing.

[0178] The transmit power is the transmit power of the i-th IoT device in time slot t, and the transmit power of the g-th gateway in time slot t.

[0179] Computing resource allocation refers to the computing resources (e.g., the CPU capacity of the satellite for processing tasks) that the s-th satellite allocates to the i-th IoT device in time slot t.

[0180] Channel gain is the channel gain from the g-th gateway to the s-th satellite in the q-th sub-band. Channel gain reflects the quality of wireless communication.

[0181] The task completion delay is the actual completion delay of the task of the i-th IoT device in the previous time slot (t-1), reflecting the actual performance of task unloading and processing in the past.

[0182] By constructing a policy network and the model required for its pre-training, the problem of task offloading and resource allocation in multi-satellite and multi-gateway scenarios is effectively solved. When the policy network is used to make offloading decisions, the system latency and device power consumption can be significantly reduced.

[0183] In one alternative implementation, the task latency of the i-th IoT device in local computing mode and the task latency in satellite computing mode are determined by the following steps:

[0184] Step S201: Based on the first binary decision variable, determine whether the i-th IoT device chooses to offload the computing task to the satellite through the gateway in time slot t.

[0185] Based on the aforementioned first binary decision variable, it is determined whether the i-th IoT device chooses to offload the computing task to the satellite via the gateway in time slot t. If the first binary decision variable indicates "do not offload", the task is processed locally by the device; if it indicates "offload", the device enters satellite computing mode and determines the target gateway to be uploaded to and the target satellite to which the target gateway needs to access.

[0186] Step S202: If the computing task is not selected to be offloaded to the satellite through the gateway, the task latency of the i-th IoT device in local computing mode is calculated based on the local computing capability of the i-th IoT device and the task data volume of the un-offloaded computing task, and the task latency of the i-th IoT device in local computing mode is used as the actual completion latency of the task.

[0187] Without selecting unloading, the local task latency is calculated based on the following parameters: task data size, the size of the task data generated by device i in time slot t (bits); computation density, the number of computation cycles required per bit of task data (cycles / bit); local computing power, the available computing power of device i in time slot t (cycles / second).

[0188] Specifically, the total computational requirement of the task is obtained by multiplying the task data volume by the computational density. Dividing the total computational requirement by the local computing capacity yields the task latency in the local computing mode. This local computing mode latency is directly used as the actual task completion latency.

[0189] The formula for calculating local latency is as follows:

[0190]

[0191] The amount of data for the task (in bits); Calculate density; For device i's local computing capabilities; To unload the decision factor, the value is set to 1 during local calculation.

[0192] Step S203: If the task is selected to be unloaded through the gateway, the target gateway to which the i-th IoT device is connected and the target satellite corresponding to the target gateway are determined according to the second binary decision variable.

[0193] In the case of selecting offloading, the target gateway and target satellite are determined according to the second binary decision variable. As mentioned above, the target gateway g is the gateway selected by device i, chosen from multiple candidate gateways covering the device; the target satellite s is the satellite corresponding to the target gateway, chosen from multiple candidate satellites covering the gateway. The target gateway establishes a connection with the target satellite through a specified sub-frequency band q to ensure no signal interference.

[0194] Step S204: Determine the communication delay from the i-th IoT device to the target gateway based on the transmission rate from the i-th IoT device to the target gateway, and determine the total communication delay from the target gateway to the target satellite based on the transmission rate from the target gateway to the target satellite.

[0195] Based on the transmission rate from device i to the target gateway (determined by parameters such as bandwidth, spreading factor, and coding rate), the transmission delay from device i to gateway is obtained by dividing the task data volume by the transmission rate. Since the distance between device i and gateway is short, the propagation delay can be ignored; therefore, the total communication delay is the transmission delay.

[0196] The communication latency from the IoT device to the gateway is shown in the following formula:

[0197]

[0198] The formula for transmission rate is:

[0199]

[0200] In the formula, For bandwidth; For equipment With gateway Spreading factor; This represents the coding rate.

[0201] The total communication delay from the gateway to the satellite is shown in the following formula:

[0202]

[0203] The transmission delay is:

[0204]

[0205] The formula for transmission rate is:

[0206]

[0207] In the formula, For Ka-band sub-band bandwidth; This is the signal-to-interference-to-noise ratio.

[0208] The propagation delay is:

[0209]

[0210] In the formula, Let S be the orbital altitude of satellite S and gateway G; It is the speed of light.

[0211] Step S205: Based on the communication delay from the i-th IoT device to the target gateway and the total communication delay from the target gateway to the target satellite, the computation processing delay of the computing task offloaded by the target satellite to the i-th IoT device is added to obtain the task delay of the i-th IoT device in satellite computing mode, and the task delay of the i-th IoT device in satellite computing mode is used as the actual completion delay of the task.

[0212] The communication delay from gateway to satellite includes transmission delay and propagation delay. Transmission delay is obtained by dividing the amount of mission data by the transmission rate from the target gateway to the target satellite. Propagation delay is calculated based on the round-trip distance between the target satellite and the target gateway (dynamically calculated from the satellite's orbital altitude and position) and the speed of light.

[0213] Adding the transmission delay to the propagation delay gives the total communication delay from the gateway to the satellite.

[0214] The formula for the total time delay in satellite computing mode is as follows:

[0215]

[0216] The satellite computing and processing latency is as follows:

[0217]

[0218] In the formula, Indicates whether device i processes the task via satellite s; The computing resources allocated to device i for satellite s.

[0219] The observation space of the multi-agent deep reinforcement learning model is defined by the observation information of the i-th IoT device in time slot t. The observation information includes: the amount of task data, computation density and maximum tolerable latency of the i-th IoT device in time slot t, the transmit power of the i-th IoT device in time slot t, the transmit power of the g-th gateway in time slot t, the computational resources allocated to the i-th IoT device by the s-th satellite, the channel gain from the g-th gateway to the s-th satellite in the q-th sub-frequency band, and the actual task completion latency of the i-th IoT device in time slot t-1.

[0220] After receiving the mission data, the target satellite multiplies the amount of mission data by the computational density based on the computing resources allocated to device i to obtain the total computational requirement, and then divides it by the allocated computing resources to obtain the satellite's computational processing delay.

[0221] The action space of the multi-agent deep reinforcement learning model is defined by a first binary decision variable, which determines whether the i-th IoT device unloads its computing task to the satellite through the gateway in time slot t, and a second binary decision variable, which determines whether the g-th gateway accesses the s-th satellite through the q-th sub-frequency band.

[0222] The total task latency in satellite computing mode is obtained by superimposing the communication latency from the device to the gateway, the communication latency from the gateway to the satellite, and the satellite computing processing latency, and is used as the actual task completion latency.

[0223] Through the above steps, the present invention dynamically selects the optimal unloading path and accurately calculates the latency based on task attributes and network status, thereby achieving a balance between latency and energy consumption between local processing and satellite unloading, ultimately improving the task processing efficiency of IoT devices in remote areas and the overall system effectiveness.

[0224] In one optional implementation, determining the communication delay from the i-th IoT device to the target gateway based on the transmission rate from the i-th IoT device to the target gateway includes:

[0225] Based on the amount of task data of the computing task offloaded by the i-th IoT device and the transmission rate of the target gateway, the transmission delay from the i-th IoT device to the target gateway is determined.

[0226] Based on the communication link parameters between the device and the target gateway (including bandwidth, spreading factor, coding rate, etc.), the maximum effective transmission rate of the current link is determined. The total amount of data from the computational task to be offloaded is divided by the transmission rate of the target gateway to obtain the communication delay from the device to the gateway. This process only involves the data transmission time in the channel; since the distance between the device and the gateway is relatively short, the propagation delay between them is negligible.

[0227] Based on the transmission rate from the target gateway to the target satellite, the total communication delay from the target gateway to the target satellite is determined, including:

[0228] Step S301: Based on the amount of task data of the computing tasks offloaded by the target gateway and the transmission rate of the target satellite, determine the transmission delay from the target gateway to the target satellite.

[0229] Based on the total amount of computational task data to be forwarded to the satellite by the gateway (consistent with the original task data amount of the device), and combined with the current transmission rate of the satellite link (determined by factors such as wireless channel quality and modulation method), the transmission delay is calculated. The transmission delay reflects the actual time spent transmitting data in the satellite communication channel.

[0230] Step S302: Determine the propagation delay from the target gateway to the target satellite based on the signal transmission distance and the speed of light between the target gateway and the target satellite.

[0231] The spatial straight-line distance between the target gateway and the target satellite is obtained in real time (affected by dynamic changes in satellite orbital altitude and position). The distance value is multiplied by 2 (considering the uplink and downlink round-trip paths of the signal) and then divided by the speed of light to obtain the propagation delay from the target gateway to the target satellite.

[0232] Step S303: The sum of the transmission delay from the target gateway to the target satellite and the propagation delay from the target gateway to the target satellite is determined as the total communication delay from the target gateway to the target satellite.

[0233] The transmission delay calculated in step S301 is added to the propagation delay calculated in step S302 to obtain the total communication delay from the gateway to the satellite, which reflects the overall time consumption from the start of data transmission to the completion of satellite reception.

[0234] In one optional implementation, calculating the task latency of the i-th IoT device in local computing mode based on the local computing power of the i-th IoT device and the amount of task data of the unloaded computing tasks includes:

[0235] Step S401: Obtain the task data volume, computation density, and local computing power of the computation task that is not unloaded in time slot t for the i-th IoT device. The computation density represents the number of computation cycles required per bit of task data.

[0236] The amount of unloaded task data refers to the total amount of task data that IoT devices do not choose to offload to satellites within time slot t and need to be processed locally. It is determined by the size of sensor data (such as images, temperature, etc.) collected by IoT devices.

[0237] Computational density represents the number of computation cycles required per bit of task data, and is determined by the task type and algorithm complexity (for example, the computational density of image analysis tasks is higher than that of temperature monitoring tasks).

[0238] Local computing power refers to the local processor performance of an IoT device within time slot t, measured by the number of computing cycles that can be completed per unit time (cycles / second), and is affected by the device's hardware performance and current load status.

[0239] Step S402: Determine the total computational requirements of the task based on the amount of task data and computational density.

[0240] Multiply the task data volume by the computation density to obtain the total number of computation cycles required to complete the task.

[0241] Step S403: Divide the total computing requirements by the local computing power to obtain the task latency of the i-th IoT device in local computing mode.

[0242] Divide the total computing demand by the local computing capacity to get the latency of the device processing the task locally. For example, assuming the total computing demand is 500,000 CPU cycles and the local computing capacity is 1,000 cycles / second, the latency is 500,000 ÷ 1,000 = 500 seconds.

[0243] According to the definition of the reward function, the delay calculation result must meet the maximum tolerable delay constraint of the task; otherwise, the task timeout penalty mechanism will be triggered.

[0244] In one alternative implementation, the task power consumption of the i-th IoT device in local computing mode and in satellite computing mode are determined by the following steps:

[0245] Step S501: Based on the first binary decision variable, determine whether the i-th IoT device chooses to offload the computing task to the satellite through the gateway in time slot t.

[0246] The implementation method of step S501 is similar to that of step S201, and will not be described again here.

[0247] Step S502: If the computing task is not offloaded to the satellite through the gateway, calculate the task energy consumption of the i-th IoT device in the local computing mode based on the local computing capability of the i-th IoT device, the task latency in the local computing mode, and the energy consumption coefficient.

[0248] Local computing power consumption is: In the formula, The energy efficiency coefficient of the chip architecture; The local computing power of device i in time slot t; Calculate latency locally.

[0249] Calculated This represents the task energy consumption of the i-th IoT device in local computing mode.

[0250] Step S503: If the computing task is selected to be offloaded to the satellite via the gateway, the task energy consumption of the i-th IoT device in satellite computing mode is calculated based on the transmit power of the i-th IoT device in time slot t and the communication delay from the i-th IoT device to the target gateway.

[0251]

[0252] In satellite computing mode, the energy consumption of IoT devices only includes the communication energy consumption for transmitting data to the gateway:

[0253] In the formula, Let i be the transmit power of IoT device i in time slot t. The communication latency from IoT device i to the target gateway.

[0254] In one alternative implementation, the utility maximization problem is constructed according to the following steps:

[0255] With the goal of maximizing the average utility of all IoT devices in a single time slot, and under the conditions of satisfying the total satellite computing resource constraints and the maximum transmission power constraints of each IoT device and gateway, the utility maximization problem is established. The utility maximization problem includes the following variables: the first binary decision variable, the second binary decision variable, the transmission power allocation variable, and the computing resource allocation variable.

[0256] To maximize the average utility of all IoT devices within a single time slot, a joint optimization problem (i.e., a utility maximization problem) is constructed. The utility comprehensively considers the reduction of task processing latency and the saving of device energy consumption. As mentioned earlier, the utility function is defined as a weighted combination of latency gain and energy consumption gain. The latency gain is the difference between local computation latency and satellite offload latency, reflecting the contribution of latency optimization; the energy consumption gain is the difference between local computation energy consumption and satellite offload energy consumption, reflecting the contribution of energy consumption optimization. The two weighting coefficients represent the importance of latency and energy consumption, respectively.

[0257] The average utility is obtained by summing the utility of all devices and dividing it by the total number of devices and the total number of time slots. This average utility is then used as the optimization objective.

[0258] The utility maximization problem requires joint optimization of the following variables: the first binary decision variable, the second binary decision variable, the transmission power allocation variable (including the transmission power of device i and the transmission power of gateway g), and the computing resource allocation variable (representing the computing power allocated to device i by satellite s).

[0259] The optimization problem must meet the following hard constraints:

[0260] First, the total computing resources allocated to all devices by satellite s must not exceed its maximum computing capacity; second, the transmission power of device i must not exceed its maximum allowable power; third, the transmission power of gateway g must not exceed its maximum allowable power; fourth, the actual completion delay of the mission must not exceed its maximum tolerable delay.

[0261] The above objectives and constraints are integrated into a mathematical optimization model, which achieves the global optimal allocation of task offloading decisions, network access selection, power and computing resources by maximizing average utility.

[0262] In one alternative implementation, the convex optimization model is configured to independently solve multiple resource allocation subproblems of the utility maximization problem, the resource allocation subproblems including a first subproblem, a second subproblem, and a third subproblem.

[0263] The method further includes:

[0264] The utility maximization problem is decomposed into: a first sub-problem for optimizing the transmission power from the IoT device to the target gateway, a second sub-problem for optimizing the transmission power from the target gateway to the target satellite, and a third sub-problem for optimizing the allocation of computing resources to the target satellite.

[0265] Due to the high dimensionality and nonlinearity of the utility maximization problem, direct solutions are extremely complex. This invention employs a hierarchical optimization strategy, decomposing the main problem into three independently solvable convex optimization subproblems: the first subproblem (satellite computing resource allocation P1), the second subproblem (device-to-gateway transmission power optimization P2), and the third subproblem (gateway-to-satellite transmission power optimization P3).

[0266] As shown in Figure 1, under the premise of fixed decisions, subproblems P1, P2, and P3 are solved sequentially to update the power and resource allocation variables. Furthermore, the resource allocation results are fed back to the outer model to adjust subsequent decisions, forming a closed-loop optimization.

[0267] In one optional implementation, under the premise of the offloading decision scheme and the network access selection scheme, at least one of the following is determined using a convex optimization model: the computing resources allocated by the s-th satellite to the i-th IoT device, the transmit power of the i-th IoT device in time slot t, and the transmit power of the g-th gateway in time slot t, including:

[0268] Step S601: Model the first subproblem as a linear programming problem, and use the interior point method to solve for the transmission power of the i-th IoT device in time slot t under the constraint of the maximum transmission power of the i-th IoT device.

[0269] The goal of the first sub-problem is to allocate computing resources to each IoT device according to the principle of fairness, under the constraint of total satellite resources.

[0270] The first subproblem (P1) is modeled as a convex optimization problem, with the following mathematical form:

[0271]

[0272] The constraint of the first subproblem is that the computing resources allocated to each device must be greater than zero, that is:

[0273]

[0274] And the total computing resources of the satellite do not exceed its maximum capacity, that is:

[0275]

[0276] The first subproblem is solved using the proportional allocation formula, and the optimal resource allocation formula is derived using the Lagrange multiplier method:

[0277]

[0278] The purpose of solving the first subproblem using the above method is to ensure that resource allocation is proportional to the square root of the equipment's computational needs and weight coefficients, thereby ensuring that high-demand equipment receives more resources while avoiding resource monopoly.

[0279] In this embodiment, the convex optimization model simplifies the solution of the resource allocation problem, while the proportional allocation takes into account both efficiency and fairness.

[0280] Step S602: Model the second subproblem as a convex optimization problem, and use a convex optimization tool to solve for the transmit power of the g-th gateway in time slot t under the constraint of the maximum transmit power of the g-th gateway.

[0281] The goal of the second subproblem is to minimize the energy consumption of the device due to task offloading while satisfying the device's maximum transmit power constraint.

[0282] The second subproblem (P2) is constructed as a linear programming problem, with the following mathematical form:

[0283]

[0284] The constraint of the second subproblem is that the transmission power of IoT devices must not exceed their maximum allowable value, that is:

[0285]

[0286] By introducing a constraint function to incorporate the constraints into the objective function, the problem is transformed into an unconstrained optimization problem, and then the optimal solution is approximated through iteration. The interior-point solver in Python's SciPy library can be used, leveraging its efficient handling of linear programming problems to quickly obtain the optimal power allocation. .

[0287] In this embodiment, the linear programming model guarantees the global optimality of the solution, while the interior point method has high computational efficiency when dealing with large-scale constraints.

[0288] Step S603: For the third sub-problem, based on the total amount of task data of all tasks offloaded to the s-th satellite and the satellite resource capacity of the s-th satellite, a convex optimization model is constructed. The computing resources of the s-th satellite are allocated proportionally to each IoT device offloaded to the s-th satellite, and the computing resources allocated by the s-th satellite to the i-th IoT device are determined.

[0289] The third subproblem (P3) is modeled as a convex optimization problem, with the following mathematical form:

[0290]

[0291]

[0292] The constraint of the third subproblem is: the gateway's transmit power must not exceed its maximum allowable value, that is:

[0293]

[0294] By leveraging CVXPY's convex optimization modeling capabilities, the problem can be transformed into a standard convex form, and then the embedded solver (such as ECOS or SCS) can be called for numerical solution.

[0295] In this embodiment, convex optimization guarantees the existence of a unique globally optimal solution to the problem, and the CVXPY tool simplifies the implementation process of complex models.

[0296] Step S604: Solve the first subproblem, the second subproblem, and the third subproblem independently to obtain the closed solution of the first subproblem, and update the solutions of the second subproblem and the third subproblem iteratively.

[0297] After the outer model (multi-agent deep reinforcement learning) determines the offloading and access decisions, the inner model solves subproblems P1, P2, and P3 independently in sequence. Each subproblem is optimized independently while keeping other variables fixed, reducing the overall complexity. The solution to the current subproblem is fed back to the outer model to update the training data of the policy network. The outer model adjusts subsequent offloading decisions based on the feedback, forming a closed-loop optimization process of decision-making, resource allocation, and re-decision-making. In this embodiment, the objective function of the first subproblem is a convex function, and the constraints are linear constraints, making it a convex optimization problem overall. Therefore, the first subproblem can be solved directly using the Lagrange multiplier method to obtain a closed-loop solution. The second and third subproblems, the former being a linear programming problem requiring interior-point solution (requiring iterative solution) and the latter a convex optimization problem requiring numerical solution (also requiring iterative solution), are different.

[0298] Step S605: Under the condition of satisfying the preset iteration conditions, the optimal solutions to the second subproblem and the third subproblem are obtained.

[0299] Define the threshold for the change of the objective function If the difference between the objective functions of two consecutive iterations is less than If so, then it is considered convergent.

[0300] For example: ,in For the first Average utility over several iterations.

[0301] Set the maximum number of iterations To prevent infinite loops, for either the second or third subproblem, when any termination condition is met, output the current optimal solution to that subproblem: device transmit power. or gateway transmit power .

[0302] Step S606: Based on the closed solution of the first subproblem, and the optimal solutions of the second and third subproblems, obtain the computing resources allocated by the s-th satellite to the i-th IoT device, the transmission power of the i-th IoT device in time slot t, and the transmission power of the g-th gateway in time slot t.

[0303] Integrate the closed-form solution of the first subproblem, as well as the optimal solutions of the second and third subproblems, and extract the solution from the solution of P1. Distribute computing power proportionally and fairly, and extract it from the solution of P2. To ensure that the equipment completes data transmission while maintaining energy efficiency, and to extract data from the solution of P3. Optimize gateway power consumption and avoid channel interference.

[0304] Through collaborative optimization, the overall three-layer architecture achieves a balance between latency, energy consumption, and resource utilization, ultimately realizing the effectiveness of IoT task processing in remote areas.

[0305] Figure 3 is a structural block diagram of a joint optimization device for offloading tasks from IoT devices to satellites according to an embodiment of the present invention. As shown in Figure 3, it is applied to a three-layer offloading system involving multiple IoT devices, multiple gateways, and multiple satellites. In the three-layer offloading system, a single IoT device is simultaneously covered by multiple gateways, and a single gateway is simultaneously covered by multiple satellites. The device includes:

[0306] The data acquisition module 701 is used to collect current environmental data through multiple IoT devices. The current environmental data includes at least one of the following: image, temperature, and wind speed.

[0307] The acquisition module 702 is used to acquire the observation information of the i-th IoT device in time slot t;

[0308] The input module 703 is used to input the observation information of the i-th IoT device in time slot t into the policy network deployed locally by the i-th IoT device, and generate an offloading decision scheme and a network access selection scheme. The offloading decision scheme indicates whether the i-th IoT device offloads the computing task to the satellite through the gateway in time slot t, and the network access selection scheme indicates whether the g-th gateway accesses the s-th satellite through the q-th sub-frequency band.

[0309] The convex optimization module 704 is used to determine at least one of the following using a convex optimization model, under the premise of the offloading decision scheme and the network access selection scheme: the amount of computing resources allocated by the s-th satellite to the i-th IoT device, the transmission power of the i-th IoT device in time slot t, and the transmission power of the g-th gateway in time slot t.

[0310] In one optional implementation, the policy network locally deployed by the i-th IoT device is obtained through training a multi-agent deep reinforcement learning model; the device further includes a multi-agent deep reinforcement learning model definition module, wherein the reward function of the multi-agent deep reinforcement learning model includes a device utility reward and a task timeout penalty, the device utility reward being determined based on the task latency and energy consumption of the i-th IoT device in local computing mode, and the task latency and energy consumption in satellite computing mode; the task timeout penalty being determined based on the maximum tolerable latency of the task and the actual completion latency of the task;

[0311] The observation space of the multi-agent deep reinforcement learning model is defined by the observation information of the i-th IoT device in time slot t. The observation information includes: the amount of task data, computation density and maximum tolerable latency of the i-th IoT device in time slot t, the transmit power of the i-th IoT device in time slot t, the transmit power of the g-th gateway in time slot t, the computational resources allocated to the i-th IoT device by the s-th satellite, the channel gain from the g-th gateway to the s-th satellite in the q-th sub-frequency band, and the actual task completion latency of the i-th IoT device in time slot t-1.

[0312] The action space of the multi-agent deep reinforcement learning model is defined by a first binary decision variable, which determines whether the i-th IoT device unloads its computing task to the satellite through the gateway in time slot t, and a second binary decision variable, which determines whether the g-th gateway accesses the s-th satellite through the q-th sub-frequency band.

[0313] In one alternative embodiment, the device further includes:

[0314] The first judgment module is used to determine, based on the first binary decision variable, whether the i-th IoT device chooses to offload the computing task to the satellite through the gateway in time slot t;

[0315] The first computing module is used to calculate the task latency of the i-th IoT device in local computing mode based on the local computing capability of the i-th IoT device and the task data volume of the un-unloaded computing task when the computing task is not selected to be unloaded to the satellite through the gateway, and to use the task latency of the i-th IoT device in local computing mode as the actual completion latency of the task.

[0316] The second calculation module is used to determine the target gateway to which the i-th IoT device is connected and the target satellite corresponding to the target gateway, based on the second binary decision variable, when the task is selected to be unloaded through the gateway.

[0317] The third calculation module is used to determine the communication delay from the i-th IoT device to the target gateway based on the transmission rate from the i-th IoT device to the target gateway, and to determine the total communication delay from the target gateway to the target satellite based on the transmission rate from the target gateway to the target satellite.

[0318] The fourth calculation module is used to obtain the task delay of the i-th IoT device in satellite computing mode based on the communication delay from the i-th IoT device to the target gateway, the total communication delay from the target gateway to the target satellite, and the calculation processing delay of the computing task offloaded by the target satellite to the i-th IoT device. The task delay of the i-th IoT device in satellite computing mode is then used as the actual completion delay of the task.

[0319] In one optional implementation, the third computing module includes:

[0320] The first calculation submodule is used to determine the transmission delay from the i-th IoT device to the target gateway based on the amount of task data of the calculation task offloaded by the i-th IoT device and the transmission rate of the target gateway.

[0321] The third calculation module also includes:

[0322] The first calculation submodule is used to determine the transmission delay from the target gateway to the target satellite based on the amount of task data of the calculation tasks offloaded by the target gateway and the transmission rate of the target satellite.

[0323] The second calculation submodule is used to determine the propagation delay from the target gateway to the target satellite based on the signal transmission distance and the speed of light between the target gateway and the target satellite;

[0324] The third calculation submodule is used to determine the total communication delay from the target gateway to the target satellite by summing the transmission delay from the target gateway to the target satellite and the propagation delay from the target gateway to the target satellite.

[0325] In one optional implementation, the first computing module includes:

[0326] The fourth calculation submodule is used to obtain the task data volume, calculation density and local computing power of the calculation task of the i-th IoT device that is not unloaded in time slot t, wherein the calculation density represents the number of calculation cycles required per bit of task data.

[0327] The fifth calculation submodule is used to determine the total computational requirements of the task based on the amount of task data and the computational density.

[0328] The sixth calculation submodule is used to divide the total calculation requirement by the local computing capability to obtain the task latency of the i-th IoT device in local computing mode.

[0329] In one alternative embodiment, the device further includes:

[0330] The second judgment module is used to determine, based on the first binary decision variable, whether the i-th IoT device chooses to offload the computing task to the satellite through the gateway in time slot t;

[0331] The fifth computing module is used to calculate the task energy consumption of the i-th IoT device in local computing mode, based on the local computing capability of the i-th IoT device, the task latency in local computing mode, and the energy consumption coefficient, when the computing task is not selected to be offloaded to the satellite through the gateway.

[0332] The sixth calculation module is used to calculate the task energy consumption of the i-th IoT device in satellite computing mode, based on the transmit power of the i-th IoT device in time slot t and the communication delay from the i-th IoT device to the target gateway, when the computing task is selected to be offloaded to the satellite through the gateway.

[0333] In one alternative embodiment, the device further includes:

[0334] A module is established to address the utility maximization problem, which aims to maximize the average utility of all IoT devices in a single time slot, while satisfying the constraints of total satellite computing resources and the maximum transmission power constraints of each IoT device and gateway. The utility maximization problem includes the following variables: a first binary decision variable, a second binary decision variable, a transmission power allocation variable, and a computing resource allocation variable.

[0335] In one optional implementation, the convex optimization model is configured to independently solve multiple resource allocation subproblems of the utility maximization problem, the resource allocation subproblems including a first subproblem, a second subproblem, and a third subproblem; the apparatus further includes:

[0336] The decomposition module is used to decompose the utility maximization problem into: a first sub-problem for optimizing the transmission power from the IoT device to the target gateway, a second sub-problem for optimizing the transmission power from the target gateway to the target satellite, and a third sub-problem for optimizing the allocation of computing resources of the target satellite.

[0337] In one optional implementation, the convex optimization module includes:

[0338] The first convex optimization submodule is used to construct a convex optimization model for the first subproblem based on the total amount of task data of all tasks offloaded to the s-th satellite and the satellite resource capacity of the s-th satellite. The model allocates the computing resources of the s-th satellite to each IoT device offloaded to the s-th satellite in proportion, and determines the computing resources allocated by the s-th satellite to the i-th IoT device.

[0339] The second convex optimization submodule is used to model the second subproblem as a linear programming problem and use the interior point method to solve for the transmission power of the i-th IoT device in time slot t under the constraint of the maximum transmission power of the i-th IoT device.

[0340] The third convex optimization submodule is used to model the third subproblem as a convex optimization problem and use convex optimization tools to solve for the transmit power of the g-th gateway in time slot t under the constraint of the maximum transmit power of the g-th gateway.

[0341] The fourth convex optimization submodule is used to independently solve the first subproblem, the second subproblem, and the third subproblem, obtain the closed solution of the first subproblem, and update the solutions of the second subproblem and the third subproblem through iteration;

[0342] The fifth convex optimization submodule is used to obtain the optimal solutions to the second subproblem and the third subproblem under the condition of satisfying the preset iteration conditions;

[0343] The sixth convex optimization submodule is used to obtain the computing resources allocated by the s-th satellite to the i-th IoT device, the transmission power of the i-th IoT device in time slot t, and the transmission power of the g-th gateway in time slot t based on the closed solution of the first subproblem and the optimal solutions of the second and third subproblems.

[0344] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, apparatus, electronic devices, and storage media. Therefore, embodiments of the present invention can take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware aspects. Furthermore, embodiments of the present invention can take the form of a computer program product embodied on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0345] Embodiments of the present invention are described with reference to flowchart illustrations and / or block diagrams of methods and apparatus according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, create means for implementing the functions specified in one or more flowchart illustrations and / or one or more block diagrams. These computer program instructions may also be stored in a computer-readable storage medium capable of directing a computer or other programmable data processing terminal device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means that implement the functions specified in one or more flowchart illustrations and / or one or more block diagrams. These computer program instructions may also be loaded onto a computer or other programmable data processing terminal equipment to cause a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, thereby providing steps for implementing the functions specified in one or more flowcharts and / or one or more block diagrams.

[0346] Although preferred embodiments of the present invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present invention.

[0347] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the term "comprising" or any other variations thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element. The above provides a detailed description of a joint optimization method and apparatus for offloading tasks from IoT devices to satellites provided by the present invention. Specific examples have been used to illustrate the principles and implementation methods of the invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention; at the same time, for those skilled in the art, based on the ideas of the present invention, there will be changes in specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A joint optimization method for offloading tasks from IoT devices to satellites, characterized in that, A three-layer offloading system is applied to multiple IoT devices, multiple gateways, and multiple satellites. In this system, a single IoT device is simultaneously covered by multiple gateways, and a single gateway is simultaneously covered by multiple satellites. The method includes: collecting current environmental data through multiple IoT devices, where the current environmental data includes at least one of the following: image, temperature, and wind speed; acquiring the observation information of the i-th IoT device in time slot t; inputting the observation information of the i-th IoT device in time slot t into the policy network locally deployed by the i-th IoT device to generate an offloading decision scheme and a network access selection scheme, wherein the offloading decision scheme characterizes the... The system determines whether the i-th IoT device offloads its computing task to the satellite via a gateway in time slot t, and whether the g-th gateway accesses the s-th satellite via the q-th sub-frequency band. The policy network locally deployed by the i-th IoT device is trained using a multi-agent deep reinforcement learning model. This model, as an outer layer, models each IoT device as an agent, independently generating the offloading decision scheme and the network access selection scheme. The observation information serves as the input to the multi-agent deep reinforcement learning model, adjusting the output of the policy network in real time. The action space of the multi-agent deep reinforcement learning model is defined by a first binary decision variable regarding whether the i-th IoT device offloads its computing task to the satellite via a gateway in time slot t, and a second binary decision variable regarding whether the g-th gateway accesses the s-th satellite via the q-th sub-band. Under the premise of the offloading decision scheme and the network access selection scheme, a convex optimization model is used to determine at least one of the following: the amount of computing resources allocated by the s-th satellite to the i-th IoT device, the transmit power of the i-th IoT device in time slot t, and the transmit power of the g-th gateway in time slot t. The convex optimization model is configured as follows: The system independently solves multiple resource allocation subproblems to maximize utility, which are nested within the outer model as an inner model. These subproblems include one for optimizing the transmit power from the target gateway to the target satellite to reduce uplink energy consumption. After determining offloading and access decisions, the inner model sequentially and independently solves each subproblem, optimizing each subproblem individually while keeping other variables fixed. The solution to the current subproblem is then fed back to the outer model. The outer model then feeds back the resource allocation results determined by the inner model to adjust subsequent offloading decisions.

2. The method according to claim 1, characterized in that, The policy network locally deployed by the i-th IoT device is obtained through training a multi-agent deep reinforcement learning model. The reward function of the multi-agent deep reinforcement learning model includes a device utility reward and a task timeout penalty. The device utility reward is determined based on the task latency and energy consumption of the i-th IoT device in local computing mode and in satellite computing mode. The task timeout penalty is determined based on the maximum tolerable latency and the actual completion latency of the task. The observation space of the multi-agent deep reinforcement learning model is defined by the observation information of the i-th IoT device in time slot t. The observation information includes: the observation information of the i-th IoT device... The task data volume, computational density, and maximum tolerable latency in time slot t; the transmit power of the i-th IoT device in time slot t; the transmit power of the g-th gateway in time slot t; the computational resources allocated to the i-th IoT device by the s-th satellite; the channel gain from the g-th gateway to the s-th satellite in the q-th sub-frequency band; and the actual task completion latency of the i-th IoT device in time slot t-1; the action space of the multi-agent deep reinforcement learning model is defined by a first binary decision variable (whether the i-th IoT device unloads the computational task to the satellite through the gateway in time slot t) and a second binary decision variable (whether the g-th gateway accesses the s-th satellite through the q-th sub-frequency band).

3. The method according to claim 2, characterized in that, The task latency of the i-th IoT device in local computing mode and in satellite computing mode are determined by the following steps: Based on the first binary decision variable, it is determined whether the i-th IoT device chooses to offload the computing task to the satellite through the gateway in time slot t. Without selecting to offload computing tasks to the satellite via the gateway, the task latency of the i-th IoT device in local computing mode is calculated based on the local computing capability of the i-th IoT device and the task data volume of the un-offloaded computing tasks, and the task latency of the i-th IoT device in local computing mode is used as the actual task completion latency. When choosing to offload the task through a gateway, the target gateway to which the i-th IoT device is connected and the target satellite corresponding to the target gateway are determined according to the second binary decision variable; the communication delay from the i-th IoT device to the target gateway is determined according to the transmission rate from the i-th IoT device to the target gateway; and the total communication delay from the target gateway to the target satellite is determined according to the transmission rate from the target gateway to the target satellite. Based on the communication latency from the i-th IoT device to the target gateway, the total communication latency from the target gateway to the target satellite, and the computation processing latency of the computing tasks offloaded by the target satellite to the i-th IoT device, the task latency of the i-th IoT device in satellite computing mode is obtained, and the task latency of the i-th IoT device in satellite computing mode is taken as the actual task completion latency.

4. The method according to claim 3, characterized in that, The communication delay from the i-th IoT device to the target gateway is determined based on the transmission rate from the i-th IoT device to the target gateway, including: determining the transmission delay from the i-th IoT device to the target gateway based on the task data volume of the computing tasks offloaded by the i-th IoT device and the transmission rate of the target gateway; the total communication delay from the target gateway to the target satellite is determined based on the transmission rate from the target gateway to the target satellite, including: determining the transmission delay from the target gateway to the target satellite based on the task data volume of the computing tasks offloaded by the target gateway and the transmission rate of the target satellite; the propagation delay from the target gateway to the target satellite is determined based on the signal transmission distance and the speed of light between the target gateway and the target satellite; the sum of the transmission delay from the target gateway to the target satellite and the propagation delay from the target gateway to the target satellite is determined as the total communication delay from the target gateway to the target satellite.

5. The method according to claim 3, characterized in that, The step of calculating the task latency of the i-th IoT device in local computing mode based on the local computing power of the i-th IoT device and the amount of task data of the ununloaded computing tasks includes: obtaining the amount of task data, computing density, and local computing power of the ununloaded computing tasks of the i-th IoT device in time slot t, wherein the computing density represents the number of computing cycles required per bit of task data; determining the total computing requirements of the task based on the amount of task data and the computing density; and dividing the total computing requirements by the local computing power to obtain the task latency of the i-th IoT device in local computing mode.

6. The method according to claim 2, characterized in that, The task energy consumption of the i-th IoT device in local computing mode and in satellite computing mode are determined by the following steps: Based on the first binary decision variable, it is determined whether the i-th IoT device chooses to offload the computing task to the satellite through the gateway in time slot t. Without selecting to offload computing tasks to the satellite via a gateway, the task energy consumption of the i-th IoT device in local computing mode is calculated based on the local computing capability of the i-th IoT device, the task latency in local computing mode, and the energy consumption coefficient. When choosing to offload computing tasks to a satellite via a gateway, the task energy consumption of the i-th IoT device in satellite computing mode is calculated based on the transmit power of the i-th IoT device in time slot t and the communication delay from the i-th IoT device to the target gateway.

7. The method according to claim 2, characterized in that, The utility maximization problem is constructed according to the following steps: with the objective of maximizing the average utility of all IoT devices in a single time slot, the utility maximization problem is established under the conditions of satisfying the total satellite computing resource constraints and the maximum transmission power constraints of each IoT device and gateway. The utility maximization problem includes the following variables: the first binary decision variable, the second binary decision variable, the transmission power allocation variable, and the computing resource allocation variable.

8. The method according to claim 7, characterized in that, The convex optimization model is configured to independently solve multiple resource allocation subproblems of the utility maximization problem, the resource allocation subproblems including a first subproblem, a second subproblem, and a third subproblem; the method further includes: decomposing the utility maximization problem into: a first subproblem for optimizing the transmission power from the IoT device to the target gateway, a second subproblem for optimizing the transmission power from the target gateway to the target satellite, and a third subproblem for optimizing the allocation of computing resources for the target satellite.

9. The method according to claim 8, characterized in that, Under the premise of the offloading decision scheme and the network access selection scheme, a convex optimization model is used to determine at least one of the following: the computing resources allocated by the s-th satellite to the i-th IoT device, the transmit power of the i-th IoT device in time slot t, and the transmit power of the g-th gateway in time slot t. This includes: for the first sub-problem, based on the total task data of all tasks offloaded to the s-th satellite and the satellite resource capacity of the s-th satellite, constructing a convex optimization model, proportionally allocating the computing resources of the s-th satellite to each IoT device offloading computing tasks to the s-th satellite, and determining the computing resources allocated by the s-th satellite to the i-th IoT device; modeling the second sub-problem as a linear programming problem, and using the interior-point method to solve for the maximum transmit power constraint of the i-th IoT device, the... The transmit power of i IoT devices in time slot t is determined; the third subproblem is modeled as a convex optimization problem, and the transmit power of the g-th gateway in time slot t is solved using a convex optimization tool under the constraint of the maximum transmit power of the g-th gateway; the first, second, and third subproblems are solved independently to obtain the closed-form solution of the first subproblem, and the solutions of the second and third subproblems are updated iteratively; under the condition of satisfying the preset iteration, the optimal solutions of the second and third subproblems are obtained; based on the closed-form solution of the first subproblem and the optimal solutions of the second and third subproblems, the computing resources allocated by the s-th satellite to the i-th IoT device, the transmit power of the i-th IoT device in time slot t, and the transmit power of the g-th gateway in time slot t are obtained.

10. A joint optimization device for offloading tasks from IoT devices to satellites, characterized in that, A three-layer offloading system is applied to multiple IoT devices, multiple gateways, and multiple satellites. In this system, a single IoT device is simultaneously covered by multiple gateways, and a single gateway is simultaneously covered by multiple satellites. The device includes: a data acquisition module for acquiring current environmental data from multiple IoT devices, including at least one of the following: image, temperature, and wind speed; an acquisition module for acquiring observation information of the i-th IoT device in time slot t; and an input module for inputting the observation information of the i-th IoT device in time slot t into a locally deployed policy network of the i-th IoT device to generate an offloading decision scheme. The network access selection scheme, wherein the offloading decision scheme characterizes whether the i-th IoT device offloads the computing task to the satellite through the gateway in time slot t, and the network access selection scheme characterizes whether the g-th gateway accesses the s-th satellite through the q-th sub-frequency band; the policy network locally deployed by the i-th IoT device is obtained by training a multi-agent deep reinforcement learning model; the multi-agent deep reinforcement learning model serves as the outer model, modeling each IoT device as an agent, independently generating the offloading decision scheme and the network access selection scheme; the observation information serves as the multi-agent deep reinforcement learning model... The input of the multi-agent deep reinforcement learning model is used to adjust the output of the policy network in real time. The action space of the multi-agent deep reinforcement learning model is defined by a first binary decision variable, which determines whether the i-th IoT device offloads its computing task to the satellite through the gateway in time slot t, and a second binary decision variable, which determines whether the g-th gateway accesses the s-th satellite through the q-th sub-band. The convex optimization module is used to determine at least one of the following, based on the offloading decision scheme and the network access selection scheme: the amount of computing resources allocated by the s-th satellite to the i-th IoT device, the transmit power of the i-th IoT device in time slot t, and the... The transmit power of the g-th gateway in time slot t; the convex optimization model is configured to independently solve multiple resource allocation subproblems of the utility maximization problem, and the multiple resource allocation subproblems are nested as inner models with the outer model; the multiple subproblems include a subproblem for optimizing the transmit power from the target gateway to the target satellite to reduce uplink energy consumption; after the outer model determines the offloading and access decisions, the inner model solves multiple subproblems independently in sequence, each subproblem is optimized separately under the premise of fixing other variables, and the solution of the current subproblem is fed back to the outer model, and the outer model adjusts subsequent decisions according to the feedback.