Network resource allocation method and apparatus, electronic device, and storage medium
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-22
- Publication Date
- 2026-08-11
AI Technical Summary
然而,空天地一体化网络的资源调度是强制计算任务在单一物理域执行,无法适配多个物理域之间的协同需求
[0017]As can be seen from the above description, the network resource allocation method, apparatus, electronic device, and storage medium provided in this disclosure are as follows: When any of the air-based controller, space-based controller, and ground-based controller receives a target computing task, the physical domain where the controller receiving the target computing task is located is taken as the current physical domain. Using a pre-trained target action network in the current physical domain, the target physical domain for executing the target computing task is determined based on the target computing resources required by the target computing task and the available resources in each physical domain. In this way, each controller can independently determine the target physical domain for executing the target computing task, quickly and accurately determining the target physical domain based on the real-time available resources in each physical domain, achieving rapid response. When the target physical domain contains other physical domains, the target computing task is allocated to other physical domains besides the current physical domain for execution in those other physical domains. Thus, when it is determined that the target computing task needs to be executed in another physical domain, the target computing task can be quickly scheduled to be executed in that other physical domain, fulfilling the collaborative needs between various physical domains. When the target physical domain contains the current physical domain, the currently allocated resources of the current physical domain are determined, and the target computing task is executed in the current physical domain using the currently allocated resources. In this way, when it is determined that the target computing task needs to be executed in the target physical domain, the target physical domain can quickly and accurately determine the currently allocated resources, and thus use the currently allocated resources to execute the target computing task in the target physical domain.
Smart Images

Figure CN121419018B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of data processing technology, and in particular to a method, apparatus, electronic device and storage medium for network resource allocation. Background Technology
[0002] As a core architecture of 6G, the integrated air-space-ground network achieves full-domain coverage by integrating satellites, drones, and ground facilities, supporting low-latency scenarios such as vehicle-to-everything (V2X) and disaster remote sensing. However, the resource scheduling of the integrated air-space-ground network forces computing tasks to be executed in a single physical domain, which cannot adapt to the collaborative needs between multiple physical domains.
[0003] In view of this, how to adapt to the collaborative needs between multiple physical domains when scheduling computing tasks has become an urgent technical problem to be solved. Summary of the Invention
[0004] In view of this, the purpose of this disclosure is to provide a network resource allocation method, apparatus, electronic device and storage medium to solve or partially solve the above-mentioned technical problems.
[0005] Based on the above objectives, the first aspect of this disclosure proposes a network resource allocation method, applied to a network resource allocation device, the device comprising: a space-based controller disposed in a space-based physical domain, a space-based controller disposed in a space-based physical domain, and a ground-based controller disposed in a ground-based physical domain; the method comprising:
[0006] In response to determining that any one of the airborne controller, the space-based controller, and the ground-based controller has received a target computing task, the physical domain where the controller that received the target computing task is located is taken as the current physical domain;
[0007] Using the pre-trained target action network in the current physical domain, the target physical domain for executing the target computing task is determined based on the target computing resources required for the target computing task and the available resources in each physical domain;
[0008] In response to determining that the target physical domain contains other physical domains, the target computing task is assigned to one of the other physical domains for execution in that other physical domain; wherein, the other physical domain is a physical domain other than the current physical domain;
[0009] In response to determining that the target physical domain contains the current physical domain, the currently allocated resources of the current physical domain are determined, and the target computing task is executed in the current physical domain using the currently allocated resources.
[0010] Based on the same inventive concept, a second aspect of this disclosure proposes a network resource allocation device, the device comprising: an airborne controller disposed in an airborne physical domain, a spaceborne controller disposed in a space-based physical domain, and a ground-based controller disposed in a ground-based physical domain.
[0011] The current physical domain determination module is configured to, in response to determining that any one of the airborne controller, the space-based controller, and the ground-based controller has received a target computing task, take the physical domain where the controller that received the target computing task is located as the current physical domain;
[0012] The target physical domain determination module is configured to use a pre-trained target action network in the current physical domain to determine the target physical domain for executing the target computing task, based on the target computing resources required for the target computing task and the available resources in each physical domain.
[0013] The first allocation module is configured to allocate the target computing task to the other physical domain in response to determining that the target physical domain contains other physical domains, so that the target computing task can be executed in the other physical domains; wherein, the other physical domains are physical domains other than the current physical domain;
[0014] The second allocation module is configured to, in response to determining that the target physical domain contains the current physical domain, determine the currently allocated resources of the current physical domain and execute the target computing task in the current physical domain using the currently allocated resources.
[0015] Based on the same inventive concept, a third aspect of this disclosure proposes an electronic device including a memory, a processor, and a computer program stored in the memory and executable by the processor, wherein the processor implements the method described above when executing the computer program.
[0016] Based on the same inventive concept, a fourth aspect of this disclosure provides a non-transitory computer-readable storage medium that stores computer instructions for causing a computer to perform the methods described above.
[0017] As can be seen from the above description, the network resource allocation method, apparatus, electronic device, and storage medium provided in this disclosure are as follows: When any of the air-based controller, space-based controller, and ground-based controller receives a target computing task, the physical domain where the controller receiving the target computing task is located is taken as the current physical domain. Using a pre-trained target action network in the current physical domain, the target physical domain for executing the target computing task is determined based on the target computing resources required by the target computing task and the available resources in each physical domain. In this way, each controller can independently determine the target physical domain for executing the target computing task, quickly and accurately determining the target physical domain based on the real-time available resources in each physical domain, achieving rapid response. When the target physical domain contains other physical domains, the target computing task is allocated to other physical domains besides the current physical domain for execution in those other physical domains. Thus, when it is determined that the target computing task needs to be executed in another physical domain, the target computing task can be quickly scheduled to be executed in that other physical domain, fulfilling the collaborative needs between various physical domains. When the target physical domain contains the current physical domain, the currently allocated resources of the current physical domain are determined, and the target computing task is executed in the current physical domain using the currently allocated resources. In this way, when it is determined that the target computing task needs to be executed in the target physical domain, the target physical domain can quickly and accurately determine the currently allocated resources, and thus use the currently allocated resources to execute the target computing task in the target physical domain. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in this disclosure or related technologies, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a flowchart of a network resource allocation method according to an embodiment of the present disclosure;
[0020] Figure 2 This is a schematic diagram of an integrated air-space-ground network scenario according to an embodiment of this disclosure;
[0021] Figure 3 This is a schematic diagram of the task scheduling and resource allocation structure based on MADDPG according to an embodiment of this disclosure;
[0022] Figure 4 This is a schematic diagram of the structure of a network resource allocation device according to an embodiment of the present disclosure;
[0023] Figure 5 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present disclosure. Detailed Implementation
[0024] To make the objectives, technical solutions, and advantages of this disclosure clearer, the following detailed description is provided in conjunction with specific embodiments and the accompanying drawings.
[0025] It should be noted that, unless otherwise defined, the technical or scientific terms used in the embodiments of this disclosure should have the ordinary meaning understood by one of ordinary skill in the art to which this disclosure pertains. The terms "first," "second," and similar terms used in the embodiments of this disclosure do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0026] Based on the background technology description, the Space-Air-Ground Integrated Network (SAGIN), as the core architecture of the 6th Generation Mobile Communication Technology (6G), achieves full-area coverage by integrating satellites, drones, and ground facilities, supporting low-latency scenarios such as vehicle-to-everything (V2X) and disaster remote sensing. However, the resource scheduling of the SAGIN faces three major challenges: resource heterogeneity (limited computing power on space-based networks, energy constraints on air-based networks, and ample bandwidth on ground-based networks), highly dynamic topology (high-speed movement of satellites / drones causes time-varying links), and diverse tasks (high-definition rendering requires computing power, and autonomous driving requires real-time performance). Among current optimization methods, traditional heuristic algorithms rely on preset rules, making them difficult to adapt to dynamic environments and prone to getting trapped in local optima; centralized deep reinforcement learning schemes incur communication overhead due to global state collection, failing to meet millisecond-level response requirements such as vehicle control; while distributed multi-agent strategies (e.g., Multi-Agent Proximal Policy Optimization, or MAPPO) improve real-time performance, but the discrete action space limits resource allocation accuracy and does not explicitly handle cross-domain resource contention, resulting in a task conflict rate exceeding 19.8%.
[0027] Current research attempts to alleviate these problems through hierarchical control, such as deploying domain controllers for local resource management. However, the heterogeneity of space-based, airborne, and ground-based systems leads to low efficiency in cross-domain collaboration: satellites have limited coverage time due to orbital motion, UAVs face significant energy constraints, and ground computing centers suffer from uneven load distribution. Scheduling schemes based on game theory or independent deep reinforcement learning (DRL) (e.g., Dueling Double Deep Q-Network, D3QN) partially optimize latency, but they cause a surge in energy consumption (exhibiting over 245kJ in experiments) by ignoring resource contention, and cannot support fine-grained scheduling such as task fragmentation. Especially when multiple domains compete for scarce spectrum or storage, the lack of a collaborative mechanism will lead to an overall performance degradation of more than 28%.
[0028] While deep reinforcement learning offers a new path for dynamic optimization, the current MAPPO framework still has fundamental limitations: first, discrete scheduling decisions force tasks to execute in a single domain, failing to adapt to wide-domain collaborative requirements; second, distributed training lacks a global competition awareness mechanism, leading to frequent resource conflicts. In contrast, the centralized training architecture and continuous action space characteristics of Multi-Agent Deep Deterministic Policy Gradient (MADDPG) naturally align with SAGIN's cross-domain resource allocation needs, but current research has not yet applied it to solve the joint optimization problem of latency, energy consumption, and conflicts in space-air-ground scenarios.
[0029] As mentioned above, how to adapt to the collaborative needs between multiple physical domains when scheduling computing tasks has become an important research problem.
[0030] Based on the above description, such as Figure 1 As shown, the network resource allocation method proposed in this embodiment is applied to a network resource allocation device, which includes: a space-based controller disposed in the space-based physical domain, a space-based controller disposed in the space-based physical domain, and a ground-based controller disposed in the ground-based physical domain; the method includes:
[0031] Step 101: In response to determining that any one of the air-based controller, the space-based controller, and the ground-based controller has received a target computing task, the physical domain where the controller that received the target computing task is located is taken as the current physical domain.
[0032] Step 102: Using the pre-trained target action network in the current physical domain, determine the target physical domain for executing the target computing task based on the target computing resources required for the target computing task and the available resources in each physical domain.
[0033] Step 103: In response to determining that the target physical domain contains other physical domains, the target computing task is assigned to the other physical domains for execution in the other physical domains; wherein, the other physical domains are physical domains other than the current physical domain.
[0034] Step 104: In response to determining that the target physical domain contains the current physical domain, determine the currently allocated resources of the current physical domain, and execute the target computing task in the current physical domain using the currently allocated resources.
[0035] In practice, Figure 2 This is a schematic diagram of an integrated air-space-ground network scenario according to an embodiment of this disclosure. For example... Figure 2 As shown, the design goal of the SAGIN architecture is to provide computing service support for various application scenarios with different needs through seamless global coverage and collaborative computing across the three physical domains of airspace, space, and ground. The SAGIN architecture comprises three core components: airspace network, space-based network, and ground-based network.
[0036] At the ground-based network layer, the focus is on three types of devices capable of handling computing tasks: IoT devices, wired network devices, and dedicated computing devices. Given the limited coverage of 5G / 6G cellular networks in remote areas, IoT devices are widely deployed in these regions to handle smaller computing tasks such as video surveillance and the fusion of image and sound sensor data. Leveraging the SAGIN architecture, IoT devices can achieve globally scalable connectivity and collaborate with other domains to perform computing tasks. Computationally intensive tasks can be offloaded to cloud servers or terrestrial network computing centers via drones, satellites, or relays, thereby enhancing the overall network's computing power.
[0037] The terrestrial wired network primarily consists of routers, switches, gateways, and other devices forming the core network, mainly responsible for hop-by-hop forwarding of computing tasks. Computing devices are typically deployed at the network edge, especially in large computing centers. The servers in these centers contain one or more processing modules, such as a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), and a Field Programmable Gate Array (FPGA), responsible for real-time task processing. User-generated computing tasks are transmitted to these computing devices via network devices (e.g., routers and switches). Notably, by endowing routers with computing capabilities, they can perform both forwarding and computing functions. When computing tasks flow through these computing-capable routers, the tasks can be processed locally or forwarded to other devices. Therefore, for some small computing tasks, the computation can be completed on the forwarding path and the result returned to the user.
[0038] At the airborne network level, flying drones act as edge servers, providing low-latency edge caching and computing services to ground users. For example, lightweight drones with lightweight AI platforms can schedule deep learning inference tasks to be executed on their own lightweight AI platform. Each drone typically flies along a fixed trajectory, serving a specific area.
[0039] At the space-based network level, in areas with low user density and limited network resources (e.g., covered only by space-based networks), quality of experience can be achieved through satellite communication. Quality of Experience (QoE) and Quality of Service (QoS) requirements. Meanwhile, satellites generate numerous computational tasks during operation, such as Earth observation data processing, fault diagnosis, and management. However, Low Earth Orbit (LEO) satellites, limited by their computing power and onboard power supply, typically can only perform lightweight computational tasks. In contrast, Geostationary Earth Orbit (GEO) and Medium Earth Orbit (MEO) satellites have more abundant computing resources and solar power. Therefore, if GEO satellites have sufficient resources, computational tasks generated by MEO satellites can be scheduled to GEO satellites for processing. Furthermore, for services that cannot be covered by terrestrial networks and have high timeliness requirements, directly scheduling tasks to satellites can meet latency requirements. The embodiments disclosed in this disclosure consider a scenario where multiple satellites achieve full regional coverage.
[0040] SAGIN's control architecture divides the entire system into space-based, airborne, and ground-based domains, each deploying its own space-based, airborne, and ground-based controllers. These controllers are primarily responsible for information aggregation, task scheduling, and resource management and allocation within their respective physical domains. When a smart application device generates a computing task, the corresponding domain controller determines which physical domain the task will be scheduled to execute in and how many resources will be allocated to each physical domain. Specifically, ground-based controllers can be deployed in ground relay stations or base stations; airborne controllers are deployed inside UAVs; and space-based controllers can be deployed on LEO satellites or satellite ground stations.
[0041] The control process first utilizes novel in-band telemetry technology to flexibly sense resource information in three domains, including available computing resources, storage resources, bandwidth resources, spectrum resources, and remaining energy. This resource information serves as the state input for a deep reinforcement learning (DRL) model. Each domain controller is treated as an agent. The agent analyzes and learns based on the current state and makes decisions (actions). Subsequently, the environment provides reward signals for the decisions, guiding the agent to adjust subsequent decisions. By continuously repeating the process of state perception, decision-making, reward feedback, and learning, dynamic management and optimization of the network are achieved. The core issue is that the domain controller needs to dynamically make two key decisions: which physical domain to schedule the computing task for execution, and how many resources to allocate to the local physical domain (if the task is executed locally) to execute the computing task.
[0042] This disclosure provides a method for scheduling space-air-ground integrated computing power network resources based on multi-agent deep deterministic policy gradient (MADDPG). Its core lies in constructing a hierarchical collaborative architecture and designing an intelligent decision-making mechanism in a continuous action space.
[0043] In practice, the controller is first deployed on a low Earth orbit satellite or satellite ground station at the space-based network layer to coordinate onboard computing resources and limited energy. At the air-based network layer, the controller is integrated into the UAV swarm to dynamically manage the computing power and communication links of the flight edge nodes. At the ground-based network layer, the controller is set up at the ground computing center or base station to coordinate server resource pools and high-bandwidth transmission facilities.
[0044] In the above embodiments, when any of the airborne, spaceborne, and ground-based controllers receives a target computing task, the physical domain where the controller receiving the target computing task is located is designated as the current physical domain. Utilizing a pre-trained target action network in the current physical domain, the target physical domain for executing the target computing task is determined based on the target computing resources required by the target computing task and the available resources in each physical domain. In this way, each controller can independently determine the target physical domain for executing the target computing task, quickly and accurately determining the target physical domain based on the real-time available resources in each physical domain, achieving rapid response. If the target physical domain contains other physical domains, the target computing task is allocated to other physical domains besides the current physical domain for execution in those other physical domains. Thus, when it is determined that the target computing task needs to be executed in another physical domain, it can be quickly scheduled to that other physical domain for execution, fulfilling the collaborative needs between different physical domains. If the target physical domain contains the current physical domain, the currently allocated resources of the current physical domain are determined, and the target computing task is executed in the current physical domain using these currently allocated resources. In this way, when it is determined that the target computing task needs to be executed in the target physical domain, the target physical domain can quickly and accurately determine the currently allocated resources, and thus use the currently allocated resources to execute the target computing task in the target physical domain.
[0045] In some embodiments, the pre-training process of the target action network includes:
[0046] Step 105: Obtain the sample computing task and use the initial action network to determine the resource allocation method based on the sample computing resources required by the sample computing task and the available resources in each physical domain.
[0047] In practical implementation, at the system modeling level, the SAGIN architecture deploys independent domain controllers in its space-based, airborne, and ground-based networks. These controllers achieve globally optimal task scheduling and resource allocation through a collaborative mechanism. For the multi-domain collaborative decision-making problem, this embodiment employs a multi-agent deep deterministic policy gradient (MADDPG) algorithm, modeling each domain controller as an independent agent. These agents only need to perceive and process resource status and environmental information within their own domain to autonomously make decisions and execute corresponding operations. This distributed architecture effectively reduces the computational and storage pressure on individual control nodes while improving computational efficiency during the model training phase.
[0048] Step 106: Determine the total system cost when performing the sample calculation task according to the resource allocation method, and determine the reward function based on the total system cost.
[0049] In specific implementation, the transmission latency and computation latency in each physical domain are determined when performing the sample computation task according to the resource allocation method, and the total latency in each physical domain is determined based on the transmission latency and the computation latency. The transmission energy consumption and computation energy consumption in each physical domain are determined when performing the sample computation task according to the resource allocation method, and the total energy consumption in each physical domain is determined based on the transmission energy consumption and computation energy consumption. Based on the total latency and total energy consumption, the total system cost when performing the sample computation task according to the resource allocation method is determined, and the reciprocal of the total system cost is used as the reward function.
[0050] Step 107: Obtain the status of space-based resources in the space-based physical domain, the status of space-based resources in the space-based physical domain, and the status of ground-based resources in the ground-based physical domain; determine the global status based on the status of space-based resources, the status of space-based resources, and the status of ground-based resources.
[0051] In practical implementation, by introducing a policy sharing mechanism and a joint learning framework, each agent can achieve real-time synchronization of network parameters and construct a global perspective to support cross-domain collaborative decision-making. Based on the above design, this disclosure proposes a task scheduling and resource allocation strategy based on the MADDPG algorithm. By optimizing task response latency and device energy consumption indicators, it ultimately forms a strategy as follows: Figure 3 The collaborative computing architecture shown.
[0052] Each domain controller acquires multi-dimensional status information within the domain in real time through a novel in-band telemetry technology, including computing resources. t (unit: GHz), storage resources t (Unit: MB), bandwidth resources b t (in MHz) and remaining available energy e t (Unit: kJ) is used to form a local observation state vector, which is then uploaded to the central processing unit via a low-latency channel.
[0053] The task scheduling and resource allocation problem is modeled as a typical multi-agent Markov decision process, represented by tuples (S, A, P, R), where each domain controller is an agent, S is the global state space shared by all agents i∈{g,a,p}, A represents the joint action space of all agents, and a i Let A represent the decision action of agent i, P represent the state transition function, and R be the function for calculating the global reward. At the current time t, each agent, according to policy π... i (a i |s i t Choose one decision action a i The common goal of all intelligent agents is to minimize the system cost.
[0054] The global state s at time t t It can be represented as follows:
[0055]
[0056] in, This provides information related to resource status and sample computation tasks in the ground-based physical domain. This provides information related to resource status and sample computation tasks in the space-based physical domain. This provides information related to resource status and sample computation tasks in the space-based physical domain. This is sensed by each domain controller, including computing resources, storage resources, bandwidth resources, and available energy. A n This represents information related to the sample computation tasks that require scheduling decisions at the current time t. These represent the computing resources of the ground-based physical domain, the air-based physical domain, and the space-based physical domain, respectively, i.e., the ability to perform sample computing tasks; These represent the storage resources of the ground-based physical domain, the air-based physical domain, and the space-based physical domain, respectively, i.e., the device's ability to store sample computation task data. These represent the bandwidth resources of the ground-based physical domain, the air-based physical domain, and the space-based physical domain, respectively. In other words, the computing device's ability to communicate with other domains or other local devices determines its data transmission rate and capacity. These represent the available energy in the ground-based physical domain, air-based physical domain, and space-based physical domain, respectively, which is the power of the computing equipment. This information can reflect whether the equipment can support the completion of the sample computing task during decision-making.
[0057] Step 108: Using the initial value network, determine the value function based on the resource allocation method, the reward function, and the global state. Update the network parameters of the initial action network by maximizing the value function to obtain the target action network.
[0058] In practice, the current state value and the next state value are determined from the global state, and the current state value, resource allocation method, reward function and next state value are stored as tuple data in the experience replay pool.
[0059] The first loss function of the initial value network is determined based on the reward function and the next value function corresponding to the next state value. The target value network is obtained by updating the initial value network based on the first loss function.
[0060] The temporal difference error is determined based on the current value function corresponding to the current state value and the next value function corresponding to the next state value. The generalization advantage estimation parameters are then determined based on the temporal difference error. The second loss function of the initial action network is determined based on the generalization advantage estimation parameters. This second loss function is then backpropagated to the initial action network, and the network parameters of the initial action network are updated according to the second loss function to obtain the target action network.
[0061] The above scheme utilizes an initial value network to determine a value function based on resource allocation, reward function, and global state. The target action network is obtained by updating its parameters by maximizing this value function. This ensures that the resource allocation method output by the trained target action network maximizes the value function, thus achieving optimal resource allocation.
[0062] In some embodiments, step 106 includes:
[0063] Step 1061: Determine the transmission delay and computation delay in each physical domain when performing the sample computation task according to the resource allocation method, and determine the total delay in each physical domain based on the transmission delay and the computation delay.
[0064] Step 1062: Determine the transmission energy consumption and computing energy consumption in each physical domain when performing the sample computing task according to the resource allocation method, and determine the total energy consumption in each physical domain based on the transmission energy consumption and the computing energy consumption.
[0065] Step 1063: Based on the total latency and the total energy consumption, determine the total system cost when executing the sample calculation task according to the resource allocation method, and use the reciprocal of the total system cost as the reward function.
[0066] In practical implementation, the optimization objective is to minimize the combined overhead of total system latency and energy consumption. To this end, a dedicated latency and energy consumption model was established. The model considers the differences in computing power among different nodes, which directly leads to different computational latencies when executing the same task. When a task is scheduled for processing, its data needs to be transmitted to the target node, resulting in transmission latency; this latency is significantly affected by factors such as channel quality and allocated bandwidth resources. It is worth noting that the queuing latency of data packets during forwarding is usually much smaller than the computational and transmission latency, and therefore can be ignored in the model. Furthermore, different task scheduling and resource allocation decisions will also lead to varying energy consumption from different devices. These established models lay the theoretical foundation for the design of subsequent algorithms and the evaluation of system performance.
[0067] The above scheme determines the total system cost when executing sample computing tasks according to resource allocation based on total latency and total energy consumption. This enables joint optimization of application devices' scheduling decisions, transmission latency, and execution energy consumption for sample computing tasks. By using the reciprocal of the total system cost as the reward function, the sample computing tasks can be completed in a way that minimizes total latency and total energy consumption.
[0068] In some embodiments, step 1061 includes:
[0069] Step 1061A: Determine the data transfer rate between the application device and each physical domain when performing the sample computation task according to the resource allocation method.
[0070]
[0071] in, The data transmission rate between the application device n and each physical domain i. p is the link bandwidth resource allocated for the application device n to communicate with each physical domain i. n The transmission power of the task is calculated for the sample. The channel gain is defined as the communication gain between the application device n and each physical domain i. σ is the distance from the application device n to the computing nodes in each physical domain i. i θ represents the channel background noise when the application device n communicates with each physical domain i, and θ is the distance attenuation factor.
[0072] In practice, in the SAGIN architecture, intelligent application devices (i.e., application device n) connect to terrestrial networks, airborne networks (sky), and spaceborne networks (space) through three wireless communication methods.
[0073] Specifically, for any application device n, the data transmission rate between application device n and the terrestrial network (i.e., the ground-based physical domain) is expressed as:
[0074]
[0075] Where g represents the ground physical domain, The data transmission rate between application device n and the ground-based physical domain g. p is the link bandwidth resource allocated for application device n to communicate with the ground physical domain g. n The transmission power of the sample calculation task is calculated. The channel gain is used when the application device n communicates with the ground-based physical domain g. σ is the distance from the application device n to the computing node in the ground-based physical domain g. g θ represents the background noise of the channel when the application device n communicates with the ground physical domain g, and θ is the distance attenuation factor.
[0076] Similarly, the data transmission rates between the application device n and the space-based network (i.e., the space-based physical domain) and the space-based network (i.e., the space-based physical domain) can be determined separately, and expressed as follows:
[0077]
[0078] Where 'a' represents the space-based physical domain, The data transmission rate between application device n and space-based physical domain a. p is the link bandwidth resource allocated for application device n to communicate with space-based physical domain a. n The transmission power of the sample calculation task is calculated. The channel gain when application device n communicates with the space-based physical domain a. σ is the distance from the application device n to the computing node in the space-based physical domain a. a Let θ be the channel background noise when the application device n communicates with the space-based physical domain a, and let θ be the distance attenuation factor.
[0079]
[0080] Where p represents the space-based physical domain, The data transmission rate between application device n and space-based physical domain p. The link bandwidth resources allocated for application device n to communicate with space-based physical domain p, p n The transmission power of the sample calculation task is calculated. The channel gain is used when application device n communicates with the space-based physical domain p. σ is the distance from application device b to the computing node in the space-based physical domain p. p Let θ be the channel background noise when the application device n communicates with the space-based physical domain p, and let θ be the distance attenuation factor.
[0081] Step 1061B: Determine the transmission delay of the sample computation task to each physical domain based on the data transmission rate.
[0082]
[0083] in, The transmission delay for the sample task to be transmitted to each physical domain i is calculated. x is the data transmission rate between the application device n and each physical domain i. n The amount of input data for the sample calculation task.
[0084] In practical implementation, transmission latency refers to the time consumed by a computation task to be transmitted to each physical domain via a wireless link. Specifically, when a sample computation task is scheduled to a ground-based, airborne, or space-based network within the SAGIN architecture, the transmission latency of the sample computation task can be characterized by a quantization model. This latency difference mainly stems from the differences in physical layer parameters such as transmission distance, channel conditions, and link bandwidth between different physical domains, and requires modeling and analysis in conjunction with the specific network topology and communication protocol.
[0085]
[0086] Where g represents the ground physical domain, The transmission delay of the sample computation task to the ground-based physical domain g is given. x represents the data transmission rate between the application device n and the ground-based physical domain g. n The amount of input data for the sample calculation task.
[0087]
[0088] Where 'a' represents the space-based physical domain, The transmission delay of the sample computation task to the space-based physical domain a is given. x represents the data transmission rate between application device n and space-based physical domain a. n The amount of input data for the sample calculation task.
[0089]
[0090] Where p represents the space-based physical domain, The transmission delay of the sample computation task to the space-based physical domain p is given. x represents the data transmission rate between application device n and space-based physical domain p. n The amount of input data for the sample calculation task.
[0091] Step 1061C: Determine the sample allocation resources in each physical domain when executing the sample computing task according to the resource allocation method, and determine the computing latency of the sample computing task in each physical domain based on the sample allocation resources.
[0092]
[0093] in, c represents the computation latency of the sample computation task performed in each physical domain i. n The total computing resources for the sample computation task. Allocate resources for the sample computation task in each physical domain i.
[0094] In practical implementation, computation latency refers to the time required for a computation task to complete execution on the computing device. In the SAGIN architecture, the sample computation tasks generated by the vehicle have cross-domain execution capabilities and can be processed in ground-based networks, air-based networks, or space-based networks. The computation latency of these three physical domains can be quantified and analyzed using the following formulas.
[0095]
[0096] Where g represents the ground physical domain, c represents the computational latency of the sample computation task performed in the ground-based physical domain g. n The total computational resources for the sample computation task. Allocate resources for the samples in the ground-based physical domain g for the sample computation task.
[0097]
[0098] Where 'a' represents the space-based physical domain, The computational latency of the sample computation task performed in the space-based physical domain a, c n The total computational resources for the sample computation task. Allocate resources for the samples in the space-based physical domain a for the sample computation task.
[0099]
[0100] Where p represents the space-based physical domain, c represents the computational latency of the sample computation task performed in the space-based physical domain p. n The total computational resources for the sample computation task. Allocate resources for the samples in the space-based physical domain p for the sample computation task.
[0101] During the computation task scheduling process of the SAGIN architecture, when the computation task of application device n is assigned to any physical domain among air-based, space-based, and ground-based networks for execution, sufficient storage space must be pre-configured for the computation task. If the storage resources are insufficient, even if enough computing resources have been allocated for the computation task, the execution of the computation task may still be interrupted or fail due to data storage bottlenecks.
[0102] Step 1061D: Based on the transmission delay and the computation delay, determine the total delay in each physical domain when executing the sample computation task according to the resource allocation method.
[0103]
[0104] in, The total latency in each physical domain i when performing the sample computation task according to the resource allocation method, The transmission delay for the sample task to be transmitted to each physical domain i is calculated. The computational latency for the sample computation task to be performed in each physical domain i.
[0105] In practical implementation, based on the above analysis, the total latency of the sample computation task An, from its generation to its final execution and return of results, consists of three parts: transmission latency, computation latency, and reception latency. Transmission latency corresponds to the time it takes for the sample computation task data to be transmitted to each physical domain via the wireless link; computation latency reflects the processing time required for the computing device to execute the sample computation task; and reception latency refers to the transmission delay of the computation result from the computation domain back to the user terminal. Given that the amount of data in the computation result is usually much smaller than the original sample computation task data, the transmission latency of the return process can be ignored. Therefore, the total latency of the sample computation task An can be simplified as the sum of the transmission latency and the computation latency, which can be expressed mathematically as:
[0106]
[0107] Where g represents the ground physical domain, The total latency in the ground-based physical domain g when performing sample computation tasks according to the resource allocation method. The transmission delay of the sample computation task to the ground-based physical domain g is given. The computational delay for the sample computation task performed in the ground-based physical domain g.
[0108]
[0109] Where 'a' represents the space-based physical domain, The total latency in the space-based physical domain a when performing sample computation tasks according to the resource allocation method. The transmission delay of the sample computation task to the space-based physical domain a is given. The computational latency for the sample computation task performed in the space-based physical domain a.
[0110]
[0111] Where p represents the space-based physical domain, The total latency in the space-based physical domain p when performing sample computation tasks according to the resource allocation method. The transmission delay of the sample computation task to the space-based physical domain p is given. The computational latency of the sample computation task performed in the space-based physical domain p.
[0112] The above scheme determines the total latency in each physical domain when performing sample computation tasks according to the resource allocation method, based on transmission latency and computation latency. By summing the transmission latency and computation latency, the total latency in each physical domain when performing sample computation tasks according to the resource allocation method can be accurately obtained.
[0113] In some embodiments, step 1062 includes:
[0114] Step 1062A: Determine the computational energy consumption and transmission energy consumption in each physical domain when performing the sample computation task according to the resource allocation method.
[0115] Step 1062B: Based on the computational energy consumption and the transmission energy consumption, determine the total energy consumption in each physical domain when performing the sample computation task according to the resource allocation method.
[0116]
[0117] in, This represents the total energy consumption in each physical domain i when performing the sample computation task according to the described resource allocation method. The computational energy consumption of the sample computation task performed in each physical domain i. The transmission energy consumption for the sample task to be transmitted to each physical domain i is calculated.
[0118] In practical implementation, within the integrated air-space-ground network architecture, application devices in the air-based, space-based, and ground-based physical domains all consume energy when performing computational tasks. To address the processor energy consumption modeling problem, energy consumption is determined through quantification expressions. Taking the Central Processing Unit (CPU) as an example, its energy consumption mainly depends on two core indicators: the effective capacitance coefficient k (this parameter is determined by the processor chip's physical architecture) and the number of computation cycles C required per unit bit of data. For the terminal device initiating the computation request, its energy consumption is primarily driven by the wireless transmission energy consumption of the sample computation task data. Therefore, the system energy consumption generated when the sample computation task is scheduled across the three physical domains is expressed as follows:
[0119]
[0120] Where g represents the ground physical domain, The total energy consumption in the ground-based physical domain g when performing sample computation tasks according to the resource allocation method. The computational energy consumption for the sample computation task performed in the ground-based physical domain g. The energy consumption for transmitting the sample computation task to the ground-based physical domain g.
[0121]
[0122] Where 'a' represents the space-based physical domain, The total energy consumption in the space-based physical domain a when performing sample computation tasks according to the resource allocation method. The computational energy consumption of the sample computation task performed in the space-based physical domain a. The energy consumption for transmitting the sample computation task to the space-based physical domain a.
[0123]
[0124] Where p represents the space-based physical domain, The total energy consumption in the space-based physical domain p when performing sample computation tasks according to the resource allocation method. The computational energy consumption of the sample computation task performed in the space-based physical domain p. The energy consumption for transmitting the sample computation task to the space-based physical domain p.
[0125] The above scheme determines the total energy consumption in each physical domain when performing sample computation tasks according to the resource allocation method, based on computational and transmission energy consumption. By summing the computational and transmission energy consumption, the total energy consumption in each physical domain when performing sample computation tasks according to the resource allocation method can be accurately obtained.
[0126] In some embodiments, step 1063 includes:
[0127] Step 1063A: Based on the total latency and the total energy consumption, determine the total system cost when executing the sample computation task according to the resource allocation method.
[0128]
[0129] Among them, C t The total system cost is defined as the resource allocation method used to execute the sample computation task, where N is the number of sample computation tasks and β is the weighting coefficient. This refers to the scheduling position when the sample computation task is executed according to the described resource allocation method. The total latency in each physical domain i when performing the sample computation task according to the resource allocation method, The total energy consumption in each physical domain i when performing the sample computation task according to the resource allocation method.
[0130] In specific implementation, for the sample computation task An, this embodiment of the disclosure stipulates that the sample computation task An cannot be simultaneously scheduled to be executed in multiple physical domains. Therefore, at the current time t, for the task scheduling decision (i.e., resource allocation method), the scheduling position is defined. In the formula, It is a Boolean value indicating whether to schedule the sample computation task An to the i-th physical domain, where g, a, and p represent the ground-based network, air-based network, and space-based network, respectively.
[0131] The objective of this disclosure is to jointly optimize the scheduling decision, task transmission latency, and execution energy consumption of the sample computing task for application device n, so as to complete the sample computing task in a way that minimizes the combined overhead of total latency and total energy consumption.
[0132] Define the total system cost at the current time t:
[0133]
[0134] Where i∈{g,a,p}, C t The total system cost is defined as the resource allocation method used to execute the sample computation task, where N is the total number of sample computation tasks to be processed, and β is the weighting coefficient. This refers to the scheduling position when the sample computation task is executed according to the described resource allocation method. This represents the total latency in each physical domain i when performing sample computation tasks according to the resource allocation method. This represents the total energy consumption in each physical domain i when performing sample computation tasks according to the resource allocation method.
[0135] In the formula, N serves as the upper limit of the outer summation symbol, representing the summation over all sample computation tasks. N ensures that the formula covers every sample computation task in the system. By summing from sample computation task 1 to sample computation task N, the total system cost integrates the total latency and total energy consumption of all sample computation tasks, achieving global optimization. β is used to balance the relative importance of total latency and total energy consumption in the total system cost, with a value ranging from [0,1]. When β is close to 1, the total latency contributes more to the total system cost (optimization objective biases towards low latency); when β is close to 0, the total energy consumption contributes more to the total system cost (optimization objective biases towards low energy consumption).
[0136] Therefore, the optimization problem is transformed into minimizing the total system cost:
[0137] minC t
[0138]
[0139] Where, minC t The objective function is to minimize the total system cost.
[0140] Step 1063B: Use the reciprocal of the total system cost as the reward function.
[0141]
[0142] Where, r t For the reward function, C t The total system cost when performing the sample computation task according to the resource allocation method.
[0143] In practical implementation, the core optimization objective of this embodiment is to minimize the total latency and total energy consumption of the entire sample computation task processing process. At the implementation level, each domain controller follows the principle of distributed optimization, independently optimizing the total latency and total energy consumption only for sample computation tasks scheduled for execution within its local domain. To achieve this objective, the reward function design must possess dual guidance: it must guide the agent to form the optimal decision-making strategy through autonomous learning, and it must ensure that the decision-making behavior directly targets the joint optimization of sample computation task processing latency and device energy consumption. Therefore, the reward function is set as follows:
[0144]
[0145] The above scheme determines the total system cost when executing sample computation tasks according to resource allocation based on total latency and total energy consumption. This allows for joint optimization of the application device's scheduling decisions, transmission latency, and execution energy consumption for sample computation tasks, minimizing both total latency and total energy consumption. The reciprocal of the total system cost is used as the reward function, enabling the training of the initial value network and initial action network based on this reward function.
[0146] In some embodiments, step 108 includes:
[0147] Step 108A: Determine the current state value and the next state value from the global state, and store the current state value, the resource allocation method, the reward function, and the next state value as tuple data of the current time in the experience replay pool.
[0148] Step 108B: Determine the first loss function of the initial value network based on the reward function and the next value function corresponding to the next state value.
[0149]
[0150] in, Let the first loss function of the initial value network be . For the reward function, The next state value s t+1 The corresponding next value function.
[0151] Step 108C: Update the initial value network based on the first loss function to obtain the target value network.
[0152] Step 108D: Determine the time difference error based on the current value function corresponding to the current state value and the next value function corresponding to the next state value.
[0153]
[0154] Where, δ t The time difference error is... Let γ be the reward function, and γ be the discount factor. The next state value s t+1 The corresponding next value function, The current state value s t The corresponding current value function.
[0155] Step 108E: Determine the generalization advantage estimation parameters based on the time difference error.
[0156]
[0157] in, δ is the generalization advantage estimation parameter. t Let be the time difference error, γ be the discount factor, and λ be the smoothing parameter.
[0158] Step 108F: Determine the second loss function of the initial action network based on the generalization advantage estimation parameters.
[0159]
[0160] Among them, L CLIP (θ i ) is the second loss function of the initial action network. The average of the tuple data at all times in the experience replay pool. To limit parameters, The generalization advantage estimation parameters are as follows: Let ε be the pruning function, ε be a preset positive number, and δ be the coefficient of the entropy of the resource allocation method. The resource allocation method The entropy.
[0161] Step 108G: Backpropagate the second loss function to the initial action network, and update the network parameters of the initial action network according to the second loss function to obtain the target action network.
[0162] (1) Algorithm Principle
[0163] The task scheduling and resource allocation algorithm based on the Multi-Agent Deep Deterministic Policy Gradient (MADDPG) is a multi-agent reinforcement learning algorithm specifically designed for the SAGIN integrated space-air-ground computing network. Through a framework of centralized training and distributed execution, it enables collaborative decision-making among multiple domain controllers (space-based, air-based, and ground-based) to optimize network resource utilization and improve task processing efficiency.
[0164] The MADDPG algorithm, as a multi-agent extension of the DDPG algorithm, inherits its ability to handle continuous action spaces. By introducing a centralized training and decentralized execution (CTDE) mechanism, it effectively solves the instability and credit allocation problems in multi-agent environments. In the SAGIN scenario, each domain controller is considered an agent with an independent Actor-Critic architecture. The Actor network (action network) is responsible for generating actions based on the current local state (e.g., task scheduling decisions, resource allocation strategies), while the Critic network (value network) uses the global state (including the states and actions of other agents) to evaluate the value of the current action, thereby guiding policy updates.
[0165] (2) Network Structure
[0166] In the MADDPG-based task scheduling and resource allocation algorithm, network architecture design is one of the core aspects of algorithm implementation. This algorithm equips each agent (i.e., space-based, air-based, and ground-based domain controllers) with an independent Actor-Critic architecture to achieve efficient decision-making and evaluation.
[0167] Actor network structure:
[0168] Actor networks are used to generate continuous actions based on the current local state, such as task scheduling priorities and resource allocation ratios. The input layer of an actor network receives local state information of the agent, which includes, but is not limited to, the specific requirements of the sample computation task (e.g., task type, data volume, computational load), the real-time state of network resources (e.g., bandwidth resources, storage resources, computing power), and other relevant information (e.g., task deadline, priority). To fully capture the features of the state information, the input layer typically contains multiple neurons, each corresponding to a specific dimension of the state information.
[0169] Next are the hidden layers. Actor networks may contain multiple hidden layers to increase the network's expressive power. Each hidden layer consists of multiple neurons connected by weights. The number of hidden layers and the number of neurons per layer can be adjusted according to the needs of the specific task. Generally, more hidden layers and more neurons can improve the network's expressive power, but may also increase the risk of overfitting. Therefore, in practical applications, it is necessary to determine the most suitable network structure through experiments.
[0170] In hidden layers, activation functions are typically used to introduce non-linearity, enabling the network to learn more complex patterns. Commonly used activation functions include ReLU (Rectified Linear Unit), Sigmoid, and Tanh. The ReLU function is widely used in deep learning due to its advantages such as simple computation and less pronounced gradient vanishing problem.
[0171] Finally, there's the output layer, which generates consecutive actions in the Actor network. The number of neurons in the output layer corresponds to the dimension of the action space. For example, if task scheduling priority and resource allocation ratio are two independent consecutive actions, then the output layer will contain two neurons, one for each action. The activation function of the output layer is typically determined based on the range of the action space. If the action space is bounded (e.g., resource allocation ratio is between 0 and 1), then a sigmoid or tanh function can be used to constrain the output within a reasonable range.
[0172] Critic network structure:
[0173] Critic networks are used to evaluate the Q-value of a current state-action pair, that is, to assess the long-term benefit of an agent taking a certain action in the current state. Unlike actor networks, Critic networks take into account both a global state and global actions as input. The global state is a vector composed of the states of all agents, containing complete information about the entire environment. The global actions are a vector composed of the actions of all agents, reflecting the decisions made by all agents.
[0174] The input layer of a Critic network therefore consists of two parts: one corresponding to the global state and the other to the global actions. These two parts may need to be concatenated or fused before entering the input layer so that the Critic network can consider both state and action information simultaneously.
[0175] Similar to Actor networks, Critic networks may also contain multiple hidden layers to increase the network's expressive power. The number of hidden layers and the number of neurons in each layer can also be adjusted according to the needs of the specific task. Activation functions are also used in the hidden layers to introduce non-linearity.
[0176] The output layer of a Critic network generates a scalar value, the Q-value of the current state-action pair. This Q-value reflects the agent's expected long-term reward after taking a certain action in the current state. The output layer typically does not use an activation function (or uses a linear activation function) so that it can directly output the numerical value of the Q-value.
[0177] Network update mechanism:
[0178] To train the Actor and Critic networks, the MADDPG algorithm employs an experience replay pool and a centralized training mechanism. During training, all agents interact with the SAGIN environment and store the experience of each interaction (current state, action, reward, and next state) in a shared experience replay pool.
[0179] During network training, a batch of experiences is randomly drawn from the experience replay pool to update the Critic and Actor networks. The Critic network updates its parameters by minimizing the mean squared error loss function to approximate the true Q-value. The first loss function calculates the difference between the Q-value predicted by the Critic network and the target Q-value, and minimizes this difference using a gradient descent algorithm.
[0180] The Actor network updates its parameters by maximizing the Q-value output of the Critic network. Specifically, the Actor network uses a deterministic policy gradient method, updating its network parameters based on the Q-value gradient provided by the Critic network, thereby generating a better action policy.
[0181] Target network:
[0182] To improve the stability of the training process, the MADDPG algorithm also introduces target Actor networks and target Critic networks. The parameters of these target networks are gradually updated to approximate the parameters of the main network through a soft update method, thereby reducing oscillations and instability during training. The soft update formula ensures smooth changes in the target network parameters, contributing to maintaining the stability of the training process.
[0183] (3) Training process
[0184] In the initial phase of training, each agent interacts with the environment according to the current resource allocation method (determined by the Actor network). The agent first observes the current environmental state, which contains all information relevant to the agent, such as the specific requirements of the computational task and the real-time status of network resources.
[0185] Based on the observed state, the agent uses an Actor network to generate an action (i.e., a resource allocation method). This action is continuous and may represent task scheduling priority, resource allocation ratio, etc. The Actor network maps the state to the action space through a deep neural network, providing decision support for the agent.
[0186] After the action is performed, the environment will transition to a new state and return a reward signal. This reward signal indicates whether the action was good or bad, and is an important basis for the agent's learning. For example, if an action leads to efficient task processing or reasonable resource allocation, the agent will receive a positive reward; conversely, if an action leads to task delays or resource waste, the agent will receive a negative reward.
[0187] The agent stores the experience of this interaction (current state, action, reward, and next state) in an experience replay pool. The experience replay pool is a buffer for storing the agent's interaction experience with the environment, breaking down correlations between data and improving training efficiency. By storing a large number of experience samples, the algorithm can update network parameters more stably, avoiding training fluctuations caused by the similarity between consecutive samples.
[0188] During training, the algorithm periodically draws a batch of experience samples from the experience replay pool to update the Critic and Actor networks. This random sampling method increases the diversity of samples and improves the network's generalization ability.
[0189] For the Critic network, the goal is to accurately evaluate the Q-value of the current state-action pair. The algorithm uses a target network to compute a target Q-value, which is calculated based on the next state and all possible actions the agent can take. By minimizing the mean squared error between the predicted Q-value and the target Q-value, the algorithm updates the parameters of the Critic network, gradually approximating the true Q-value function. This process is implemented using gradient descent, ensuring that the Critic network provides accurate action value assessments.
[0190] For Actor networks, the goal is to generate actions that maximize the Q-value. The algorithm uses the Q-value gradient provided by the Critic network to update the parameters of the Actor network. Specifically, the algorithm calculates the gradient of the Q-value relative to the action and backpropagates this gradient into the Actor network to update its parameters. In this way, the Actor network can learn strategies to generate better actions, thereby improving the agent's decision-making ability.
[0191] To improve the stability of the training process, the MADDPG algorithm also introduces a target network mechanism. The target network is a copy of the main network, and its parameters are gradually updated to approximate the main network parameters through soft updates. This soft update method can reduce oscillations and instability during training, ensuring that the algorithm can stably converge to the optimal policy.
[0192] To improve training stability, the MADDPG algorithm introduces a target Actor network and a target Critic network. The parameters of the target network are gradually updated to approximate the parameters of the main network through soft updates, reducing oscillations during training.
[0193] In the MADDPG algorithm of this disclosure, each domain controller is considered an independent agent. Taking the ground-based controller as an example, this paper describes how to use the MADDPG algorithm to solve the computational task scheduling and resource allocation problem in the SAGIN scenario. The principles of the space-based controller and the airborne controller are similar. Each agent contains two networks, i.e., the parameters are θ. i The Actor network and parameters are as follows The Critic network. At the current time t, the ground controller dynamically senses the resource information of the local area. (i.e., the current state value at the current moment), and set the current state value The relevant information for the sample computation task is input into the Actor network. The Actor network generates the action probability distribution. That is: to acquire task scheduling and resource allocation strategies. Ground-based intelligent agents then make execution decisions. After (i.e., resource allocation method), instant rewards are obtained. (i.e., the reward function) and the next state value at the next time step. And and The data is stored in the experience replay pool. Meanwhile, the Critic network generates the current value function based on the set of local state observations of all agents at time t. Based on the next state value Generate the next value function Therefore, the first loss function of the Critic network is obtained:
[0194]
[0195] By cutting To obtain a new strategy to optimize the objective function, and thus optimize the Actor network, the algorithm can achieve better performance. Therefore, the parameters θ of the Actor network are updated by maximizing the objective function. i :
[0196]
[0197] in, This represents the updated policy parameter θ. i The ratio of the policy parameter θi,old before the update takes the value [1-ε, 1+ε], thus ensuring the update range; Indicating in strategy The entropy of task scheduling and resource allocation decisions output by sensing the status of ground-based network resources is δ; It is calculated using the Generalized Advantage Estimation (GAE). The calculation formula is as follows:
[0198]
[0199] in, It is the current value function at the current time t; γ is the discount factor, and λ is the smoothing parameter. The second loss function L... CLIP (θ i The backpropagation is performed to the Actor network to update the parameters.
[0200] The MADDPG algorithm employs a "centralized training-distributed execution" paradigm, where a global Critic network and a distributed Actor network collaborate to achieve optimization decisions. The Critic network receives the fused global state s. t =s g ,s a ,s p and combined actions a t =a g ,a a ,a p The system outputs a cross-domain collaborative value function Q(s,a|\phi), which comprehensively evaluates the long-term impact of latency, energy consumption, and resource conflicts. Each domain Actor network generates deterministic action outputs based on its local real-time state, innovatively designed as a dual-channel continuous vector: the first channel outputs a flexible scheduling probability vector w = [w...]. g ,w a ,w p ]∈[0,1] 2 Satisfying the probability normalization constraint ∑ i∈g,a,p w i =1, supports the fragmented scheduling of computational tasks to multi-domain parallel execution (e.g., splitting a high-definition remote sensing task into 60% ground-based rendering, 30% space-based preprocessing, and 10% space-based data verification); the second channel outputs a resource allocation ratio vector d. i =[c i ,o i ,b i ]∈[0,1] 3 To satisfy the total resource constraint c i +o i +b i ≤1, dynamically allocate computing cores, storage space and spectrum resources.
[0201] To effectively suppress cross-domain resource contention, this disclosure innovatively introduces a contention-aware penalty mechanism into the reward function:
[0202]
[0203] Where, d i This refers to the resource allocation ratio of physical domain i: di =[c i ,o i ,b i (Percentage of computing resources, storage resources, and bandwidth resources); |d i ∩d j | refers to the resource allocation overlap, taken as d. i With d j Sum of minimum values in each dimension This refers to the total amount of resources allocated, taking the value d. i With d j The sum of the maximum values of each dimension; This represents the total latency (including transmission latency) of the sample computation task An in physical domain i. Computation delay ), For the corresponding energy consumption (including transmission energy consumption) Computational energy consumption ), β∈[0,1] is the time delay weight factor (preferably 0.7); the penalty term is calculated using Jaccard similarity. The overlap of resource allocation is quantified (e.g., the conflict coefficient is 0.8 when two domains simultaneously request 80% of the public spectrum), and the penalty intensity is dynamically adjusted by λ (empirical value 0.2).
[0204] The training process employs a dual-network architecture and experience replay technique: the Actor network updates parameters via gradients using a deterministic policy.
[0205]
[0206] θ i Let J be the trainable parameters (weight matrix, bias vector) of Actor network i; J is the objective function of Actor network, which is to maximize the expected cumulative reward. N is the batch size sampled from the experience replay pool (e.g., N = 64); k is the k-th sample in the batch; s (k) The global state s = s for the k-th sample g ,s a ,s p ;a (k) Joint action for the k-th sample It is a scalar function, referring to the Q-value of the Critic network output (which evaluates the long-term value of state-action pairs); It is the gradient vector, referring to the action a of Q on agent i. i gradient (guiding a) i How to improve to increase the Q value); Refers to the policy function (input state s) of Actor network i. i Output action a i ); For the policy function with respect to parameter θ i gradient (guiding θ) i How to update to optimize actions).
[0207] This formula transforms the Q-value gradient into the parameter update direction using the chain rule:
[0208] 1. Indicates directions for improvement in the action (e.g., increasing bandwidth allocation can improve the Q value);
[0209] 2. The action improvement is translated into network parameter updates, ultimately enabling the Actor network to generate actions with higher Q values (e.g., better scheduling decisions w and resource allocation d). i ).
[0210] The Critic network optimizes value assessment by minimizing temporal difference error:
[0211]
[0212] These are the trainable parameters of the Critic network; Here is the loss function for Critic (which needs to be minimized); r (k) s′ represents the immediate reward for the k-th sample; γ represents the future reward decay rate (γ∈[0,1], preferably 0.95); (k) The next state s for the k-th sample t+1 ;μ′(s′ (k) |θ′) is the policy function of the target Actor network (the parameter θ′ is a delayed copy of the main Actor network); The target Q-value (parameter) output by the target Critic network. It is a delayed replica of the main Critic network.
[0213] The formula aims to minimize the mean square error between the predicted Q-value and the target Q-value.
[0214] 1. Predicted Q-value: The Critic network's estimate of the current state-action pair;
[0215] 2. Target Q-value: A more stable estimate (calculated using the target network), updated via gradient descent, making the Critic network's evaluation more accurate (experiments show Q-value prediction error <5%).
[0216] The target network parameters are gradually synchronized through soft updates θ′←τθ+(1-τ)θ′ (τ=0.01). The exploration mechanism employs an Ornstein-Uhlenbeck stochastic process to inject relevant noise into the actions, with a noise attenuation coefficient set to 0.15 to balance exploration and utilization. After training, each domain controller can independently execute optimization strategies, generating scheduling decisions and resource allocation schemes based on real-time status, achieving millisecond-level response.
[0217] In some embodiments, the method further includes:
[0218] Step 105: In response to determining that the target computing task meets the partitioning conditions, the target computing task is divided into multiple subtasks.
[0219] Step 106: Based on the scheduling probability vector output by the target action network, the multiple subtasks are assigned to multiple physical domains contained in the target physical domain for parallel execution.
[0220] Step 107: Receive the subtask execution results returned by each physical domain, and perform fusion processing on the subtask execution results to obtain the execution result of the target computing task.
[0221] In specific implementation, the task sharding and multi-domain parallel execution mechanism includes: based on the flexible scheduling probability vector generated by the aforementioned MADDPG algorithm, this disclosure further provides a task sharding and multi-domain parallel execution mechanism to realize the collaborative processing of a single computing task across multiple physical domains and maximize the utilization of heterogeneous computing resources of the integrated air-space-ground network.
[0222] 1. Task partitioning strategy:
[0223] When the target computation task An has high computational complexity, large data volume, or contains sub-modules that can be processed in parallel, the controller of the current physical domain (such as the ground controller) will initiate the task sharding procedure. The sharding strategy is dynamically adjusted according to the task type:
[0224] Data parallel sharding: For tasks such as image rendering and large-scale sensor data fusion, the input data x is sharded. n Divided into multiple data blocks Each data block can be processed independently.
[0225] Functional parallel partitioning: For a task pipeline that includes multiple stages such as preprocessing, feature extraction, and decision reasoning, it is divided into multiple subtasks according to functional modules. The sharding process must ensure that the data dependencies between subtasks are satisfied and construct a directed acyclic graph (DAG) to describe their execution order.
[0226] 2. Subtask allocation based on scheduling vectors:
[0227] The scheduling probability vector w output by the Actor network in the MADDPG algorithm directly determines the allocation ratio of subtasks after partitioning. The controller dynamically schedules the subtask set to various physical domains based on w.
[0228] The proportion of subtasks allocated to the ground-based physical domain is approximately w. g .
[0229] The proportion of subtasks allocated to the space-based physical domain is approximately w. a .
[0230] The proportion of sub-tasks allocated to the space-based physics domain is approximately w. p .
[0231] In specific allocation, the real-time resource status of each domain (such as...) needs to be considered. The algorithm prioritizes assigning computationally intensive subtasks to ground-based nodes with sufficient computing power, assigning low-latency subtasks to nearby space-based nodes, and assigning global perception subtasks to space-based nodes.
[0232] 3. Parallel execution and result fusion:
[0233] After each subtask is assigned to the target physical domain, it is executed in parallel by the computing nodes within that domain. Local, airborne, and spaceborne controllers are responsible for monitoring the execution status of the subtasks within their respective domains. Once all subtasks have been completed, a designated master control domain (usually the task initiation domain or a ground-based controller) is responsible for collecting the intermediate processing results from each domain, fusing and summarizing the results according to the task DAG, and finally generating the complete output of task An, which is then returned to the application device.
[0234] 4. Enhancement of the system reward function:
[0235] To incentivize the effective use of the sharding parallel mechanism, this embodiment enhances the reward function by introducing a sharding efficiency reward term:
[0236]
[0237] Where η is the slice reward coefficient, and ParallelGain is the parallel gain, which is calculated as follows: ParallelGain = (T sequential -T parallel ) / T sequential T sequential Let T be the theoretical time for the task to execute within a single fastest domain. sequential -T parallel The actual time for parallel execution of the slices is represented by 0. This reward encourages the agent to learn the slice and scheduling strategies that maximize the reduction of the total task processing time.
[0238] The above solution breaks through the limitation of "a task can only be executed in one domain" in traditional methods, and realizes cross-domain collaboration at the task granularity level, which significantly improves system throughput and resource utilization. It is especially suitable for computationally intensive and time-sensitive scenarios such as high-definition remote sensing image processing, wide-area Internet of Things data fusion, and emergency disaster response.
[0239] Through the above-described scheme, this disclosed embodiment overcomes the limitations of traditional discrete decision-making: the flexible scheduling vector increases the task fragmentation execution ratio to 63% (compared to 0% in MAPPO), the conflict penalty mechanism reduces the resource contention rate from 19.8% to 12.4%, and the cross-domain collaboration of centralized Critic stabilizes the latency within 2.1 seconds (a 34.4% improvement over the baseline). The design is specifically adapted to wide-area collaborative scenarios in 6G networks, such as cross-domain fusion processing of satellite remote sensing data and UAV imagery in disaster emergency response, or hierarchical computation offloading of high-precision positioning tasks in vehicle-to-everything (V2X) networks.
[0240] 1. This disclosure significantly optimizes the resource utilization efficiency of the integrated air-space-ground network through the continuous action space design and contention-aware mechanism of the MADDPG algorithm. Compared to traditional discrete decision-making methods (e.g., Boolean scheduling in MAPPO), the flexible scheduling vector allows task fragmentation for multi-domain parallel processing, with a measured task fragmentation ratio as high as 63% (compared to 0% in MAPPO). This enables full utilization of idle computing power from satellites, UAVs, and ground servers, improving resource utilization by 31.2%. Simultaneously, the resource conflict penalty term based on Jaccard similarity explicitly suppresses cross-domain contention, reducing the conflict rate from a baseline of 19.8% to 12.4% (a 38% reduction) in spectrum overlap scenarios, effectively avoiding transmission interruptions or computational failures caused by resource contention. For example, in a typical scenario where satellites and UAVs share a 20MHz spectrum, conflict events are reduced by 52 times per hour, ensuring highly reliable transmission of disaster remote sensing data streams.
[0241] 2. Based on a cross-domain collaborative mechanism using a centralized Critic network, this solution achieves a breakthrough balance between latency and energy consumption. In terms of latency optimization: By prioritizing sensitive tasks such as vehicle control with a dynamic weight (0.7), and by reducing single-domain load accumulation through segmented scheduling of the joint action space, the measured average task latency is reduced to 2.1 seconds (a 34.4% improvement compared to MAPPO's 3.2 seconds). Specifically, the latency for heavy tasks in the ground-based domain is reduced by 41% (from 3.5 seconds to 2.1 seconds), and the response speed for light tasks in the air-based domain is improved by 28% (from 1.8 seconds to 1.3 seconds). In terms of energy consumption control: The continuous adjustment capability of the resource allocation ratio vector matches the real-time load of nodes, avoiding energy waste caused by over-allocation; reduced conflicts and synchronization lower retransmission energy consumption, resulting in a total system energy consumption of 182.3 kJ (a 13.5% decrease compared to MAPPO's 210.7 kJ). Especially in UAV energy-constrained scenarios, the single-unit endurance is extended by 17%, meeting the long-term monitoring needs of remote areas.
[0242] This disclosed embodiment breaks through the traditional paradigm of "discrete decision-making-independent optimization" and achieves a four-dimensional improvement in resource utilization, latency, energy consumption, and robustness through continuous collaborative scheduling, laying a technical foundation for the large-scale deployment of integrated air-space-ground networks in key scenarios such as emergency communication and intelligent transportation.
[0243] The above scheme determines the first loss function of the initial value network based on the reward function and the next value function corresponding to the next state value. The target value network is obtained by updating the initial value network based on the first loss function. The temporal difference error is determined based on the current value function corresponding to the current state value and the next value function corresponding to the next state value. The generalization advantage estimation parameters are determined based on the temporal difference error, and the second loss function of the initial action network is determined based on the generalization advantage estimation parameters. The second loss function is backpropagated to the initial action network, and the network parameters of the initial action network are updated based on the second loss function to obtain the target action network. This improves the stability of the training process and ensures that the algorithm can stably converge to the optimal policy.
[0244] In the above embodiments, when any of the airborne, spaceborne, and ground-based controllers receives a target computing task, the physical domain where the controller receiving the target computing task is located is designated as the current physical domain. Utilizing a pre-trained target action network in the current physical domain, the target physical domain for executing the target computing task is determined based on the target computing resources required by the target computing task and the available resources in each physical domain. In this way, each controller can independently determine the target physical domain for executing the target computing task, quickly and accurately determining the target physical domain based on the real-time available resources in each physical domain, achieving rapid response. If the target physical domain contains other physical domains, the target computing task is allocated to other physical domains besides the current physical domain for execution in those other physical domains. Thus, when it is determined that the target computing task needs to be executed in another physical domain, it can be quickly scheduled to that other physical domain for execution, fulfilling the collaborative needs between different physical domains. If the target physical domain contains the current physical domain, the currently allocated resources of the current physical domain are determined, and the target computing task is executed in the current physical domain using these currently allocated resources. In this way, when it is determined that the target computing task needs to be executed in the target physical domain, the target physical domain can quickly and accurately determine the currently allocated resources, and thus use the currently allocated resources to execute the target computing task in the target physical domain.
[0245] It should be noted that the method of this disclosure embodiment can be executed by a single device, such as a computer or server. The method of this embodiment can also be applied to a distributed scenario, where multiple devices cooperate to complete the task. In such a distributed scenario, one of these devices may execute only one or more steps of the method of this disclosure embodiment, and the multiple devices will interact with each other to complete the method described.
[0246] It should be noted that the above description describes some embodiments of this disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in a different order than that shown in the above embodiments and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0247] Based on the same inventive concept, corresponding to any of the above-described embodiments, this disclosure also provides a network resource allocation device.
[0248] refer to Figure 4 The network resource allocation device includes: an airborne controller located in the airborne physical domain, a spaceborne controller located in the space-based physical domain, and a ground-based controller located in the ground-based physical domain.
[0249] The current physical domain determination module 401 is configured to, in response to determining that any one of the air-based controller, the space-based controller, and the ground-based controller has received a target computing task, take the physical domain where the controller that received the target computing task is located as the current physical domain.
[0250] The target physical domain determination module 402 is configured to use a pre-trained target action network in the current physical domain to determine the target physical domain for executing the target computing task, based on the target computing resources required for the target computing task and the available resources in each physical domain.
[0251] The first allocation module 403 is configured to allocate the target computing task to the other physical domain in response to determining that the target physical domain contains other physical domains, so that the target computing task can be executed in the other physical domains; wherein, the other physical domains are physical domains other than the current physical domain;
[0252] The second allocation module 404 is configured to, in response to determining that the target physical domain contains the current physical domain, determine the currently allocated resources of the current physical domain and execute the target computing task in the current physical domain using the currently allocated resources.
[0253] In some embodiments, the apparatus further includes a training module, the training module comprising:
[0254] The resource allocation method determination unit is configured to acquire sample computing tasks and use the initial action network to determine the resource allocation method based on the sample computing resources required by the sample computing tasks and the available resources in each physical domain.
[0255] The reward function determination unit is configured to determine the total system cost when the sample calculation task is executed according to the resource allocation method, and to determine the reward function based on the total system cost;
[0256] The global state determination unit is configured to acquire the state of space-based resources in the space-based physical domain, the state of space-based resources in the space-based physical domain, and the state of ground-based resources in the ground-based physical domain, and determine the global state based on the state of space-based resources, the state of space-based resources, and the state of ground-based resources;
[0257] The update processing unit is configured to use the initial value network to determine a value function based on the resource allocation method, the reward function, and the global state, and to update the network parameters of the initial action network by maximizing the value function to obtain the target action network.
[0258] In some embodiments, the reward function determination unit includes:
[0259] The total latency determination subunit is configured to determine the transmission latency and computation latency in each physical domain when the sample computation task is executed according to the resource allocation method, and to determine the total latency in each physical domain based on the transmission latency and the computation latency;
[0260] The total energy consumption determination subunit is configured to determine the transmission energy consumption and computing energy consumption in each physical domain when the sample computing task is executed according to the resource allocation method, and to determine the total energy consumption in each physical domain based on the transmission energy consumption and the computing energy consumption.
[0261] The reward function determination subunit is configured to determine the total system cost when performing the sample calculation task according to the resource allocation method based on the total latency and the total energy consumption, and use the reciprocal of the total system cost as the reward function.
[0262] In some embodiments, the total delay determination subunit is specifically configured as follows:
[0263] Determine the data transfer rate between the application device and each physical domain when performing the sample computation task according to the described resource allocation method.
[0264]
[0265] in, The data transmission rate between the application device n and each physical domain i. p is the link bandwidth resource allocated for the application device n to communicate with each physical domain i. n The transmission power of the task is calculated for the sample. The channel gain is defined as the communication gain between the application device n and each physical domain i. σ is the distance from the application device n to the computing nodes in each physical domain i. i θ represents the background noise of the channel when the application device n communicates with each physical domain i, and θ is the distance attenuation factor.
[0266] The transmission delay of the sample computation task to each physical domain is determined based on the data transmission rate.
[0267]
[0268] in, The transmission delay for the sample task to be transmitted to each physical domain i is calculated. x is the data transmission rate between the application device n and each physical domain i. nThe amount of input data for the sample calculation task;
[0269] When executing the sample computation task according to the described resource allocation method, determine the sample allocation resources in each physical domain, and determine the computation latency of the sample computation task in each physical domain based on the sample allocation resources.
[0270]
[0271] in, c represents the computation latency of the sample computation task performed in each physical domain i. n The total computing resources for the sample computation task. Allocate resources for the sample computation task in each physical domain i;
[0272] Based on the transmission delay and the computation delay, determine the total delay in each physical domain when executing the sample computation task according to the resource allocation method.
[0273]
[0274] in, The total latency in each physical domain i when performing the sample computation task according to the resource allocation method, The transmission delay for the sample task to be transmitted to each physical domain i is calculated. The computational latency for the sample computation task to be performed in each physical domain i.
[0275] In some embodiments, the reward function determines the sub-unit, specifically configured as follows:
[0276] Based on the total latency and the total energy consumption, determine the total system cost when executing the sample computation task according to the resource allocation method.
[0277]
[0278] Among them, C t The total system cost is defined as the resource allocation method used to execute the sample computation task, where N is the number of sample computation tasks and β is the weighting coefficient. This refers to the scheduling position when the sample computation task is executed according to the described resource allocation method. The total latency in each physical domain i when performing the sample computation task according to the resource allocation method, The total energy consumption in each physical domain i when performing the sample computation task according to the resource allocation method;
[0279] The reciprocal of the total system cost is used as the reward function.
[0280]
[0281] Where, r t For the reward function, C t The total system cost when performing the sample computation task according to the resource allocation method.
[0282] In some embodiments, the update processing unit includes:
[0283] The storage subunit is configured to determine the current state value at the current moment and the next state value at the next moment from the global state, and store the current state value, the resource allocation method, the reward function and the next state value as tuple data at the current moment into the experience replay pool;
[0284] The first loss function determination sub-unit is configured to determine the first loss function of the initial value network based on the reward function and the next value function corresponding to the next state value.
[0285]
[0286] in, Let the first loss function of the initial value network be . For the reward function, The next state value s t+1 The corresponding next value function;
[0287] The first update processing subunit is configured to update the initial value network based on the first loss function to obtain the target value network.
[0288] The entropy coefficient determination subunit is configured to determine the time difference error based on the current value function corresponding to the current state value and the next value function corresponding to the next state value.
[0289]
[0290] Where, δ t The time difference error is... Let γ be the reward function, and γ be the discount factor. The next state value s t+1 The corresponding next value function, The current state value s t The corresponding current value function;
[0291] The generalization advantage estimation parameter determination subunit is configured to determine the generalization advantage estimation parameters based on the time difference error.
[0292]
[0293] in, δ is the generalization advantage estimation parameter. t The time difference error is denoted as γ, where γ is the discount factor and λ is the smoothing parameter.
[0294] The second loss function determination subunit is configured to determine the second loss function of the initial action network based on the generalization advantage estimation parameters.
[0295]
[0296] Among them, L CLIP (θ i ) is the second loss function of the initial action network. The average of the tuple data at all times in the experience replay pool. To limit parameters, The generalization advantage estimation parameters are as follows: Let ε be the pruning function, ε be a preset positive number, and δ be the coefficient of the entropy of the resource allocation method. The resource allocation method Entropy;
[0297] The second update processing subunit is configured to backpropagate the second loss function to the initial action network and update the network parameters of the initial action network according to the second loss function to obtain the target action network.
[0298] In some embodiments, the apparatus further includes:
[0299] The task partitioning module is configured to divide the target computing task into multiple subtasks in response to determining that the target computing task meets the partitioning conditions.
[0300] The subtask allocation module is configured to allocate the multiple subtasks to multiple physical domains contained in the target physical domain for parallel execution based on the scheduling probability vector output by the target action network;
[0301] The execution result generation module is configured to receive the subtask execution results returned by each physical domain, and to perform fusion processing on the subtask execution results to obtain the execution result of the target computing task.
[0302] For ease of description, the above apparatus is described in terms of its functions, divided into various modules. Of course, in implementing this disclosure, the functions of each module can be implemented in one or more software and / or hardware.
[0303] The apparatus of the above embodiments is used to implement the corresponding network resource allocation method in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0304] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, this disclosure also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the network resource allocation method described in any of the above embodiments.
[0305] Figure 5 This embodiment illustrates a more specific hardware structure of an electronic device, which may include a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, memory 1020, input / output interface 1030, and communication interface 1040 are interconnected internally via the bus 1050.
[0306] The processor 1010 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0307] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program code is stored in the memory 1020 and is called and executed by the processor 1010.
[0308] The input / output interface 1030 is used to connect input / output modules to realize information input and output. Input / output modules can be configured as components within the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Input devices may include keyboards, mice, touchscreens, microphones, various sensors, etc., while output devices may include displays, speakers, vibrators, indicator lights, etc.
[0309] The communication interface 1040 is used to connect a communication module (not shown in the figure) to enable communication between this device and other devices. The communication module can communicate via wired means (such as USB (Universal Serial Bus), network cable, etc.) or wireless means (such as mobile network, WIFI (Wireless Fidelity), Bluetooth, etc.).
[0310] Bus 1050 includes a pathway for transmitting information between various components of the device, such as processor 1010, memory 1020, input / output interface 1030, and communication interface 1040.
[0311] It should be noted that although the above-described device only shows the processor 1010, memory 1020, input / output interface 1030, communication interface 1040, and bus 1050, in specific implementations, the device may also include other components necessary for normal operation. Furthermore, those skilled in the art will understand that the above-described device may only include the components necessary for implementing the embodiments of this specification, and not necessarily all the components shown in the figures.
[0312] The electronic devices described above are used to implement the corresponding network resource allocation methods in any of the foregoing embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0313] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, this disclosure also provides a non-transitory computer-readable storage medium storing computer instructions for causing the computer to execute the network resource allocation method as described in any of the above embodiments.
[0314] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.
[0315] The computer instructions stored in the storage medium of the above embodiments are used to cause the computer to execute the network resource allocation method as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0316] Based on the same inventive concept, corresponding to any of the above embodiments, this application also provides a computer program product, including computer program instructions. When the computer program instructions are run on a computer, the computer executes the network resource allocation method as described in any of the above embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0317] It is understood that before using the technical solutions of the various embodiments in this disclosure, users will be informed of the type, scope of use, and usage scenarios of the personal information involved in an appropriate manner, and user authorization will be obtained.
[0318] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose, based on the prompt message, whether to provide personal information to the software or hardware such as electronic devices, applications, servers, or storage media performing the operations of this disclosed technical solution.
[0319] As an optional but not limited implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0320] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0321] Those skilled in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of this disclosure is limited to these examples; within the framework of this disclosure, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of the embodiments of this disclosure as described above, which are not provided in detail for the sake of brevity.
[0322] Additionally, to simplify the description and discussion, and to avoid obscuring the embodiments of this disclosure, the provided drawings may or may not show well-known power / ground connections to integrated circuit (IC) chips and other components. Furthermore, the apparatus may be shown in block diagram form to avoid obscuring the embodiments of this disclosure, and this also takes into account the fact that the details of implementation of these block diagram apparatuses are highly dependent on the platform on which the embodiments of this disclosure will be implemented (i.e., these details should be fully understood by those skilled in the art). While specific details (e.g., circuitry) have been set forth to describe exemplary embodiments of this disclosure, it will be apparent to those skilled in the art that the embodiments of this disclosure may be implemented without these specific details or with variations thereof. Therefore, these descriptions should be considered illustrative rather than restrictive.
[0323] Although this disclosure has been described in conjunction with specific embodiments thereof, many substitutions, modifications, and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed.
[0324] This disclosure is intended to cover all such substitutions, modifications, and variations that fall within the broad scope of this disclosure. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this disclosure should be included within the protection scope of this disclosure.
Claims
1. A method for allocating network resources, characterized in that, An application to a network resource allocation device, the device comprising: a space-based controller disposed in a space-based physical domain, a space-based controller disposed in a space-based physical domain, and a ground-based controller disposed in a ground-based physical domain; the method comprising: In response to determining that any one of the airborne controller, the space-based controller, and the ground-based controller has received a target computing task, the physical domain where the controller that received the target computing task is located is taken as the current physical domain; Using the pre-trained target action network in the current physical domain, the target physical domain for executing the target computing task is determined based on the target computing resources required for the target computing task and the available resources in each physical domain; In response to determining that the target physical domain contains other physical domains, the target computing task is assigned to one of the other physical domains for execution in that other physical domain; wherein, the other physical domain is a physical domain other than the current physical domain; In response to determining that the target physical domain contains the current physical domain, the currently allocated resources of the current physical domain are determined, and the target computing task is executed in the current physical domain using the currently allocated resources; The pre-training process of the target action network includes: Obtain sample computing tasks, and use the initial action network to determine the resource allocation method based on the sample computing resources required by the sample computing tasks and the available resources in each physical domain; Determine the total system cost when executing the sample computation task according to the resource allocation method, and determine the reward function based on the total system cost; The status of space-based resources in the space-based physical domain, the status of space-based resources in the space-based physical domain, and the status of ground-based resources in the ground-based physical domain are obtained, and the global status is determined based on the status of space-based resources, the status of space-based resources, and the status of ground-based resources. The initial value network is used to determine the value function based on the resource allocation method, the reward function, and the global state. The network parameters of the initial action network are updated by maximizing the value function to obtain the target action network. The process of determining a value function using an initial value network based on the resource allocation method, the reward function, and the global state, and then updating the network parameters of the initial action network by maximizing the value function to obtain the target action network includes: The current state value and the next state value are determined from the global state, and the current state value, the resource allocation method, the reward function, and the next state value are stored as tuple data in the experience replay pool. The first loss function of the initial value network is determined based on the reward function and the next value function corresponding to the next state value. in, Let the first loss function of the initial value network be . For the reward function, The next state value The corresponding next value function; The target value network is obtained by updating the initial value network based on the first loss function; The time difference error is determined based on the current value function corresponding to the current state value and the next value function corresponding to the next state value. in, The time difference error is... For the reward function, As a discount factor, The next state value The corresponding next value function, The current state value The corresponding current value function; The generalization advantage estimation parameters are determined based on the time difference error. in, The generalization advantage estimation parameters are as follows: The time difference error is... As a discount factor, For smoothing parameters; The second loss function of the initial action network is determined based on the generalization advantage estimation parameters. in, The second loss function of the initial action network is... The average of the tuple data at all times in the experience replay pool. To limit parameters, The generalization advantage estimation parameters are as follows: For the clipping function, As a preset positive number, Let be the coefficient of entropy for the resource allocation method. The resource allocation method Entropy; The second loss function is backpropagated to the initial action network, and the network parameters of the initial action network are updated according to the second loss function to obtain the target action network.
2. The method according to claim 1, characterized in that, The determination of the total system cost when executing the sample computation task according to the resource allocation method, and the determination of the reward function based on the total system cost, includes: Determine the transmission latency and computation latency in each physical domain when executing the sample computation task according to the resource allocation method, and determine the total latency in each physical domain based on the transmission latency and the computation latency; Determine the transmission energy consumption and computing energy consumption in each physical domain when executing the sample computing task according to the resource allocation method, and determine the total energy consumption in each physical domain based on the transmission energy consumption and the computing energy consumption; Based on the total latency and the total energy consumption, the total system cost when executing the sample calculation task according to the resource allocation method is determined, and the reciprocal of the total system cost is used as the reward function.
3. The method according to claim 2, characterized in that, The step of determining the transmission latency and computation latency in each physical domain when executing the sample computation task according to the resource allocation method, and determining the total latency in each physical domain based on the transmission latency and the computation latency, includes: Determine the data transfer rate between the application device and each physical domain when performing the sample computation task according to the described resource allocation method. in, For the application device With each physical domain Data transfer rate between For the application device With each physical domain The link bandwidth resources allocated during communication. This represents the logarithmic operation with base 2. The transmission power of the task is calculated for the sample. For the application device With each physical domain Channel gain during communication The application device is To each physical domain The distance between the computing nodes For the application device With each physical domain Channel background noise during communication This is the distance attenuation factor; The transmission delay of the sample computation task to each physical domain is determined based on the data transmission rate. in, The sample computation task is transmitted to each physical domain. Transmission delay, For the application device With each physical domain Data transfer rate between The amount of input data for the sample calculation task; When executing the sample computation task according to the described resource allocation method, determine the sample allocation resources in each physical domain, and determine the computation latency of the sample computation task in each physical domain based on the sample allocation resources. in, The computational task for the samples is performed in each physical domain. The computational latency performed in the process, The total computing resources for the sample computation task. The computational task for the samples is performed in each physical domain. Resource allocation for samples; Based on the transmission delay and the computation delay, determine the total delay in each physical domain when executing the sample computation task according to the resource allocation method. in, In order to perform the sample computation task in accordance with the resource allocation method in each physical domain Total latency in the process, The sample computation task is transmitted to each physical domain. Transmission delay, The computational task for the samples is performed in each physical domain. The computational latency during execution.
4. The method according to claim 2, characterized in that, The step of determining the total system cost when executing the sample computation task according to the resource allocation method based on the total latency and the total energy consumption, and using the reciprocal of the total system cost as the reward function, includes: Based on the total latency and the total energy consumption, determine the total system cost when executing the sample computation task according to the resource allocation method. in, This represents the total system cost when performing the sample computation task according to the described resource allocation method. Calculate the number of tasks for the sample. These are the weighting coefficients. This refers to the scheduling position when the sample computation task is executed according to the described resource allocation method. In order to perform the sample computation task in accordance with the resource allocation method in each physical domain Total latency in the process, In order to perform the sample computation task in accordance with the resource allocation method in each physical domain Total energy consumption; The reciprocal of the total system cost is used as the reward function. in, For the reward function, The total system cost when performing the sample computation task according to the resource allocation method.
5. The method according to claim 1, characterized in that, The method further includes: In response to determining that the target computing task meets the partitioning conditions, the target computing task is divided into multiple subtasks; Based on the scheduling probability vector output by the target action network, the multiple subtasks are assigned to multiple physical domains contained in the target physical domain for parallel execution; The system receives the subtask execution results returned by each physical domain and performs fusion processing on the subtask execution results to obtain the execution result of the target computing task.
6. A network resource allocation device, characterized in that, The device includes: an airborne controller disposed in the airborne physical domain, a spaceborne controller disposed in the space-based physical domain, and a ground-based controller disposed in the ground-based physical domain. The current physical domain determination module is configured to, in response to determining that any one of the airborne controller, the space-based controller, and the ground-based controller has received a target computing task, take the physical domain where the controller that received the target computing task is located as the current physical domain; The target physical domain determination module is configured to use a pre-trained target action network in the current physical domain to determine the target physical domain for executing the target computing task, based on the target computing resources required for the target computing task and the available resources in each physical domain. The first allocation module is configured to allocate the target computing task to the other physical domain in response to determining that the target physical domain contains other physical domains, so that the target computing task can be executed in the other physical domains; wherein, the other physical domains are physical domains other than the current physical domain; The second allocation module is configured to, in response to determining that the target physical domain contains the current physical domain, determine the currently allocated resources of the current physical domain and execute the target computing task in the current physical domain using the currently allocated resources; The device further includes a training module, the training module comprising: The resource allocation method determination unit is configured to acquire sample computing tasks and use the initial action network to determine the resource allocation method based on the sample computing resources required by the sample computing tasks and the available resources in each physical domain. The reward function determination unit is configured to determine the total system cost when the sample calculation task is executed according to the resource allocation method, and to determine the reward function based on the total system cost; The global state determination unit is configured to acquire the state of space-based resources in the space-based physical domain, the state of space-based resources in the space-based physical domain, and the state of ground-based resources in the ground-based physical domain, and determine the global state based on the state of space-based resources, the state of space-based resources, and the state of ground-based resources; The update processing unit is configured to use the initial value network to determine a value function based on the resource allocation method, the reward function, and the global state, and update the network parameters of the initial action network by maximizing the value function to obtain the target action network; The update processing unit includes: The storage subunit is configured to determine the current state value at the current moment and the next state value at the next moment from the global state, and store the current state value, the resource allocation method, the reward function and the next state value as tuple data at the current moment into the experience replay pool; The first loss function determination sub-unit is configured to determine the first loss function of the initial value network based on the reward function and the next value function corresponding to the next state value. in, Let the first loss function of the initial value network be . For the reward function, The next state value The corresponding next value function; The first update processing subunit is configured to update the initial value network based on the first loss function to obtain the target value network. The entropy coefficient determination subunit is configured to determine the time difference error based on the current value function corresponding to the current state value and the next value function corresponding to the next state value. in, The time difference error is... For the reward function, As a discount factor, The next state value The corresponding next value function, The current state value The corresponding current value function; The generalization advantage estimation parameter determination subunit is configured to determine the generalization advantage estimation parameters based on the time difference error. in, The generalization advantage estimation parameters are as follows: The time difference error is... As a discount factor, For smoothing parameters; The second loss function determination subunit is configured to determine the second loss function of the initial action network based on the generalization advantage estimation parameters. in, The second loss function of the initial action network is... The average of the tuple data at all times in the experience replay pool. To limit parameters, The generalization advantage estimation parameters are as follows: For the clipping function, As a preset positive number, Let be the coefficient of entropy for the resource allocation method. The resource allocation method Entropy; The second update processing subunit is configured to backpropagate the second loss function to the initial action network and update the network parameters of the initial action network according to the second loss function to obtain the target action network.
7. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor, when executing the program, implements the method as claimed in any one of claims 1 to 5.
8. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium stores computer instructions for causing a computer to perform the method described in any one of claims 1 to 5.
Citation Information
Patent Citations
Slice-based collaborative task unloading method in air-space-ground integrated Internet of Vehicles
CN116193396A
Task unloading and resource allocation method, equipment, medium and system
CN116582836A