A risk-aware resource scheduling method

By constructing an integrated air-space-ground network model and utilizing CVaR and dual-critic network optimization for resource scheduling, the problem of uneven resource allocation in the integrated air-space-ground network was solved, achieving efficient computing resource management and risk control, and improving service quality.

CN119383661BActive Publication Date: 2025-10-28CHONGQING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411653420.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-19
Publication Date
2025-10-28
Estimated Expiration
2044-11-19

AI Technical Summary

Technical Problem

In integrated air-space-ground networks, traditional computing task scheduling methods are difficult to effectively balance computing resource allocation and risk management, leading to network congestion, reduced service quality, and an inability to meet the diverse needs of users.

Method used

We construct communication, latency, and energy consumption models for an integrated air-space-ground network, utilize Conditional Risk Value (CVaR) and dual-critic network to assess costs and risks, and optimize resource scheduling strategies through deep neural networks to achieve optimal task offloading and resource allocation.

Benefits of technology

It significantly improved service quality, balanced costs and risks, optimized resource allocation, and met the diverse needs of users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119383661B_ABST
    Figure CN119383661B_ABST
Patent Text Reader

Abstract

This invention relates to a risk-aware resource scheduling method, belonging to the field of radio communication technology. The method includes: constructing communication models, latency models, and energy consumption models for the space-based, air-based, and ground-based networks in an integrated air-space-ground network; formulating the problem as an uncertain integer nonlinear optimization problem under the energy constraints of the UAV, and using the conditional risk value CVaR to measure and check the risk; using a dual-critic network to evaluate costs and risks, with each network independently evaluating the action-value function, and then performing a comprehensive evaluation; implementing optimal task offloading and resource scheduling strategies through parameterization and time-series difference updates via a deep neural network; reducing the impact of the inaccuracy of individual networks on strategy updates while improving the stability and convergence of learning; the risk-aware resource scheduling method proposed in this invention can meet diverse user needs and solve the high-risk problems caused by traditional decision-making.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of communication technology and relates to a resource scheduling method based on risk perception. Background Technology

[0002] With the development of mobile communication technology, applications running on user devices are becoming increasingly diverse and complex, leading to more frequent computing demands. These mobile applications often require substantial computing resources and consume significant energy during operation. However, current user devices are typically constrained by their limited computing resources and battery capacity, making it difficult to meet the applications' demands for low latency and high computing power. Users can offload computing tasks to nearby ground base stations, which not only reduces costs and latency but also conserves their own resources. However, simply relying on offloading tasks to ground base stations makes it difficult to ensure that task processing achieves ideal performance levels. Ground networks centered around base stations often face overload issues, potentially leading to network congestion, slow processing, and degraded service quality. Drones and satellites are considered powerful supplements to enhance ground capabilities. Satellite communication serves as an effective means of providing global communication coverage. To reduce transmission latency and resource consumption and provide reliable services to users, Mobile Edge Computing (MEC) technology has emerged. MEC servers can be deployed at Low Earth Orbit Satellites (LEO) to provide computing services to users. Drones, characterized by flexible deployment and agile management, can act as relays to transmit data or provide communication and computing resources on demand. Integrating these three elements constitutes an integrated air-space-ground network. This integrated network is a heterogeneous, multi-dimensional network; the fusion of multiple networks makes its structure extremely complex, and its resources exhibit diverse characteristics. Due to the highly dynamic and random nature of computational task arrival, traditional reinforcement learning typically focuses on maximizing cumulative rewards. However, in some cases, considering only reward maximization may lead to high-risk decisions. Therefore, scheduling computational tasks in an integrated air-space-ground network still faces significant challenges. Designing an effective resource allocation strategy is therefore crucial, as it can both meet diverse user needs and address the high-risk problems arising from traditional decision-making. Summary of the Invention

[0003] To achieve the above objectives, the present invention provides the following technical solution:

[0004] A risk-aware-based resource scheduling method includes the following steps:

[0005] S1: In the integrated air-space-ground network, construct communication models, latency models, and energy consumption models for the space-based network, air-based network, and ground-based network respectively;

[0006] S2: Under the energy constraints of the UAV, the problem is formulated as an integer nonlinear optimization problem with uncertainty, and the conditional risk value CVaR is used to measure the inspection risk;

[0007] S3: Utilize a dual-critic network to evaluate costs and risks, and achieve optimal task offloading and resource scheduling strategies through parameterization and temporal difference (TD) updates using deep neural networks (DNN).

[0008] Furthermore, S1 includes the following steps:

[0009] In ground-to-air communication, the channel gain for user k and UAV m is defined as:

[0010]

[0011] Where α0 is the channel gain at a reference distance d = 1m, q(i+1) is the position of the UAV after moving m in the next time slot i, and p k (i) represents the position of user k in time slot i, H is the fixed flight altitude of UAV m, and ||q(i+1)-p k (i)|| 2 +H 2 This represents the Euclidean distance between user k and drone m;

[0012] The data rate unloaded onto drone m is expressed as:

[0013]

[0014] Among them, B u For the available spectrum bandwidth from user k to drone m, η k To determine the percentage of bandwidth allocated to user k, we need to calculate ∑η k <1, P u For the user's transmit power on the uplink, g k (i) represents the channel gain. For noise power, f k (i) Check if there is a blockage between user k and drone m. If there is a blockage, add an extra P. NLOS noise power, P NLOS Indicates additional noise power;

[0015] In terrestrial / airborne-space communication, the channel gain from user k to the satellite is determined by the fixed distance D between the satellite and the ground and the satellite's computing power f. LEO The calculated data rate from the ground or drone (m) to the satellite is expressed as:

[0016]

[0017] in, For the data rate from ground or drone m to satellite, B l P represents the communication bandwidth to the satellite. l This refers to the transmit power on this link. Noise power;

[0018] set up This represents the partial uninstallation decision made by user k. The percentage of data that user k unloads onto drone m. The percentage of user k that is unloaded to low-Earth orbit satellites;

[0019] The partial uninstallation decision for user k is set as follows:

[0020]

[0021] in, The percentage of data that user k unloads onto drone m. The percentage of user k that is unloaded to low-Earth orbit satellites.

[0022] Furthermore, S1 also includes the following steps:

[0023] Based on task size b k (i) Number of cycles required per bit s k (i) The user's computing power f UE Locally calculated proportion Get the local computation latency of user k

[0024] The computing resources allocated to user k by drone m in each time slot are represented as follows:

[0025]

[0026] Where κ1 is the energy consumption coefficient of UAV m, p k (i) represents the computing resources allocated to user k by drone m, expressed as:

[0027] p k (i)=μ k P

[0028] Where, μ k The proportion of resource allocation, ∑μ k <1;

[0029] Computational tasks from unloading to the drone Number of cycles required per bit (s) k (i) The computing power f of the droneUAV,k The computational delay D for unloading the m-part of the task onto the drone was calculated. UAV,k (i);

[0030] Priority factor, expressed as:

[0031]

[0032] Among them, t k The priority factor is the time when the task was generated, and t is the current time. The smaller the priority factor, the higher the priority of the task, so that it can be unloaded and processed earlier.

[0033] Furthermore, S1 also includes the following steps:

[0034] Energy consumption transmitted to drone m and calculate energy consumption They are represented as follows:

[0035]

[0036]

[0037] Depend on and The sum gives the local computation latency of user k.

[0038] The flight energy consumption E of drone m fly (i) is represented as:

[0039] E fly (i)=φ||ν(i)|| 2

[0040] Here, ν(i) includes the UAV's flight speed and flight angle. A modulo operation ||ν(i)|| is performed to obtain the flight speed, which is in the range [0, ν]. max ] takes a value from, ν max φ represents the maximum flight speed; β(i) represents the flight angle, β(i)∈[0,2π]; φ represents the relevant parameters of the UAV.

[0041] Furthermore, S2 includes the following steps:

[0042] drone m-track Represented as:

[0043]

[0044] Uninstallation strategy Λ, represented as:

[0045]

[0046] Resource allocation strategy Φ is represented as:

[0047]

[0048] Latency D(i) is defined as the maximum value of the latency of each part of the processing task, expressed as:

[0049]

[0050] The problem can be modeled as follows:

[0051]

[0052]

[0053]

[0054]

[0055]

[0056]

[0057] Where C1 represents the range of values ​​for the unloading ratio; C2 represents the resource allocation to all users not exceeding the total resources; C3 determines whether the latency tolerance of each user k is met; C4 represents the drone's movement range not exceeding the predetermined range; C5 represents the drone's flight energy consumption, computing energy consumption, and transmission energy consumption within the time slot not exceeding the drone's energy capacity constraint, E b This represents the energy capacity of the drone;

[0058] Using Markov decision-making, the cost c is obtained, expressed as:

[0059]

[0060] Where ζ is the penalty for exceeding the time delay constraint, o k (i) is used to determine whether the task latency tolerance of each user k is met;

[0061] The current scheduling problem can be described as a constrained Markov decision process, expressed as:

[0062]

[0063] stC1,C2,C3,C4

[0064]

[0065] The Conditional Value at Risk (CVaR) is used as a measure of risk, which measures the average loss under the worst-case scenario at a certain confidence level.

[0066] Let X be a random variable representing the amount of energy consumed by the drone exceeding its energy capacity limit:

[0067]

[0068] in, This refers to the actual energy consumed by the drone, E. b X represents the drone's energy capacity; when X > 0, it indicates that the drone's energy consumption has exceeded the limit.

[0069] Set the confidence level α = 0.05; for a given α, calculate the Value at Risk (VaR) to satisfy P(X ≤ -VaR) = α; find all data points X that satisfy X > -VaR. j (j=1,2,...,m), calculate CVaR, expressed as:

[0070]

[0071] We take CVaR as the main component of the risk function; let the risk function R(X) = CVaR α (X).

[0072] Furthermore, S3 includes the following steps:

[0073] The expected long-term discount cost value function V(s) is expressed as:

[0074]

[0075] in, This represents the discount factor in the current time slot, used to measure the importance of future rewards relative to current rewards;

[0076] Define the Q-value function to further evaluate the expected long-run discount cost obtained under the current state-action pair, expressed as:

[0077]

[0078] With the goal of minimizing cost, select the function with the minimum Q value Q. * (s,a) = minQ(s,a) as the optimal Q value; simultaneously, take action a. * Obtain the optimal Q value; the Q-value function is updated according to the Temporal Difference (TD) equation, expressed as:

[0079]

[0080] Where η represents the learning rate, which is used to determine the magnitude of each update;

[0081] Parameterization is performed using Deep Neural Networks (DNNs), and the loss function is expressed as:

[0082]

[0083] in, This represents the Q-value function based on DNN. For TD error;

[0084] The risk function for the portion exceeding the drone's energy capacity is expressed as:

[0085]

[0086] Wherein, Ω is the defined set of error states, indicating that the drone's energy consumption exceeds its energy capacity.

[0087] Furthermore, S3 also includes the following steps:

[0088] Value function of expected long-term discount risk Represented as:

[0089]

[0090] The expected long-term discount risk estimated using the Q-value function is expressed as:

[0091]

[0092] Perform a greedy action To obtain the optimal Q value

[0093] Suppose the optimal Q values ​​obtained under two different objectives are Q(s) i ,a i )and By adding a dynamically updated weighting coefficient δ between the two to achieve balance, a new Q-value function is obtained, expressed as:

[0094]

[0095] Furthermore, S3 also includes the following steps:

[0096] Initialize parameterized cost network Risk Network and the action network μ(s|θ) μ Initialization parameters Experience replay Weights δ and step size Δ; based on μ(s|θ) μ Choose a i Storage experience (s) i ,a i ,c i ,r i ,s i+1 )arrive middle;

[0097] Based on the target value y j To update, it is indicated as:

[0098]

[0099] Calculate the network loss L c , represented as:

[0100]

[0101] Obtain the optimal strategy action Represented as:

[0102]

[0103] The weights are updated simultaneously with the parameters through a soft update, expressed as follows:

[0104]

[0105] if Then δ←δ+Δ, otherwise δ←δ-Δ.

[0106] Furthermore, the maximum value λ of the unloading ratio max The unloading ratio is constrained by a value of 0.5, expressed as follows:

[0107]

[0108] The beneficial effects of this invention are as follows:

[0109] Since link availability and task arrival are highly dynamic, the focus is on minimizing the average processing latency of all tasks. In traditional integrated air-space-ground networks, communication, latency, and energy consumption models are constructed for each of the three network layers, taking into account their respective characteristics. To address the instability caused by dynamic task arrival and optimize the bias of estimation results caused by a single target, an uncertain integer nonlinear optimization problem is proposed under the energy constraint of the UAV, and a conditional risk value (CVaR) is introduced to measure the risk. A dual-critic network structure is designed to balance cost and risk, and a risk-aware resource scheduling method is proposed. The method provided in this invention can achieve a significant improvement in service quality and achieve a good trade-off between cost and risk.

[0110] Other advantages of the invention: The objectives and features will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from an examination of the following, or may be learned from the practice of the invention. Attached Figure Description

[0111] To make the objectives, technical solutions, and advantages of the present invention clearer, the preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings, wherein:

[0112] Figure 1 This is a network system framework diagram according to an embodiment of the present invention;

[0113] Figure 2 These are the steps of the risk-aware resource scheduling method according to an embodiment of the present invention;

[0114] Figure 3 This is a flowchart illustrating the problem-solving method of the risk-aware resource scheduling method according to an embodiment of the present invention. Detailed Implementation

[0115] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0116] Please see Figure 1 This is a network system framework diagram according to an embodiment of the present invention; please refer to [link / reference]. Figure 2The steps of the risk-aware resource scheduling method according to an embodiment of the present invention are described below; please refer to [link to relevant documentation]. Figure 3 This is a flowchart illustrating the problem-solving method of the risk-aware resource scheduling approach according to an embodiment of the present invention. Combined with... Figures 1 to 3 This invention specifically describes a risk-aware resource scheduling method, which includes the following steps:

[0117] S1: In the integrated air-space-ground network, communication models, latency models, and energy consumption models of the three-layer network are constructed respectively based on their respective characteristics.

[0118] S1.1: The system under consideration consists of K users, M drones, and one LEO satellite. Users k = {1, 2, ..., K}, and drones m = {1, 2, ..., M}. Assume the satellite consistently covers the region within a period T, dividing the time into I time slots, then T = {1, 2, ..., I}. The positions of user k and drone m in each time slot are represented by p. k (i), q(i). The task generated by the user in time slot i is represented as a tuple. Where b k (i) represents the size of the task generated by user k in time slot i. s represents the latency tolerance of all tasks generated by the same user k. k (i) represents the number of cycles per bit required to process the task.

[0119] S1.2: The channel gain for user k and drone m is defined as follows:

[0120]

[0121] Where α0 is the channel gain at a reference distance d = 1m, q(i+1) is the position of the UAV after moving m in the next time slot i, and p k (i) represents the position of user k in time slot i, H is the fixed flight altitude of UAV m, and ||q(i+1)-p k (i)|| 2 +H 2 This represents the Euclidean distance between user k and drone m.

[0122] Data rate unloaded to drone m Represented as:

[0123]

[0124] Among them, B u For the available spectrum bandwidth from user k to drone m, η k To determine the percentage of bandwidth allocated to user k, we need to calculate ∑η k <1, Pu For the user's transmit power on the uplink, g k (i) represents the channel gain. For noise power, f k (i) is used to indicate whether there is a blockage between user k and drone m. If there is a blockage, add P. NLOS noise power, P NLOS This indicates additional noise power.

[0125] Ground / Air-Space Communication: The channel gain from user k to the satellite can be determined by the fixed distance D between the satellite and the ground and the satellite's computing power f. LEO The calculated data rate unloaded to LEO is expressed as:

[0126]

[0127] in, For the data rate from ground or drone m to satellite, B l P represents the communication bandwidth to the satellite. l This refers to the transmit power on this link. This represents noise power.

[0128] S1.3: Set the partial offloading decision for user k, represented as:

[0129]

[0130] in, The percentage of data that user k unloads onto drone m. This represents the percentage of the data that user k will unload to low-Earth orbit satellites. The maximum unloading percentage is set to 0.5, which is expressed as:

[0131]

[0132] S1.4: Local computation latency of user k Based on task size b k (i) Number of cycles required per bit s k (i) The user's computing power f UE Locally calculated proportion The calculation yields the following result, which is expressed as:

[0133]

[0134] The computing resources allocated to user k by drone m in each time slot are represented as follows:

[0135]

[0136] Where κ1 is the energy consumption coefficient of UAV m, p k(i) represents the computing resources allocated to user k by drone m, expressed as:

[0137] p k (i)=μ k P

[0138] Where, μ k The proportion of resources allocated must satisfy ∑μ k <1.

[0139] Computational tasks from unloading to the drone Number of cycles required per bit (s) k (i) The computing power f of the drone UAV,k The computational delay D for unloading the m-part of the task onto the drone was calculated. UAV,k (i) is represented as:

[0140]

[0141] The task transmission delay from user k to drone m is represented as follows: The computational delay for user k to offload the satellite portion of the task is expressed as:

[0142] Due to the long distance between the ground / airborne satellite and the satellite, the data propagation delay d must also be considered during transmission. s Considering the task's tolerable latency, task generation time, and priority factor.

[0143] S1.5: Energy consumption transmitted to drone m and calculate energy consumption They are represented as follows:

[0144]

[0145]

[0146] User k's energy consumption E k (i), by and The sum is obtained by adding them together.

[0147] Energy consumption for transmission to satellite Transmission power P from ground to satellite l The user's energy consumption per CPU cycle e loc Calculated.

[0148] The flight energy consumption of drone m is expressed as:

[0149] E fly (i)=φ||ν(i)|| 2

[0150] Here, ν(i) includes the UAV's flight speed and flight angle. A modulo operation ||ν(i)|| is performed to obtain the flight speed, which is in the range [0, ν]. max ] takes a value from, ν max It represents the maximum flight speed; β(i) represents the flight angle, β(i)∈[0,2π]; φ represents the relevant parameters of the UAV, which are determined by the UAV's payload M and the fixed flight time t. fly The calculation yields the following result, which is expressed as:

[0151] φ=0.5M UAV t fly

[0152] S2: Under the energy constraints of UAVs, an integer nonlinear optimization problem with uncertainty is proposed, which uses Conditional Value at Risk (CVaR) to measure the inspection risk;

[0153] S2.1: The drone's m-trajectory is The uninstallation strategy is Resource allocation strategy Latency D(i) is defined as the maximum value of the latency of each part of the processing task, expressed as:

[0154]

[0155] The problem can be modeled as follows:

[0156]

[0157] Where C1 represents the range of values ​​for the unloading ratio; C2 indicates that the resource allocation to all users cannot exceed the total resource amount; C3 is used to determine whether the task latency tolerance of each user k is met; C4 indicates that the UAV's movement range cannot exceed the predetermined range; C5 indicates that the UAV's flight energy consumption, computing energy consumption, and transmission energy consumption within all time slots cannot exceed the UAV's energy capacity constraint, E b This represents the energy capacity of the drone.

[0158] S2.2: Define a Markov Decision Plugin (MDP) as a tuple. At time slot i, the description of the system state s i , represented as:

[0159] s i ={E[i],Q[i],P[i],D remain [i],f K [i]}

[0160] In time slot i, an action a is performed based on the current state.i , represented as:

[0161] a i =(β(i),ν(i),λ(i))

[0162] Where β(i), ν(i), and λ(i) represent the UAV's flight direction, UAV's moving speed, and unloading ratio, respectively.

[0163] Considering the penalty for exceeding the task latency requirement, the cost c is expressed as:

[0164]

[0165] Where ζ is the penalty for exceeding the time delay constraint, o k (i) is used to determine whether the task latency tolerance for each user k is met.

[0166] The task scheduling problem based on MDP, which minimizes latency, is represented as:

[0167]

[0168] stC1,C2,C3,C4

[0169]

[0170] Here, problem P2 represents the expected average cost. Energy consumption is not a component of the cost function of problem P2, requiring the design of a method suitable for handling Constrained Markov Decision Process (CMDP) problems. The Conditional Value at Risk (CVaR) is introduced as a measure of risk, used to measure the average loss under the worst-case scenario at a certain confidence level.

[0171] First, let X be a random variable representing the drone's energy consumption exceeding its energy capacity limit:

[0172]

[0173] in, This refers to the actual energy consumed by the drone, E. b This represents the drone's energy capacity. When X > 0, it indicates that the drone's energy consumption has exceeded the limit.

[0174] Based on the acceptable level of risk for the drone mission, an appropriate confidence level α is selected. We aim for the drone to operate normally in 95% of cases without energy overruns, so α is set to 0.05.

[0175] To calculate the Value at Risk (VaR) for a given confidence level α, the value of P(X ≤ -VaR) must satisfy α. Assume there are n historical or simulated energy consumption data points X1, X2, ..., Xn from various drone mission scenarios. n The energy consumption data is sorted, and VaR is approximately expressed as follows, based on the confidence level α: This represents the maximum possible loss boundary at confidence level α when the drone's energy consumption exceeds the capacity limit.

[0176] Once VaR is determined, CVaR is calculated to represent the average of these losses when the loss exceeds VaR. All data points X satisfying X > -VaR are then found. j (j=1,2,...,m), calculate CVaR, and obtain:

[0177]

[0178] We take CVaR as a major component of the risk function. Let the risk function be R(X) = CVaR. α (X).

[0179] S3: A dual-critic network structure was designed, and a risk-aware resource scheduling method was proposed.

[0180] Discount cost models can be used to consider the consequences of long-term decisions and are used to calculate state-value functions or action-value functions. Let V(s) represent the value function of expected long-term discount costs, expressed as:

[0181]

[0182] in, This represents the discount factor in the current time slot, used to measure the importance of future rewards relative to current rewards.

[0183] Based on the above equation, we define a Q-value function to further evaluate the expected long-term discount cost obtained under the current state-action pair:

[0184]

[0185] With the goal of minimizing cost, select the function with the minimum Q value Q. * (s,a) = minQ(s,a) is taken as the optimal Q value. Simultaneously, action a is taken. * Obtain the optimal Q value. The Q-value function is updated based on the Temporal Difference (TD) equation:

[0186]

[0187] Here, η represents the learning rate, which determines the magnitude of each update.

[0188] Parameterization is performed using Deep Neural Networks (DNNs), and the loss function is expressed as:

[0189]

[0190] in, This represents the Q-value function based on DNN. This is the TD error.

[0191] The risk function for the portion exceeding the drone's energy capacity is expressed as:

[0192]

[0193] Here, Ω is the defined set of error states, representing the drone's energy consumption exceeding its energy capacity.

[0194] Similarly, a value function that represents the expected long-term discount risk can be obtained. Represented as:

[0195]

[0196] The expected long-term discount risk estimated using the Q-value function is expressed as:

[0197]

[0198] Perform a greedy action To obtain the optimal Q value

[0199] Suppose the optimal Q values ​​obtained under two different objectives are Q(s) i ,a i )and By adding a dynamically updated weighting coefficient δ between the two to achieve balance, a new Q-value function is obtained, expressed as:

[0200]

[0201] Initialize parameterized cost network Risk Network and the action network μ(s|θ) μ ), initialization parameters Experience replay Weights δ and step size Δ. Based on μ(s|θ) μ Choose a i Storage experience (s) i ,a i ,c i,r i ,s i+1 )arrive middle.

[0202] Based on the target value y j To update, it is indicated as:

[0203]

[0204] Calculate the network loss L c , represented as:

[0205]

[0206] Obtain the optimal strategy action Represented as:

[0207]

[0208] The weights are updated simultaneously with the parameters through a soft update, expressed as follows:

[0209]

[0210] if Then δ←δ+Δ, otherwise δ←δ-Δ. This yields the optimal task unloading and resource scheduling method. This invention introduces a dual-critic network structure for learning, where each network independently evaluates the action-value function, and then a comprehensive evaluation is performed. This reduces the impact of inaccuracies in individual networks on policy updates and minimizes overestimation. It also improves the stability and convergence of the learning process.

[0211] This invention designs a risk-aware resource scheduling method. First, in traditional integrated air-space-ground networks, communication, latency, and energy consumption models are constructed for each of the three network layers (space-based, air-based, and ground-based networks) based on their respective characteristics. To address the instability caused by dynamic mission arrival and to optimize the bias in estimation results caused by a single target, an uncertain integer nonlinear optimization problem is proposed under the energy constraint of the UAV, using Conditional Value at Risk (CVaR) to measure and assess risk. A dual-critic network structure is designed to balance cost and risk, and a risk-aware resource scheduling method is proposed. The method provided by this invention can achieve a significant improvement in service quality and achieves a good trade-off between cost and risk.

[0212] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A resource scheduling method based on risk perception, characterized in that, Includes the following steps: S1: In the integrated air-space-ground network, construct communication models, latency models, and energy consumption models for the space-based network, air-based network, and ground-based network respectively; S2: Under the energy constraints of the UAV, the problem is formulated as an integer nonlinear optimization problem with uncertainty, and the conditional risk value CVaR is used to measure the inspection risk; S3: Utilize a dual-critic network to evaluate costs and risks, and achieve optimal task offloading and resource scheduling strategies through parameterization and time-series differential TD updates via deep neural network (DNN). S2 includes the following steps: The trajectory Q of the drone m is represented as: Uninstallation strategy Λ, represented as: Resource allocation strategy Φ is represented as: Latency D(i) is defined as the maximum value of the latency of each part of the processing task, expressed as: The problem can be modeled as follows: Where C1 represents the range of values ​​for the unloading ratio; C2 represents the resource allocation to all users not exceeding the total resources; C3 determines whether the latency tolerance of each user k is met; C4 represents the drone's movement range not exceeding the predetermined range; C5 represents the drone's flight energy consumption, computing energy consumption, and transmission energy consumption within the time slot not exceeding the drone's energy capacity constraint, E b This represents the energy capacity of the drone; Using Markov decision-making, the cost c is obtained, expressed as: Where ζ is the penalty for exceeding the time delay constraint, o k (i) is used to determine whether the task latency tolerance of each user k is met; The current scheduling problem can be described as a constrained Markov decision process, expressed as: stC1,C2,C3,C4 The risk condition value (CVaR) is used as a measure of risk to measure the average loss under the worst-case scenario at a certain confidence level. Let X be a random variable representing the amount of energy consumed by the drone exceeding its energy capacity limit: in, This refers to the actual energy consumed by the drone, E. b X represents the drone's energy capacity; when X > 0, it indicates that the drone's energy consumption has exceeded the limit. Set the confidence level α = 0.05; for a given α, calculate the risk value VaR, which must satisfy P(X ≤ -VaR) = α; find all data points X that satisfy X > -VaR. j (j=1,2,...,m), calculate CVaR, expressed as: We take CVaR as the main component of the risk function; let the risk function R(X) = CVaR α (X).

2. The resource scheduling method based on risk perception according to claim 1, characterized in that, The S1 Includes the following steps: In ground-to-air communication, the channel gain for user k and UAV m is defined as: Where α0 is the channel gain at a reference distance d = 1m, q(i+1) is the position of the UAV after moving m in the next time slot i, and p k (i) represents the position of user k in time slot i, H is the fixed flight altitude of UAV m, and ||q(i+1)-p k (i)|| 2 +H 2 This represents the Euclidean distance between user k and drone m; The data rate unloaded onto drone m is expressed as: Among them, B u For the available spectrum bandwidth from user k to drone m, η k To determine the percentage of bandwidth allocated to user k, we need to calculate ∑η k <1, P u For the user's transmit power on the uplink, g k (i) represents the channel gain. For noise power, f k (i) Check if there is a blockage between user k and drone m. If there is a blockage, add an extra P. NLOS noise power, P NLOS Indicates additional noise power; In terrestrial / airborne-space communication, the channel gain from user k to the satellite is determined by the fixed distance D between the satellite and the ground and the satellite's computing power f. LEO The calculated data rate from the ground or drone (m) to the satellite is expressed as: in, For the data rate from ground or drone m to satellite, B l P represents the communication bandwidth to the satellite. l This refers to the transmit power on this link. Noise power; set up This represents the partial uninstallation decision made by user k. The percentage of data that user k unloads onto drone m. The percentage of user k that is unloaded to low-Earth orbit satellites; The partial uninstallation decision for user k is set as follows: in, The percentage of data that user k unloads onto drone m. The percentage of user k that is unloaded to low-Earth orbit satellites.

3. The resource scheduling method based on risk perception according to claim 2, characterized in that, S1 further includes the following steps: Based on task size b k (i) Number of cycles required per bit s k (i) The user's computing power f UE Locally calculated proportion Get the local computation latency of user k The computing resources allocated to user k by drone m in each time slot are represented as follows: Where κ1 is the energy consumption coefficient of UAV m, p k (i) represents the computing resources allocated to user k by drone m, expressed as: p k (i)=μ k P Where, μ k The proportion of resource allocation, ∑μk<1; Computational tasks from unloading to the drone Number of cycles required per bit (s) k (i) The computing power f of the drone UAV,k The computational delay D for unloading the m-part of the task onto the drone was calculated. UAV,k (i); Priority factor, expressed as: Among them, t k The priority factor is the time when the task was generated, and t is the current time. The smaller the priority factor, the higher the priority of the task, so that it can be unloaded and processed earlier.

4. The resource scheduling method based on risk perception according to claim 3, characterized in that, S1 further includes the following steps: Energy consumption transmitted to drone m and calculate energy consumption They are represented as follows: Depend on and The sum gives the local computation latency of user k. The flight energy consumption E of drone m fly (i) is represented as: E fly (i)=φ||ν(i)|| 2 Here, ν(i) includes the UAV's flight speed and flight angle. A modulo operation ||ν(i)|| is performed to obtain the flight speed, which is in the range [0, ν]. max ] takes a value from, ν max φ represents the maximum flight speed; β(i) represents the flight angle, β(i)∈[0,2π]; φ represents the relevant parameters of the UAV.

5. A resource scheduling method based on risk perception according to claim 1, characterized in that, S3 includes the following steps: The expected long-term discount cost value function V(s) is expressed as: in, This represents the discount factor in the current time slot, used to measure the importance of future rewards relative to current rewards; Define the Q-value function to further evaluate the expected long-run discount cost obtained under the current state-action pair, expressed as: With the goal of minimizing cost, select the function with the minimum Q value Q. * (s,a) = minQ(s,a) as the optimal Q value; simultaneously, take action a. * Obtain the optimal Q value; the Q-value function is updated according to the time-difference TD equation, expressed as: Where η represents the learning rate, which is used to determine the magnitude of each update; Parameterization is performed using a deep neural network (DNN), and the loss function is expressed as: in, This represents the Q-value function based on DNN. For TD error; The risk function for the portion exceeding the drone's energy capacity is expressed as: Wherein, Ω is the defined set of error states, indicating that the drone's energy consumption exceeds its energy capacity.

6. A resource scheduling method based on risk perception according to claim 5, characterized in that, S3 further includes the following steps: Value function of expected long-term discount risk Represented as: The expected long-term discount risk estimated using the Q-value function is expressed as: Perform a greedy action To obtain the optimal Q value The optimal Q values ​​obtained under the two different objectives are Q(s) i ,a i )and By adding a dynamically updated weighting coefficient δ between the two to achieve balance, a new Q-value function is obtained, expressed as:

7. A resource scheduling method based on risk perception according to claim 6, characterized in that, S3 further includes the following steps: Initialize parameterized cost network Risk Network and the action network μ(s|θ) μ Initialization parameters Experience replay D, weight δ, and step size Δ; based on μ(s|θ) μ Choose a i Storage experience (s) i ,a i ,c i ,r i ,s i+1 ) to D; Based on the target value y j To update, it is indicated as: Calculate the network loss L c , represented as: Obtain the optimal strategy action Represented as: The weights are updated simultaneously with the parameters through a soft update, expressed as follows: if Then δ←δ+Δ, otherwise δ←δ-Δ.

8. A resource scheduling method based on risk perception according to claim 2, characterized in that, The maximum value λ of the unloading ratio max The unloading ratio is constrained by a value of 0.5, expressed as follows:

Citation Information

Patent Citations

  • Online resource scheduling method for guaranteeing service quality

    CN116489226A

  • Satellite edge computing task unloading and resource allocation method based on deep reinforcement learning

    CN118250750A