UAV edge network control methods, UAVs and computer program products

By constructing the POMDP model and using the RSAC algorithm, the problems of flight trajectory planning and resource allocation of UAV MEC systems in dynamic environments were solved, enabling efficient communication of UAVs in complex terrain and post-disaster areas.

CN120711015BActive Publication Date: 2025-10-28XIAN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511202825.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-27
Publication Date
2025-10-28
Estimated Expiration
2045-08-27

AI Technical Summary

Technical Problem

How to accurately plan the flight trajectory of UAVs and rationally allocate computing resources in a dynamically changing environment, especially the communication needs in complex terrain, remote areas and disaster-damaged areas, has not been effectively addressed.

Method used

A POMDP model is constructed, which combines user mobility model, task generation model and environmental observability constraints. The RSAC algorithm is used for decision-making, and the environment transition probability is modeled by LSTM to generate simulation experience in order to minimize the computational energy consumption of user equipment and maximize the uplink rate, thereby determining the flight trajectory and computational resource allocation of the UAV.

Benefits of technology

It enables precise planning of UAV flight trajectories and rational allocation of computing resources in dynamic environments, improving the communication efficiency and coverage of user equipment and reducing network latency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120711015B_ABST
    Figure CN120711015B_ABST
Patent Text Reader

Abstract

This application provides a UAV edge network control method, a UAV, and a computer program product, relating to the field of wireless communication technology. The method includes: loading a UAV edge network scenario model; loading a user terminal mobility model and a task generation model; loading a UAV flight model and a coverage model; loading a communication model between the UAV and the user equipment; loading a predefined energy consumption model and objective function; loading the seven-tuple of the POMDP model; acquiring observation information of the observation space in the current time slot; determining target decision action information based on the observation information of the observation space in the current time slot, the seven-tuple, the Dyna environment model, and the RSAC algorithm, with the goal of minimizing the computational energy consumption of the user equipment and maximizing the uplink rate of the user equipment; and controlling the UAV edge network in the current time slot based on the target decision action information. This method is used to provide a UAV edge network control scheme.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of wireless communication technology, and in particular to a method for controlling an edge network of a drone, a drone, and a computer program product. Background Technology

[0002] Unmanned aerial vehicle (UAV) communication networks have gained widespread attention and application in the communications field in recent years due to their significant advantages such as high flexibility, rapid deployment, and wide coverage. This technology overcomes the communication limitations of traditional terrestrial networks in complex terrain, remote areas, and after disasters, further improving the user's communication experience, reducing network latency, and effectively compensating for the deficiencies of ground infrastructure, providing an important supplement to modern communication networks.

[0003] With the rapid development of Mobile Edge Computing (MEC) technology, computing resources can be deployed closer to user devices at the edge. Drone-based MEC systems, by carrying lightweight MEC servers on drones, provide efficient task offloading and computing services to ground-based user devices. In drone-based MEC systems, the maneuverability of drones allows them to dynamically adjust their positions, covering areas that are difficult for traditional networks to reach, such as mountainous areas, oceans, or disaster-stricken regions, thus demonstrating great potential in emergency communications, rural network coverage, and smart city construction.

[0004] Therefore, the need for a drone MEC system is to accurately plan the drone's flight trajectory and rationally allocate computing resources in a dynamically changing environment. Summary of the Invention

[0005] Based on this, this application provides an edge network control method for unmanned aerial vehicles (UAVs), an UAV, and a computer program product, which can accurately plan the flight trajectory of the UAV and rationally allocate computing resources.

[0006] Firstly, this application provides a method for controlling an unmanned aerial vehicle (UAV) edge network. The method includes: loading a predefined UAV edge network scenario model; the UAV edge network scenario model includes a UAV, multiple user devices, a discrete-time system, movement boundary coordinates, and flight hovering protocol information; the UAV carries an MEC server to provide computing services to multiple user devices; loading a predefined user terminal mobility model and task generation model; the user terminal mobility model represents the speed and position of the user terminal in each time slot; the task generation model represents the task generation rules of the user terminal; loading a predefined UAV flight model and coverage model; the flight model represents the flight parameters of the UAV; the coverage model represents the coverage area of ​​the UAV; loading a predefined communication model between the UAV and user devices; the communication model determines the uplink rate of the user devices; and loading a predefined energy consumption model and objective function; the energy consumption model represents the user's... The computational energy consumption of the device; the objective function is a function that aims to minimize the computational energy consumption of the user device and maximize the uplink rate of the user device; loading the predefined POMDP model's seven-tuple; the seven-tuple includes the state space, observation space, action space, state transition probability, observation transition probability, reward function, and discount factor of the UAV edge network system; the state space includes complete information of the UAV edge network system in each time slot; the observation space includes the observation information that the UAV can observe in each time slot; the action space includes the decision action information of the UAV in each time slot; obtaining the observation information of the observation space in the current time slot; based on the observation information of the observation space in the current time slot, the seven-tuple, the Dyna environment model, and the RSAC algorithm, determining the target decision action information with the goal of minimizing the computational energy consumption of the user device and maximizing the uplink rate of the user device; controlling the UAV edge network in the current time slot based on the target decision action information.

[0007] The UAV edge network control method provided in this application jointly considers user mobility models, task arrival models, and partially observable environmental constraints to construct a POMDP. Subsequently, the RSAC algorithm is used to make decisions based on historical information; LSTM is used to model environmental transition probabilities, generating simulation experience and accelerating RSAC algorithm convergence. Finally, target decision action information is determined with the objectives of minimizing user computational energy consumption and maximizing user uplink speed in the network scenario. Based on this target decision action information, the UAV edge network in the current time slot is controlled, thus providing a UAV edge network control scheme that accurately plans the UAV's flight trajectory and rationally allocates computational resources in a dynamically changing environment.

[0008] Optionally, the discrete-time system includes a set of time slots; the set of time slots includes multiple time slots of equal length; each time slot includes a flight duration and a hovering duration; the UAV flies at a constant speed during the flight duration and does not connect to user equipment; the hovering duration includes multiple sub-time slots, each sub-time slot is used to be allocated to the corresponding user equipment for task offloading; the flight duration is less than or equal to the time slot duration; the movement boundary coordinates include the boundary coordinates of the rectangular region.

[0009] Optionally, in the user terminal mobile model, the user equipment exist The speed of the time slot satisfies the following formula:

[0010] ;

[0011] in, Indicates user equipment exist The speed of the time slot; Indicates user equipment exist The speed of the time slot; Indicates Gaussian white noise; Indicates memory level; Indicates asymptotic average velocity; Indicates asymptotic standard deviation;

[0012] User equipment exist The location of the time slot satisfies the following formula:

[0013]

[0014] in, Indicates user equipment exist The location of the time slot; Indicates user equipment exist The location of the time slot; Indicates the duration of the time slot;

[0015] In the task generation model, each user device at the beginning of each time slot is assigned a preset probability. Received fixed data size The task; each user device maintains an infinitely large task queue. If a received task is not processed or unloaded to the drone immediately, the task will accumulate in the task queue. The total number of tasks in the drone edge network system is fixed at M.

[0016] Optionally, in the flight model, the drone operates at a fixed altitude. At a speed of meters Uniform flight It is a positive number; The horizontal coordinates of the UAV at the beginning of the time slot are ,and , ; For drones The flight angle of the time slot, and drones exist The position update at the beginning of the time slot satisfies the following formula:

[0017]

[0018] in, and Indicates drone exist The horizontal coordinate at the beginning of the time slot; Indicates that drones are in Flight duration of time slots; Indicates drone Flight speed;

[0019] In the coverage model, Indicates drone exist The set of user equipment within the coverage area of ​​the time slot, and , , It represents a collection of multiple user devices. Indicates user terminal With drones The two-dimensional distance between them Indicates the maximum coverage radius of the drone. , This indicates the maximum elevation angle between the drone and the user's destination.

[0020] Optionally, in the communication model, Time-slot drones With ground user terminal The uplink transmission rate between nodes satisfies the following formula:

[0021]

[0022] in, express Time-slot drones With ground user terminal Inter-link uplink transmission rate; Indicates bandwidth; Indicates the transmit power of the user equipment; Indicates Gaussian white noise; express Time-slot drones With ground user terminal Overall channel gain between;

[0023]

[0024] in, express Time-slot drones With ground user terminal Line-of-sight (LOS) channel probability; express Time-slot drones With ground user terminal LOS channel gain between; express Time-slot drones With ground user terminal Non-line-of-sight NLOS channel probability; express Time-slot drones With ground user terminal Inter-NLOS channel gain;

[0025]

[0026]

[0027]

[0028]

[0029] in, and Represents environmentally relevant constants; Represents the speed of light; Indicates the carrier frequency of the wireless channel; express Time-slot drones With ground user terminal The three-dimensional Euclidean distance between them ; Add a loss factor to NLOS; This represents the path loss exponent of the LOS channel.

[0030] Optionally, in the energy consumption model, express Time slots allocated to observable user equipment The sub-time slot, and ; Determine which user to assign sub-slot Afterwards, the amount of data unloaded Time slot User equipment Offloading energy consumption during transmission tasks , Indicates the transmit power of user equipment; time slot User equipment Locally computed workload , This indicates the fixed frequency of the processor in the user terminal; A time slot represents the number of processor cycles required to process a unit of bit. User equipment Energy consumption for processing local computing tasks , Indicates the energy consumption coefficient. Indicates time slot duration; User equipment exist The workload of a time slot is expressed by the following formula:

[0031]

[0032] in, Indicates user equipment exist The workload in a time slot; Indicates user equipment exist The workload in a time slot; Indicates the size of the task's data; Indicates time slot The workload of drone calculations; This represents the preset probability that each user device will receive a task at the beginning of each time slot;

[0033] The objective function is expressed by the following formula:

[0034]

[0035] in, express Time-slot drones With ground user terminal Inter-link uplink transmission rate; Indicates time slot User equipment Energy consumption for processing local computing tasks; This indicates the flight angle constraints of the drone. Indicates that drones are in The flight angle of the time slot; Indicates a single sub-slot constraint. express Time slots allocated to observable user equipment The sub-time slot, express Drone flight duration within the time slot; This means that the sum of the flight duration and all sub-slots equals the slot duration; This indicates a constraint on the total amount of tasks. This represents the total number of tasks generated by all user devices throughout the entire task cycle. Indicates an indicator variable. ; This indicates the amount of work the drone calculates throughout the entire mission cycle. The workload of local computation on user devices The sum is greater than or equal to M.

[0036] Optionally, the complete information includes: the current two-dimensional position of the UAV, the two-dimensional positions of all user devices, the task queues of all user devices, and the channel gain between all user devices and the UAV;

[0037] The observation information includes: the current two-dimensional position of the UAV, the two-dimensional position of the observable user equipment, the task queue of the observable user equipment, and the channel gain between the observable user equipment and the UAV;

[0038] Decision-making information includes UAV flight time, UAV flight angle, and sub-time slot division information of all observable user equipment;

[0039] The reward function satisfies the following relationship:

[0040]

[0041] in, This represents the reward function value; express Time-slot drones With ground user terminal Inter-link uplink transmission rate; Indicates the weighting coefficient; Indicates time slot User equipment Energy consumption during offloading of transmission tasks; Indicates time slot User equipment Energy consumption for processing local computing tasks.

[0042] Optionally, based on the observation information of the observation space in the current time slot, the seven-tuple, the Dyna environment model, and the RSAC algorithm, with the goal of minimizing the computational energy consumption of the user equipment and maximizing the uplink rate of the user equipment, target decision-making action information is determined, including:

[0043] Initialize the network parameters of the Actor network Critic network parameters and Dyna environment model parameters And set the target Critic network parameters. and Initialize the real experience replay pool and virtual experience pool All are empty sets;

[0044] Iterate through the set maximum number of rounds, and initialize the hidden state of the LSTM network at the beginning of each round;

[0045] Within each round, the maximum number of steps is further iterated, and each step is based on the current observation information. and historical observation information in hidden state Use strategy Select decision action information and in the information of decision-making actions Then, the immediate reward is obtained based on the reward function. and observation information for the next time slot ;

[0046] After each round, the current trajectory data will be... Store in the real experience pool ;

[0047] from Medium sampling Groups, each group Each strip is [length missing] Sub-trajectories, training the Dyna environment model Then, update using the following loss function. :

[0048]

[0049] in, Represents the loss function of the Dyna environment model; The first prediction made by the Dyna environment model represents the Sub-trajectories in time slots Observed values; Indicates the Sub-trajectories in time slots The actual observed value; Represents the square of the L2 norm; The first prediction made by the Dyna environment model represents the Sub-trajectories in time slots The reward value; Indicates the Sub-trajectories in time slots The actual reward value;

[0050] Initialize the hidden state of the Dyna environment model Iterate through each round Each step is based on the current observation information. Use strategy Select decision action information ,according to , and hidden state Historical observation information and historical decision-making information Predicting instant reward value and observation information for the next time slot ;

[0051] Generate simulated trajectory Store in the model experience pool ;

[0052] from and Medium sampling Groups, each group Each strip is [length missing] Training the Actor network and Critic network using sub-trajectories Each time, the loss function is updated using the following methods respectively. , , and The converged Actor and Critic networks are obtained:

[0053]

[0054] in, express The loss value; This represents the reward value output by the first Critic network. Indicates the discount factor; Indicates time slot Observational information; Indicates time slot Information on the selection and decision-making actions; Represents the time slots of the outputs of the two target Critic networks. The reward value of the information on the choice and decision-making action; Represents the entropy-temperature factor;

[0055]

[0056] in, express The loss value; This represents the reward value output by the second Critic network;

[0057]

[0058] in, express The loss value;

[0059]

[0060] in, express The loss value; Represents the target entropy. , Indicates the dimension of the action space; network parameters of the soft-update target Critic network: ; , This represents the updated network parameters of the first target Critic network; Indicates the soft update coefficient; This represents the network parameters before the first target Critic network update; The second objective is to update the network parameters of the Critic network. This represents the network parameters before the Critic network update for the second objective;

[0061] Based on the observation information of the observation space in the current time slot and the converged Actor network, the target decision action information is determined.

[0062] Secondly, this application provides an edge network control device for unmanned aerial vehicles (UAVs), which includes various functional modules for the method described in the first aspect above.

[0063] Thirdly, this application provides a drone including a processor and a memory; the memory stores processor-executable instructions; when the processor is configured to execute the instructions, the drone performs the method described in the first aspect above.

[0064] Fourthly, this application provides a computer program product comprising: computer instructions that, when executed in a drone, cause the drone to perform the methods described in the first aspect, thereby implementing the methods described in the first aspect.

[0065] Fifthly, this application provides a readable storage medium comprising: software instructions; when the software instructions are executed in a drone, causing the drone to perform the method described in the first aspect above.

[0066] The beneficial effects of the second to fifth aspects mentioned above can be referred to the first aspect, and will not be repeated here. Attached Figure Description

[0067] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0068] Figure 1 This is a schematic diagram of the application architecture of the UAV edge network control method provided in the embodiments of this application;

[0069] Figure 2 A flowchart illustrating the UAV edge network control method provided in this application embodiment;

[0070] Figure 3 This is a schematic diagram of the composition of the UAV edge network control device provided in the embodiments of this application;

[0071] Figure 4 This is a schematic diagram of the composition of the drone provided in an embodiment of this application. Detailed Implementation

[0072] Hereinafter, the terms "first," "second," and "third," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined with "first," "second," or "third," etc., may explicitly or implicitly include one or more of that feature.

[0073] Unmanned aerial vehicle (UAV) communication networks have gained widespread attention and application in the communications field in recent years due to their significant advantages such as high flexibility, rapid deployment, and wide coverage. This technology overcomes the communication limitations of traditional terrestrial networks in complex terrain, remote areas, and after disasters, further improving the user's communication experience, reducing network latency, and effectively compensating for the deficiencies of ground infrastructure, providing an important supplement to modern communication networks.

[0074] With the rapid development of MEC technology, computing resources can be moved closer to the edge, closer to user equipment. Drone MEC systems, by mounting lightweight MEC servers on drones, provide efficient task offloading and computing services to ground user equipment. In drone MEC systems, the maneuverability of drones allows them to dynamically adjust their positions, covering areas that are difficult for traditional networks to reach, such as mountainous areas, oceans, or disaster-stricken regions, thus demonstrating great potential in emergency communications, rural network coverage, and smart city construction.

[0075] Therefore, the need for a drone MEC system is to accurately plan the drone's flight trajectory and rationally allocate computing resources in a dynamically changing environment.

[0076] Based on this, embodiments of this application provide an edge network control method for unmanned aerial vehicles (UAVs), an UAV, and a computer program product, which can accurately plan the flight trajectory of the UAV and rationally allocate computing resources.

[0077] The following description is provided in conjunction with the accompanying drawings.

[0078] Figure 1 This is a schematic diagram of the application architecture of the UAV edge network control method provided in the embodiments of this application. Figure 1 As shown, the application architecture includes: drone 100 and multiple user devices 200.

[0079] The drone 100 can be a fixed-wing drone, a swooping drone, an unmanned airship, a paragliding drone, a flapping-wing drone, a satellite drone, a light drone, a small drone, a medium-sized drone, a large drone, an ultra-low-altitude drone, a low-altitude drone, a centrally controlled drone, a high-altitude drone, an ultra-high-altitude drone, etc. This application does not limit the specific type of drone 100.

[0080] The drone 100 can be equipped with an MEC server, which can provide edge computing services to the user equipment 200.

[0081] In some embodiments, the drone 100 can also control the drone edge network. The specific process can be referred to the drone edge network control method provided in the following method embodiments, and will not be repeated here.

[0082] User equipment 200 can be a smartphone, tablet computer, laptop computer, industrial Internet of Things (IoT) terminal, emergency rescue terminal, power / traffic inspection terminal, smartwatch, smart bracelet, or mobile medical terminal, etc. This application embodiment does not limit the specific form of user equipment 200.

[0083] User equipment 200 can perform edge computing through the MEC server in drone 100.

[0084] As an example, user equipment 200 can acquire computing tasks and offload them to drone 100 for computation. The specific process can be found in related technical documents and will not be elaborated upon here.

[0085] The execution subject of the UAV edge network control method provided in this application embodiment is a UAV edge network control device. This UAV edge network control device can be a UAV (e.g., the UAV 100 described above). Optionally, the UAV edge network control device can also be a processor (e.g., a central processing unit (CPU)) in the aforementioned UAV; or, the UAV edge network control device can also be a software system or platform in the aforementioned UAV; or, the UAV edge network control device can also be a functional module in the aforementioned UAV used to perform UAV edge network control, etc. This application embodiment does not impose any limitations on these aspects.

[0086] The following describes the UAV edge network control method provided in the embodiments of this application.

[0087] Figure 2 This is a flowchart illustrating the UAV edge network control method provided in an embodiment of this application. Figure 2 As shown, the method includes the following steps:

[0088] S101. Load the predefined drone edge network scenario model.

[0089] The UAV edge network scenario model includes a UAV, multiple user devices, a discrete-time system, mobile boundary coordinates, and flight hovering protocol information. The UAV carries an MEC server to provide computing services to the multiple user devices.

[0090] As an example, it can be defined Indicates drone, express A collection of user devices.

[0091] As an example, a discrete-time system may include a set of time slots. The time slot set includes multiple time slots of equal length, with a time slot duration of [missing information]. Each time slot includes flight duration. The hovering duration; the drone flies at a constant speed during the flight duration and does not connect to user equipment. During the hovering duration, the drone hovers, and user equipment can connect to the drone and offload tasks using Time Division Multiplexing (TDM) mode. The hovering duration includes multiple sub-time slots, each allocated to a corresponding user equipment for task offloading; the flight duration is less than or equal to the time slot duration; the movement boundary coordinates include the boundary coordinates of the rectangular area.

[0092] S102, Load the predefined user terminal mobile model and task generation model.

[0093] The user terminal mobility model represents the speed and position of the user terminal in each time slot. The task generation model represents the task generation rules of the user terminal.

[0094] As an example, in the user terminal mobile model, the user equipment exist The speed of the time slot satisfies the following formula:

[0095] Formula (1)

[0096] in, Indicates user equipment exist The speed of the time slot; Indicates user equipment exist The speed of the time slot; Indicates Gaussian white noise; Indicates memory level; Indicates asymptotic average velocity; Indicates asymptotic standard deviation;

[0097] User equipment exist The location of the time slot satisfies the following formula:

[0098] Formula (2)

[0099] in, Indicates user equipment exist The location of the time slot; Indicates user equipment exist The location of the time slot; Indicates the duration of the time slot. Must meet the following conditions: , Ensure that the user equipment location remains within the rectangular area.

[0100] As an example, in the task generation model, each user device at the beginning of each time slot is assigned a preset probability. Received fixed data size The task; each user device maintains an infinitely large task queue. If a received task is not processed or unloaded to the drone immediately, the task will accumulate in the task queue. The total number of tasks in the drone edge network system is fixed at M.

[0101] S103, Load the predefined flight model and coverage model of the UAV.

[0102] The flight model represents the drone's flight parameters. The coverage model represents the drone's coverage area.

[0103] As an example, in the flight model, the drone is at a fixed altitude. At a speed of meters Uniform flight It is a positive number; The horizontal coordinates of the UAV at the beginning of the time slot are ,and , ; For drones The flight angle of the time slot, and drones exist The position update at the beginning of the time slot satisfies the following formula:

[0104] Formula (3)

[0105] in, and Indicates drone exist The horizontal coordinate at the beginning of the time slot; Indicates that drones are in Flight duration of time slots; Indicates drone Flight speed.

[0106] As an example, in the coverage model, Indicates drone exist The set of user equipment within the coverage area of ​​the time slot, and , , It represents a collection of multiple user devices. Indicates user terminal With drones The two-dimensional distance between them Indicates the maximum coverage radius of the drone. , This represents the maximum elevation angle between the drone and the user's destination. The drone's flight needs to cover different areas to serve all user devices; considering user device mobility, trajectory planning needs to optimize flight time. and angle In addition, it optimizes sub-slot allocation, reduces computing power consumption, and increases uplink speed.

[0107] S104, Load the predefined communication model between the UAV and the user equipment.

[0108] The communication model is used to determine the uplink rate of the user equipment.

[0109] As an example, the communication model may include an elevation angle, a key parameter of the communication link between the UAV and the user equipment. The elevation angle determines whether the signal propagation path is easily affected by obstacles; a higher elevation angle generally means fewer obstacles, thus increasing the probability of a line-of-sight (LOS) channel. In this application, the UAV altitude is fixed at [value missing]. The horizontal coordinate is Ground user terminal Location is ,definition for Time-slot drones With ground user terminal The angle of elevation between them.

[0110] As an example, in the communication model, Time-slot drones With ground user terminal The uplink transmission rate between nodes satisfies the following formula:

[0111] Formula (4)

[0112] in, express Time-slot drones With ground user terminal Inter-link uplink transmission rate; Indicates bandwidth; Indicates the transmit power of the user equipment; Indicates Gaussian white noise; express Time-slot drones With ground user terminal Overall channel gain between;

[0113] Formula (5)

[0114] in, express Time-slot drones With ground user terminal Line-of-sight (LOS) channel probability; express Time-slot drones With ground user terminal LOS channel gain between; express Time-slot drones With ground user terminal Non-line-of-sight NLOS channel probability; express Time-slot drones With ground user terminal Inter-NLOS channel gain;

[0115] ;Formula (6)

[0116] ;Formula (7)

[0117] ;Formula (8)

[0118] ;Formula (9)

[0119] in, and Represents environmentally relevant constants; Represents the speed of light; Indicates the carrier frequency of the wireless channel; express Time-slot drones With ground user terminal The three-dimensional Euclidean distance between them ; Add a loss factor to NLOS; This represents the path loss exponent of the LOS channel, used to simulate signal attenuation caused by obstacles.

[0120] S105, Load the predefined energy consumption model and objective function.

[0121] The energy consumption model is used to represent the computing energy consumption of user equipment, and the objective function is a function that aims to minimize the computing energy consumption of user equipment and maximize the uplink speed of user equipment.

[0122] As an example, in energy consumption models, express Time slots allocated to observable user equipment The sub-time slot, and ; Determine which user to assign sub-slot Afterwards, the amount of data unloaded Time slot User equipment Offloading energy consumption during transmission tasks , Indicates the transmit power of user equipment; time slot User equipment Locally computed workload , This indicates the fixed frequency of the processor in the user terminal; A time slot represents the number of processor cycles required to process a unit of bit. User equipment Energy consumption for processing local computing tasks , Indicates the energy consumption coefficient. Indicates time slot duration; User equipment exist The workload of a time slot is expressed by the following formula:

[0123] Formula (10)

[0124] in, Indicates user equipment exist The workload in a time slot; Indicates user equipment exist The workload in a time slot; Indicates the size of the task's data; Indicates time slot The workload of drone calculations; This represents the preset probability that each user device will receive a task at the beginning of each time slot.

[0125] As an example, the optimization objective is to jointly optimize the drone trajectory: the drone's flight time in each time slot. Flight angle Selection and sub-slots for each user equipment The allocation of resources is designed to maximize user uplink speeds and minimize user computing power consumption. The objective function can be expressed as follows:

[0126]

[0127] in, express Time-slot drones With ground user terminal Inter-link uplink transmission rate; Indicates time slot User equipment Energy consumption for processing local computing tasks; This indicates the flight angle constraints of the drone. Indicates that drones are in The flight angle of the time slot; Indicates a single sub-slot constraint. express Time slots allocated to observable user equipment The sub-time slot, express Drone flight duration within the time slot; This means that the sum of the flight duration and all sub-slots equals the slot duration; This indicates a constraint on the total amount of tasks. This represents the total number of tasks generated by all user devices throughout the entire task cycle. Indicates an indicator variable. ; This indicates the amount of work the drone calculates throughout the entire mission cycle. The workload of local computation on user devices The sum is greater than or equal to M.

[0128] S106. Load the seven-tuple of the predefined Partially Observable Markov Decision Process (POMDP) ​​model.

[0129] The seven-tuple includes the state space, observation space, action space, state transition probability, observation transition probability, reward function, and discount factor of the UAV edge network system. The state space contains complete information about the UAV edge network system at each time slot. The observation space contains the observation information that the UAV can observe at each time slot. The action space contains the decision-making action information of the UAV at each time slot.

[0130] As an example, the seven-tuple of the POMDP model can be represented as: .in, For state space, For the action space, For observation space, For reward function (or instant reward). As a discount factor, For environmental transition probability, The POMDP framework effectively models the interaction between UAVs and user equipment, especially when the location of user equipment, task queue, and channel state are partially unknown. Its advantage lies in its ability to handle incomplete information and infer hidden states by combining historical observations and predictive models, thereby making adaptive decisions.

[0131] As an example, state space It can contain complete information about the drone's edge network system in each time slot. This complete information includes the drone's current two-dimensional position. Two-dimensional location of all user devices Task queues for all user devices and channel gain between all user equipment and drones time slot The internal state space consists of complete information about the drone and all ground users, as follows: .

[0132] As an example, the observation space is a subset of the state space, containing only the observational information that the UAV can observe in each time slot. The state space within the system consists of observation information from the UAV and observable user equipment, and is represented as: , This represents the current two-dimensional position of the drone. For the observable two-dimensional position of user equipment, For observable user equipment task queues, For observable channel gain between user equipment and drones.

[0133] As an example, action space The decision-making action information of the UAV in each time slot is defined, including the UAV flight time. Flight angle Time sub-slot partitioning information for all observable user equipment Specifically, it is expressed as: .

[0134] As an example, the reward function is a core component of POMDP, reflecting the system's optimization objective. The objective of this device is to maximize the user device's uplink speed while minimizing its computational energy consumption. Directly solving for this objective function during optimization presents certain limitations. First, the objective function tends towards infinity as the denominator approaches zero. Second, mainstream optimization algorithms (gradient descent, reinforcement learning, etc.) are better suited to the addition / subtraction form than the ratio form. Third, energy efficiency optimization is essentially a trade-off between performance metrics and energy consumption, and the subtraction form more intuitively reflects this characteristic. Therefore, the original ratio maximization problem can be transformed into a subtraction-based optimization problem.

[0135] First, for convenience, the rate term in the original objective function is removed. Recorded as Energy consumption items Recorded as ,in , This represents the set of feasible solutions. A penalty term is introduced based on the original objective function. If and only if Then, the original problem, namely maximizing the uplink speed of user devices and minimizing the computing power consumption of user devices, can be transformed into:

[0136] Formula (12)

[0137] in, It is a coefficient that adjusts the rate of change and the energy consumption weight. This is the optimal solution to the original problem. The detailed proof is as follows:

[0138] Necessity: Let This is the optimal solution to the original problem. For any ,have:

[0139] Formula (13)

[0140] Therefore, it can be proven Too The optimal solution.

[0141] Sufficiency: Let yes The optimal solution, and For any ,have:

[0142]

[0143] Formula (14)

[0144] Therefore, it can be proven It is also the optimal solution to the original problem.

[0145] The above reasoning can prove that when When the optimal solution in the subtraction form is also the optimal solution to the original problem, then, with the correct choice... In this case, the original optimization problem can be equivalently transformed into maximizing... Based on this, the reward of this application can be defined as:

[0146] Formula (15)

[0147] in, This represents the reward function value; express Time-slot drones With ground user terminal Inter-link uplink transmission rate; Indicates the weighting coefficient; Indicates time slot User equipment Energy consumption during offloading of transmission tasks; Indicates time slot User equipment Energy consumption for processing local computing tasks.

[0148] S107. Obtain observation information of the observation space in the current time slot.

[0149] S108. Based on the observation information of the observation space under the current time slot, the seven-tuple, the Dyna environment model, and the Recurrent Soft Actor-Critic (RSAC) algorithm, the target decision action information is determined with the goal of minimizing the computational energy consumption of the user equipment and maximizing the uplink rate of the user equipment.

[0150] As an example, the above S108 may specifically include the following steps:

[0151] Step 1: Initialize the network parameters of the Actor network Critic network parameters and Dyna environment model parameters And set the target Critic network parameters. and Initialize the real experience replay pool and virtual experience pool All are empty sets.

[0152] Step 2: Iterate through the set maximum number of rounds, and initialize the hidden state of the Long Short-Term Memory (LSTM) network at the beginning of each round.

[0153] Step 3: Within each round, further iterate through the maximum number of steps, and at each step, based on the current observation information. and historical observation information in hidden state Use strategy Select decision action information and in the information of decision-making actions Then, the immediate reward is obtained based on the reward function. and observation information for the next time slot .

[0154] Step 4: After each round, record the current trajectory data. Store in the real experience pool .

[0155] Step 5, from Medium sampling Groups, each group Each strip is [length missing] Sub-trajectories, training the Dyna environment model Then, update using the following loss function. :

[0156] Formula (16)

[0157] in, Represents the loss function of the Dyna environment model; The first prediction made by the Dyna environment model represents the Sub-trajectories in time slots Observed values; Indicates the first Sub-trajectories in time slots The actual observed value; Represents the square of the L2 norm; The first prediction made by the Dyna environment model represents the Sub-trajectories in time slots The reward value; Indicates the first Sub-trajectories in time slots The actual reward value;

[0158] Step 6: Initialize the hidden state of the Dyna environment model Iterate through each round Each step is based on the current observation information. Use strategy Select decision action information ,according to , and hidden state Historical observation information and historical decision-making information Predicting instant reward value and observation information for the next time slot .

[0159] Step 7: Generate simulated trajectory Store in the model experience pool .

[0160] Step 8, from and Medium sampling Groups, each group Each strip is [length missing] Training the Actor network and Critic network using sub-trajectories Each time, the loss function is updated using the following methods respectively. , , and The converged Actor and Critic networks are obtained:

[0161] Formula (17)

[0162] in, express The loss value; This represents the reward value output by the first Critic network. Indicates the discount factor; Indicates time slot Observational information; Indicates time slot Information on the selection and decision-making actions; Represents the time slots of the outputs of the two target Critic networks. The reward value of the information on the choice and decision-making action; This represents the entropy-temperature factor.

[0163] Formula (18)

[0164] in, express The loss value; This represents the reward value output by the second Critic network.

[0165] Formula (19)

[0166] in, express The loss value.

[0167] Formula (20)

[0168] in, express The loss value; Represents the target entropy. , Indicates the dimension of the action space; network parameters of the soft-update target Critic network: ; , This represents the updated network parameters of the first target Critic network; Indicates the soft update coefficient; This represents the network parameters before the first target Critic network update; The second objective is to update the network parameters of the Critic network. This represents the network parameters before the second objective Critic network was updated.

[0169] Step 9: Based on the observation information of the observation space in the current time slot and the converged Actor network, determine the target decision action information.

[0170] As an example, step 9 may specifically include: inputting the observation information of the observation space in the current time slot into the converged Actor network to obtain the target decision action information output by the converged Actor network.

[0171] S109. Control the UAV edge network in the current time slot based on target decision action information.

[0172] As an example, as described above, the decision action information may include the UAV flight time, the UAV flight angle, and the sub-time slot allocation information of all observable user equipment. In this case, S109 may specifically include: planning or controlling the flight trajectory of the UAV based on the UAV flight time and UAV flight angle in the target decision action information, and allocating computing resources based on the sub-time slot allocation information of all observable user equipment.

[0173] The UAV edge network control method provided in this application embodiment jointly considers the user mobility model, task arrival model, and partially observable environmental constraints to construct a POMDP. Subsequently, the RSAC algorithm is used to make decisions based on historical information; LSTM is used to model the environment transition probability, generating simulation experience and accelerating the convergence of the RSAC algorithm. Finally, target decision action information is determined with the objectives of minimizing user computational energy consumption and maximizing user uplink speed in the network scenario. Based on this target decision action information, the UAV edge network in the current time slot is controlled, thus providing a UAV edge network control scheme that accurately plans the UAV's flight trajectory and rationally allocates computational resources in a dynamically changing environment.

[0174] The above primarily describes the solutions provided by the embodiments of this application from a methodological perspective. To achieve the above functions, the UAV edge network control device includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should readily recognize that, in conjunction with the units and algorithm steps of the various examples described in the embodiments disclosed herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Experts may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0175] In an exemplary embodiment, this application also provides an unmanned aerial vehicle (UAV) edge network control device. Figure 3 This is a schematic diagram illustrating the composition of a drone edge network control device provided in an embodiment of this application. Figure 3 As shown, the device includes: a loading module 301, an acquisition module 302, and a processing module 303.

[0176] Loading module 301 is used to load a predefined UAV edge network scenario model. The UAV edge network scenario model includes a UAV, multiple user devices, a discrete-time system, mobile boundary coordinates, and flight hovering protocol information. The UAV carries an MEC server to provide computing services to multiple user devices. It also loads predefined user terminal mobility models and task generation models. The user terminal mobility model represents the speed and position of the user terminal in each time slot; the task generation model represents the task generation rules of the user terminal. Furthermore, it loads predefined UAV flight models and coverage models. The flight model represents the UAV's flight parameters; the coverage model represents the UAV's coverage area. Finally, it loads predefined UAV and user... The device's communication model is used to determine the uplink rate of the user device. A predefined energy consumption model and objective function are loaded. The energy consumption model represents the computational energy consumption of the user device. The objective function is a function that aims to minimize the computational energy consumption of the user device and maximize its uplink rate. A predefined seven-tuple of the POMDP model is loaded. The seven-tuple includes the state space, observation space, action space, state transition probability, observation transition probability, reward function, and discount factor of the UAV edge network system. The state space includes complete information about the UAV edge network system in each time slot. The observation space includes the observation information that the UAV can observe in each time slot. The action space includes the decision-making action information of the UAV in each time slot.

[0177] The acquisition module 302 is used to acquire observation information of the observation space under the current time slot.

[0178] The processing module 303 is used to determine target decision action information based on the observation information of the observation space in the current time slot, the seven-tuple, the Dyna environment model, and the RSAC algorithm, with the goal of minimizing the computational energy consumption of the user equipment and maximizing the uplink rate of the user equipment; and to control the UAV edge network in the current time slot based on the target decision action information.

[0179] In some possible embodiments, the discrete-time system includes a set of time slots; the set of time slots includes multiple time slots of equal length; each time slot includes a flight duration and a hovering duration; the UAV flies at a constant speed during the flight duration without connecting to user equipment; the hovering duration includes multiple sub-time slots, each sub-time slot being allocated to a corresponding user equipment for task offloading; the flight duration is less than or equal to the time slot duration; the movement boundary coordinates include the boundary coordinates of a rectangular region.

[0180] In some possible embodiments, in the user terminal mobility model, the user equipment exist The speed of the time slot satisfies the following formula:

[0181] ;

[0182] in, Indicates user equipment exist The speed of the time slot; Indicates user equipment exist The speed of the time slot; Indicates Gaussian white noise; Indicates memory level; Indicates asymptotic average velocity; Indicates asymptotic standard deviation;

[0183] User equipment exist The location of the time slot satisfies the following formula:

[0184]

[0185] in, Indicates user equipment exist The location of the time slot; Indicates user equipment exist The location of the time slot; Indicates the duration of the time slot;

[0186] In the task generation model, each user device at the beginning of each time slot is assigned a preset probability. Received fixed data size The task; each user device maintains an infinitely large task queue. If a received task is not processed or unloaded to the drone immediately, the task will accumulate in the task queue. The total number of tasks in the drone edge network system is fixed at M.

[0187] In some possible embodiments, in the flight model, the drone is at a fixed altitude of At a speed of meters Uniform flight It is a positive number; The horizontal coordinates of the UAV at the beginning of the time slot are ,and , ; For drones The flight angle of the time slot, and drones exist The position update at the beginning of the time slot satisfies the following formula:

[0188]

[0189] in, and Indicates drone exist The horizontal coordinate at the beginning of the time slot; Indicates that drones are in Flight duration of time slots; Indicates drone Flight speed;

[0190] In the coverage model, Indicates drone exist The set of user equipment within the coverage area of ​​the time slot, and , , It represents a collection of multiple user devices. Indicates user terminal With drones The two-dimensional distance between them Indicates the maximum coverage radius of the drone. , This indicates the maximum elevation angle between the drone and the user's destination.

[0191] In some possible embodiments, in the communication model, Time-slot drones With ground user terminal The uplink transmission rate between nodes satisfies the following formula:

[0192]

[0193] in, express Time-slot drones With ground user terminal Inter-link uplink transmission rate; Indicates bandwidth; Indicates the transmit power of the user equipment; Indicates Gaussian white noise; express Time-slot drones With ground user terminal Overall channel gain between;

[0194]

[0195] in, express Time-slot drones With ground user terminal Line-of-sight (LOS) channel probability; express Time-slot drones With ground user terminal LOS channel gain between; express Time-slot drones With ground user terminal Non-line-of-sight NLOS channel probability; express Time-slot drones With ground user terminal Inter-NLOS channel gain;

[0196]

[0197]

[0198]

[0199]

[0200] in, and Represents environmentally relevant constants; Represents the speed of light; Indicates the carrier frequency of the wireless channel; express Time-slot drones With ground user terminal The three-dimensional Euclidean distance between them ; Add a loss factor to NLOS; This represents the path loss exponent of the LOS channel.

[0201] In some possible embodiments, in the energy consumption model, express Time slots allocated to observable user equipment The sub-time slot, and ; Determine which user to assign sub-slot Afterwards, the amount of data unloaded Time slot User equipment Offloading energy consumption during transmission tasks , Indicates the transmit power of user equipment; time slot User equipment Locally computed workload , This indicates the fixed frequency of the processor in the user terminal; A time slot represents the number of processor cycles required to process a unit of bit. User equipment Energy consumption for processing local computing tasks , Indicates the energy consumption coefficient. Indicates time slot duration; User equipment exist The workload of a time slot is expressed by the following formula:

[0202]

[0203] in, Indicates user equipment exist The workload in a time slot; Indicates user equipment exist The workload in a time slot; Indicates the size of the task's data; Indicates time slot The workload of drone calculations; This represents the preset probability that each user device will receive a task at the beginning of each time slot;

[0204] The objective function is expressed by the following formula:

[0205]

[0206] in, express Time-slot drones With ground user terminal Inter-link uplink transmission rate; Indicates time slot User equipment Energy consumption for processing local computing tasks; This indicates the flight angle constraints of the drone. Indicates that drones are in The flight angle of the time slot; Indicates a single sub-slot constraint. express Time slots allocated to observable user equipment The sub-time slot, express Drone flight duration within the time slot; This means that the sum of the flight duration and all sub-slots equals the slot duration; This indicates a constraint on the total amount of tasks. This represents the total number of tasks generated by all user devices throughout the entire task cycle. Indicates an indicator variable. ; This indicates the amount of work the drone calculates throughout the entire mission cycle. The workload of local computation on user devices The sum is greater than or equal to M.

[0207] In some possible embodiments, complete information includes: the current 2D position of the UAV, the 2D positions of all user devices, the task queues of all user devices, and the channel gain between all user devices and the UAV; observation information includes: the current 2D position of the UAV, the 2D positions of observable user devices, the task queues of observable user devices, and the channel gain between observable user devices and the UAV; decision action information includes the UAV flight time, the UAV flight angle, and the sub-time slot division information of all observable user devices; the reward function satisfies the following relationship:

[0208]

[0209] in, This represents the reward function value; express Time-slot drones With ground user terminal Inter-link uplink transmission rate; Indicates the weighting coefficient; Indicates time slot User equipment Energy consumption during offloading of transmission tasks; Indicates time slot User equipment Energy consumption for processing local computing tasks.

[0210] In some possible embodiments, processing module 303 is specifically used to initialize the network parameters of the Actor network. Critic network parameters and Dyna environment model parameters And set the target Critic network parameters. and Initialize the real experience replay pool and virtual experience pool All are empty sets; iterate through the set maximum number of rounds, initializing the hidden state of the LSTM network at the start of each round; within each round, iterate through the maximum number of steps, and at each step, iterate based on the current observation information. and historical observation information in hidden state Use strategy Select decision action information and in the information of decision-making actions Then, the immediate reward is obtained based on the reward function. and observation information for the next time slot After each round, the current trajectory data will be... Store in the real experience pool ;from Medium sampling Groups, each group Each strip is [length missing] Sub-trajectories, training the Dyna environment model Then, update using the following loss function. :

[0211]

[0212] in, Represents the loss function of the Dyna environment model; The first prediction made by the Dyna environment model represents the Sub-trajectories in time slots Observed values; Indicates the Sub-trajectories in time slots The actual observed value; Represents the square of the L2 norm; The first prediction made by the Dyna environment model represents the Sub-trajectories in time slots The reward value; Indicates the first Sub-trajectories in time slots The actual reward value;

[0213] Initialize the hidden state of the Dyna environment model Iterate through each round Each step is based on the current observation information. Use strategy Select decision action information ,according to , and hidden state Historical observation information and historical decision-making information Predicting instant reward value and observation information for the next time slot Generate simulated trajectory Store in the model experience pool ;from and Medium sampling Groups, each group Each strip is [length missing] Training the Actor network and Critic network using sub-trajectories Each time, the loss function is updated using the following methods respectively. , , and The converged Actor and Critic networks are obtained:

[0214]

[0215] in, express The loss value; This represents the reward value output by the first Critic network. Indicates the discount factor; Indicates time slot Observational information; Indicates time slot Information on the selection and decision-making actions; Represents the time slots of the outputs of the two target Critic networks. The reward value of the information on the choice and decision-making action; Represents the entropy-temperature factor;

[0216]

[0217] in, express The loss value; This represents the reward value output by the second Critic network;

[0218]

[0219] in, express The loss value;

[0220]

[0221] in, express The loss value; Represents the target entropy. , Indicates the dimension of the action space; network parameters of the soft-update target Critic network: ; , This represents the updated network parameters of the first target Critic network; Indicates the soft update coefficient; This represents the network parameters before the first target Critic network update; The second objective is to update the network parameters of the Critic network. This represents the network parameters before the Critic network update for the second objective;

[0222] Based on the observation information of the observation space in the current time slot and the converged Actor network, the target decision action information is determined.

[0223] It should be noted that Figure 3 The module division shown is illustrative and represents only one logical functional division; in actual implementation, other division methods are possible. For example, two or more functions can be integrated into a single processing module. These integrated modules can be implemented in hardware or as software functional units.

[0224] In an exemplary embodiment, this application also provides a drone. Figure 4 This is a schematic diagram illustrating the composition of a drone provided in an embodiment of this application. Figure 4 As shown, the drone includes a processor 402, a communication interface 403, and a bus 404. As an example, the drone may also include a memory 401.

[0225] Processor 402 may implement or execute various exemplary logic blocks, modules, and circuits described in conjunction with this disclosure. Processor 402 may be a central processing unit, a general-purpose processor, a digital signal processor, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It may implement or execute various exemplary logic blocks, modules, and circuits described in conjunction with this disclosure. Processor 402 may also be a combination that implements computational functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0226] Communication interface 403 is used to connect to other devices via a communication network. This communication network can be Ethernet, wireless access network, wireless local area network (WLAN), etc.

[0227] The memory 401 may be a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, or electrically erasable programmable read-only memory (EEPROM), disk storage medium or other magnetic storage device, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but is not limited thereto.

[0228] As one possible implementation, the memory 401 can exist independently of the processor 402. The memory 401 can be connected to the processor 402 via a bus 404 and is used to store instructions or program code. When the processor 402 calls and executes the instructions or program code stored in the memory 401, it can implement the UAV edge network control method provided in this embodiment of the present disclosure.

[0229] In another possible implementation, the memory 401 can also be integrated with the processor 402.

[0230] Bus 404 can be an extended industry standard architecture (EISA) bus, etc. Bus 404 can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 4 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0231] In an exemplary embodiment, this application also provides a readable storage medium including software instructions that, when run on a drone, cause the drone to perform any of the methods provided in the above embodiments.

[0232] In an exemplary embodiment, this application also provides a computer program product containing computer execution instructions, which, when run on a drone, causes the drone to perform any of the methods provided in the above embodiments.

[0233] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented using software programs, implementation can be, in whole or in part, in the form of a computer program product. This computer program product includes one or more computer-executable instructions. When these computer-executable instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer-executable instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer-executable instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device containing one or more servers, data centers, etc., that can be integrated with the medium. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state disks, SSDs).

[0234] Although this application has been described herein in conjunction with various embodiments, those skilled in the art, by reviewing the accompanying drawings, disclosure, and appended claims, will understand and implement other variations of the disclosed embodiments in carrying out the claimed application. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude multiple instances. A single processor or other unit can implement several functions listed in the claims. While different dependent claims may recite certain measures, this does not mean that these measures cannot be combined to produce good results.

[0235] Although this application has been described in conjunction with specific features and embodiments, it is obvious that various modifications and combinations can be made thereto without departing from the spirit and scope of this application. Accordingly, this specification and drawings are merely exemplary illustrations of this application as defined by the appended claims, and are considered to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from the spirit and scope of this application. Thus, if such modifications and modifications of this application fall within the scope of the claims of this application and their equivalents, this application is also intended to include such modifications and modifications.

[0236] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for controlling an unmanned aerial vehicle (UAV) via an edge network, characterized in that, The method includes: Load a predefined UAV edge network scenario model; the UAV edge network scenario model includes a UAV, multiple user devices, a discrete-time system, mobile boundary coordinates, and flight hovering protocol information; the UAV carries a mobile edge computing (MEC) server to provide computing services to the multiple user devices; Load a predefined user terminal mobility model and a task generation model; the user terminal mobility model is used to represent the speed and position of the user terminal in each time slot; the task generation model is used to represent the task generation rules of the user terminal. Load predefined UAV flight models and coverage models; the flight model represents the UAV's flight parameters; the coverage model represents the UAV's coverage area. Load a predefined communication model between the UAV and the user equipment; the communication model is used to determine the uplink rate of the user equipment. Load a predefined energy consumption model and objective function; the energy consumption model is used to represent the computing energy consumption of the user equipment; the objective function is a function that aims to minimize the computing energy consumption of the user equipment and maximize the uplink rate of the user equipment. Load a predefined septum of a partially observable Markov decision process (POMDP) ​​model; the septum includes the state space, observation space, action space, state transition probability, observation transition probability, reward function, and discount factor of the UAV edge network system; the state space includes complete information of the UAV edge network system in each time slot; the observation space includes the observation information that the UAV can observe in each time slot; the action space includes the decision action information of the UAV in each time slot. Obtain observation information of the observation space in the current time slot; Based on the observation information of the observation space under the current time slot, the seven-tuple, the Dyna environment model, and the cyclic flexible actor-critic RSAC algorithm, the target decision action information is determined with the goal of minimizing the computational energy consumption of the user equipment and maximizing the uplink rate of the user equipment. The UAV edge network in the current time slot is controlled based on the target decision action information.

2. The method according to claim 1, characterized in that, The discrete-time system includes a set of time slots; the set of time slots includes multiple time slots of equal length; each time slot includes a flight duration and a hovering duration; the UAV flies at a constant speed within the flight duration and does not connect to user equipment; the hovering duration includes multiple sub-time slots, each sub-time slot is used to be allocated to the corresponding user equipment for task offloading; the flight duration is less than or equal to the time slot duration; the movement boundary coordinates include rectangular region boundary coordinates.

3. The method according to claim 1, characterized in that, In the user terminal mobility model, the user equipment exist The speed of the time slot satisfies the following formula: ; in, Indicates user equipment exist The speed of the time slot; Indicates user equipment exist The speed of the time slot; Indicates Gaussian white noise; Indicates memory level; Indicates asymptotic average velocity; Indicates asymptotic standard deviation; User equipment exist The location of the time slot satisfies the following formula: ; in, Indicates user equipment exist The location of the time slot; Indicates user equipment exist The location of the time slot; Indicates the duration of the time slot; In the task generation model, each user device at the beginning of each time slot is assigned a preset probability. Received fixed data size The task; each user device maintains an infinitely large task queue. If a received task is not processed or unloaded to the drone immediately, the task will accumulate in the task queue. The total number of tasks in the drone edge network system is fixed at M.

4. The method according to claim 1, characterized in that, In the flight model, the drone is at a fixed altitude. At a speed of meters Uniform flight It is a positive number; The horizontal coordinates of the UAV at the beginning of the time slot are ,and , ; For drones The flight angle of the time slot, and ; drones exist The position update at the beginning of the time slot satisfies the following formula: ; in, and Indicates drone exist The horizontal coordinate at the beginning of the time slot; Indicates that drones are in Flight duration of a time slot; Indicates drone Flight speed; In the coverage model Indicates drone exist The set of user equipment within the coverage area of ​​the time slot, and , , This represents the collection of the multiple user equipment. Indicates user terminal With drones The two-dimensional distance between them Indicates the maximum coverage radius of the drone. , This indicates the maximum elevation angle between the drone and the user's destination.

5. The method according to claim 1, characterized in that, In the communication model Time-slot drones With ground user terminal The uplink transmission rate between nodes satisfies the following formula: ; in, express Time-slot drones With ground user terminal Inter-link uplink transmission rate; Indicates bandwidth; Indicates the transmit power of the user equipment; Indicates Gaussian white noise; express Time-slot drones With ground user terminal Overall channel gain between; ; in, express Time-slot drones With ground user terminal Line-of-sight (LOS) channel probability; express Time-slot drones With ground user terminal LOS channel gain between; express Time-slot drones With ground user terminal Non-line-of-sight NLOS channel probability; express Time-slot drones With ground user terminal Inter-NLOS channel gain; ; ; ; ; in, and Represents environmentally relevant constants; Represents the speed of light; Indicates the carrier frequency of the wireless channel; express Time-slot drones With ground user terminal The three-dimensional Euclidean distance between them ; Add a loss factor to NLOS; This represents the path loss exponent of the LOS channel.

6. The method according to claim 1, characterized in that, In the energy consumption model express Time slots allocated to observable user equipment The sub-time slot, and ; Determine which user to assign sub-slot Afterwards, unload the amount of data. ; Time slot User equipment Offloading energy consumption during transmission tasks , Indicates the transmit power of user equipment; time slot User equipment Locally computed workload , This indicates the fixed frequency of the processor in the user terminal; A time slot represents the number of processor cycles required to process a unit of bit. User equipment Energy consumption for processing local computing tasks , Indicates the energy consumption coefficient. Indicates time slot duration; User equipment exist The workload of a time slot is expressed by the following formula: ; in, Indicates user equipment exist The workload in a time slot; Indicates user equipment exist The workload in a time slot; Indicates the size of the task's data; Indicates time slot The workload of drone calculations; This represents the preset probability that each user device will receive a task at the beginning of each time slot; The objective function is expressed by the following formula: in, express Time-slot drones With ground user terminal Inter-link uplink transmission rate; Indicates time slot User equipment Energy consumption for processing local computing tasks; This indicates the flight angle constraints of the drone. Indicates that drones are in The flight angle of the time slot; Indicates a single sub-slot constraint. express Time slots allocated to observable user equipment The sub-time slot, express Drone flight duration within the time slot; This means that the sum of the flight duration and all sub-slots equals the slot duration; This indicates a constraint on the total amount of tasks. This represents the total number of tasks generated by all user devices throughout the entire task cycle. Indicates an indicator variable. ; This indicates the amount of work the drone calculates throughout the entire mission cycle. The workload of local computation on user devices The sum is greater than or equal to M.

7. The method according to claim 1, characterized in that, The complete information includes: the current two-dimensional position of the UAV, the two-dimensional positions of all user devices, the task queues of all user devices, and the channel gain between all user devices and the UAV; The observation information includes: the current two-dimensional position of the UAV, the two-dimensional position of the observable user equipment, the task queue of the observable user equipment, and the channel gain between the observable user equipment and the UAV; The decision-making action information includes the UAV flight time, UAV flight angle, and sub-time slot division information of all observable user equipment; The reward function satisfies the following relationship: ; in, This represents the reward function value; express Time-slot drones With ground user terminal Inter-link uplink transmission rate; Indicates the weighting coefficient; Indicates time slot User equipment Energy consumption during offloading of transmission tasks; Indicates time slot User equipment Energy consumption for processing local computing tasks.

8. The method according to claim 1, characterized in that, Based on the observation information of the observation space under the current time slot, the seven-tuple, the Dyna environment model, and the RSAC algorithm, with the goal of minimizing the computational energy consumption of the user equipment and maximizing the uplink rate of the user equipment, the target decision action information is determined, including: Initialize the network parameters of the Actor network Critic network parameters and Dyna environment model parameters And set the target Critic network parameters. and Initialize the real experience replay pool and virtual experience pool All are empty sets; Iterate through the set maximum number of rounds, and initialize the hidden state of the Long Short-Term Memory (LSTM) network at the beginning of each round; Within each round, the maximum number of steps is further iterated, and each step is based on the current observation information. and the historical observation information in the hidden state Use strategy Select decision action information and in the information of decision-making actions Then, an immediate reward is obtained based on the aforementioned reward function. and observation information for the next time slot ; After each round, the current trajectory data will be... Store in the real experience pool ; from Medium sampling Groups, each group Each strip is [length missing] Sub-trajectories, training the Dyna environment model Then, update using the following loss function. : in, Represents the loss function of the Dyna environment model; The first prediction made by the Dyna environment model represents the Sub-trajectories in time slots Observed values; Indicates the first Sub-trajectories in time slots The actual observed value; Represents the square of the L2 norm; The first prediction made by the Dyna environment model represents the Sub-trajectories in time slots The reward value; Indicates the first Sub-trajectories in time slots The actual reward value; Initialize the hidden state of the Dyna environment model Iterate through each round Each step is based on the current observation information. Use strategy Select decision action information ,according to , and hidden state Historical observation information and historical decision-making information Predicting instant reward value and observation information for the next time slot ; Generate simulated trajectory Store in the model experience pool ; from and Medium sampling Groups, each group Each strip is [length missing] Training the Actor network and Critic network using sub-trajectories Each time, the loss function is updated using the following methods respectively. , , and The converged Actor and Critic networks are obtained: in, express The loss value; This represents the reward value output by the first Critic network. Indicates the discount factor; Indicates time slot Observational information; Indicates time slot Information on the selection and decision-making actions; Represents the time slots of the outputs of the two target Critic networks. The reward value of the information on the choice and decision-making action; Represents the entropy-temperature factor; in, express The loss value; This represents the reward value output by the second Critic network; in, express The loss value; in, express The loss value; Represents the target entropy. , Indicates the dimension of the action space; network parameters of the soft-update target Critic network: ; , This represents the updated network parameters of the first target Critic network; Indicates the soft update coefficient; This represents the network parameters before the first target Critic network update; The second objective is to update the network parameters of the Critic network. This represents the network parameters before the Critic network update for the second objective; Based on the observation information of the observation space in the current time slot and the converged Actor network, the target decision action information is determined.

9. A drone, characterized in that, include: Processor and memory; The memory stores instructions that the processor can execute; When the processor is configured to execute the instructions, it causes the drone to perform the method as described in any one of claims 1-8.

10. A computer program product, characterized in that, include: Computer instructions; When the computer instructions are executed in the drone, the drone causes the drone to perform the method as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Flight control and calculation unloading method and system for multi-unmanned aerial vehicle mobile edge calculation

    CN115454527A

  • Unmanned aerial vehicle auxiliary calculation migration method based on depth deterministic strategy gradient

    CN115640131A