Autonomous intelligent Internet of Things driven unmanned aerial vehicle cluster self-organization method

Through the autonomous intelligent IoT-driven drone cluster self-organization method, combined with stochastic game theory and edge-cloud collaboration framework, the drone cluster trajectory and communication topology are optimized, and the network topology changes and energy consumption imbalance of traditional drone clusters in dynamic environments are solved, and efficient independent decision-making and resource optimization are achieved.

CN120523232AActive Publication Date: 2025-08-22NANJING UNIV OF INFORMATION SCI & TECH

Patent Information

Application Number
CN202510997261.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-18
Publication Date
2025-08-22
Estimated Expiration
2045-07-18

AI Technical Summary

Technical Problem

The trajectory optimization of traditional drone groups is difficult to cope with network topology changes, energy consumption imbalance and collaborative decision-making problems in dynamic environments, and existing drone edge computing solutions are difficult to meet real-time and security needs.

Method used

Adopting an autonomous intelligent IoT-driven drone cluster self-organization method, combining stochastic game theory with edge-cloud collaboration framework, through deep reinforcement learning algorithms, the drone cluster trajectory and communication topology are optimized, and independent decision-making and resource optimization are achieved.

Benefits of technology

It realizes independent decision-making in complex dynamic communication scenarios, reduces communication delay and energy consumption, and improves system communication performance and task completion rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120523232A_ABST
    Figure CN120523232A_ABST
Patent Text Reader

Abstract

The invention discloses an unmanned aerial vehicle cluster self-organizing method driven by an autonomous intelligent Internet of Things, belongs to the technical field of unmanned aerial vehicle cluster cooperative control and wireless network communication, and provides a large-scale unmanned aerial vehicle cluster self-organizing networking and dynamic trajectory optimization scheme by deeply fusing a random game theory and an edge-cloud cooperative framework. According to the scheme, a deep reinforcement learning algorithm is adopted to minimize the total time delay of global task completion and the total energy consumption of an unmanned aerial vehicle cluster and realize autonomous decision-making in a complex dynamic communication scene; compared with other methods, the task-driven semantic communication is introduced, and the transmission load is effectively reduced and the overall efficiency of a communication system is improved by extracting task related characteristics and further compressing data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of drone swarm collaborative control and wireless network communication, and specifically relates to a drone swarm self-organization method driven by an autonomous intelligent Internet of Things. Background Art

[0002] With the rapid development of 6G and IoT technologies, edge computing platforms based on unmanned aerial vehicles (UAVs) have become a crucial solution for meeting the demands of low latency, high bandwidth, and reliability. In complex scenarios such as smart transportation, signal coverage, and emergency communications, ground-based terminals often face real-time constraints and energy consumption constraints, while network load and transmission latency struggle to meet the demands of large-scale, multi-tasking concurrency. Therefore, using drone swarms as mobile edge nodes not only allows computing resources to be deployed closer to users, reducing communication and task processing latency, but also leverages the high maneuverability and deployability of drones to flexibly adjust their locations based on user distribution and task requirements, providing reliable communication services.

[0003] However, traditional drone swarm trajectory optimization relies primarily on pre-defined algorithms and static parameter adjustments, making it difficult to cope with network topology changes, uneven energy consumption, and collaborative decision-making challenges in dynamic environments. Furthermore, existing drone edge computing solutions often rely on bit-level communication and centralized scheduling. Due to data transmission delays and insufficient computing resources, traditional centralized approaches struggle to meet real-time and security requirements.

[0004] The Autonomous Intelligent Internet of Things (Agentic AIoT) empowers IoT devices with autonomous decision-making and proactive interaction, enabling them to adapt locally to the real-time environment. Existing technologies have yet to combine the autonomous agent concept of Agentic AIoT with edge computing for drone swarm trajectory optimization. Consequently, a systematic solution is urgently needed that can respond to environmental changes in real time, enable large-scale drone swarm communication and information sharing, and combine local autonomy with global collaboration. Summary of the Invention

[0005] In response to the shortcomings of the existing technology, the present invention proposes an autonomous intelligent IoT-driven drone swarm self-organization method. Based on the Agentic AIoT model, it deeply integrates random game theory with the edge-cloud collaborative framework and adopts a deep reinforcement learning algorithm. It aims to minimize the total delay of global task completion and the total energy consumption of the drone swarm, and realize autonomous decision-making in complex dynamic communication scenarios.

[0006] The present invention adopts the following technical solutions to solve the above technical problems:

[0007] An autonomous intelligent Internet of Things-driven drone cluster self-organization method includes a drone cluster assisted wireless communication system, wherein the drone cluster assisted wireless communication system includes a ground user layer, a drone edge layer and a cloud platform layer; wherein the ground user layer includes Single-antenna user equipment, each single-antenna user equipment is equipped with a semantic encoder and decoder for generating computing tasks, extracting task semantic features and transmitting them to the nearest UAV; the UAV edge layer includes A UAV equipped with a mobile edge server, each with computing, communication, and perception capabilities, deploys distributed MASAC agents to collaboratively process user tasks; the cloud platform layer is responsible for offline pre-training of the MASAC model and periodically delivering global weight parameters; It includes the following three steps: Step 1: Initialization and semantic extraction: The cloud platform completes offline pre-training of the MASAC model before the mission begins and distributes the critic and actor network parameters to each UAV periodically or under event triggering. After generating a computational task, the ground user extracts only the key semantic features of the task, compresses them using a local semantic encoder, and transmits them to the nearest UAV. Step 2: Each UAV runs the MASAC agent locally based on its current state and the received mission semantics. It uses a random game framework to optimize the associations between UAVs, generate a neighbor set, and form a robust distributed communication topology. Step 3: Synchronize edge task execution with event-triggered strategies: Each UAV executes the collected task subset on the local edge server and feeds back the execution results to the local MASAC agent in real time for the next generation of decision iteration. When the local network parameter increment exceeds the threshold, the parameter difference is broadcast to neighbors to reduce communication overhead.

[0008] As a further preferred embodiment of the autonomous intelligent Internet of Things-driven UAV cluster self-organization method of the present invention, in step 1, the communication environment adopts a three-dimensional Cartesian coordinate system, the UAV flies horizontally at a fixed altitude, and the ground users are randomly distributed; it is assumed that within each time slot, the terminal users are static.

[0009] As a further preferred solution of the autonomous intelligent IoT-driven drone swarm self-organization method of the present invention, in step 2, MASAC performs local parallel learning of trajectories and computational scheduling, while also embedding game gain design; the maximum entropy strategy is used to improve exploration efficiency, enabling a balance between exploration and utilization in the strategy evaluation and strategy improvement stages.

[0010] As a further preferred solution of the autonomous intelligent Internet of Things-driven drone cluster self-organization method of the present invention, in step 3, the execution result includes optimizing the drone trajectory and adjusting the user association between drones, thereby minimizing the communication delay and drone energy consumption, and ultimately improving the communication performance and task completion rate of the entire system.

[0011] Compared with the prior art, the present invention adopts the above technical solution and has the following technical effects:

[0012] 1. The present invention proposes a self-organizing method for drone clusters driven by an autonomous intelligent Internet of Things. Based on the AgenticAIoT model, it deeply integrates random game theory with the edge-cloud collaborative framework, and proposes a self-organizing networking and dynamic trajectory optimization scheme for large-scale drone clusters. The scheme adopts a deep reinforcement learning algorithm to minimize the total delay in global task completion and the total energy consumption of the drone cluster, and realizes autonomous decision-making in complex dynamic communication scenarios. Compared with other inventions, the present invention introduces task-driven semantic communication, which further compresses data by extracting task-related features, effectively reduces the transmission load, and improves the overall efficiency of the communication system.

[0013] 2. The MASAC algorithm proposed in this paper not only focuses on maximizing the cumulative reward, but also takes policy entropy as one of the key optimization objectives; the algorithm aims to maximize the weighted sum of cumulative reward and policy entropy, and by introducing policy randomness, it encourages the intelligent agent to have a higher exploration ability when selecting actions. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without paying any creative work.

[0015] Figure 1 This is the application scenario of the present invention, which consists of three layers: ground user layer, drone layer and cloud platform layer;

[0016] Figure 2 is a flow chart of the MASAC algorithm deployed on a UAV of the present invention;

[0017] Figure 3 is an inventive flow chart of the present invention;

[0018] Figure 4 This is a comparison of the convergence performance of the MASAC algorithm used in the present invention and other multi-agent reinforcement learning algorithms (MAPPO, MADDPG) under the system model of the present invention. DETAILED DESCRIPTION

[0019] The technical solution of the present invention is further described in detail below with reference to the accompanying drawings:

[0020] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention. The present invention is described in detail below based on the drawings and preferred embodiments. The purpose and effect of the present invention will become more clear. It should be understood that the specific embodiments described here are only used to explain the present invention and are not used to limit the present invention.

[0021] Based on the Agentic AIoT concept, this paper proposes a method for drone swarm edge computing and trajectory optimization that uses semantic communication-driven distributed multi-agent reinforcement learning and stochastic game-based collaborative optimization. The system, consisting of a ground user layer, a drone edge layer, and a cloud platform layer, leverages the advantages of semantic communication, MASAC local decision-making, and stochastic game-based optimization to achieve an organic fusion of global collaboration and local autonomy.

[0022] The communication environment between ground users and a swarm of drones consists of a collection of ground users and a swarm of drones. In this environment, a ground user can only be served by one UAV within a given time step, whereas a single UAV can serve multiple ground users within the same time step. The communication model between drones and ground users uses a line-of-sight (LOS) channel probability model. The air-ground channel contains both LOS and NLOS links, each with its own connection probability. A specific function is used to model and analyze the LOS transmission probability.

[0023] like Figure 3 As shown, the optimization method specifically includes the following three steps:

[0024] (1) System initialization and semantic extraction. The cloud platform completes offline pre-training of the MASAC model before the mission begins, and sends the critic and actor network parameters to each UAV periodically or on an event-triggered basis. After generating a computational task, the ground user extracts only the key semantic features of the task, compresses them using a local semantic encoder, and transmits them to the nearest UAV.

[0025] (2) Each UAV runs the MASAC agent locally based on its current state and the received task semantics, and combines the random game framework to optimize the association between UAVs, generate a neighbor set, and form a robust distributed communication topology.

[0026] (3) Edge task execution is synchronized with event-triggered strategies. Each UAV executes a subset of the collected tasks on the local edge server and feeds the execution results back to the local MASAC agent in real time for the next generation of decision iteration. When the local network parameter increment exceeds a threshold, the parameter difference is broadcast to neighbors to reduce communication overhead.

[0027] Furthermore, the environment in step (1) includes the following features: the environment adopts a three-dimensional Cartesian coordinate system, the UAV flies horizontally at a fixed altitude, and the ground users are randomly distributed. Considering that each time slot interval is very short, it can be assumed that the end users are static within each time slot.

[0028] Furthermore, in step (2), MASAC performs local parallel learning of trajectories and computational scheduling, while also embedding game payoff design. The core idea of ​​the proposed MASAC algorithm is to improve exploration efficiency through the maximum entropy strategy, so that exploration and exploitation can be balanced in the strategy evaluation and strategy improvement stages.

[0029] Furthermore, the execution result of step (3) includes optimizing the drone trajectory and adjusting the user association between drones,

[0030] This minimizes communication delay and drone energy consumption, ultimately improving the communication performance and mission completion rate of the entire system.

[0031] System model: such as Figure 1 As shown, the UAV cluster assisted wireless communication system includes a ground user layer, a UAV edge layer and a cloud platform layer; wherein the ground user layer includes Single-antenna user equipment, each single-antenna user equipment is equipped with a semantic encoder and decoder for generating computing tasks, extracting task semantic features and transmitting them to the nearest UAV; the UAV edge layer includes A UAV equipped with a mobile edge server, each UAV has computing, communication and perception capabilities, deploys distributed MASAC agents, and collaboratively processes user tasks; the cloud platform layer is responsible for offline pre-training of the MASAC model and periodically delivering global weight parameters.

[0032] The entire flight cycle of the UAV is , the total flight time Divided into equal intervals time slots, and the length of each time slot is ;Use the three-dimensional Cartesian coordinate system to model the flight state of the UAV, and set the flight altitude of the UAV as a fixed value ; in the In each time slot, the position coordinates of the UAV are expressed as ; The drone's position update follows a specific formula in, It's the drone. The moving distance at the moment quantifies the movement amplitude of the UAV in that time step; It's the drone in The flight direction at any moment is the yaw angle, which accurately indicates the flight direction of the drone;

[0033] In order to ensure the flight safety of drones and avoid collisions between drones, this patent also introduces a safety distance constraint, stipulating that any two drones and The Euclidean distance between in, is a pre-set safety distance threshold; Communication model: Each user is equipped with a task encoding processing unit to extract the semantics of the task as a task representation; In the time slot, ground users The original task size generated is , the task encoding ratio is , which is expressed as the ratio between the optimized transmission content size and the original task size during task encoding; for users , the time required to extract the task representation from the original task is in, is a function of the task encoding ratio, indicating that the user The additional computational load required to extract lightweight task representations during task encoding; Represents a user the computational power available during task encoding; remember for Moment User The transmission power, For users With UAV The channel bandwidth, For users With UAV Therefore, the user With UAV The transmission rate between is expressed as: in, The same-frequency interference caused by other drones and ground users, is the noise power.

[0034] Furthermore, users With UAV The data transmission delay is expressed as ; Markov decision process: In deep reinforcement learning applications, the Markov decision process is defined as the basic framework for solving sequential decision problems in a random environment. A Markov decision process usually consists of five parts: state space , action space , state transition probability function, reward function, discount factor At each decision moment, the agent chooses an action based on a policy that defines the probability distribution of actions chosen in a given state as: ; After each action is executed, the agent transitions to the next state according to the state transition probability function and is given an immediate reward. This state transition and reward feedback process guides policy optimization. The agent adjusts its strategy through iterative interaction with the environment to gradually approach the optimal strategy. Stochastic Game Framework: To further improve communication reliability and resource utilization efficiency in drone swarms, we extend the Markov decision process model to a multi-agent stochastic game. By integrating the stochastic game framework with MASAC, we adaptively establish reliable connections at the game level while enabling drones to autonomously optimize their trajectories and edge computing. The random game model is represented by the quintuple ;in Represented as a collection of drones, Represented as the set of all states, Indicates the The action space of the drones, the joint action space is , Represents the state probability transfer function, which is composed of flight dynamics and link random failure probability composition, Indicates the Immediate benefits of a drone; Global State Concatenate the vectors of all local states: , local state Defined as a triple ,in Indicates the The spatial position of the UAV, Indicates the The remaining battery power of the UAV, The link after combining large-scale path loss and small-scale fading normalization is quality indicators; In the random game model, UAV in time slot The compound action should contain four parts of decision in, Indicates that the UAV is in the current time slot displacement vector within, neighbor subset Indicates the UAV in this time slot Selected communication partners, power allocation For each Link transmission power and bandwidth allocation For each Link bandwidth resources; In the MASAC algorithm, the policy function of each agent is Output these four types of decisions simultaneously, so that the selection of neighbor sets is coordinated with objectives such as drone trajectory and communication delay to achieve optimal results; No. UAVs to select neighbor sets The immediate benefit is composed of communication capacity benefit, power cost and bandwidth cost in, For Link The link capacity, that is, the maximum transmission rate of the link; By combining this game gain with MASAC’s trajectory and instant rewards, a complete integrated reward can be formed to drive the Critic and Actor networks to make decisions. MASAC algorithm: Figure 2 As shown, under the MASAC algorithm framework, for the above composite states and actions, the MASAC agent Instant rewards It consists of three parts: game income, task delay and total energy consumption; among them, the delay includes task extraction delay , transmission delay And calculate delay ; Energy consumption is determined by flight energy consumption , calculate energy consumption And communication energy consumption Composition; Agent Instant rewards It is defined as follows: in, and They are respectively the weights of delay and energy consumption, which can be adjusted according to system requirements; In order to improve the sampling rate, the MASAC algorithm adopts a dual network structure, including an actor network and two critic networks. is the Critic network parameter, is the target Critic network parameter, is the Actor network parameter; the critic network estimates the Q value by minimizing the temporal difference loss, and the loss function is expressed as in is the target Q value, For experience replay cache, Current status Execute an action Then transfer to the next state; at the target Q value In the entropy regularization term, the temperature parameter To control the weight of entropy in the optimization process and balance the relationship between exploration and utilization; expressed as in, For the next state According to the policy function The next action of sampling; In the policy improvement phase, the MASAC algorithm uses the policy gradient method to optimize the policy, and the update goal is to minimize the policy network loss function: ; This loss takes into account both maximum entropy and maximum value, encouraging random exploration and high-reward actions, thereby updating the Actor network parameters In order to save communication overhead in distributed deployment, an event-triggered synchronization mechanism is introduced: after each Critic or Actor network update, the UAV Calculation parameters and decision increments Only when Exceeding the threshold When the local network or location changes are considered significant enough to notify neighbors broadcast and , the broadcast only contains the incremental vector, so the message size is much smaller than the complete parameter and is more lightweight; other UAV Upon receiving any neighbor After the increment, it will be weightedly fused with the old parameters retained by itself: in, is the channel gain; is the normalized weight; For the fusion rate, control the ratio of new and old information.

[0035] The proposed drone cluster self-organization method in the Agentic AIoT mode has significant advantages in optimizing system performance through the random game and MASAC algorithm. It significantly reduces the communication load while maintaining strategy consistency, optimizes the drone cluster trajectory by minimizing task delay and total energy consumption, and improves the overall communication performance of the system. Figure 4 As shown in the figure, compared with the traditional multi-agent reinforcement learning algorithms MAPPO and MADDPG, it has higher real-time and robustness, faster convergence speed and larger reward value.

[0036] Those skilled in the art will understand that the above descriptions are merely preferred embodiments of the invention and are not intended to limit the invention. Although the invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art will still be able to modify the technical solutions described in the aforementioned embodiments or substitute equivalents for some of the technical features. Any modifications, equivalent substitutions, etc. made within the spirit and principles of the invention shall be included within the scope of protection of the invention. All technical features in this embodiment may be freely combined according to actual needs.

[0037] Finally, it should be noted that the above is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or make equivalent substitutions for some of the technical features therein. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A self-organizing method for drone swarms driven by an autonomous intelligent Internet of Things, characterized by: It includes a UAV cluster-assisted wireless communication system, which includes a ground user layer, a UAV edge layer and a cloud platform layer; wherein, The ground user layer includes Single-antenna user devices, each equipped with a semantic encoder and decoder, are used to generate computing tasks, extract task semantic features, and transmit them to nearby drones; The drone edge layer includes A UAV equipped with a mobile edge server, each with computing, communication, and perception capabilities, deploys a distributed MASAC agent to collaboratively process user tasks; The cloud platform layer is responsible for offline pre-training of the MASAC model and periodically delivering global weight parameters; It includes the following three steps: Step 1: Initialization and semantic extraction: The cloud platform completes offline pre-training of the MASAC model before the mission begins and distributes the critic and actor network parameters to each UAV periodically or under event triggering. After generating a computational task, the ground user extracts only the key semantic features of the task, compresses them using a local semantic encoder, and transmits them to the nearest UAV. Step 2: Each UAV runs the MASAC agent locally based on its current state and the received mission semantics. It uses a random game framework to optimize the associations between UAVs, generate a neighbor set, and form a robust distributed communication topology. Step 3: Synchronize edge task execution with event-triggered strategies: Each UAV executes the collected task subset on the local edge server and feeds back the execution results to the local MASAC agent in real time for the next generation of decision iteration. When the local network parameter increment exceeds the threshold, the parameter difference is broadcast to neighbors to reduce communication overhead.

2. The method for self-organizing a drone swarm driven by an autonomous intelligent Internet of Things according to claim 1, characterized in that: In step 1, the communication environment adopts a three-dimensional Cartesian coordinate system, the UAV flies horizontally at a fixed altitude, and the ground users are randomly distributed; it is assumed that in each time slot, the terminal users are static.

3. The autonomous intelligent Internet of Things-driven drone swarm self-organization method according to claim 1, characterized in that: In step 2, MASAC performs local parallel learning of trajectories and computational scheduling, while also embedding game-benefit design. It improves exploration efficiency through the maximum entropy strategy, enabling a balance between exploration and exploitation during the strategy evaluation and improvement stages.

4. The autonomous intelligent Internet of Things-driven drone swarm self-organization method according to claim 1, characterized in that: In step 3, the execution results include optimizing the UAV trajectory and adjusting the user associations between UAVs, thereby minimizing the communication delay and UAV energy consumption, and ultimately improving the communication performance and task completion rate of the entire system.

5. The autonomous intelligent Internet of Things-driven drone swarm self-organization method according to claim 1, characterized in that: The entire flight cycle of the UAV is , the total flight time Divided into equal intervals time slots, and the length of each time slot is ;Use the three-dimensional Cartesian coordinate system to model the flight state of the UAV, and set the flight altitude of the UAV as a fixed value ; in the In each time slot, the position coordinates of the UAV are expressed as ; The drone's position update follows a specific formula in, It's the drone in The moving distance at the moment quantifies the movement amplitude of the UAV in that time step; It's the drone. The flight direction at the moment is the yaw angle; At the same time, a safety distance constraint is introduced, stipulating that any two drones and The Euclidean distance between in, is a pre-set safety distance threshold; Communication model: Each user is equipped with a task encoding processing unit to extract the semantics of the task as a task representation; In the time slot, ground users The original task size generated is , the task encoding ratio is , which is expressed as the ratio between the optimized transmission content size and the original task size during task encoding; for users , the time required to extract the task representation from the original task is in, is a function of the task encoding ratio, indicating that the user The additional computational load required to extract lightweight task representations during task encoding; Represents a user the computational power available during task encoding; remember for Moment User The transmission power, For users With UAV The channel bandwidth, For users With UAV Channel gain of user With UAV The transmission rate between is expressed as: in, The same-frequency interference caused by other drones and ground users, is the noise power; Furthermore, users With UAV The data transmission delay is expressed as ; Markov decision process: In deep reinforcement learning applications, the Markov decision process is defined as the basic framework for solving sequential decision problems in a random environment. A Markov decision process usually consists of five parts: state space , action space , state transition probability function, reward function, discount factor At each decision moment, the agent chooses an action based on a policy that defines the probability distribution of actions chosen in a given state as: ; After each action is executed, the agent transitions to the next state according to the state transition probability function and is given an immediate reward. This state transition and reward feedback process guides policy optimization. The agent adjusts its strategy through iterative interaction with the environment to gradually approach the optimal strategy. Stochastic Game Framework: Based on the Markov decision process model, it is extended to multi-agent stochastic games. By integrating the stochastic game framework with MASAC, it can adaptively establish reliable connections at the game level while enabling drones to autonomously optimize their trajectories and edge computing. The random game model is represented by the quintuple ;in Represented as a collection of drones, Represented as the set of all states, Indicates the The action space of the drones, the joint action space is , Represents the state probability transfer function, which is composed of flight dynamics and link random failure probability composition, Indicates the Immediate benefits of a drone; Global State Concatenate the vectors of all local states: , local state Defined as a triple ,in Indicates the The spatial position of the UAV, Indicates the The remaining battery power of the UAV, The link after combining large-scale path loss and small-scale fading normalization is quality indicators; In the random game model, UAV in time slot The compound action should contain four parts of decision in, Indicates that the UAV is in the current time slot displacement vector within, neighbor subset Indicates the UAV in this time slot Selected communication partners, power allocation For each Link transmission power and bandwidth allocation For each Link bandwidth resources; In the MASAC algorithm, the policy function of each agent is Output these four types of decisions simultaneously, so that the selection of neighbor sets is coordinated with objectives such as drone trajectory and communication delay to achieve optimal results; No. UAVs to select neighbor sets The immediate benefit is composed of communication capacity benefit, power cost and bandwidth cost in, For Link The link capacity, that is, the maximum transmission rate of the link; By combining this game gain with MASAC’s trajectory and instant rewards, a complete integrated reward can be formed to drive the Critic and Actor networks to make decisions. MASAC algorithm: Under the MASAC algorithm framework, for the above composite states and actions, the MASAC agent Instant rewards It consists of three parts: game income, task delay and total energy consumption; among them, the delay includes task extraction delay , transmission delay And calculate delay ; Energy consumption is determined by flight energy consumption , calculate energy consumption And communication energy consumption Composition; Agent Instant rewards It is defined as follows: in, and They are respectively the weights of delay and energy consumption, which can be adjusted according to system requirements; In order to improve the sampling rate, the MASAC algorithm adopts a dual network structure, including an actor network and two critic networks. is the Critic network parameter, is the target Critic network parameter, is the Actor network parameter; the critic network estimates the Q value by minimizing the temporal difference loss, and the loss function is expressed as in is the target Q value, For experience replay cache, Current status Execute an action Then transfer to the next state; at the target Q value In the entropy regularization term, the temperature parameter is introduced To control the weight of entropy in the optimization process and balance the relationship between exploration and utilization; expressed as in, For the next state According to the policy function The next action of sampling; In the policy improvement phase, the MASAC algorithm uses the policy gradient method to optimize the policy, and the update goal is to minimize the policy network loss function: ; This loss takes into account both maximum entropy and maximum value, encouraging random exploration and high-reward actions, thereby updating the Actor network parameters In order to save communication overhead in distributed deployment, an event-triggered synchronization mechanism is introduced: after each Critic or Actor network update, the UAV Calculation parameters and decision increments Only when Exceeding the threshold When the local network or location changes are considered significant enough to notify neighbors broadcast and , the broadcast only contains the incremental vector, so the message size is much smaller than the complete parameter and is more lightweight; other UAV Upon receiving any neighbor After the increment, it will be weightedly fused with the old parameters retained by itself: in, is the channel gain; is the normalized weight; To control the ratio of new and old information, the fusion rate is set.

Citation Information

Patent Citations

  • Unmanned aerial vehicle data acquisition trajectory and user association joint optimization method based on reinforcement learning in wireless network

    CN115616906A

  • Distributed multi-unmanned aerial vehicle relay network coverage method

    CN116017479A

  • Multi-machine collaborative task scheduling method for aerospace network edge computing scene

    CN119759587A

  • Communication network optimization method for air RIS attitude change

    CN120018159A

  • Multi-agent unmanned aerial vehicle cluster collaborative search method based on semantic communication technology

    CN120143879A

Cited By

  • Unmanned aerial vehicle communication system and method based on narrowband Internet of Things

    CN121690337A

  • Edge cluster self-organization and reconstruction method based on deep reinforcement learning

    CN122119756A

  • An edge cluster self-organization and reconstruction method based on deep reinforcement learning

    CN122119756B