Unmanned aerial vehicle group congestion obstacle avoidance control method
Through real-time perception and multi-agent collaboration, the drone paths and strategies are optimized, and the problems of drone groups avoid obstacles and energy utilization in complex environments are solved, achieving efficient and safe task completion.
Patent Information
- Application Number
- CN202411918391.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-25
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-25
AI Technical Summary
How to effectively avoid congestion and collisions in complex and dynamic environments, ensure the smooth progress of flight missions, and ensure the efficient use of energy is a key issue.
A method for avoiding obstacles for drone groups is proposed, including real-time perception of the drone status and its neighbor's environmental information, and implementing congestion detection and strategic solutions based on improved consensus algorithms. Multi-agent methods are used to optimize the drone path, consider fuel efficiency, task completion and safety, avoid collisions through dynamic formations in a congested environment, and dynamically monitor the energy state of each drone, and perform energy replenishment or task handover if necessary.
Effectively respond to dynamic environmental changes, optimize drone flight paths, avoid collisions, reduce energy consumption, ensure smooth completion of tasks, and improve system efficiency and robustness.
Smart Images

Figure CN119937627A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of unmanned aerial vehicle (UAV) technology, and in particular relates to a method for controlling UAV swarm congestion and obstacle avoidance. Background Technology
[0002] With the continuous development of drone technology, drone swarms are playing an important role in various fields, such as environmental monitoring, disaster relief, traffic monitoring, and agricultural spraying. However, drone swarms face several challenges when performing collaborative tasks, especially in complex and dynamic environments. How to effectively avoid congestion and collisions, ensure the smooth progress of flight missions, and ensure the efficient use of energy is a key issue.
[0003] Traditional UAV obstacle avoidance control methods typically rely on centralized algorithms or single-agent decision-making. This can lead to a sharp increase in computational complexity in multi-UAV collaborative operations, making it difficult to achieve real-time and efficient decision-making. Furthermore, in complex environments, the dynamic changes of obstacles and other aircraft require obstacle avoidance strategies to be highly adaptable and flexible, which traditional methods often struggle to handle.
[0004] To address these issues, distributed control methods based on multi-agent cooperation have been widely applied in recent years. These methods avoid the bottleneck of centralized computing and improve system scalability and real-time performance through information sharing and collaborative decision-making among UAVs. Meanwhile, with the increasing diversification of UAV applications, energy constraints have become another significant limiting factor. The limited battery life and mission requirements of UAVs force researchers to consider energy consumption in flight path planning and mission allocation. Optimizing energy use while ensuring mission completion is a crucial research direction in current UAV swarm control. Summary of the Invention
[0005] In view of this, the objective of this invention is to propose a method for controlling congestion and obstacle avoidance in a drone swarm, comprising the following steps:
[0006] Step 1: Real-time sensing of the drone's status and the environmental information of its neighbors;
[0007] Step 2: Implement congestion detection and strategy resolution among drones based on the improved consensus algorithm;
[0008] Step 3: Utilize a multi-agent approach to optimize the drone's path, taking into account fuel efficiency, mission completion, and safety.
[0009] Step 4: In congested environments, avoid collisions through dynamic formation to ensure overall formation stability;
[0010] Step 5: Dynamically monitor the energy status of each UAV and replenish energy or hand over tasks when necessary.
[0011] Specifically, the real-time sensing of the drone's status and its neighbor's environmental information includes the following steps:
[0012] Status acquisition: UAV i The system acquires its own position, velocity, and energy state information in real time through sensors, which is represented as a state vector s. i (t).
[0013] Neighbor identification: within the communication range R c Within, neighbor sets are identified through broadcast communication.
[0014] Status sharing: Broadcast the drone's status information to neighboring drones.
[0015] Specifically, the congestion detection includes the following steps:
[0016] At each time step, collision detection is performed first to predict each pair of drones that may collide.
[0017] Calculate the time to shortest distance for each pair of drones, and measure the time it takes for two drones to reach their closest point: Where, p i and p j : respectively U drone i and U j Position vector, v i and v j : respectively U drone i and U j The velocity vector when TCPA ij A value greater than 0 indicates that the two drones may approach a certain minimum distance in the future;
[0018] Calculate the shortest distance for each pair of drones, and measure the minimum distance between two drones to reach their closest point:
[0019] CPA ij =‖p i +TCPA ij ·v i -(p j +TCPA ij ·v j )‖
[0020] If CPA ij <d safe If d, then a potential conflict is considered to exist between the two drones, where d safe The set safe distance threshold;
[0021] Each pair of drones with potential conflict (U iU j The system calculates the conflict level based on the time taken to reach the shortest distance and the shortest distance itself, and assigns a resolution priority to each pair of conflicts. Conflict level calculation: Among them, w1 and w2 are weighting coefficients used to adjust the impact of time to shortest distance and shortest distance on conflict level. The higher the conflict level, the greater the necessity of prioritizing its resolution.
[0022] Specifically, the strategy solution includes the following steps:
[0023] Each drone that detects a potential conflict generates multiple candidate avoidance strategies, each candidate strategy u i,k Includes the following information: speed adjustment Δv i,k Change speed to avoid collision, adjust heading Δθ i,k Change the heading angle and adjust the altitude Δh i,k Change flight altitude;
[0024] The goal of the candidate strategies is to reduce collision risk and minimize energy consumption C. energy (u i,k );
[0025] Each drone broadcasts its set of candidate strategies to its neighboring drones. Enable all potentially conflicting drones to share strategy information;
[0026] Consensus variable definition and initialization: Each drone defines a consensus variable y. i (t) represents the avoidance strategy selection at the current time step t. In the initial state, the candidate strategy with the minimum energy consumption is selected as the initial value:
[0027] At each time step, each drone updates its consensus variables based on information from its neighbors. The goal is to gradually reach consensus across the entire drone swarm, using the Laplace consensus algorithm to update these variables. Where ∈ represents the learning rate, controlling the step size for each update, and y j (t) represents the neighboring drone U j Consensus variables at time step t;
[0028] When consensus variable y i When (t) converges to a stable value, it is considered that the drone swarm has reached a consensus, and each drone executes the final avoidance strategy based on this consensus variable, with position updates: p i (t+1)=p i (t)+v i (t)Δt, velocity update: v i (t+1)=vi (t)+Δv i,k Energy Renewal: E i (t+1)=E i (t)-C energy (u i,k ), p i (t) represents the unmanned aerial vehicle U i The three-dimensional position of the UAV at time t, v i (t) represents the unmanned aerial vehicle U i The velocity vector of the UAV at time t, E i (t) represents the unmanned aerial vehicle U i The energy state of the drone at time t.
[0029] The method of using a multi-agent approach to optimize the path of a drone, taking into account fuel efficiency, mission completion, and safety, includes the following steps:
[0030] A multi-objective reward function was designed: R i (t)=w1R fuel (t)+w2R safety (t)+w3R task (t), where w1, w2, w3 are the weights of each objective, determining the importance of each objective in path optimization, R fuel (t) represents the fuel efficiency bonus, which is based on energy consumption; the less energy consumed, the greater the bonus. R safety (t) represents the safety bonus, based on the minimum distance to obstacles or other drones; the greater the distance, the higher the bonus. R task (t) represents the task completion reward, based on the degree to which the task objective is achieved;
[0031] Each drone U i The state space is defined as a vector containing its current position, velocity, energy, and state information of the surrounding environment. Where, p i (t) represents the unmanned aerial vehicle U i The three-dimensional position vector, v, at time t. i (t) represents the unmanned aerial vehicle U i velocity vector, E i (t) represents the unmanned aerial vehicle U i The remaining energy, This provides obstacle information, indicating the nearest obstacle and its location. This is mission information, indicating the current progress of the drone's mission;
[0032] Design Action Network Used according to state s i (t) Generate the optimal action a i(t), the goal is to maximize the cumulative reward: Where, θ i These are parameters of the actor network;
[0033] Design Reviewers Network V i (s i ), used to estimate state s i The value function of , i.e., the expected reward in a given state;
[0034] Design the advantage function A i (t), which measures the quality of the currently chosen action relative to the average level, is calculated using the formula: A i (t)=R i (t)+γV i (s i (t+1))-V i (s i (t)), where γ is the discount factor, and the actor adjusts the strategy through the advantage function;
[0035] The parameters of the actor network are updated according to the gradient of the advantage function to implement the policy, while the commentator network is updated by minimizing the mean squared error (MSE) to update the value function.
[0036] Once training is complete, the optimized strategy is used for path planning. At each time step, the drone determines its path based on its current state s. i (t) Generate the optimal action a through the actor network. i (t), after executing this action, the drone will update its position, speed and energy status.
[0037] The multi-agent method employs a parameter sharing mechanism and a policy communication mode, and includes the following steps:
[0038] Status information sharing: Each drone shares its status information s at each time step t. i (t) It broadcasts to its neighbors, and the status information is transmitted to the neighboring drones. Used to adjust their strategy choices;
[0039] Sharing of motion information, each U drone i According to state s i (t) generates an action a i (t), that is, it makes flight decisions based on the current environmental state, and the drone transmits its selected action a via wireless communication. i (t) is passed to its neighboring drone This allows neighboring drones to adjust their strategies and share experiences;
[0040] Rewards and experience sharing, each drone based on its selected actions.i (t) and environmental feedback receive immediate rewards R i (t), the drone will send its reward R i (t) along with the action taken and the state, is stored as an experience quadruple: ε i (t)=(s i (t),a i (t),R i (t),s i (t+1) is broadcast to other neighboring drones so that they can use this information to update their strategies;
[0041] At each time step t, each drone sends its experience quadruple to the global experience pool, where the experiences of all drones are aggregated. A batch of experiences is randomly drawn from the experience pool for training, updating the policy network parameters of all drones. Where γ is the discount factor, which controls the impact of future rewards, and α is the learning rate, which adjusts the step size of each parameter update;
[0042] Each drone generates new control inputs based on a shared policy network and adjusts them according to the policies of other drones. Each agent updates its policy after each update. Broadcast to all neighboring drones via wireless communication After receiving the policy, the neighboring drone updates its local policy network, enabling it to make decisions based on shared knowledge.
[0043] Specifically, the dynamic monitoring of the energy status of each UAV, and the replenishment of energy or handover of tasks when necessary, includes the following steps:
[0044] When the energy of a certain drone is lower than the set threshold E threshold At that time, the task transfer mechanism will be activated;
[0045] Select a suitable drone to take over the mission, and use path planning algorithms to calculate the mission takeover path;
[0046] The drones were replaced to join the formation, maintain formation stability, and adjust the path to avoid conflict;
[0047] When the drone's energy is insufficient, a replenishment mechanism and task handover ensure the successful completion of the mission.
[0048] The beneficial effects of this invention are as follows: To effectively cope with dynamic environmental changes, this technical solution provides a drone swarm obstacle avoidance control method based on multi-agent cooperation, reinforcement learning, and distributed consensus. It aims to solve the problem of how to optimize flight paths, avoid collisions, reduce energy consumption, and ensure successful mission completion in congested environments. Through task transfer and path adjustment mechanisms, it ensures that when a local drone is unable to continue its mission due to insufficient energy, the mission can be promptly reassigned to other drones, thereby avoiding mission interruption and improving the efficiency and robustness of the entire system. Attached Figure Description
[0049] Figure 1 A flowchart illustrating the overall process of a real-time obstacle avoidance UAV trajectory planning method is presented. Detailed Implementation
[0050] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this invention, and not all embodiments. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.
[0051] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0052] like Figure 1 As shown in the figure, this embodiment proposes a method for controlling congestion and obstacle avoidance in a drone swarm, including the following steps:
[0053] Step 1: Real-time sensing of the drone's status and the environmental information of its neighbors;
[0054] Step 2: Implement congestion detection and strategy resolution among drones based on the improved consensus algorithm;
[0055] Step 3: Utilize a multi-agent approach to optimize the drone's path, taking into account fuel efficiency, mission completion, and safety.
[0056] Step 4: In congested environments, avoid collisions through dynamic formation to ensure overall formation stability;
[0057] Step 5: Dynamically monitor the energy status of each UAV and replenish energy or hand over tasks when necessary.
[0058] Specifically, the real-time sensing of the drone's status and its neighbor's environmental information includes the following steps:
[0059] Status acquisition: UAV iThe system acquires its own position, velocity, and energy state information in real time through sensors, which is represented as a state vector s. i (t).
[0060] Neighbor identification: within the communication range R c Within, neighbor sets are identified through broadcast communication.
[0061] Status sharing: Broadcast the drone's status information to neighboring drones.
[0062] Specifically, the congestion detection includes the following steps:
[0063] At each time step, collision detection is performed first to predict each pair of drones that may collide.
[0064] Calculate the time to shortest distance for each pair of drones, and measure the time it takes for two drones to reach their closest point: Where, p i and p j : respectively U drone i and U j Position vector, v i and v j : respectively U drone i and U j The velocity vector when TCPA ij A value greater than 0 indicates that the two drones may approach a certain minimum distance in the future;
[0065] Calculate the shortest distance for each pair of drones, and measure the minimum distance between two drones to reach their closest point:
[0066] CPA ij =‖p i +TCPA ij ·v i -(p j +TCPA ij ·v j )‖
[0067] If CPA ij <d safe If d, then a potential conflict is considered to exist between the two drones, where d safe The set safe distance threshold;
[0068] Each pair of drones with potential conflict (U i U j The system calculates the conflict level based on the time taken to reach the shortest distance and the shortest distance itself, and assigns a resolution priority to each pair of conflicts. Conflict level calculation: Among them, w1 and w2 are weighting coefficients used to adjust the impact of time to shortest distance and shortest distance on conflict level. The higher the conflict level, the greater the necessity of prioritizing its resolution.
[0069] Specifically, the strategy solution includes the following steps:
[0070] Each drone that detects a potential conflict generates multiple candidate avoidance strategies, each candidate strategy u i,k Includes the following information: speed adjustment Δv i,k Change speed to avoid collision, adjust heading Δθ i,k Change the heading angle and adjust the altitude Δh i,k Change flight altitude;
[0071] The goal of the candidate strategies is to reduce collision risk and minimize energy consumption C. energy (u i,k );
[0072] Each drone broadcasts its set of candidate strategies to its neighboring drones. Enable all potentially conflicting drones to share strategy information;
[0073] Consensus variable definition and initialization: Each drone defines a consensus variable y. i (t) represents the avoidance strategy selection at the current time step t. In the initial state, the candidate strategy with the minimum energy consumption is selected as the initial value:
[0074] At each time step, each drone updates its consensus variables based on information from its neighbors. The goal is to gradually reach consensus across the entire drone swarm, using the Laplace consensus algorithm to update these variables. Where ∈ represents the learning rate, controlling the step size for each update, and y j (t) represents the neighboring drone U j Consensus variables at time step t;
[0075] When consensus variable y i When (t) converges to a stable value, it is considered that the drone swarm has reached a consensus, and each drone executes the final avoidance strategy based on this consensus variable, with position updates: p i (t+1)=p i (t)+v i (t)Δt, velocity update: v i (t+1)=v i (t)+Δv i,k Energy Renewal: E i (t+1)=E i (t)-C energy(u i,k ), p i (t) represents the unmanned aerial vehicle U i The three-dimensional position of the UAV at time t, v i (t) represents the unmanned aerial vehicle U i The velocity vector of the UAV at time t, E i (t) represents the unmanned aerial vehicle U i The energy state of the drone at time t.
[0076] The method of using a multi-agent approach to optimize the path of a drone, taking into account fuel efficiency, mission completion, and safety, includes the following steps:
[0077] A multi-objective reward function was designed: R i (t)=w1R fuel (t)+w2R safety (t)+w3R task (t), where w1, w2, w3 are the weights of each objective, determining the importance of each objective in path optimization, R fuel (t) represents the fuel efficiency bonus, which is based on energy consumption; the less energy consumed, the greater the bonus. R safety (t) represents the safety bonus, based on the minimum distance to obstacles or other drones; the greater the distance, the higher the bonus. R task (t) represents the task completion reward, based on the degree to which the task objective is achieved;
[0078] Each drone U i The state space is defined as a vector containing its current position, velocity, energy, and state information of the surrounding environment. Where, p i (t) represents the unmanned aerial vehicle U i The three-dimensional position vector, v, at time t. i (t) represents the unmanned aerial vehicle U i velocity vector, E i (t) represents the unmanned aerial vehicle U i The remaining energy, This provides obstacle information, indicating the nearest obstacle and its location. This is mission information, indicating the current progress of the drone's mission;
[0079] Design Action Network Used according to state s i (t) Generate the optimal action a i (t), the goal is to maximize the cumulative reward: Where, θ i These are parameters of the actor network;
[0080] Design Reviewers Network Vi (s i ), used to estimate state s i The value function of , i.e., the expected reward in a given state;
[0081] Design the advantage function A i (t), which measures the quality of the currently chosen action relative to the average level, is calculated using the formula: A i (t)=R i (t)+γV i (s i (t+1))-V i (s i (t)), where γ is the discount factor, and the actor adjusts the strategy through the advantage function;
[0082] The parameters of the actor network are updated according to the gradient of the advantage function to implement the policy, while the commentator network is updated by minimizing the mean squared error (MSE) to update the value function.
[0083] Once training is complete, the optimized strategy is used for path planning. At each time step, the drone determines its path based on its current state s. i (t) Generate the optimal action a through the actor network. i (t), after executing this action, the drone will update its position, speed and energy status.
[0084] Specifically, the multi-agent algorithm optimizes paths based on a multi-objective reward function, effectively balancing energy efficiency, safety, and task completion. During training, agents collaborate through shared strategies, ensuring the efficient operation of the drone swarm in complex environments and enabling real-time adjustments based on environmental and mission requirements.
[0085] In energy-constrained UAV congestion avoidance control methods, parameter sharing mechanisms and policy communication enable UAVs to learn from the policies of other UAVs during training, making the policies of each agent (i.e., UAV) more general and effective. This approach improves the overall cooperation of multi-UAV systems, reduces conflicts between individuals, and accelerates task completion.
[0086] The multi-agent method employs a parameter sharing mechanism and a policy communication mode, and includes the following steps:
[0087] Status information sharing: Each drone shares its status information s at each time step t. i (t) It broadcasts to its neighbors, and the status information is transmitted to the neighboring drones. Used to adjust their strategy choices;
[0088] Sharing of motion information, each U dronei According to state s i (t) generates an action a i (t), that is, it makes flight decisions based on the current environmental state, and the drone transmits its selected action a via wireless communication. i (t) is passed to its neighboring drone This allows neighboring drones to adjust their strategies and share experiences;
[0089] Rewards and experience sharing, each drone based on its selected actions. i (t) and environmental feedback receive immediate rewards R i (t), the drone will send its reward R i (t) along with the action taken and the state, is stored as an experience quadruple: ε i (t)=(s i (t),a i (t),R i (t),s i (t+1) is broadcast to other neighboring drones so that they can use this information to update their strategies;
[0090] At each time step t, each drone sends its experience quadruple to the global experience pool, where the experiences of all drones are aggregated. A batch of experiences is randomly drawn from the experience pool for training, updating the policy network parameters of all drones. Where γ is the discount factor, which controls the impact of future rewards, and α is the learning rate, which adjusts the step size of each parameter update;
[0091] Each drone generates new control inputs based on a shared policy network and adjusts them according to the policies of other drones. Each agent updates its policy after each update. Broadcast to all neighboring drones via wireless communication After receiving the policy, the neighboring drone updates its local policy network, enabling it to make decisions based on shared knowledge.
[0092] In energy-constrained drone congestion avoidance control methods, task transfer and path adjustment are crucial, especially when a drone's energy is insufficient to complete its current task. Task transfer mechanisms ensure task continuity and, by replanning paths, allocate tasks to other drones. This process involves multiple aspects, including energy perception, task allocation, path planning, and formation adjustment, ensuring that the drone swarm can still complete tasks efficiently and collaboratively under energy constraints.
[0093] Specifically, the dynamic monitoring of the energy status of each UAV, and the replenishment of energy or handover of tasks when necessary, includes the following steps:
[0094] When the energy of a certain drone is lower than the set threshold E threshold At that time, the task transfer mechanism will be activated;
[0095] Select a suitable drone to take over the mission, and use path planning algorithms to calculate the mission takeover path;
[0096] The drones were replaced to join the formation, maintain formation stability, and adjust the path to avoid conflict;
[0097] When the drone's energy is insufficient, a replenishment mechanism and task handover ensure the successful completion of the mission.
[0098] As used herein, the term "preferred" is meant as an example, illustration, or illustration. Any aspect or design described herein as "preferred" need not be construed as being more advantageous than other aspects or designs. Rather, the use of the term "preferred" is intended to present the concept in a specific manner. As used in this application, the term "or" is intended to mean an inclusive "or" rather than an exclusionary "or." That is, unless otherwise specified or clear from the context, "X uses A or B" naturally includes either of the permutations. That is, if X uses A; X uses B; or X uses both A and B, then "X uses A or B" is satisfied in any of the foregoing examples.
[0099] Furthermore, although this disclosure has been shown and described with respect to one or more implementations, equivalent variations and modifications will occur to those skilled in the art based on a reading and understanding of this specification and the accompanying drawings. This disclosure includes all such modifications and variations and is limited only by the scope of the appended claims. In particular, with respect to the various functions performed by the aforementioned components (e.g., elements, etc.), the terminology used to describe such components is intended to correspond to any component (unless otherwise indicated) that performs the specified function of said component (e.g., is functionally equivalent to it), even if structurally not equivalent to the disclosed structure performing the functions in the exemplary implementations of this disclosure shown herein. Moreover, although specific features of this disclosure have been disclosed with respect to only one of several implementations, such features may be combined with one or more features of other implementations that may be desirable and advantageous for a given or particular application. Furthermore, with regard to the use of the terms “comprising,” “having,” “containing,” or variations thereof in the Detailed Description or claims, such terms are intended to be included in a manner similar to the term “including.”
[0100] The functional units in this invention embodiment can be integrated into a processing module, or each unit can exist physically separately, or multiple units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. The aforementioned devices or systems can execute the storage methods in the corresponding method embodiments.
[0101] In summary, the above embodiments are one implementation of the present invention, but the implementation of the present invention is not limited to the embodiments described above. Any changes, modifications, substitutions, combinations, or simplifications made that deviate from the spirit and principle of the present invention should be considered equivalent substitutions and are included within the protection scope of the present invention.
Claims
1. A method for controlling a group of drones to avoid obstacles during congestion, characterized in that: The steps include: Step 1: Real-time perception of the drone status and its neighbors’ environmental information; Step 2: Implement congestion detection and strategy resolution between drones based on the improved consensus algorithm; Step 3: Use a multi-agent approach to optimize the UAV’s path, taking into account fuel efficiency, mission completion, and safety. Step 4: Avoid collisions through dynamic formation in congested environments to ensure the stability of the overall formation; Step 5: Dynamically monitor the energy status of each drone and perform energy replenishment or task handover when necessary; The real-time perception of the status of the drone and the environmental information of its neighbors includes the following steps: Status collection: UAV i The sensors are used to obtain the position, speed, and energy status information in real time, which is expressed as a state vector s i (t); Neighbor identification: within the communication range R c In the network, the neighbor set is identified through broadcast communication State sharing: Broadcast the drone’s state information to neighboring drones.
2. The method for controlling a drone swarm from congestion to avoiding obstacles according to claim 1, characterized in that: The congestion detection comprises the following steps: At each time step, conflict detection is first performed to predict each pair of UAVs that may collide; Calculate the time to the shortest distance TCPA for each pair of drones ij , measures the time it takes for two drones to reach their closest point, Among them, p i and p j UAV i and U j The position vector, v i and v j UAV i and U j The velocity vector of TCPA ij When >0, it means that the two drones may approach a certain minimum distance in the future; Calculate the shortest distance CPA for each pair of drones ij , measuring the minimum distance between two drones to reach the closest point: CPA ij =‖p i +TCPA ij ·v i -(p j +TCPA ij ·v j )‖ If CPA ij <d safe , then it is considered that there is a potential conflict between the two drones, where d safe is the set safety distance threshold; Each pair of potentially conflicting UAVs (U i ,U j ) will calculate the conflict level based on its time to the shortest distance and the shortest distance, and set a resolution priority for each pair of conflicts. The conflict level calculation is: Among them, w1 and w2 are weight coefficients, which are used to adjust the impact of time to the shortest distance and the shortest distance on the conflict level. The higher the conflict level, the greater the necessity of prioritizing resolution.
3. The method for controlling a drone group from congestion to avoiding obstacles according to claim 2, characterized in that: The strategic solution includes the following steps: Each UAV that detects a potential conflict will generate multiple candidate avoidance strategies, each candidate strategy u i,k Contains the following information: Speed adjustment Δv i,k , change speed to avoid collision, and adjust heading Δθ i,k , change the heading angle, adjust the height Δh i,k , change the flight altitude; The goal of the candidate strategy is to reduce the risk of collision and minimize the energy consumption C energy (u i,k ); Each drone will broadcast its candidate strategy set to neighboring drones. Have all potentially conflicting drones share strategic information; Definition and initialization of consensus variables. Each drone defines a consensus variable y i (t), represents the avoidance strategy selection at the current time step t, y i (t) is the neighboring drone U i The consensus variable at time step t, in the initial state, selects the candidate strategy with the minimum energy consumption as the initial value: At each time step, each drone will update the consensus variable based on the information of its neighbors. The goal is to gradually reach a consensus among the entire drone swarm and use the Laplace consensus algorithm to update the consensus variable. When the consensus variable y i When (t) converges to a stable value, it is considered that the drone group has reached a consensus, and each drone executes the final avoidance strategy according to the consensus variable, and the position is updated: p i (t+1)=p i (t)+v i (t)Δt, speed update: v i (t+1)=v i (t)+Δv i,k , Energy Update: E i (t+1)=E i (t)-C energy (u i,k ), p i (t) is the UAV U i The three-dimensional position of the drone at time t, v i (t) is the UAV U i The velocity vector of the drone at time t, E i (t) is the UAV U i The energy state of the drone at time t, Δt is the time change value.
4. The method for controlling a drone group from congestion to avoiding obstacles according to claim 3 is characterized in that: The multi-agent approach to optimizing the path of a drone, taking into account fuel efficiency, mission completion, and safety, includes the following steps: A multi-objective reward function is designed: R i (t) = w1R fuel (t)+w2R safety (t)+w3R task (t), where w1, w2, and w3 are the weights of each target, which determine the importance of each target in path optimization. fuel (t) represents the fuel efficiency reward, which is based on energy consumption. The less energy consumption, the greater the reward. R safety (t) represents the safety reward, which is based on the minimum distance to obstacles or other drones. The greater the distance, the higher the reward. task (t) represents the task completion reward, which is based on the degree of achievement of the task objectives; Each drone i The state space of is defined as a vector containing the position, velocity, energy and state information of the surrounding environment at time t. Among them, p i (t) represents the drone U i The three-dimensional position at time t, v i (t) represents the drone U i The velocity vector at time t, E i (t) represents the drone U i The remaining energy at time t, is the obstacle information, indicating the nearest obstacle and its location. It is the mission information, indicating the progress of the current mission of the drone; Designing a network of actors Used according to the state i (t) Generate the optimal action a i (t), the goal is to maximize the cumulative reward: Among them, θ i are the parameters of the actor network; Design Reviewer Network V i (s i ), used to estimate the state s i The value function of , which is the expected reward in a given state; Design advantage function A i (t) measures the quality of the currently selected action relative to the average level, and is calculated as: i (t) = R i (t)+γ(V i (s i (t+1))-V i (s i (t))), γ is the discount factor, and the actor adjusts the strategy through the advantage function; The parameters of the actor network are updated strategically according to the gradient of the advantage function, and the update of the critic network: the value function is updated by minimizing the mean square error; After the training is completed, the optimized strategy is used for path planning. At each time step, the drone moves according to its current state s i (t) Generate the optimal action a through the actor network i (t), after executing this action, the drone will update its position, velocity, and energy status.
5. The method for controlling a drone group from congestion to avoiding obstacles according to claim 4, characterized in that: The multi-agent method adopts a parameter sharing mechanism and a strategic communication mode, and includes the following steps: Sharing of state information: Each drone will share its own state information s at each time step t i (t) Broadcast to its neighbors, the status information is passed to the neighboring drones the choice of strategies for adjusting them; Sharing of action information, each drone U i According to the status i (t) Generate an action a i (t), that is, it makes flight decisions based on the current environmental state, and the drone communicates its selected action a via wireless communication. i (t) is passed to its neighboring drones This allows neighboring drones to adjust their strategies and share experiences; Reward and experience sharing, each drone chooses an action a i (t) and environmental feedback to get the immediate reward R i (t), the drone will send its reward R i (t) is stored as an experience quadruple together with the action taken and the state: i (t)=(s i (t),a i (t),R i (t),s i (t+1)), is broadcasted to other neighboring drones so that they can use this information for strategy updates; At each time step t, each drone sends its experience quadruple to the global experience pool, aggregates the experience of all drones in the experience pool, randomly extracts batches of experience from the experience pool for training, and updates the policy network parameters of all drones; Each drone generates new control inputs based on the shared policy network and adjusts them according to the policies of other drones. After each update, each agent sets the policy π θi Broadcast to all neighboring drones via wireless communication After receiving the strategy, the neighboring drone updates the local strategy network and can make decisions based on the shared knowledge.
6. The method for controlling a drone swarm from congestion to avoiding obstacles according to claim 5, characterized in that: The dynamic monitoring of the energy status of each drone and performing energy replenishment or task handover when necessary include the following steps: When the energy of a drone is lower than the set threshold E threshold When the task transfer mechanism is started; Select the appropriate drone to take over the task and use the path planning algorithm to calculate the task takeover path; Replacing drones to join the formation, keeping the formation stable, and adjusting paths to avoid conflicts; When the drone is running low on energy, a replenishment mechanism and task handover are used to ensure smooth completion of the mission.
Citation Information
Patent Citations
Multi-unmanned aerial vehicle cooperative city perception method based on block chain
CN115903911A
Unmanned aerial vehicle cooperative dynamic task allocation method based on extended consistency packet algorithm
CN116610144A
Unmanned aerial vehicle system based on improved intelligent potential field and consensus protocol and obstacle avoidance algorithm thereof
CN118915820A
Cited By
Shared unmanned aerial vehicle intelligent scheduling method and system based on Internet of Things
CN120782160A