Unmanned aerial vehicle formation conflict avoidance control method and system, and medium

By using a Steinberg game architecture and a single-network ADP algorithm, the problems of trajectory tracking, formation stability, and obstacle avoidance safety of UAV formations in dense obstacle environments were solved, and distributed collaborative control of formation stability and collision avoidance safety was achieved.

CN121806930APending Publication Date: 2026-04-07BEIHANG UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-07
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Traditional drone formation control methods struggle to simultaneously ensure formation trajectory tracking accuracy, formation stability, and distributed obstacle avoidance safety in dense obstacle environments.

Method used

By introducing the Steinberg game architecture and combining the virtual leader-follower strategy, the obstacle avoidance penalty mechanism based on the speed obstacle method, and the single-network ADP algorithm, a distributed optimal cooperative control framework is constructed. The hierarchical optimization objective function is defined by the virtual leader and followers, and the Hamilton-Jacobi-Bellman equation is solved online using the single neural network adaptive dynamic programming algorithm to generate the distributed control law.

Benefits of technology

It achieves stability and collision avoidance safety of UAV formations in obstacle environments, improves the real-time performance and robustness of the formations, reduces communication and computing burden, and ensures decentralized collaborative control of the formations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121806930A_ABST
    Figure CN121806930A_ABST
Patent Text Reader

Abstract

The invention provides an unmanned aerial vehicle formation conflict avoidance control method and system and a medium, and the method comprises the steps: introducing a virtual leader-follower-based Stackelberg game architecture, enabling a virtual leader unmanned aerial vehicle to adaptively adjust the trajectory according to the formation state, and enabling the follower unmanned aerial vehicle to maintain the stability of the formation while completing a set tracking task; constructing an obstacle avoidance penalty term based on a speed obstacle method, and embedding the obstacle avoidance penalty term into an optimization objective function of the follower unmanned aerial vehicle to ensure that the follower unmanned aerial vehicle realizes collision avoidance in an obstacle environment; and respectively solving Hamiltonian-Jacobian-Bellman equations corresponding to the leader unmanned aerial vehicle and the follower unmanned aerial vehicle by adopting a single-network ADP algorithm, and generating a controller in real time. According to the invention, the problems of safe obstacle avoidance and tracking control of a multi-unmanned aerial vehicle system with formation stability requirements are effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of unmanned aerial vehicles, and particularly provides a method and system for collision avoidance control of unmanned aerial vehicle formation and a medium. BACKGROUND

[0002] The unmanned aerial vehicle formation has been widely applied in the fields of military reconnaissance, cluster logistics and wide-area area inspection due to its excellent cooperative operation and flexible deployment capability. However, the cooperative operation of the multi-unmanned aerial vehicle system in the dense obstacle environment and the requirement of formation stability make it difficult for the traditional centralized or decentralized control method to simultaneously guarantee the formation trajectory tracking accuracy, the stability of the formation and the distributed obstacle avoidance safety.

[0003] In recent years, as a data-driven adaptive control solution method, the adaptive dynamic programming (ADP) provides a new way for realizing the optimal cooperative control of the multi-unmanned aerial vehicle system. However, the existing research still has deficiencies in the aspects of the design of the formation control architecture, the cooperative integration of the distributed obstacle constraint processing and the online learning mechanism. Therefore, the Steinberg game architecture is introduced, the virtual leader-following strategy, the obstacle avoidance punishment mechanism based on the speed obstacle method and the single-network ADP algorithm are fused, and a distributed optimal cooperative control framework capable of considering the formation stability, the trajectory tracking and the collision avoidance safety is constructed, which has important theoretical significance and engineering application value. SUMMARY

[0004] In order to solve the problems in the prior art, the present application is proposed to provide a solution or partial solution to the above problems.

[0005] In a first aspect, the present application provides a UAV formation collision avoidance control method, comprising the following steps: establishing a Stahlberg game decision architecture between a virtual leader UAV and at least one follower UAV, wherein the virtual leader serves as an upper decision maker, and its decision information is broadcast to all followers; based on the Stahlberg game architecture, a hierarchical optimization objective function is defined for the virtual leader and each follower respectively; wherein the optimization objective function of the virtual leader is configured to include an evaluation item for the trajectory tracking performance of the virtual leader itself, and a formation maintenance cost item related to the formation state of the follower; by minimizing the objective function, the trajectory optimization of the virtual leader itself and the overall stability of the formation are realized; the optimization objective function of each follower is configured to include evaluation items for the trajectory tracking performance of the follower itself and the formation maintenance performance, and an adaptive obstacle avoidance penalty item related to the distance from the obstacle generated based on the speed obstacle method; by minimizing the objective function, the tracking of the leader, the formation maintenance and the safe obstacle avoidance are realized; a single neural network adaptive dynamic programming algorithm is used to solve the Hamilton-Jacobi-Bellman equation corresponding to the optimization objective function of the virtual leader and each follower in parallel online, and a distributed control law is generated and executed in real time according to the solving result.

[0006] Preferably, the optimization objective function of the virtual leader is specifically: wherein γ is a decay coefficient, is the state vector of the leader, and u0 is the control input of the leader, , are positive diagonal weight matrices respectively; is the Lagrange multiplier related to the i-th follower, and V i is the value function of the i-th follower; is the state vector of the leader, wherein the acquisition process of the state vector of the leader includes: under the Stahlberg game architecture, the ideal reference position and the ideal reference speed of the leader UAV are defined as and respectively; the ideal state vector of the leader UAV is defined as: wherein ; the leader augmented dynamic equation considering the tracking performance and the formation stability is defined as: wherein, , is the dynamic vector related to , , , is a system gain matrix.

[0007] Preferably, the value function of the i-th follower is specifically: The specific objective function for optimizing a follower drone is as follows: in, It is the attenuation coefficient. , , , , The diagonal weight matrix is ​​positive definite. Let u be the augmented state vector of the i-th follower. i u is the control input for the i-th follower. j u0 is the control input for the j-th follower, and u0 is the control input for the leader. For the first An adaptive obstacle avoidance penalty term for each obstacle; Wherein, the augmented state vector of the i-th follower The acquisition process includes: Under the Steinberg game framework, for each follower drone, the control objective is: , , ,in, For the first The ideal relative positions of the follower drones with respect to the leader drone; the state equations for the leader drone and the follower drones are defined as follows: , The first one that balances formation and tracking performance The augmented state equation for a follower is defined as: ,in, , , , , This is the system gain matrix.

[0008] Preferably, the adaptive obstacle avoidance penalty term Based on the speed barrier method, the calculation formula is as follows: in, This represents the adaptive safety factor function. Indicates the first The location of the obstacle Indicates the first The position status of a follower drone It is a scheduling function. express hour, The value of , where .

[0009] Preferably, the optimal values of the optimization objective function of the leader and the optimization objective function of the follower are approximated by constructing individual neural networks, and the weights of the neural networks are updated online using gradient descent method whose direction is determined by the corresponding Bellman error. Preferably, the optimal control law of the i-th follower is , which is obtained by solving the optimal value of the optimization objective function of the follower

[0010] and the corresponding Hamiltonian function, and is expressed as: where, is the Hamiltonian function of the i-th follower, u i is the control input of the i-th follower, is the control input of other followers except the i-th follower, u0 is the control input of the leader, is a positive definite diagonal weight matrix. . Preferably, the optimal control law of the virtual leader is , which is obtained by solving the optimal value of the optimization objective function of the leader and the corresponding Hamiltonian function, and is expressed as:

[0011] where, , , , , u0 is the control input of the leader, u i , is the control input of the i-th follower, is a positive definite diagonal weight matrix.

[0012] ​​​In a second aspect, the present application provides a UAV formation cooperative control system based on a Stahberg game, which comprises: a system for controlling a formation comprising a virtual leader UAV and at least one follower UAV, the system comprising: an architecture building module for establishing a Stahberg game decision architecture between the virtual leader UAV and each of the follower UAVs, wherein the virtual leader serves as an upper decision maker and is configured with a broadcasting unit for broadcasting its decision information to all followers; an optimization objective function defining module connected to the architecture building module, for defining a hierarchical optimization objective function for the virtual leader and each of the followers based on the Stahberg game architecture; wherein the optimization objective function defined for the virtual leader is configured to comprise an evaluation item for the trajectory tracking performance of the virtual leader itself and a formation maintenance cost item related to the formation state of the followers; and by minimizing the objective function, the trajectory optimization of the virtual leader and the overall stability of the formation are achieved; wherein the optimization objective function defined for each of the followers is configured to comprise evaluation items for the trajectory tracking performance of the follower itself and the formation maintenance performance, and an adaptive obstacle avoidance penalty item related to the distance to obstacles generated based on the speed obstacle method; and by minimizing the objective function, the tracking of the leader, the formation maintenance and the safe obstacle avoidance are achieved; and a distributed solving and control module connected to the optimization objective function defining module, the distributed solving and control module comprising local solvers deployed on the virtual leader and each of the followers, each of the local solvers being configured to: adopt a single neural network adaptive dynamic programming algorithm to solve a Hamilton-Jacobi-Bellman equation corresponding to the own optimization objective function online, and generate and execute control instructions in real time according to the solving result.

[0013] In a third aspect, the present application provides a computer readable storage medium, wherein a plurality of program codes are stored, the program codes being adapted to be loaded and run by a processor to execute the UAV formation conflict avoidance control method.

[0014] In a fourth aspect, the present application provides a UAV formation system comprising a virtual leader UAV and at least one follower UAV, wherein: the virtual leader UAV and each of the follower UAVs are each loaded with the UAV formation cooperative control system based on a Stahberg game; and the virtual leader UAV and each of the follower UAVs are connected to each other through a communication network to run the Stahberg game decision architecture in the control system and achieve distributed cooperative control.

[0015] The low-altitude unmanned aerial vehicle formation conflict avoidance control method based on the Stanberger game architecture has the beneficial effects that: for the formation stability control problem of the multi-unmanned aerial vehicle formation, a virtual leader-follower-based Stanberger game architecture is introduced, a virtual leader unmanned aerial vehicle can adaptively adjust a trajectory according to a formation state, so that a follower unmanned aerial vehicle can complete a given tracking task while keeping the stability of the formation; an obstacle avoidance penalty term is constructed based on a speed obstacle method and is embedded into a value function of the follower unmanned aerial vehicle, so as to guarantee that the follower unmanned aerial vehicle realizes collision avoidance in an obstacle environment; a single-network ADP algorithm is adopted to solve a Hamilton-Jacobi-Bellman equation corresponding to the leader and the follower unmanned aerial vehicles respectively, so as to generate a controller in real time Further, the virtual leader unmanned aerial vehicle is regarded as an intelligent agent capable of adaptively adjusting a trajectory, instead of a fixed reference trajectory, and a state equation of the virtual leader unmanned aerial vehicle is: The state equation not only contains an error term relative to an ideal reference trajectory, but also includes a formation error term relative to the follower, and the design form of the state equation gives the virtual leader unmanned aerial vehicle the ability to adaptively adjust the trajectory. BRIEF DESCRIPTION OF DRAWINGS

[0016] The disclosed content of the present application will become more apparent with reference to the drawings. It is easy for those skilled in the art to understand that the drawings are only for the purpose of illustration, and are not intended to constitute a limitation on the scope of protection of the present application. In addition, similar numbers in the drawings are used to represent similar components, wherein: Figure 1 A flowchart of a low-altitude unmanned aerial vehicle formation conflict avoidance control method based on a Stanberger game architecture according to an embodiment of the present application; Figure 2 A three-dimensional trajectory diagram of a virtual leader and a follower unmanned aerial vehicle in a top view according to an embodiment of the present application; Figure 3 A three-dimensional trajectory diagram of a virtual leader and a follower unmanned aerial vehicle in a side view according to an embodiment of the present application; Figure 4 A formation error diagram of a follower relative to a virtual leader according to an embodiment of the present application; Figure 5 A formation error diagram of a follower relative to a virtual leader according to an embodiment of the present application; Figure 6 A three-dimensional trajectory diagram of a virtual leader and an ideal reference trajectory according to an embodiment of the present application; Figure 7 A tracking error diagram of a virtual leader relative to an ideal reference trajectory according to an embodiment of the present application; Figure 8 A weight convergence diagram of a virtual leader according to an embodiment of the present application; Figure 9 follower of one embodiment of the present invention weight convergence graph of Figure 10 follower of one embodiment of the present invention weight convergence graph of DETAILED DESCRIPTION

[0017] Some embodiments of the present application will be described below with reference to the accompanying drawings. Those skilled in the art will understand that these embodiments are only used to explain the technical principles of the present application, and are not intended to limit the protection scope of the present application.

[0018] The application discloses a UAV formation conflict avoidance control method, comprising the following steps: establishing a Stenberger game decision framework between a virtual leader UAV and at least one follower UAV, wherein the virtual leader serves as an upper decision maker, and decision information of the virtual leader is broadcast to all followers; defining a hierarchical optimization objective function for the virtual leader and each follower based on the Stenberger game framework; wherein the optimization objective function of the virtual leader is configured to contain an evaluation item of trajectory tracking performance of the virtual leader itself and a formation maintenance cost item related to the formation state of the follower; trajectory optimization of the virtual leader and overall stability of the formation are realized by minimizing the objective function; the optimization objective function of each follower is configured to contain an evaluation item of trajectory tracking performance of the follower itself and formation maintenance performance, and an adaptive obstacle avoidance penalty item related to the distance from the obstacle generated based on the speed obstacle method; tracking of the leader, formation maintenance and safe obstacle avoidance are realized by minimizing the objective function; a single neural network adaptive dynamic programming algorithm is adopted to solve Hamilton-Jacobi-Bellman equations corresponding to the optimization objective functions of the virtual leader and each follower in parallel online, and a distributed control law is generated and executed in real time according to the solving result.

[0019] In the embodiment, the virtual leader UAV: as an upper decision maker, is not an entity UAV, but a preset formation reference center, and its decision information (such as a reference trajectory and ideal formation configuration parameters) is broadcast to all follower UAVs in real time to dominate the overall movement direction of the formation; the follower UAV: as a lower decision maker, receives the broadcast information of the virtual leader, and perceives the surrounding obstacles and adjacent UAV states, and optimizes the movement strategy under the constraint of the decision of the leader; the virtual leader first optimizes its own decision with the overall stability of the formation as the target, and then the follower optimizes the local strategy with the trajectory tracking and obstacle avoidance safety of the follower as the target based on the decision of the leader, to form a global-local collaborative decision closed loop, and avoid the movement conflict of the UAVs in the formation.

[0020] Different optimization objective functions are designed for different decision-making positions of the virtual leader and followers, and the hierarchical objective functions are defined as follows: The optimization objective function of the virtual leader is configured as a self-trajectory tracking performance evaluation item + a formation keeping cost item, and by minimizing the function, the dual goals of "self-trajectory accuracy" and "overall stability of the formation" are achieved: The self-trajectory tracking performance evaluation item is used to measure the error of the virtual leader in tracking the preset global reference trajectory, for example, the sum of squares of deviations of the ideal reference position and the ideal reference speed of the virtual leader from the corresponding parameters of the preset reference trajectory, and the higher the weight of this item, the higher the requirement for trajectory tracking accuracy.

[0021] The formation keeping cost item is used to measure the relative position deviation of all follower drones from the virtual leader, which is directly related to the ideal configuration of the formation, such as fixed spacing and geometric shape, for example, defined as the sum of squares of deviations of the ideal relative distance and the actual relative distance between each follower and the leader, and the higher the weight of this item, the higher the accuracy of the formation configuration; by minimizing the weighted sum of the above two cost items, the virtual leader will actively adjust its motion state while tracking the global reference trajectory, guiding the followers to maintain the preset formation configuration and avoiding the collapse of the formation.

[0022] The optimization objective function of each follower is configured as a self-trajectory tracking performance evaluation item + a formation keeping performance evaluation item + an adaptive obstacle avoidance penalty item, and by minimizing the function, the cooperative control of "tracking the leader, maintaining the formation, and safely avoiding obstacles" is achieved: The self-trajectory tracking performance evaluation item measures the error of the follower in tracking the broadcast trajectory of the virtual leader, for example, the sum of squares of deviations of the position / velocity of the follower from the corresponding parameters of the ideal trajectory of the leader; The formation keeping performance evaluation item measures the relative position deviation of the follower from the adjacent drones, ensuring that there is no collision between the drones in the formation, for example, defined as the sum of squares of deviations of the ideal spacing and the actual spacing between the follower and the adjacent drones; The adaptive obstacle avoidance penalty item based on the speed barrier method is negatively related to the real-time distance of the follower to the obstacle, the closer the distance, the larger the penalty item value, and the farther the distance, the penalty item tends to 0; the introduction of this item makes the follower actively avoid obstacles when optimizing the motion strategy, avoiding entering the collision risk area; by minimizing the weighted sum of the three cost items, the follower will prioritize obstacle avoidance while tracking the leader and maintaining the formation configuration, achieving conflict-free cooperation between formation movement and obstacle avoidance.

[0023] The HJB equation corresponding to the optimization objective function of the virtual leader and each follower is constructed, the input of the equation is the real-time state of the UAV (such as position, speed, formation relative state, obstacle distance, etc.), and the output is the optimal control cost; a single neural network (such as a feedforward neural network, a convolutional neural network) is used as a value function approximator, the input of the neural network is the real-time state of the UAV, and the output is the approximate solution (i.e. the optimal value function) of the HJB equation; compared with a multi-neural network structure, the single neural network has fewer parameters and higher calculation efficiency, and is suitable for online real-time solving; the virtual leader and each follower independently and in parallel execute the training and solving process of the neural network. Without a centralized computing center, each UAV only needs to update the neural network parameters based on the state information perceived by itself and the broadcast information received, to approximate the optimal solution of the HJB equation; this method greatly reduces the communication and calculation burden of the formation and improves the real-time performance and robustness of the system; according to the solving result of the HJB equation, the optimal control law (such as speed adjustment amount, attitude angle control amount, etc.) of each UAV is derived in real time. The control law of the virtual leader is used to optimize the reference trajectory of itself, and the control law of each follower is used to adjust the motion state of itself; since the control law is independently generated by each UAV, it belongs to a distributed control law, and the decentralized cooperative control of the formation can be realized.

[0024] Of course, the parameters involved in the above steps (such as the weight coefficients of each term in the objective function, the number of layers and the number of neurons of the neural network, the solving iteration step length of the HJB equation, etc.) can be flexibly set according to the scale of the UAV formation, the flight speed, the obstacle density and other actual application scenarios, as long as the core control objectives of formation trajectory tracking, formation keeping and conflict avoidance can be achieved.

[0025] In one possible control method, as shown in Figures 1-10 The present application provides a UAV formation conflict avoidance control method, comprising the following steps: S1, a virtual leader-based UAV formation mathematical model is established.

[0026] For each UAV, its mathematical model can be constructed as a second-order dynamic system. The second-order dynamic model of the virtual leader UAV can be expressed as: , , wherein, is the position state of the leader UAV, is the speed state of the leader UAV, is the control input of the leader UAV.

[0027] Similarly, for each follower UAV, its dynamic model can be expressed as: , , where, is the index of the follower UAV, is the position state of the th follower UAV, is the velocity state of the th follower UAV, is the control input of the th follower UAV.

[0028] S2, establish the state equation of the follower UAV based on the Steinberg game architecture.

[0029] Under the Steinberg game architecture, for each follower UAV, the control objective is: , , , where, is the ideal relative position of the th follower UAV relative to the leader UAV; Define the state equation of the leader UAV and the follower UAV as: , ; Define the formation error of the th follower UAV as: ; Define the augmented vector of the th follower UAV as: ; Define the neighbor tracking error of the th follower UAV and other UAVs as: where, ; The th follower augmented state equation considering formation and tracking performance is defined as: where, , , , , , .

[0030] S3, establish the leader UAV state equation based on the stanberg game architecture. Under the stanberg game architecture, for the leader UAV, its control objective is: , , , Where, and are the ideal reference position and ideal reference velocity of the leader UAV, respectively; The ideal state vector of the leader UAV is defined as: Where, ; The leader augmented dynamic equation considering both tracking performance and formation stability is defined as: Where, , is the dynamic vector related to , , , .

[0031] S4, calculate the obstacle avoidance penalty term .

[0032] To achieve obstacle avoidance control, consider defining the position of the th obstacle as: Where, ; The obstacle avoidance penalty term is: Where, the adaptive safety coefficient function is defined as: Where, is the distance between the th follower UAV and the th obstacle, is the radius of the th obstacle, is the region radius for the th obstacle to take obstacle avoidance maneuver, is a scaling coefficient, and the scheduling function is defined as: Where, is the scaling coefficient.

[0033] S5, define the follower value function based on the Stacberg game and give the optimal control law. For the i-th follower UAV, its corresponding value function based on the Stacberg game is: where, is the decay coefficient, , , , are positive definite diagonal weight matrices, respectively; The optimal value function for the i-th follower UAV is: Its corresponding Hamiltonian function is: Its corresponding Hamilton-Jacobi-Bellman equation is: According to the necessary condition of the optimality principle, the optimal control law for the i-th follower UAV is: S6, define the leader value function based on the Stacberg game and give the optimal control law. Based on the Lagrange multiplier method and considering the constraint condition of the follower value function, consider , as the constraint factor of the leader value function.

[0034] The derivative with respect to time is: The leader value function based on the Stacberg game is defined as: where, , are positive definite diagonal weight matrices, respectively; is the Lagrange multiplier, in order to ensure the stability of the Lagrange term, rewrite the Lagrange term as: where, is a scaling factor, further, the optimal leader value function can be written as: The corresponding Hamiltonian function is: The Lagrange multiplier can be expressed as: ​​​ The optimal control law for the leader is: where, .

[0035] S7, update the neural network weights of the virtual leader-follower UAV formation and compute the control law. For each follower UAV, use the neural network to approximate the optimal value function , define the neural network structure for the follower UAVs as: where, is the weight vector of the th follower UAV, is the activation function vector of the th follower UAV, is the obstacle avoidance penalty term, and: where, , is a very small positive number. The nominal control law for the th follower UAV is given by: where, is the nominal control law for the leader UAV.

[0036] The approximate Bellman equation for the th follower UAV is given by: The Bellman error for the th follower UAV is: where, . Update the weights using gradient descent: where, , is the corresponding learning rate.

[0037] Similarly, for the leader UAV, use the neural network to approximate the optimal value function , define the neural network structure for the leader UAV as: where, is the weight vector of the leader UAV, is the activation function vector of the leader UAV. The nominal control law of the leader UAV is given as: The approximate Bellman equation of the leader UAV is given as: The Bellman error of the leader UAV is defined as: where The weights are updated by gradient descent method as: where , is the corresponding learning rate; thus, the nominal control law calculated according to the weights is finally adopted as the control law of the virtual leader-follower UAV formation.

[0038] In one possible UAV formation collision avoidance control method, the UAV parameters are configured as follows: The initial position and initial velocity of the leader are: , ; the initial position and initial velocity of the follower are: , ; the initial position and initial velocity of the follower are: , ; the ideal reference trajectory of the leader is: ; the ideal relative position of the follower relative to the leader is: ; the ideal relative position of the follower relative to the leader is: ; The environment is configured as follows: The center coordinates of the obstacle are: , and the radius is ; the center coordinates of the obstacle 2 are: , and the radius is .

[0039] The parameter selection is as follows: The value function weights of the leader are: , ; the value function weights of the follower are: , , ; the value function decay factor is defined as: ; the related parameter selection of the other obstacle avoidance penalty term is: , , ; the leader's neural network activation function vector is selected as: The follower's neural network activation function vector is selected as: wherein, represents the i-th element of the leader's state vector , represents the i-th element of the follower's state vector , ; the neural network weight learning rates of the leader and the follower are defined as: , , it is assumed that the UAVs can always detect obstacles during the flight mission. The results of

[0040] Figure 2 and Figure 3 show that this method can enable the follower UAV to complete the tracking of the leader UAV while avoiding obstacles; in addition, Figure 4 and Figure 5 show that the formation error of the follower UAV relative to the leader UAV gradually converges to a small neighborhood near the origin. Figure 6 The results of Figure 7 show that, under the Steinberg game architecture, the leader UAV will make trajectory adjustments while taking into account the formation self-adaptation of the UAV formation; in addition, Figures 7-9 shows that the tracking error of the leader UAV relative to the ideal reference trajectory will gradually converge to a small neighborhood near the origin. The results of

[0041] show that the weights of the neural networks corresponding to the leader UAV and the follower UAV quickly converge to their optimal values. The application discloses a UAV formation cooperative control system based on Steinberg game. The UAV formation cooperative control system based on Steinberg game mainly comprises a region and surface division module, an integral equation construction module, an equation adaptation module, a matrix equation construction module, a matrix solution module and a scattering characteristic calculation module. In some embodiments, one or more of the region and surface division module, the integral equation construction module, the equation adaptation module, the matrix equation construction module, the matrix solution module and the scattering characteristic calculation module can be combined together to form one module. In some embodiments, the region and surface division module can be configured to perform the program of step S1. The integral equation construction module can be configured to perform the program of step S2. The equation adaptation module can be configured to perform the program of step S3. The matrix equation construction module performs the program of step S4. The matrix solution module performs the program of step S5. The scattering characteristic calculation module performs the program of step S6. In one embodiment, the description of the specific implementation function can refer to steps S1-S6.

[0042] The UAV formation cooperative control system based on Steinberg game described above is used to perform the UAV formation conflict avoidance control method embodiment, and the technical principles, the technical problems solved and the technical effects generated are similar. Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process and related description of the UAV formation cooperative control system based on Steinberg game can refer to the content described in the UAV formation conflict avoidance control method embodiment, which will not be repeated here.

[0043] Embodiment three The application further provides a computer readable storage medium. In one computer readable storage medium embodiment according to the application, the computer readable storage medium can be configured to store the program of the UAV formation conflict avoidance control method for performing the above-mentioned method embodiment, which can be loaded and run by the processor to realize the above-mentioned UAV formation conflict avoidance control method. For the convenience of illustration, only the parts related to the embodiments of the application are shown, and the specific technical details that are not disclosed are referred to the method part of the embodiments of the application. The computer readable storage medium can be a storage device formed by various electronic devices, and optionally, the computer readable storage medium in the embodiments of the application is a non-transitory computer readable storage medium.

[0044] Embodiment four The application provides a UAV formation system, comprising a virtual leader UAV and at least one follower UAV, wherein: the virtual leader UAV and each of the follower UAVs are equipped with the UAV formation cooperative control system based on the Steinberg game; the virtual leader UAV and each of the follower UAVs are connected through a communication network to run the Steinberg game decision architecture in the control system and realize distributed cooperative control.

[0045] So far, the technical solutions of the application have been described in combination with the preferred embodiments shown in the drawings, but those skilled in the art can easily understand that the protection scope of the application is obviously not limited to these specific embodiments. Those skilled in the art can make equivalent changes or replacements to the original technical features without departing from the principles of the application, and the technical solutions after the changes or replacements will all fall within the protection scope of the application.

Claims

1. A method for conflict avoidance control in unmanned aerial vehicle (UAV) formations, characterized in that, include: Establish a Steinberg game decision-making architecture between a virtual leader drone and at least one follower drone, wherein the virtual leader acts as the upper-level decision-maker and broadcasts the virtual leader's decision information to all followers; Based on the Steinberg game architecture, hierarchical optimization objective functions are defined for the virtual leader and each follower; The optimization objective function of the virtual leader is configured to include an evaluation term for the virtual leader's own trajectory tracking performance and a formation maintenance cost term related to the formation state of the followers; by minimizing this objective function, its own trajectory optimization and overall formation stability are achieved. The optimization objective function for each follower is configured to include an evaluation term for the follower's own trajectory tracking performance and formation keeping performance, as well as an adaptive obstacle avoidance penalty term related to obstacle distance generated based on the velocity obstacle method; by minimizing this objective function, the tracking of the leader, formation keeping and safe obstacle avoidance are achieved. A single neural network adaptive dynamic programming algorithm is used to solve the Hamilton-Jacobi-Bellman equations corresponding to the optimization objective functions of the virtual leader and each of the followers in parallel online, and to generate and execute distributed control laws in real time based on the solution results.

2. The method according to claim 1, characterized in that, The specific optimization objective function of the virtual leader is as follows: Where γ is the attenuation coefficient. Let u0 be the leader's state vector, and u0 be the leader's control input. , These are positive definite diagonal weight matrices; It is the Lagrange multiplier associated with the i-th follower, V i For the first The value function of each follower; Among them, the leader's state vector The acquisition process includes: under the Steinberg game framework, defining and Let the ideal reference position and ideal reference velocity of the leader drone be respectively; the ideal state direction of the leader drone is defined as: ,in The leader augmented dynamic equation, which balances tracking performance and formation stability, is defined as follows: ,in, , Is with The relevant dynamic vectors, , , This is the system gain matrix.

3. The method according to claim 1, characterized in that, No. The specific objective function for optimizing a follower drone is as follows: in, It is the attenuation coefficient. , , , , The diagonal weight matrix is ​​positive definite. Let u be the augmented state vector of the i-th follower. i u is the control input for the i-th follower. j Let uj be the control input for the j-th follower, and u0 be the control input for the leader. For the first An adaptive obstacle avoidance penalty term for each obstacle; Wherein, the augmented state vector of the i-th follower The acquisition process includes: Under the Steinberg game framework, for each follower drone, the control objective is: , , ,in, For the first The ideal relative positions of the follower drones with respect to the leader drone; the state equations for the leader drone and the follower drones are defined as follows: , The first one that balances formation and tracking performance The augmented state equation for a follower is defined as: ,in, , , , This is the system gain matrix.

4. The method according to claim 3, characterized in that, The adaptive obstacle avoidance penalty item Based on the speed barrier method, the calculation formula is as follows: in, This represents the adaptive safety factor function. Indicates the first The location of the obstacle Indicates the first The position status of a follower drone It is a scheduling function. express hour, The value of , where .

5. The method according to claim 2 or 3, characterized in that, The optimization objective function of the leader is approximated by constructing a single neural network. optimal value and the optimization objective function of the followers optimal value The weights of the neural network are updated online using gradient descent, the direction of which is determined by the corresponding Bellman error.

6. The method according to claim 1, characterized in that, No. The optimal control law for each follower drone is: By solving the optimization objective function of the follower optimal value The corresponding Hamiltonian function is obtained and expressed as: in, The Hamiltonian function of the i-th follower, u i For the control input of the i-th follower, u0 represents the control input for followers other than follower i, and u0 represents the control input for the leader. It is a positive definite diagonal weight matrix. .

7. The method according to claim 5, characterized in that, The optimal control law corresponding to the virtual leader drone By solving the leader's optimization objective function optimal value The corresponding Hamiltonian function is obtained and expressed as: in, , , u0 is the leader's control input, u i , For the control input of the i-th follower, It is a positive definite diagonal weight matrix.

8. A drone formation cooperative control system based on Steinberg game theory, used to control a formation comprising a virtual leader drone and at least one follower drone, characterized in that, The system includes: An architecture building module is used to establish a Steinberg game decision-making architecture between the virtual leader drone and each of the follower drones, wherein the virtual leader acts as the upper-level decision-maker and is configured with a broadcasting unit to broadcast its decision information to all followers; An optimization objective function definition module, connected to the architecture construction module, is used to define hierarchical optimization objective functions for the virtual leader and each follower based on the Steinberg game architecture. The optimization objective function defined for the virtual leader is configured to include an evaluation term for the virtual leader's own trajectory tracking performance and a formation maintenance cost term related to the follower's formation state. By minimizing this objective function, the virtual leader's own trajectory optimization and overall formation stability are achieved. The optimization objective function defined for each follower is configured to include an evaluation term for the follower's own trajectory tracking performance and formation maintenance performance, and an adaptive obstacle avoidance penalty term related to obstacle distance, generated based on the speed obstacle method. By minimizing this objective function, the follower can track the leader, maintain formation, and safely avoid obstacles. A distributed solution and control module is connected to the optimization objective function definition module. The distributed solution and control module includes local solvers deployed on the virtual leader and each of the followers. Each local solver is configured to: use a single neural network adaptive dynamic programming algorithm to solve the Hamilton-Jacobi-Bellman equation corresponding to its own optimization objective function online, and generate and execute control commands in real time based on the solution results.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the UAV formation cooperative control method based on Steinberg game as described in any one of claims 1 to 7.

10. A drone formation system, characterized in that, Includes a virtual leader drone and at least one follower drone, wherein: The virtual leader drone and each of the follower drones are equipped with the drone formation cooperative control system based on Steinberg game as described in claim 8; The virtual leader drone and each of the follower drones are interconnected via a communication network to run the Steinberg game decision-making architecture in the control system and achieve distributed collaborative control.

Citation Information

Cited By

  • Distributed fault-tolerant optimal cooperative control method for multi-agent system

    CN122085710A

  • Multi-uav airspace conflict resolution and scheduling method based on counterfactual causal reasoning

    CN122135605A

  • Multi-uav airspace conflict resolution and scheduling method based on counterfactual causal reasoning

    CN122135605B