Unmanned aerial vehicle flight control method and device based on multi-alliance game fully distributed Nash equilibrium search, medium and product
By constructing a UAV dynamics model and a distributed observer, a fully distributed Nash equilibrium search for UAVs in a multi-alliance game system was achieved, solving the problem of difficulty in obtaining global information under non-cooperative alliances and improving the evolution accuracy of the system in a highly adversarial environment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIHANG UNIV
- Filing Date
- 2026-02-03
- Publication Date
- 2026-05-12
AI Technical Summary
Existing technologies cannot achieve fully distributed Nash equilibrium search in multi-alliance games, especially in non-cooperative alliances and scenarios where global information is difficult to obtain. They cannot effectively deploy decision-making algorithms, resulting in poor system performance in highly adversarial environments.
By constructing a UAV dynamics model, real-time observation status data and communication topology of the UAV are obtained. A fully distributed Nash equilibrium search algorithm based on multi-alliance game theory is adopted, and distributed observers are used for iteration until the UAV state reaches stability, thus achieving Nash equilibrium.
In scenarios where non-cooperative alliances and global information are difficult to obtain, the system improves the evolution accuracy of multi-alliance game systems, avoids dependence on global information, and enhances the system's adaptability in highly adversarial environments.
Smart Images

Figure CN122018561A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of Nash equilibrium search in multi-alliance games, and in particular to a method, device, medium and product for unmanned aerial vehicle (UAV) flight control based on fully distributed Nash equilibrium search in multi-alliance games. Background Technology
[0002] Given the high efficiency, flexibility, and difficulty in defending against swarm warfare, countermeasures against swarms have become a crucial factor in ensuring regional security. However, countermeasures based on laser and microwave weapons are dependent on the usage scenario and have limitations in terms of accuracy. Therefore, utilizing swarms to counter swarms constitutes an emerging combat method for improving swarm defense systems, and corresponding swarm countermeasure technologies have become a current research hotspot.
[0003] In swarm adversarial scenarios, different swarms attempt to evolve to their respective optimal adversarial positions, a process highly consistent with the theoretical framework of multi-alliance games. That is, all individuals are divided into multiple alliances, with competition between alliances and cooperation among individuals within an alliance to minimize the overall cost of their alliance. Currently, most research on Nash equilibrium search problems for multi-alliance games assumes that individual state information in the gradient term is known, or obtains this information through distributed estimation. However, such research only applies to cases where alliances are cooperative. Furthermore, current Nash equilibrium search algorithms for multi-alliance games all contain global information such as the number of individuals and the eigenvalues of the Laplace matrix of the communication topology, making them unsuitable for fully distributed deployment. On the one hand, in highly adversarial battlefield environments, opposing alliances cannot directly establish communication, and individuals cannot obtain the state or decision information of all individuals through distributed estimation. On the other hand, to improve the adaptability of swarm systems on the battlefield, the introduction of central nodes should be avoided as much as possible, and decision-making algorithms should be deployed in a fully distributed manner. Therefore, researching fully distributed Nash equilibrium search methods in multi-alliance game problems is one of the urgent challenges to be solved in the current engineering field.
[0004] In recent years, Nash equilibrium search algorithms for multi-alliance games have attracted the attention of many experts and scholars. However, to date, the relevant research results are only applicable to the case of cooperative alliances. Few people have ventured into the acquisition of state information and algorithm design under non-cooperative alliances.
[0005] The Nash equilibrium search methods for multi-alliance games involved in related research can be divided into three main categories. The first category treats the state information of all individuals as known information, introducing global information into the controller protocol. However, in practical applications, global information is often difficult for a single individual to obtain. In the second category, one individual is selected as the central node within each alliance, with the remaining individuals acting as subordinate nodes of that central node. Furthermore, communication is established between central nodes of different alliances, and inter-alliance interactions are represented by interactions between central nodes. Therefore, this type of method has semi-centralized characteristics, and the overall evolution of the alliance depends on its central node. When a central node is disturbed or damaged during the task, its alliance will be significantly affected. To estimate the state information of all individuals, the third category employs a distributed consensus estimation strategy. Each individual not only transmits its own state information to its neighbors but also its state estimates of other individuals. Since each individual's state estimates of other individuals can only be transmitted through communication links, and communication links can only exist between cooperative alliances, the third category is suitable for multi-alliance game problems composed of cooperative alliances. Furthermore, the feedback gain of the three types of methods mentioned above includes global information such as the Lipschitz constant, the number of individuals, and the eigenvalues of the Laplace matrix of the communication topology. These methods cannot be deployed in a fully distributed manner. In application scenarios where relevant global information is difficult to obtain, these methods cannot function effectively. Summary of the Invention
[0006] The purpose of this application is to provide a method, device, medium, and product for UAV flight control based on fully distributed Nash equilibrium search in multi-alliance game theory, which can be applied to multi-alliance game systems containing non-cooperative alliances and scenarios where relevant global information is difficult to obtain.
[0007] To achieve the above objectives, this application provides the following solution: Firstly, this application provides a UAV flight control method based on fully distributed Nash equilibrium search in a multi-alliance game, applicable to each UAV in a multi-alliance game system; the multi-alliance game system includes multiple alliances; each alliance includes multiple UAVs; the method includes: The current drone observes the state data of other drones in the multi-alliance game system, excluding the current drone, and uses this data as observation state data. Construct a dynamic model of the unmanned aerial vehicle; Obtain the communication topology of the alliance to which the current drone belongs; A fully distributed Nash equilibrium search algorithm for multi-alliance game theory is adopted. Based on the observed state data, the communication topology graph, and the UAV dynamics model, the state data of the current UAV during flight is iterated until the state data of the current UAV reaches the state stability requirement, at which point the multi-alliance game system reaches Nash equilibrium. The fully distributed Nash equilibrium search algorithm for multi-alliance game theory is based on the construction of a distributed observer. The drone is controlled to fly according to the control inputs corresponding to the state data that meet the state stability requirements.
[0008] In a second aspect, this application provides a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the UAV flight control method based on fully distributed Nash equilibrium search of multi-alliance game as described above.
[0009] Thirdly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the UAV flight control method based on fully distributed Nash equilibrium search of multi-alliance game as described above.
[0010] Fourthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the UAV flight control method based on fully distributed Nash equilibrium search of multi-alliance game as described above.
[0011] According to the specific embodiments provided in this application, this application has the following technical effects: This application provides a UAV flight control method, device, medium, and product based on a fully distributed Nash equilibrium search in a multi-alliance game. By constructing a UAV dynamics model, it obtains the observed state data of other UAVs in the multi-alliance game system and the communication topology of the current UAV's alliance. For cooperative and non-cooperative alliances in the multi-alliance game system, it utilizes the observed state data and the communication between individuals within the current UAV's alliance (i.e., the communication topology) to enable the current UAV to acquire the state data of all other UAVs. A fully distributed Nash equilibrium search algorithm for multi-alliance games is employed. Based on the observed state data, communication topology, and UAV dynamics model, the state data of the current UAV during flight is iterated. Since the iteration process uses observed state data of other UAVs, when the current UAV's state data reaches the state stability requirement, it can be determined that the state data of other UAVs in the multi-alliance game system also reach the state stability requirement. Therefore, when the current UAV's state data reaches the state stability requirement, the multi-alliance game system reaches a Nash equilibrium. This application designs a Nash equilibrium search method based on a distributed observer in a fully distributed manner, avoiding the introduction of global information such as the eigenvalues of the communication topology Laplace matrix, the total number of individuals in the system, the strong convexity coefficients of the cost function, and the Lipschitz constant. This enables the fully distributed Nash equilibrium search algorithm for multi-alliance games to be applied to multi-alliance game scenarios where global information is difficult to obtain, thereby improving the accuracy of the evolution of multi-alliance game systems to Nash equilibrium. Attached Figure Description
[0012] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 This is a flowchart of a UAV flight control method based on fully distributed Nash equilibrium search in a multi-alliance game according to an embodiment of this application; Figure 2 A communication topology diagram of a multi-alliance game system provided in an embodiment of this application; Figure 3 This is a schematic diagram of the initial spatial distribution of an individual's position provided in an embodiment of this application; Figure 4 This is a schematic diagram of the spatial distribution evolution of individuals provided in an embodiment of this application; Figure 5 A schematic diagram of the change curve of the input signal of an individual provided in an embodiment of this application; Figure 6This is a schematic diagram of the variation curve of partial derivative estimation error provided in an embodiment of this application; Figure 7 A schematic diagram of the Nash equilibrium search error variation curve provided in an embodiment of this application; Figure 8 This is a schematic diagram of the state observation error variation curve provided in an embodiment of this application; Figure 9 Adaptive parameters provided for an embodiment of this application A schematic diagram of the evolution curve; Figure 10 Adaptive parameters provided for an embodiment of this application A schematic diagram of the evolution curve; Figure 11 Adaptive parameters provided for an embodiment of this application A schematic diagram of the evolution curve; Figure 12 Adaptive parameters provided for an embodiment of this application A schematic diagram of the evolution curve; Figure 13 Adaptive parameters provided for an embodiment of this application A schematic diagram of the evolution curve; Figure 14 Adaptive parameters provided for an embodiment of this application A schematic diagram of the evolution curve; Figure 15 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation
[0014] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0015] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0016] In one exemplary embodiment, such as Figure 1 As shown, a UAV flight control method based on fully distributed Nash equilibrium search in a multi-alliance game is provided, applicable to each UAV in a multi-alliance game system. The multi-alliance game system comprises multiple alliances. Each alliance includes multiple UAVs. The method includes: Step 100: In real time, acquire the state data of other drones in the multi-alliance game system observed by the current drone, excluding the current drone, as the observation state data.
[0017] Step 200: Construct the UAV dynamics model.
[0018] Step 300: Obtain the communication topology of the alliance to which the current drone belongs.
[0019] Step 400 employs a fully distributed Nash equilibrium search algorithm for multi-alliance game theory. Based on observed state data, communication topology graph, and UAV dynamics model, the algorithm iteratively processes the state data of the current UAV during flight until the current UAV state data meets the state stability requirements, at which point the multi-alliance game system reaches Nash equilibrium. The fully distributed Nash equilibrium search algorithm for multi-alliance game theory is based on a distributed observer architecture.
[0020] Step 500: Control the UAV to fly according to the control input corresponding to the state data that meets the state stability requirements.
[0021] The communication topology graph is a connected undirected graph.
[0022] In one embodiment, a by A swarm system consisting of drones is divided into The system is configured as a multi-alliance game system, with multiple alliances. ,make ,in, For the alliance The total number of individuals in the group. In swarm adversarial games involving non-cooperative alliances (i.e., multi-alliance games), actuator failure is a common problem; therefore, the UAV dynamics model can be expressed as: .
[0023] In the formula, and express Time Alliance The Middle The position and speed of the drone express Time Alliance The Middle The actuator saturation function of a drone. for Time Alliance The Middle Ideal control input for a drone represent Time Alliance The Middle Actuator deviation fault in a drone Indicates alliance The Middle The actuator efficiency coefficient of a drone. Indicates alliance The number of drones in China.
[0024] Using hyperbolic tangent function This indicates that its physical limit is . The upper bound is . satisfy , for The lower bound. When At this time, the actuator does not experience an efficiency fault. Since an individual can only be properly controlled when the actuator's maximum output exceeds the upper limit of the deviation fault, it is necessary to... .
[0025] As an optional implementation, step 400 employs a multi-alliance game fully distributed Nash equilibrium search algorithm, which iterates the state data of the current UAV flight process based on observed state data, communication topology graph, and UAV dynamics model, including: Step 410: Acquire the current status data of the drone in real time. The status data includes position and speed.
[0026] Step 420: Initialize the current UAV state parameters. State parameters include: distributed observer parameters, partial derivative parameters, actuator bias fault compensation term parameters, and gradient term parameters.
[0027] Step 430: Determine the topology matrix based on the communication topology diagram.
[0028] Step 440: The fully distributed Nash equilibrium search algorithm of multi-alliance game is adopted to iterate the current state data of the UAV based on the current state parameters of the UAV, the observed state data, the topology matrix and the UAV dynamics model.
[0029] The process of iterating the current UAV state data using the multi-alliance game fully distributed Nash equilibrium search algorithm in step 440 includes: Step 441: Using a distributed observer, determine the state observation values based on the topology matrix, observation state data, and distributed observer parameters.
[0030] Step 442: Determine the observation location vector based on the state observation values.
[0031] Step 443: Determine the first auxiliary variable based on the observation position vector, topological matrix, and partial derivative parameters.
[0032] Step 444: Determine the actuator deviation fault compensation term parameters and gradient term parameters based on the first auxiliary variable and the current speed of the UAV.
[0033] Step 445: Determine new state data based on the current speed of the UAV, the first auxiliary variable, the actuator deviation fault compensation term parameters, the gradient term parameters, and the UAV dynamics model.
[0034] Step 446: Determine whether the state stability requirement has been met based on the state data obtained from the previous two iterations of the current UAV.
[0035] Step 447: If the state stability requirement is met, stop the iteration. If the state stability requirement is not met, determine new distributed observer parameters based on the state observations, determine new partial derivative parameters based on the first auxiliary variable, and return to step 441.
[0036] Specifically, the state stability requirement is met when the difference between the state data obtained from two consecutive iterations of the UAV is less than a set threshold.
[0037] In one implementation, in a multi-alliance game involving non-cooperative relationships, all individuals (i.e. drones) cannot directly obtain the state data of the non-cooperative alliances through communication. To address this problem, a distributed observer is designed. By utilizing the local observation behavior of different individuals on individuals in the non-cooperative alliances and the communication interaction between individuals within the alliances, each individual can obtain the state data of all other individuals in the multi-alliance game system.
[0038] Step 441 includes: Using formula Determine the derivative of the state observation. Integrate the derivative to obtain the state observation.
[0039] in, Indicates alliance The Middle The derivative of the state observations of a UAV Indicates by according to Zhang Cheng's diagonal formation, in which And the diagonal elements from top left to bottom right are respectively ; For the first The first in the alliance In a drone and multi-alliance game system, the first Adaptive parameters for inter-group observations among individual UAVs Indicates by according to Zhang Cheng's diagonal formation, Indicates alliance The Middle The first drone observation The speed of the drone ; Indicates alliance The Middle The state observation values observed by the UAV are used to determine the state observation data. This represents a column vector consisting of the positions and velocity states of all drones in a multi-alliance game system. Indicates the first The first in the alliance The drone and the first Adaptive parameters for intra-group interaction items between drones; , indicating that by the first A matrix of drones according to Zhang Cheng's diagonal block matrix, and ; , indicating that by the first Input matrix of a drone according to Zhang Cheng's diagonal block matrix, and ; Indicates alliance In the topological matrix of the first Line number Column elements, Represents a second-order identity matrix. Indicates the Kronecker product. These are standard symbolic functions.
[0040] In one implementation, the distributed observer parameters include between-group observation adaptive parameters and within-group interaction adaptive parameters. Step 447, which determines the new distributed observer parameters based on the state observations, includes: Using formula Determine the derivative of the new between-group observation adaptive parameter. Use the formula... Determine the derivatives of the new within-group interaction adaptive parameters. Integrate the derivatives to obtain the new between-group observation adaptive parameters and the new within-group interaction adaptive parameters, respectively.
[0041] in, Indicates the first The first in the alliance In a drone and multi-alliance game system, the first The derivative of the adaptive parameters of the inter-group observations among individual UAVs. Indicates alliance The Middle The first drone observation The speed of the drone Indicates the first The first in the alliance The drone is the first The state estimation vector of each UAV Indicates the first The state vector of each drone Indicates the first The input matrix of a drone, For modulo operator, It is a norm of 1; , To set parameters.
[0042] Indicates the first The first in the alliance The drone and the first The derivative of the adaptive parameters of the intra-group interaction terms among the drones. Indicates alliance In the topological matrix of the first Line number Column elements; Indicates alliance The Middle The state observation values observed by the UAV are determined based on the state observation data; , indicating that by the first Input matrix of a drone according to Zhang Cheng's diagonal block matrix.
[0043] In another exemplary embodiment of this application, a simulation is performed on the process of implementing drone flight control using a fully distributed Nash equilibrium search algorithm for multi-alliance games in a set multi-alliance game scenario involving non-cooperative alliances. The principle of the fully distributed Nash equilibrium search algorithm for multi-alliance games is mainly based on the theory of multi-alliance non-cooperative games, enabling each drone to acquire the observation state data of all other drones (including drones in cooperative and non-cooperative alliances) and determine its own state observation value through a distributed observer. Here, "fully distributed" means that the drone flight control method based on fully distributed Nash equilibrium search for multi-alliance games provided in this application can be deployed on every drone in the multi-alliance game system. Then, during the multi-alliance game process, the method is executed by the control system of each drone, causing the multi-alliance game system to reach Nash equilibrium, thereby achieving a swarm adversarial task.
[0044] Alliances in a multi-alliance game system The Taking a drone as an example, the process of implementing drone flight control using a multi-alliance game fully distributed Nash equilibrium search algorithm is as follows: 1. Set up the simulation scenario. Construct a scenario that includes an alliance. The state vector of the position of all individuals (in this embodiment, an individual refers to a drone) ,make , Indicates alliance The Middle The location of each individual. Remember the alliance. The Middle The cost function for each individual is This value can only be obtained by the first Individuals can obtain this information. Based on the communication topology within the alliance, the alliance's... The cost function is Let the feasible state space of the game be denoted as... and make Nash equilibrium A stacked vector The elements therein satisfy the formula: During simulation, the scenario must meet the following four conditions: 1) Regarding the alliance Its internal communication topology diagram This is for a connected undirected graph. Based on this, the alliance... The Laplace matrix corresponding to the communication topology diagram It is a symmetric matrix and ,in for The second smallest characteristic root.
[0045] 2) For , , Relative to time It is twice continuously differentiable. And its first derivative is... and second derivative Bounded. This condition is often used to handle average tracking problems.
[0046] 3) and It is a global Lipschitz, in which Indicates alliance The Middle The cost function of the i-th individual in the alliance Partial derivatives for each individual position, Let represent the derivative of the partial derivative with respect to time, and for any , There are positive numbers Make the formula This condition is valid. It is used to guarantee the uniqueness of the Nash equilibrium. Among them, Indicates by according to Zhang Cheng's column vector.
[0047] 4) Distributed observation mechanism satisfies ,in, For a set that includes all individuals in a multi-alliance game system, For including alliances The set of all individuals in the set. Indicated by the alliance The Middle The set of individuals observed by an individual. express exist The complement of the set, and This condition clarifies the relationship between observation behavior and state data acquisition. Under this condition, the state data of all individuals (i.e., drones) will be acquired by all drones in the alliance in the form of distributed observation, thus achieving distributed observation.
[0048] 2. Initialize the status parameters.
[0049] Set parameters , , , , , Set it to a positive number, let , ,make , , Configure according to the initial state of the simulation scenario. and Given an actuator deviation fault signal and actuator efficiency coefficient .
[0050] 3. For alliances in a multi-alliance game system as well as ,remember For the alliance The Middle Individual observations Location information of each individual. Construct vectors. , , respectively representing the alliance Let be the observation vector of all individuals in a multi-alliance game system regarding the state data of all individuals, and the observation vector of each individual in a multi-alliance game system regarding the state information of all individuals. For the alliance The Middle The velocity information of all other individuals in a multi-alliance game system observed by an individual. For the alliance The Nash equilibrium error. For In the league No. The Nash equilibrium error of an individual (equivalent to the difference in state data obtained from two consecutive iterations) is: .
[0051] Let vector , indicating the alliance The Middle The state observation vector of an individual (i.e., the state observation value, which is updated in real time based on the actual observation of the individual); , indicating the alliance The Middle The first individual observation The state observation vector of each individual ( ); , indicating the alliance The state observation vector; , represents the state observation vector of a multi-alliance game system; , representing the position and velocity vectors of all individuals; , representing the first in a multi-alliance game system The state vector of each individual.
[0052] For any alliance in a multi-alliance game system any individual (In a multi-alliance game system, any drone is identified as the current drone, and its Nash equilibrium evolution is performed until the individual...) (assuming a stable state), the designed fully distributed Nash equilibrium search algorithm for multi-alliance game is as follows: 1) Update the state observations.
[0053] To achieve distributed observation, the following is introduced Indicates the first The first in the alliance The differential equation for the distributed observer, which estimates the total state of each individual over all individuals, is as follows: .
[0054] in, This indicates the consistency between the state observed by the observer and the actual state of the observed individual. This indicates the consistency of observations of the same individual among different individuals within the alliance.
[0055] Solve the above equations in descending order. , and Then update the state observations. Adaptive parameters and (That is, the parameters of the distributed observer, which are used to update the state observations in the next iteration). , and These represent the state observation values respectively. and distributed observer parameters and The first derivative.
[0056] 2) Update the partial derivative estimate (i.e., the first auxiliary variable).
[0057] The differential equation of the designed partial derivative estimation system is: .
[0058] Based on the state observations obtained in 1) above ,from Separate the position vector , and then calculate , obtain the first auxiliary variable , Indicates alliance The Middle The estimated number of individuals The cost function of the i-th alliance is related to the i-th alliance. Estimates of the partial derivatives for each individual. Based on Determine the second auxiliary variable , express The derivative of. And according to Solve ,based on Sure ,according to renew (Used to determine the first auxiliary variable in the next iteration). And solve. and Integrate it and then update the adaptive parameters. and (For use in the next iteration for parameters) (Update). Among them, For partial derivative parameters, , These are the adaptive parameters of the sign function term and the adaptive parameters of the non-sign function term for the auxiliary system (i.e., the partial derivative estimation system), respectively. , To set parameters.
[0059] 3) Update the actuator deviation fault compensation item parameters. (i.e., the adaptive parameter of the actuator deviation fault compensation item).
[0060] according to Solve and further updates . To set parameters.
[0061] 4) Update gradient term parameters (i.e., adaptive parameters of the gradient term).
[0062] according to Solve and further updates . To set parameters.
[0063] 5) Solve for the control input.
[0064] The actual control input of the algorithm is: .in, This is a saturated input term used to compensate for individual system bias faults. Solve... Substitute it into the UAV dynamics model It updates its own state data under the guidance of control input. and .
[0065] 6) State stability determination and iterative solution.
[0066] Determine the state data obtained in the current iteration (as in step 5 above) to update your own state data. and Check if the state stability requirement is met. If the state stability requirement is met, stop iterating. If the state stability requirement is not met, repeat steps 1) to 5) until the state stability requirement is met.
[0067] In the aforementioned fully distributed Nash equilibrium search algorithm for multi-alliance games, all state parameters are designed in a distributed manner, thus avoiding global information.
[0068] In multi-alliance games involving non-cooperative relationships, individuals cannot directly obtain the state information of non-cooperative alliances through communication. This application addresses this problem by designing a distributed observer, utilizing the local observation behavior of different individuals on individuals within non-cooperative alliances and the communication interactions between individuals within the alliance, enabling each individual to obtain the state information of all other individuals. This application designs a Nash equilibrium search algorithm in a fully distributed manner, avoiding the introduction of global information such as the eigenvalues of the communication topology Laplace matrix, the total number of system individuals, the strong convexity coefficients of the cost function, and the Lipschitz constant during the estimation of partial derivatives of the alliance cost function, the distributed observer, and the solution of the actual control input. This allows the algorithm to be applied to cluster adversarial scenarios where global information is difficult to obtain. Considering the strong adversarial and interference characteristics of non-cooperative multi-alliance games, the designed algorithm can be applied to system models that include actuator deviation faults and actuator efficiency coefficients.
[0069] Furthermore, this application can be applied to fields such as multi-robot interaction, power distribution, and traffic networking. In multi-alliance game systems, drones can be replaced with robots, power equipment, or traffic participation devices, etc., while the method and process remain unchanged.
[0070] In one exemplary embodiment, the UAV flight control method based on fully distributed Nash equilibrium search of multi-alliance game theory described in the above embodiments is used to conduct a simulation experiment on a cluster system (i.e., a multi-alliance game system).
[0071] The simulation conditions are configured as follows: In a containing In a cluster system of individuals (individuals can be drones or robots, etc.), all individuals are divided into... There are several alliances, where the topology within each alliance is represented by solid lines. In each node, the first larger number represents the alliance number, and the second smaller number represents the node's index within the alliance. Observational relationships between alliances are represented by dashed lines, with the arrows indicating the direction of observation information transmission. In this communication topology, the state information of all individuals can be observed by any alliance. The corresponding inter-individual interaction relationships (i.e., the communication topology graph) are shown below. Figure 2 As shown.
[0072] In multi-alliance game systems (i.e., cluster systems), , The number of individuals in League 1 The number of individuals in League 2 The number of individuals in League 1 Considering the target withdrawal and attack / defense problem in a two-dimensional scenario, assume that targets are distributed throughout the operational area. One target evacuation point, Alliance Individuals within the alliance must reach the evacuation point and assist the target in evacuation. Because individuals within the alliance need to cooperate with each other, the alliance... There is also a maximum distance constraint between the individuals. (Alliance) As an offensive alliance plan from the alliance Simultaneous attacks from all sides, thus impacting the alliance. To achieve a comprehensive saturation attack on individuals, it is necessary to first arrive at the Alliance. The surrounding area. Alliance As a defensive league and league For cooperative relationship, with alliance As adversaries, they need to be in an alliance. The surrounding area forms a defensive network, that is, according to the alliance and Alliance The distribution reached its optimal defensive position. Furthermore, the alliance... There is also a maximum distance constraint between individuals. For the above target evacuation and attack / defense scenarios, let the evacuation point coordinates be respectively... , , The system coordinate vector is The top right corner and Each represents an individual along shaft and The coordinate components of the axes.
[0073] alliance The cost functions for individuals are respectively , , , .
[0074] alliance The cost functions for individuals are respectively , , .
[0075] alliance The cost functions for individuals are respectively , , , .
[0076] In summary, the alliance , , The cost functions are respectively , , .for , The partial derivatives of the coalition cost function satisfy the following condition. The approximate Nash equilibrium value of the system can be obtained directly through numerical calculation. .
[0077] Let the physical amplitude of the actuator be... The actuator deviation fault vector for all individuals in the system is .
[0078] alliance The actuator efficiency coefficients of the individual are respectively , , , ,alliance The actuator efficiency coefficients of the individual are respectively , , ,alliance The actuator efficiency coefficients of the individual are respectively , , , .
[0079] Given the aforementioned actuator saturation constraints and actuator deviation faults, the initial states of individuals within Alliance 1 before the simulation begins are set as follows: , , , The initial states of individuals in Alliance 2 are as follows: , , The initial states of individuals in Alliance 3 are as follows: , , , At this time, the spatial distribution of each alliance is as follows: Figure 3 As shown.
[0080] Using the routine of the UAV flight control method based on fully distributed Nash equilibrium search in multi-alliance game provided in the above embodiments, simulation results are obtained. The spatial distribution of individuals after simulation is as follows: Figure 4 As shown. In this process, all individuals evolve from the initial state to a Nash equilibrium state, and the alliance... All individuals arrived near the evacuation point and, under the influence of distance constraints, converged to... Inside the evacuation point. Alliance Individuals according to the alliance The distribution of individuals has reached the outermost region, preparing for a comprehensive strike. (Alliance) Individuals, based on the spatial distribution of the other two alliances, are in the alliance Peripheral and Alliance A defensive position is formed on the inner side. The curve of the individual's input signal change is as follows: Figure 5 As shown. Figure 5 In this context, the input signals of all individuals satisfy the actuator input saturation constraint. The variation of the partial derivative estimation error is as follows: Figure 6 As shown. Figure 6 In the middle, the estimation error of the partial derivative of the coalition cost function converges, indicating that the constructed auxiliary system for estimating the partial derivative of the coalition cost function is effective. The Nash equilibrium search error is as follows: Figure 7 As shown. Figure 7 In the above, the Nash equilibrium search error converges, indicating that all individuals in the multi-alliance game system asymptotically converge to Nash equilibrium under the designed fully distributed multi-alliance game Nash equilibrium search algorithm. The evolution of the velocity error (i.e., state observation error) in the state data is as follows: Figure 8 As shown. Figure 8 In the test, the distributed observation error converged, verifying the effectiveness of the designed distributed observer.
[0081] State parameters (i.e., adaptive parameters) , , , , as well as The change curves are as follows: Figures 9-14 As shown, all state parameters converge to finite values, thus verifying their uniform eventual boundedness.
[0082] Based on the simulation results above, under the simultaneous presence of actuator efficiency faults and bias faults, the actuator input signal satisfies the saturation constraint, and the distributed observation error, the partial derivative estimation error of the coalition cost function, and the Nash equilibrium search error converge to... The designed adaptive parameters converge to constant values. Therefore, the effectiveness of the designed distributed observer and coalition cost function partial derivative estimation system is verified. The fully distributed Nash equilibrium search algorithm for multi-coalition games can lead the system to Nash equilibrium even with input saturation constraints and actuator failures. Furthermore, the method provided in this application does not contain global information and is applicable to situations where direct communication between coalitions is not possible. It can be applied to complex game scenarios where global information is difficult to obtain directly or includes non-cooperative coalitions, avoiding the process of solving for feedback gains.
[0083] In one exemplary embodiment, a computer device is provided, which may be a server or a terminal, and its internal structure diagram may be as follows. Figure 15As shown, the computer device includes a processor, memory, input / output interfaces (I / O), and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs stored in the non-volatile storage media. The database stores data related to a UAV flight control method based on a fully distributed Nash equilibrium search in a multi-alliance game. The I / O interfaces are used for information exchange between the processor and external devices. The communication interface is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements a UAV flight control method based on a fully distributed Nash equilibrium search in a multi-alliance game.
[0084] Those skilled in the art will understand that Figure 15 The structures shown are merely block diagrams of some structures related to the present application and do not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than shown in the figures, or combine certain components, or have different component arrangements. In an exemplary embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.
[0085] In one exemplary embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.
[0086] In one exemplary embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.
[0087] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0088] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).
[0089] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, etc., and are not limited to these.
[0090] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0091] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A UAV flight control method based on fully distributed Nash equilibrium search in multi-alliance game theory, characterized in that, Each drone applied in a multi-alliance game system; The multi-alliance game system includes multiple alliances; each alliance includes multiple drones; the method includes: The current drone observes the state data of other drones in the multi-alliance game system, excluding the current drone, in real time, and uses this data as the observation state data. Construct a dynamic model of the unmanned aerial vehicle; Obtain the communication topology of the alliance to which the current drone belongs; A fully distributed Nash equilibrium search algorithm for multi-alliance game theory is adopted. Based on the observed state data, the communication topology graph, and the UAV dynamics model, the state data of the current UAV during flight is iterated until the state data of the current UAV reaches the state stability requirement, at which point the multi-alliance game system reaches Nash equilibrium. The fully distributed Nash equilibrium search algorithm for multi-alliance game theory is based on the construction of a distributed observer. The drone is controlled to fly according to the control inputs corresponding to the state data that meet the state stability requirements.
2. The UAV flight control method based on fully distributed Nash equilibrium search in multi-alliance game theory according to claim 1, characterized in that, A fully distributed Nash equilibrium search algorithm based on multi-alliance game theory is employed to iterate the state data of the current UAV flight process, based on the observed state data, the communication topology graph, and the UAV dynamics model. This includes: Real-time acquisition of the current drone status data; the status data includes position and speed; Initialize the current state parameters of the UAV; the state parameters include: distributed observer parameters, partial derivative parameters, actuator deviation fault compensation term parameters, and gradient term parameters; Determine the topology matrix based on the communication topology diagram; A fully distributed Nash equilibrium search algorithm based on multi-alliance game theory is used to iterate the state data of the current UAV based on the current UAV state parameters, the observed state data, the topology matrix, and the UAV dynamics model. The process of iterating the current UAV state data using a multi-alliance game fully distributed Nash equilibrium search algorithm includes: A distributed observer is used to determine the state observation value based on the topology matrix, the observation state data, and the parameters of the distributed observer. Determine the observation location vector based on the state observation values; The first auxiliary variable is determined based on the observed position vector, the topological matrix, and the partial derivative parameters. The parameters of the actuator deviation fault compensation term and the gradient term are determined based on the first auxiliary variable and the current speed of the UAV. New state data are determined based on the current speed of the UAV, the first auxiliary variable, the actuator deviation fault compensation term parameters, the gradient term parameters, and the UAV dynamics model. Determine whether the state stability requirement has been met based on the state data obtained from the previous two iterations of the current drone; If the state stability requirement is met, then stop iterating; If the state stability requirement is not met, new distributed observer parameters are determined based on the state observations, and new partial derivative parameters are determined based on the first auxiliary variable. Then, the process returns to the step of using a distributed observer to determine the state observations based on the topology matrix, the observed state data, and the distributed observer parameters.
3. The UAV flight control method based on fully distributed Nash equilibrium search in multi-alliance game theory according to claim 1, characterized in that, The UAV dynamics model is represented as follows: ; In the formula, and express Time Alliance The Middle The position and speed of the drone express Time Alliance The Middle The actuator saturation function of a drone. for Time Alliance The Middle Ideal control input for a drone represent Time Alliance The Middle Actuator deviation fault in a drone Indicates alliance The Middle The actuator efficiency coefficient of a drone. Indicates alliance The number of drones in China.
4. The UAV flight control method based on fully distributed Nash equilibrium search in multi-alliance game theory according to claim 2, characterized in that, Using a distributed observer, the state observation values are determined based on the topology matrix, the observed state data, and the distributed observer parameters, including: Using formula Determine the derivative of the state observation; Integrating the derivative yields the state observations; in, Indicates alliance The Middle The derivative of the state observations of a UAV Indicates by according to Zhang Cheng's diagonal formation, in which And the diagonal elements from top left to bottom right are respectively ; For the first The first in the alliance In a drone and multi-alliance game system, the first Adaptive parameters for inter-group observations among UAVs. Indicates by according to Zhang Cheng's diagonal formation, Indicates alliance The Middle The first drone observation The speed of the drone ; Indicates alliance The Middle The state observation values observed by the UAV are used to determine the state observation data. This represents a column vector consisting of the positions and velocity states of all drones in a multi-alliance game system. Indicates the first The first in the alliance The drone and the first Adaptive parameters for intra-group interaction items between drones; , indicating that by the first A matrix of drones according to Zhang Cheng's diagonal block matrix, and ; , indicating that by the first Input matrix of a drone according to Zhang Cheng's diagonal block matrix, and ; Indicates alliance In the topological matrix of the first Line number Column elements, Represents a second-order identity matrix. Indicates the Kronecker product. These are standard symbolic functions.
5. The UAV flight control method based on fully distributed Nash equilibrium search in multi-alliance game theory according to claim 2, characterized in that, Distributed observer parameters include between-group observation adaptive parameters and within-group interaction adaptive parameters; New distributed observer parameters are determined based on state observations, including: Using formula Determine the derivative of the new between-group observation adaptive parameters; Using formula Determine the derivative of the adaptive parameters for the new intragroup interaction terms; Integrating the derivatives respectively yields new adaptive parameters for the between-group observations and new adaptive parameters for the within-group interaction. in, Indicates the first The first in the alliance In a drone and multi-alliance game system, the first The derivative of the adaptive parameters of the inter-group observations among individual UAVs. Indicates alliance The Middle The first drone observation The speed of the drone Indicates the first The first in the alliance The drone is the first The state estimation vector of each UAV Indicates the first The state vector of each drone Indicates the first The input matrix of a drone, The modulo operator, It is a norm of 1; , To set parameters; Indicates the first The first in the alliance The drone and the first The derivative of the adaptive parameters of the intra-group interaction terms among the drones. Indicates alliance In the topological matrix of the first Line number Column elements; Indicates alliance The Middle The state observation values observed by the UAV are determined based on the state observation data; , indicating that by the first Input matrix of a drone according to Zhang Cheng's diagonal block matrix.
6. The UAV flight control method based on fully distributed Nash equilibrium search in multi-alliance game theory according to claim 2, characterized in that, When the difference between the state data obtained from two consecutive iterations of the drone is less than a set threshold, the state stability requirement is met.
7. The UAV flight control method based on fully distributed Nash equilibrium search in multi-alliance game theory according to claim 1, characterized in that, The communication topology is a connected undirected graph.
8. A computer device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the computer program to implement the UAV flight control method based on fully distributed Nash equilibrium search of multi-alliance game as described in any one of claims 1-7.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the UAV flight control method based on fully distributed Nash equilibrium search of multi-alliance game as described in any one of claims 1-7.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the UAV flight control method based on fully distributed Nash equilibrium search of multi-alliance game as described in any one of claims 1-7.