Urban counter-environment collaborative method for security reinforcement learning of unmanned swarm system
By combining distributed elastic observers and augmented dynamics models, the problems of network attacks and formation distortion in unmanned swarm systems under urban adversarial environments were solved, thereby improving formation performance and collaborative performance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING INST OF TECH
- Filing Date
- 2025-12-26
- Publication Date
- 2026-04-10
AI Technical Summary
In traditional urban combat environments, the formation performance of unmanned swarm systems is vulnerable to cyberattacks and formation distortion, leading to a decline in formation coordination performance.
A distributed elastic observer is used to perform cyberattack observation processing on the state of the leader vehicle and the follower unmanned vehicles. Target control data is calculated through security estimation and augmented dynamics model to achieve formation control.
It improves the formation and collaboration performance of unmanned swarm systems, maintains system stability, and preserves the stability and optimal solution of formation.
Smart Images

Figure CN121386439B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of unmanned cluster system. More particularly, the present application relates to a safety reinforcement learning coordination method of unmanned cluster system in urban confrontation environment. BACKGROUND
[0002] The traditional safety reinforcement learning coordination method of unmanned cluster system in urban confrontation environment can be understood as a formation control method of unmanned cluster system. The unmanned cluster system can also be called a networked heterogeneous vehicle cluster system, a networked multi-vehicle system or a vehicle sensor network. The unmanned cluster system is composed of a leader vehicle and multiple follower unmanned vehicles (i.e. unmanned vehicles). The leader vehicle can be a virtual leader.
[0003] The traditional safety reinforcement learning coordination method of unmanned cluster system in urban confrontation environment is to construct a dynamic model of the virtual leader vehicle and perform distributed estimation on the dynamics of the leader vehicle; construct a feedback linearization model of the unmanned vehicle; an obstacle avoidance mechanism design module realizes obstacle avoidance based on an obstacle function; a formation control module calculates the actual control input of the vehicle by using the obstacle function, the tracking error of the estimated state of the position of the unmanned vehicle itself and the feedback linearization model of the coordinate transformation and The actual control input is used to realize the formation coordination task. However, in this method, the actual control input is only related to the obstacle function, the tracking error of the estimated state of the position of the unmanned vehicle itself and the feedback linearization model of the coordinate transformation, and has no direct connection with the safety of the state of the leader vehicle. Therefore, the actual control input is easy to deviate from the optimal solution, resulting in the phenomenon of distortion or splitting of the formation, which significantly reduces the formation performance of the unmanned cluster system. In addition, the network used by the leader vehicle and all unmanned vehicles in this method has information sharing, scalability and operational flexibility. These characteristics make the network used by the leader vehicle and all unmanned vehicles vulnerable to network attacks, resulting in the inability of the formation to coordinate, which further reduces the formation performance of the unmanned cluster system. SUMMARY
[0004] The purpose of the embodiments of the present application is to provide a safety reinforcement learning coordination method of unmanned cluster system in urban confrontation environment, which can improve the formation performance of the unmanned cluster system. The embodiments of the present application mainly realize the following technical solutions:
[0005] The embodiments of the present application provide a safety reinforcement learning coordination method of unmanned cluster system in urban confrontation environment, comprising:
[0006] obtaining a system state of a leader vehicle in an unmanned cluster system, a first observed state of a distributed observer of a first follower unmanned vehicle, and a second observed state of a distributed observer of a second follower unmanned vehicle, the first follower unmanned vehicle being any one of all follower unmanned vehicles in the unmanned cluster system, and the second follower unmanned vehicle being any one of all follower unmanned vehicles except the first follower unmanned vehicle;
[0007] performing network attack observation processing on the system state of the leader vehicle, the first observed state, and the second observed state by using a distributed resilient observer, obtaining an observation result, and performing secure estimation processing on the state of the leader vehicle based on the observation result, and obtaining a secure estimation result;
[0008] when the secure estimation result is secure, performing calculation processing on the system state of the leader vehicle, the first observed state, and the second observed state by using an augmented dynamics model, and obtaining target control data;
[0009] performing formation control on the first follower unmanned vehicle based on the target control data, the formation control being used to maintain stability of a formation shape of the unmanned cluster system in a city confrontation environment.
[0010] According to an embodiment of the present application, the unmanned cluster system secure reinforcement learning collaboration method in the city confrontation environment further comprises a construction method of the distributed resilient observer, and the construction method of the distributed resilient observer comprises:
[0011] performing calculation processing on an attack strength of a denial-of-service attack of a network attack on a communication link of the first follower unmanned vehicle and the second follower unmanned vehicle based on a preset attack strength algorithm, and obtaining an observation control gain constraint value;
[0012] performing calculation processing on the observation control gain constraint value based on a preset control gain algorithm, and obtaining a target observation control gain;
[0013] constructing the distributed resilient observer based on a communication topology adjacency weight between the first follower unmanned vehicle and the second follower unmanned vehicle, the system state of the leader vehicle, and the target observation control gain.
[0014] According to an embodiment of the present application, a calculation formula of the step of performing calculation processing on an attack strength of a denial-of-service attack of a network attack on a communication link of the first follower unmanned vehicle and the second follower unmanned vehicle based on a preset attack strength algorithm, and obtaining an observation control gain constraint value is:
[0015] ;
[0016] ;
[0017] ;
[0018] wherein, is the observation control gain constraint value; is a set of communication links of the leader vehicle and all follower unmanned vehicles; denotes the i-th follower unmanned vehicle in the unmanned swarm system, i = 1, 2, …, N; denotes the first follower unmanned vehicle in the unmanned swarm system; denotes the i-th follower unmanned vehicle in the unmanned swarm system, i = 1, 2, …, N; denotes the second follower unmanned vehicle in the unmanned swarm system; is an equivalent decay rate under a denial-of-service attack on the communication link between the first follower unmanned vehicle and the second follower unmanned vehicle; is an attack intensity of the denial-of-service attack on the communication link between the first follower unmanned vehicle and the second follower unmanned vehicle; is an equivalent decay rate under no denial-of-service attack on the communication link between the first follower unmanned vehicle and the second follower unmanned vehicle; is a set of communication links under the denial-of-service attack in The symbol represents a difference set operator of a set.
[0019] According to an embodiment of the present application, the observation control gain constraint value is calculated and processed based on a preset control gain algorithm, and a calculation formula of the step of obtaining a target observation control gain is as follows:
[0020] ;
[0021] ;
[0022] wherein, is the target observation control gain; is a unit matrix; is a first constant matrix; is a transpose of ; is a positive gain constant; ; is a minimum eigenvalue function; is an information transmission matrix after the unmanned swarm system is attacked by a denial-of-service attack; is a transpose of ; ; is an operator for constructing a diagonal matrix; It is the reciprocal of the first diagonal element in the diagonal matrix; It is the first in the diagonal matrix The reciprocal of each diagonal element; It is a positive definite parameter matrix related to the control gain of the distributed elastic observer.
[0023] According to one embodiment of this application, the calculation formula for constructing the distributed resilient observer based on the communication topology adjacency weight between the first follower unmanned vehicle and the second follower unmanned vehicle, the system state of the leader vehicle, and the target observation control gain is as follows:
[0024] ;
[0025] in, This is the first observation state; These are the observation results; It is the target observation control gain; It is the set of neighbors of the first follower autonomous vehicle; It is a denial-of-service attack model; It is the communication topology adjacency weight between the first follower autonomous vehicle and the second follower autonomous vehicle; This is the second observation state; It is the expected formation error between the first follower autonomous vehicle and the second follower autonomous vehicle; It is the connection weight between the leader vehicle and the first follower autonomous vehicle; This refers to the system state of the leader vehicle; It is the expected formation error between the first follower autonomous vehicle and the leader vehicle.
[0026] According to one embodiment of this application, the calculation formula for the denial-of-service attack model is as follows:
[0027] ;
[0028] ;
[0029] in, It is the collection of communication links between the leader vehicle and all the follower autonomous vehicles; Is to The union of the time intervals during which the communication links of the first follower autonomous vehicle and the second follower autonomous vehicle are subjected to denial-of-service attacks; This is the initial moment of operation of the unmanned swarm system; It is the current moment; is a fixed attack effect allowed to the communication link of the first follower unmanned vehicle and the second follower unmanned vehicle by a denial-of-service attack; is an attack intensity of the communication link of the first follower unmanned vehicle and the second follower unmanned vehicle by a denial-of-service attack.
[0030] According to an embodiment of the present application, the step of performing a safety estimation process on the state of the leader vehicle based on the observation result to obtain a safety estimation result comprises:
[0031] In a case where the observation result is a derivative of the first observation state, it is determined that the state of the leader vehicle belongs to a safe state, and the safety estimation result is safe; otherwise, the state of the leader vehicle does not belong to a safe state, and the safety estimation result is unsafe.
[0032] According to an embodiment of the present application, the step of performing a calculation process on the system state of the leader vehicle, the first observation state and the second observation state by using an augmented dynamics model to obtain target control data comprises a calculation formula of the step:
[0033] ;
[0034] ;
[0035] ;
[0036] wherein, is the target control data; is a system state of an augmented system of the first follower unmanned vehicle; is a derivative of ; is a system matrix of the augmented system of the first follower unmanned vehicle; is an input matrix of the augmented system of the first follower unmanned vehicle; is a control gain matrix of a distributed observer of the first follower unmanned vehicle; is a transpose of ; is an augmented error term for reflecting that the augmented system of the first follower unmanned vehicle is affected by a neighborhood error of the distributed observer; is a neighborhood error term to which the observation state of the first follower unmanned vehicle is subjected.
[0037] According to one embodiment of the present application, in the case that the safety estimation result is safe, after the step of calculating and processing the system state of the leader vehicle, the first observed state and the second observed state by using the augmented dynamics model to obtain target control data, the urban confrontation environment unmanned cluster system safety reinforcement learning collaboration method further comprises:
[0038] setting an initial gain matrix;
[0039] iteratively solving and processing the initial gain matrix by using an offline policy model to obtain an offline gain matrix;
[0040] iteratively solving and processing the offline gain matrix by using an online policy model to obtain an optimized gain matrix;
[0041] cumulatively calculating and processing based on the optimized gain matrix and the system state of the augmented system of the first follower unmanned vehicle to obtain optimized control data;
[0042] controlling the formation of the first follower unmanned vehicle based on the optimized control data.
[0043] According to one embodiment of the present application, the calculation formula of the step of cumulatively calculating and processing based on the optimized gain matrix and the system state of the augmented system of the first follower unmanned vehicle to obtain optimized control data is:
[0044]
[0045] wherein, is the optimized control data; is the optimized gain matrix.
[0046] The beneficial effects of the embodiments of the present application include:
[0047] The embodiment of the application is through the network attack observation processing of the system state of the leader vehicle, the first observation state of the distributed observer of the first follower unmanned vehicle and the second observation state of the distributed observer of the second follower unmanned vehicle by using the distributed elastic observer, obtaining the observation result, and based on the observation result, the safety estimation processing of the state of the leader vehicle is carried out to obtain the safety estimation result; in the case of safety estimation result, the system state of the leader vehicle, the first observation state and the second observation state are calculated and processed by using the augmented dynamics model to obtain the target control data; the first follower unmanned vehicle is controlled based on the target control data. Since the first follower unmanned vehicle is any one of all follower unmanned vehicles in the unmanned cluster system, therefore, the unmanned cluster system can control all follower unmanned vehicles through the above steps. Compared with the prior art which is not directly related to the safety of the state of the leader vehicle, the embodiment of the application designs a distributed elastic observer, and the output result of the distributed elastic observer can be used for safety estimation of the state of the leader vehicle, so that the embodiment of the application can consider the state of the leader vehicle in the case of formation, so that the target control data (i.e. the actual control input in the prior art) approaches or reaches the optimal solution, so that the embodiment of the application can maintain the formation shape and realize optimal formation tracking, therefore, the embodiment of the application can improve the formation performance of the unmanned cluster system. Moreover, the distributed elastic observer can perform network attack elastic processing (which can also be understood as observation processing), thereby improving the cooperation performance of the unmanned cluster system, maintaining the stability of the unmanned cluster system, and further improving the formation performance of the unmanned cluster system. BRIEF DESCRIPTION OF DRAWINGS
[0048] In order to more clearly illustrate the technical solutions in the embodiments of the application or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative labor.
[0049] Figure 1 The flowchart of the safety reinforcement learning collaboration method of the unmanned cluster system in the urban confrontation environment of the application in some embodiments;
[0050] Figure 2 The flowchart of the safety reinforcement learning collaboration method of the unmanned cluster system in the urban confrontation environment of the application in some other embodiments;
[0051] Figure 3 The flowchart of the safety reinforcement learning collaboration method of the unmanned cluster system in the urban confrontation environment of the application in some other embodiments. DETAILED DESCRIPTION
[0052] In order to make the above objectives, features and advantages of the present application more clear and comprehensible, specific embodiments of the present application will be described in detail below with reference to the accompanying drawings. In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present application. It will be apparent, however, to one skilled in the art that the present application can be practiced without using some or all of these specific details. In other instances, well known process steps have not been described in detail in order to avoid obscuring the present application.
[0053] It should be noted that the terms "first", "second" and the like in the description and in the claims are used for descriptive purposes only and not to denote or imply relative importance or an ordered sequence. Therefore, a feature defined with "first", "second" can include at least one of the features, explicitly or implicitly. In the description of the present application, "a plurality of" means at least two, for example, two, three, etc., unless otherwise specifically defined.
[0054] The term "exemplary" or "for example" is used to indicate that an example is one of a possible set of alternatives. Any embodiment or design scheme described as "exemplary" or "for example" in the present application should not be construed as being more preferred or advantageous than other embodiments or design schemes. Rather, the use of "exemplary" or "for example" is merely intended to present concepts in a concrete manner.
[0055] The term "include", "comprise" or any other variant is intended to encompass non-exclusive inclusions, for example, a process, method, system, product or apparatus that comprises a list of steps or units is not necessarily limited to those steps or units that are clearly listed, but can include other steps or units that are not clearly listed or inherent to such processes, methods, products or apparatus.
[0056] The term "network attack" includes denial of service attacks and false injection attacks, and the denial of service attacks include multi-channel asynchronous denial of service attacks and synchronous denial of service attacks.
[0057] Unless otherwise defined, all technical and scientific terms used in the specification of the present application have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs. The terms used in the specification of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. The term "and / or" used in the specification of the present application includes any and all combinations of one or more related listed items.
[0058] The specific embodiments of the present application are further described below with reference to the accompanying drawings.
[0059] Reference Figure 1A flow chart of a safety reinforcement learning collaboration method of an unmanned cluster system in an urban confrontation environment is provided for an embodiment of the present application. In Figure 1 the safety reinforcement learning collaboration method of the unmanned cluster system in the urban confrontation environment comprises:
[0060] S1, obtaining a system state of a leader vehicle in an unmanned cluster system, a first observation state of a distributed observer of a first follower unmanned vehicle, and a second observation state of a distributed observer of a second follower unmanned vehicle, the first follower unmanned vehicle being any one of all follower unmanned vehicles in the unmanned cluster system, and the second follower unmanned vehicle being any one of all follower unmanned vehicles except the first follower unmanned vehicle.
[0061] The unmanned cluster system can also be referred to as a networked heterogeneous vehicle cluster system or a networked multi-vehicle system, and the unmanned cluster system is composed of a leader vehicle and multiple follower unmanned vehicles. The leader vehicle can be a virtual leader. All follower unmanned vehicles are connected through a directed communication network topology with the virtual leader as a root node and containing a directed spanning tree.
[0062] The unmanned cluster system can be applied to cooperative driving of an autonomous vehicle fleet, urban logistics cooperative transportation, emergency rescue and urban patrol, and other application scenarios.
[0063] The system state of the leader vehicle can be expressed as wherein, is the position of the leader vehicle in the two-dimensional plane direction; is the speed of the leader vehicle in the two-dimensional plane direction; is the position of the leader vehicle in the two-dimensional plane direction; is the speed of the leader vehicle in the two-dimensional plane direction; represents transposition; subscript represents the leader vehicle.
[0064] The first observation state is obtained by measuring the distributed observer of the first follower unmanned vehicle, which can be implemented by existing technologies. The first observation state includes the position of the first follower unmanned vehicle in the two-dimensional plane direction, the speed of the first follower unmanned vehicle in the two-dimensional plane direction, the position of the first follower unmanned vehicle in the two-dimensional plane direction, and the speed of the first follower unmanned vehicle in the two-dimensional plane a speed in a direction.
[0065] The second observation state is obtained by a distributed observer of the second follower unmanned vehicle, which can be realized by existing technologies. The second observation state includes a position of the second follower unmanned vehicle in a two-dimensional plane a position in a direction, a position of the second follower unmanned vehicle in a two-dimensional plane a speed in a direction, a speed of the second follower unmanned vehicle in a two-dimensional plane a position in a direction, and a position of the second follower unmanned vehicle in a two-dimensional plane a speed in a direction.
[0066] S2, a distributed resilient observer is used to perform network attack observation processing on the system state of the leader vehicle, the first observation state and the second observation state, to obtain an observation result, and to perform a security estimation processing on the state of the leader vehicle based on the observation result, to obtain a security estimation result.
[0067] The distributed resilient observer can be used to estimate the state and formation information of the leader vehicle under a multi-channel asynchronous denial of service attack, and further provide necessary formation tracking trajectory references for the networked heterogeneous vehicle cluster system through output conversion operation; the distributed resilient observer can also perform security control on the multi-channel asynchronous denial of service attack; the communication topology between the networked heterogeneous vehicle cluster system is a directed graph with the leader vehicle as a root node and containing a directed spanning tree.
[0068] Further, referring to Figure 2 illustrated, the construction method of the distributed resilient observer includes S21, S22 and S23 steps:
[0069] S21, based on a preset attack intensity algorithm, the attack intensity of the denial of service attack on the communication link of the first follower unmanned vehicle and the second follower unmanned vehicle under network attack is calculated and processed to obtain an observation control gain constraint value.
[0070] Further, the calculation formula of the S21 step is:
[0071] ;
[0072] ;
[0073] ;
[0074] wherein, is the observation control gain constraint value; is a set of communication links of the leader vehicle and all follower unmanned vehicles, that is, a directed graph communication network topology between the connected heterogeneous vehicle cluster system; refers to the first follower unmanned vehicle, that is, the first follower unmanned vehicle in the unmanned cluster system ; refers to the second follower unmanned vehicle, that is, the second follower unmanned vehicle in the unmanned cluster system ; is an equivalent attenuation rate in the case of a denial-of-service attack on the communication link between the first follower unmanned vehicle and the second follower unmanned vehicle; is the attack intensity of the denial-of-service attack on the communication link between the first follower unmanned vehicle and the second follower unmanned vehicle; is an equivalent attenuation rate in the case of no denial-of-service attack on the communication link between the first follower unmanned vehicle and the second follower unmanned vehicle; is a set of communication links subjected to a denial-of-service attack in the unmanned cluster system; The symbol represents the difference set operator of the set; represents a set of communication links not subjected to a denial-of-service attack.
[0075] The distributed resilient observer of the embodiments of the present application can realize safe estimation of the virtual leader state when the connected heterogeneous vehicle cluster system is subjected to a multi-channel asynchronous denial-of-service attack by introducing an equivalent attenuation rate, thereby providing a trajectory reference for formation tracking control of the connected heterogeneous vehicle cluster system.
[0076] The calculation formula of the S21 step can also be expressed as an inequality group based on the attack intensity of the multi-channel asynchronous denial-of-service attack signal.
[0077] S22, based on a preset control gain algorithm, the observation control gain constraint value is calculated and processed to obtain a target observation control gain.
[0078] Further, the calculation formula of the S22 step is:
[0079] ;
[0080] ;
[0081] wherein, is the target observation control gain, which can be understood as the overall control gain matrix of the distributed resilient observer; is an identity matrix, and the identity matrix is a to-be-designed parameter matrix; is a first constant matrix; is the transpose of is a positive gain constant; is a minimum eigenvalue function; is the information transfer matrix of the unmanned swarm system after being attacked by a denial-of-service attack; is the transpose of ; is an operator for constructing a diagonal matrix; is the reciprocal of the first diagonal element in the diagonal matrix; is a positive definite parameter matrix related to the distributed resilient observer control gain; is the reciprocal of the th diagonal element in the diagonal matrix. The parameter matrix to be designed can be set by those skilled in the art according to actual needs.
[0082] Further, , is the initial information transfer matrix of the unmanned swarm system before being attacked by a denial-of-service attack, is the information transfer matrix of the attacked channel in the unmanned swarm system.
[0083] Further, satisfies , is the th diagonal element.
[0084] S23, constructing the distributed resilient observer based on the communication topology adjacency weight between the first follower unmanned vehicle and the second follower unmanned vehicle, the system state of the leader vehicle and the target observation control gain.
[0085] Further, the calculation formula of the S23 step is:
[0086] ;
[0087] wherein, is the first observation state, which can also be understood as the state of the observer of the unmanned system of the first follower unmanned vehicle providing distributed estimation; is the observation result; is a target observation control gain, that is, the control gain of the distributed resilient observer; is the neighbor set of the first follower unmanned vehicle; is a denial-of-service attack model; is the communication topology adjacency weight between the first follower unmanned vehicle and the second follower unmanned vehicle; is the second observation state; It is the expected formation error between the first follower autonomous vehicle and the second follower autonomous vehicle; It is the connection weight between the leader vehicle and the first follower autonomous vehicle; This refers to the system state of the leader vehicle; This is the expected formation error between the first follower vehicle and the leader vehicle. If there is communication between the first follower vehicle and the second follower vehicle, then... Greater than zero, otherwise It equals zero. If there is communication between the first follower autonomous vehicle and the second follower autonomous vehicle, then Greater than zero, otherwise It equals zero. yes Another form of representation, namely " The set of communication links that are subject to denial-of-service attacks can be understood as a "denial-of-service attack model".
[0088] Furthermore, ; It is the expected formation error between the second follower autonomous vehicle and the leader vehicle.
[0089] Furthermore, the calculation formula for step S23 also includes: ;in, The output of the observer provides distributed estimation for the unmanned system of the first follower vehicle.
[0090] Furthermore, the calculation formula for the denial-of-service attack model is as follows:
[0091] ;
[0092] ;
[0093] in, It is the collection of communication links between the leader vehicle and all the follower autonomous vehicles; Is to The union of the time intervals during which the communication links of the first follower autonomous vehicle and the second follower autonomous vehicle are subjected to denial-of-service attacks; This is the initial moment of operation of the unmanned swarm system; It refers to the current moment; The communication link between the first follower autonomous vehicle and the second follower autonomous vehicle is subject to a fixed attack effect that allows denial-of-service attacks to occur. It is the attack strength of a denial-of-service attack on the communication link between the first follower autonomous vehicle and the second follower autonomous vehicle.
[0094] Further, the step of performing a safety estimation on the state of the leader vehicle based on the observation result to obtain a safety estimation result comprises: in a case where the observation result is a derivative of the first observation state, determining that the state of the leader vehicle belongs to a safe state, and the safety estimation result is safe; otherwise, the state of the leader vehicle does not belong to a safe state, and the safety estimation result is unsafe.
[0095] S3, in a case where the safety estimation result is safe, performing calculation processing on the system state of the leader vehicle, the first observation state and the second observation state by using an augmented dynamics model to obtain target control data.
[0096] Further, the calculation formula of the step of performing calculation processing on the system state of the leader vehicle, the first observation state and the second observation state by using an augmented dynamics model to obtain target control data is:
[0097] ;
[0098] ;
[0099] ;
[0100] wherein, is the target control data; is the system state of the augmented system of the first follower unmanned vehicle; is a derivative of ; is the system matrix of the augmented system of the first follower unmanned vehicle; is the input matrix of the augmented system of the first follower unmanned vehicle; is the control gain matrix of the distributed observer of the first follower unmanned vehicle; is the transpose of ; is an augmented error term for reflecting that the augmented system of the first follower unmanned vehicle is affected by the neighborhood error of the distributed observer; is the neighborhood error term of the observation state of the first follower unmanned vehicle. The target control data can be understood as the control input of the augmented system of the first follower unmanned vehicle. Also can be understood as neighborhood leader-follower error.
[0101] Further, . is the transpose of , is the transpose of .
[0102] , , is a vector space of dimension is a vector space of dimension is a state dimension of an augmented system of the first follower unmanned vehicle; is a control input dimension of the augmented system of the first follower unmanned vehicle. .
[0103] In some embodiments, the step of calculating the system state of the leader vehicle, the first observation state and the second observation state by using the augmented dynamics model can further comprise a calculation formula of:
[0104] ;
[0105] ;
[0106] wherein, is a platoon tracking error; ; is an output matrix of the augmented system of the first follower unmanned vehicle.
[0107] when the distributed elastic observer converges, converges to zero asymptotically, is a platoon control gain matrix to be learned, is a control gain matrix related to the feedforward term (observer), is a control gain matrix related to the feedback term (connected vehicle).
[0108] In other embodiments, .
[0109] S4, performing platoon control on the first follower unmanned vehicle based on the target control data, the platoon control being used to maintain the stability of the platoon formation of the unmanned cluster system in the urban confrontation environment.
[0110] Specifically, the target control data can be generated into a first platoon control command, and the first platoon control command can be sent to the first follower unmanned vehicle, so that the platoon controller of the first follower unmanned vehicle executes the first platoon control command to achieve the purpose of platoon. In other embodiments, the specific implementation process of the S4 step can be realized by the prior art.
[0111] The urban confrontation environment includes a high-density traffic flow environment, a human-vehicle mixed traffic environment, a vehicle network attack environment, an urban road obstacle burst environment, and an urban severe weather environment, etc.
[0112] Through the above implementation, the embodiment of the application designs a distributed elastic observer, and the output result of the distributed elastic observer can be used for safe estimation of the state of the leader vehicle, so that the embodiment of the application can consider the state of the leader vehicle in the case of formation, so that the target control data (i.e. the actual control input in the prior art) approaches or reaches the optimal solution, so that the embodiment of the application can maintain the formation shape and realize optimal formation tracking, and therefore, the embodiment of the application can improve the formation performance of the unmanned cluster system. Moreover, the distributed elastic observer can perform network attack elastic processing (which can also be understood as observation processing), thereby improving the cooperation performance of the unmanned cluster system, maintaining the stability of the unmanned cluster system operation, and further improving the formation performance of the unmanned cluster system.
[0113] In some embodiments, the embodiment of the application also defines an optimal performance index function (i.e. a reward function), and solves an optimization problem according to the optimal performance index function.
[0114] The calculation formula of the optimal performance index function is:
[0115] ;
[0116] wherein, is the base of the natural logarithm; is the discount factor of the first follower unmanned vehicle, ; is a first weight matrix, , ; is a second weight matrix, ; is an integral time variable; is the transpose of ; is the transpose of .
[0117] The calculation formula of the optimization problem is:
[0118] ;
[0119] ;
[0120] wherein, ; ; ; is Identity matrix of order.
[0121] In some embodiments, referring to FIG. 3, after the S3 step, the urban confrontation environment unmanned cluster system security reinforcement learning collaborative method further comprises S5, S6, S7, S8 and S9 steps: Figure 3
[0122] S5, setting an initial gain matrix.
[0123] The initial gain matrix can be expressed as .
[0124] Further, the initial gain matrix and the first control input of the detection noise act on the network-connected heterogeneous vehicle cluster system to enable the network-connected heterogeneous vehicle cluster system to achieve cooperative motion stability and enhance the robustness of the system to uncertainty.
[0125] The calculation formula of the first control input is .
[0126] S6, the initial gain matrix is iteratively solved by using an offline policy model to obtain an offline gain matrix.
[0127] Further, the calculation formula of the offline policy model is:
[0128] ;
[0129] ;
[0130] wherein, is the number of iterations; ; ; ; ; ; ; is the step size at the i-th iteration, ; is the maximum number of iterations; is the transpose of ;is the transpose of ; is the transpose of ; is the operation of converting a matrix into a vector; is the operation of converting a matrix into a vector; is the value matrix of the augmented system of the first follower unmanned vehicle at the k-th iteration during offline policy learning; is the value matrix of the augmented system of the first follower unmanned vehicle at the k-th iteration during offline policy learning; is the value matrix of the augmented system of the first follower unmanned vehicle at the k-th iteration during offline policy learning; is the value matrix of the augmented system of the first follower unmanned vehicle at the k-th iteration during offline policy learning; is the value matrix of the augmented system of the first follower unmanned vehicle at the k-th iteration during offline policy learning; is a quadratic coefficient matrix of the augmented system of the first follower UAV at the kth iteration when the offline policy learning is performed; is a data matrix of the augmented system of the first follower UAV at the kth iteration when the offline policy learning is performed, is an integral of , is a transpose of
[0131] The calculation formula of the offline policy model can solve the stable offline gain matrix . Embodiments of the present application take the stable as the online gain matrix .
[0132] The offline gain matrix can be understood as a stable policy. The stable policy is applied to the iterative algorithm of S7 steps.
[0133] Embodiments of the present application define a first data storage unit , a second data storage unit and a third data storage unit , and the first data storage unit, the second data storage unit and the third data storage unit are used to store offline data of system operation.
[0134] The calculation formula of the first data storage unit is: , , is a vectorization operation, is the total length of sufficient data collected in the offline policy learning stage.
[0135] The calculation formula of the second data storage unit is: , .
[0136] The calculation formula of the third data storage unit is: , , is a tensor product symbol.
[0137] Further, embodiments of the present application determine whether the offline policy value iteration algorithm (that is, the offline policy model) learns the stable policy by switching conditions. Specifically, the calculation formula of the switching condition is:
[0138] ;
[0139] wherein is a quadratic coefficient matrix of the augmented system of the first follower UAV at the kth iteration when the offline policy learning is performed; is a quadratic term coefficient matrix of the augmented system of the first follower unmanned vehicle at the kth iteration time in the offline policy value iteration algorithm; is a quadratic term coefficient matrix of the augmented system of the first follower unmanned vehicle at the kth iteration time in the offline policy value iteration algorithm; is a preset third weight matrix, used for determining whether the offline policy value iteration algorithm learns the stable policy; .
[0140] S7, using the online policy model to iteratively solve the offline gain matrix to obtain an optimized gain matrix.
[0141] Further, the calculation formula of the online policy model is:
[0142] ;
[0143] wherein, ; ; ; ; ; .
[0144] wherein, is the transpose of ; is the transpose of ; is a quadratic term coefficient matrix of the augmented system of the first follower unmanned vehicle at the kth iteration time in the online policy learning.
[0145] The calculation formula of the online policy model can solve the optimal and the optimal . The optimal is taken as the optimized gain matrix .
[0146] The embodiment of the application defines a fourth data storage unit , a fifth data storage unit and a sixth data storage unit , and the fourth data storage unit, the fifth data storage unit and the sixth data storage unit are all used to store online data of system operation.
[0147] The calculation formula of the fourth data storage unit is: wherein, , is the length of the appropriate size of data collected each time the control policy is updated in the online policy learning stage. The calculation formula of the fifth data storage unit is: , .
[0148] The calculation formula of the sixth data storage unit is: , .
[0149] The control strategy of this application (i.e., the online strategy model) is as follows: Updated once, in which, It is the integral sampling time.
[0150] Furthermore, this application will continue to include the first After the next iteration, the strategy and the second control input of the detection noise are applied to the connected heterogeneous vehicle cluster system to generate new online data.
[0151] The calculation formula for the second control input is as follows: .
[0152] S8. Based on the optimized gain matrix and the system state of the augmentation system of the first follower unmanned vehicle, perform cumulative calculation processing to obtain optimized control data.
[0153] Furthermore, the calculation formula for step S8 is as follows:
[0154] ;
[0155] in, It is the optimized control data; It is the optimized gain matrix.
[0156] In other embodiments, This can also be understood as the optimal formation controller of the unmanned swarm system. It can also be understood as the optimal feedback gain matrix.
[0157] In other embodiments, , It can be obtained by solving the Riccati equation, which is an algebraic equation for the discount factor. The solution to the discount factor algebraic Riccati equation. The optimal solution, the discount factor algebraic Riccati equation, is:
[0158] ;
[0159] in, The discount factor in the algebraic Riccati equation It is the quadratic coefficient matrix of the cost function, used to describe the state weighting or stability characteristics in optimal control.
[0160] S9. Perform formation control on the first follower unmanned vehicle based on the optimized control data.
[0161] Specifically, the optimized control data can be used to generate a second formation control command, which is then sent to the first follower unmanned vehicle (UAV) to enable its formation controller to execute the command and achieve formation. In other embodiments, step S9 can be implemented using existing technologies.
[0162] Steps S5, S6, S7, S8, and S9 can be understood as the hybrid policy reinforcement learning algorithm of this application embodiment. The use of this hybrid policy reinforcement learning algorithm enables this application embodiment to achieve safe formation coordination of heterogeneous unmanned systems by utilizing only system data from the operation of the connected heterogeneous vehicle cluster system, without needing to obtain an accurate system dynamics model. This is achieved through a paradigm of offline policy pre-learning (i.e., the offline policy model) and online policy fine-tuning to learn the optimal safe formation control strategy (i.e., the online policy model). Furthermore, the application of the offline policy model eliminates the need for an initial stable policy in traditional reinforcement learning methods, while the application of the online policy model enables the formation controller of the first follower unmanned vehicle to continuously optimize the policy in real time while reducing memory consumption, all while learning the optimal control strategy.
[0163] In some implementations, after step S7, the secure reinforcement learning collaborative method for unmanned swarm systems in an urban adversarial environment further includes:
[0164] judge Whether it is true or not, if If the condition is met, then the iteration stops. To obtain the optimal control gain matrix, As the optimized gain matrix, the optimal safe formation controller is... ,Will As the optimized control data; if If not, then The online strategy model continues to be solved iteratively, wherein, It is a small constant value greater than zero.
[0165] The optimal safe formation controller can perform optimal flexible formation control on the connected heterogeneous vehicle cluster system.
[0166] In some implementations, the secure reinforcement learning collaborative method for unmanned swarm systems in urban adversarial environments further includes: establishing a dynamic model of the leader vehicle and a dynamic model of the follower unmanned vehicles. The dynamic models of the leader vehicle and the follower unmanned vehicles can be collectively referred to as a connected heterogeneous vehicle swarm system model.
[0167] Further, the leader vehicle's dynamics model is calculated by the following formula:
[0168] ;
[0169] wherein, is a derivative of the system state of the leader vehicle; is the system state of the leader vehicle; is the output of the leader vehicle; is a second constant matrix.
[0170] Further, the follower unmanned vehicle's dynamics model is exemplified by the first follower unmanned vehicle's dynamics model, which is calculated by the following formula:
[0171] ;
[0172] ;
[0173] ;
[0174] ;
[0175] ;
[0176] ;
[0177] wherein, is a derivative of the position of the first follower unmanned vehicle in the direction of the two-dimensional plane; is a yaw angle of the first follower unmanned vehicle ; is a velocity of the first follower unmanned vehicle in the direction of the two-dimensional plane; is a velocity of the first follower unmanned vehicle in the direction of the two-dimensional plane; is a derivative of the position of the first follower unmanned vehicle in the direction of the two-dimensional plane; is a mass of the first follower unmanned vehicle ; is a derivative of ; is a yaw rate of the first follower unmanned vehicle ; is a derivative of ; is a derivative of the position of the first follower unmanned vehicle in the direction of the two-dimensional plane; is a derivative of the position of the first follower unmanned vehicle total force experienced by the first follower unmanned vehicle in the direction of the two-dimensional position of the first follower unmanned vehicle; total force experienced by the first follower unmanned vehicle in the direction of the two-dimensional position of the first follower unmanned vehicle; aerodynamic drag coefficient of the first follower unmanned vehicle; aerodynamic drag coefficient of the first follower unmanned vehicle; rolling friction coefficient of the first follower unmanned vehicle; rolling friction coefficient of the first follower unmanned vehicle; gravitational acceleration of the first follower unmanned vehicle; derivative of the gravitational acceleration of the first follower unmanned vehicle; derivative of the gravitational acceleration of the first follower unmanned vehicle; moment of inertia of the first follower unmanned vehicle; total force experienced by the first follower unmanned vehicle in the direction of the two-dimensional position of the first follower unmanned vehicle; total force experienced by the first follower unmanned vehicle in the direction of the two-dimensional position of the first follower unmanned vehicle; derivative of the total force experienced by the first follower unmanned vehicle in the direction of the two-dimensional position of the first follower unmanned vehicle; derivative of the total force experienced by the first follower unmanned vehicle in the direction of the two-dimensional position of the first follower unmanned vehicle; moment of inertia of the first follower unmanned vehicle; moment of inertia of the first follower unmanned vehicle; total moment of the yaw motion of the first follower unmanned vehicle.
[0178] Further, ; ; position of the first follower unmanned vehicle in a two-dimensional plane; position of the first follower unmanned vehicle in a two-dimensional plane; velocity of the first follower unmanned vehicle in a two-dimensional plane; transpose; vector space of dimension vector space of dimension vector space of dimension
[0179] Assuming the first follower unmanned vehicle moves with small angles and and using feedback linearization techniques, the control inputs of the first follower unmanned vehicle can be defined as and , control input of the first follower unmanned vehicle in the direction of control input of the first follower unmanned vehicle in the direction of control input of the first follower unmanned vehicle in the direction of control input of the first follower unmanned vehicle in the direction of the state vector of the first follower unmanned vehicle, then the nominal dynamic model of the first follower unmanned vehicle can be re-expressed as ; where, position of the first follower unmanned vehicle in a two-dimensional plane in the direction of position of the first follower unmanned vehicle in a two-dimensional plane in the direction of position of the first follower unmanned vehicle in a two-dimensional plane in the direction of position of the first follower unmanned vehicle in a two-dimensional plane in the direction of position of the first follower unmanned vehicle in a two-dimensional plane in the direction of a position in a direction; is a first system parameter matrix of the first follower unmanned vehicle ; ; is a second system parameter matrix of the first follower unmanned vehicle ; ; is a third system parameter matrix of the first follower unmanned vehicle ; ; ; is a system output.
[0180] It should also be noted that the embodiments of the present application can achieve optimal cooperative control of the networked heterogeneous vehicle cluster system in the case of complete unknown and multi-channel asynchronous denial of service attack, and the hybrid strategy reinforcement learning algorithm overcomes the limitations of the need for an initial stable strategy and insufficient dynamic optimization capability in existing reinforcement learning algorithms.
[0181] The communication network of the networked multi-vehicle system can still maintain the stability of the formation and the effective execution of the cooperative task when subjected to multi-channel asynchronous denial of service attack in the urban confrontation environment.
[0182] In some embodiments, the unmanned cluster system safety reinforcement learning cooperation method in the urban confrontation environment can be applied to a terminal device or a server. When the unmanned cluster system safety reinforcement learning cooperation method in the urban confrontation environment is applied to a terminal device, the terminal device includes a processor and a memory for storing a computer program, and the processor is used to call and run the computer program stored in the memory to execute the steps of the unmanned cluster system safety reinforcement learning cooperation method in the urban confrontation environment provided by the embodiments of the present application.
[0183] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when the computer program is executed, the processes of the above-mentioned embodiment methods can be included. Any reference to memory, storage, database or other medium used in the embodiments provided by the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0184] The technical features of the above embodiments can be combined without changing the basic principles of the present application. In order to make the description simple, not all possible combinations of the technical features in the above embodiments are described, but as long as the combinations of the technical features do not contradict, they should be considered as the scope of the present application.
[0185] The above embodiments only express several implementation manners of the present application, and the description is specific and detailed, but it should not be understood as a limitation on the patent scope of the application. It should be pointed out that for ordinary skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are all within the protection scope of the present application. Therefore, the patent protection scope of the present application should be subject to the appended claims.
Claims
1. A method for security reinforcement learning collaboration of unmanned swarm system in urban counter-environment, characterized in that, The method comprises the following steps: obtaining the system state of a leader vehicle in an unmanned cluster system, the first observation state of a distributed observer of a first follower unmanned vehicle, and the second observation state of a distributed observer of a second follower unmanned vehicle, the first follower unmanned vehicle being any one of all follower unmanned vehicles in the unmanned cluster system, and the second follower unmanned vehicle being any one of all follower unmanned vehicles except the first follower unmanned vehicle; performing network attack observation processing on the system state of the leader vehicle, the first observation state and the second observation state by using a distributed resilient observer, obtaining an observation result, and performing safe estimation processing on the state of the leader vehicle based on the observation result to obtain a safe estimation result; in the case that the safe estimation result is safe, performing calculation processing on the system state of the leader vehicle, the first observation state and the second observation state by using an augmented dynamics model to obtain target control data; performing formation control on the first follower unmanned vehicle based on the target control data, the formation control being used to maintain the stability of the formation shape of the unmanned cluster system in a city confrontation environment; the calculation formula of the step of performing calculation processing on the system state of the leader vehicle, the first observation state and the second observation state by using an augmented dynamics model to obtain target control data is: ; ; ; wherein, is the target control data; is the system state of the augmented system of the first follower unmanned vehicle; is a derivative of is the system matrix of the augmented system of the first follower unmanned vehicle; is the input matrix of the augmented system of the first follower unmanned vehicle; is the control gain matrix of the distributed observer of the first follower unmanned vehicle; is a transpose of is an augmented error term for reflecting that the augmented system of the first follower unmanned vehicle is constructed due to the influence of the neighborhood error of the distributed observer; is a neighborhood error term to which the observed state of the first follower unmanned vehicle is subjected. after the step of performing calculation processing on the system state of the leader vehicle, the first observation state and the second observation state by using an augmented dynamics model to obtain target control data in the case that the safe estimation result is safe, the method further comprises the following steps: setting an initial gain matrix; performing iterative solution processing on the initial gain matrix by using an offline policy model to obtain an offline gain matrix; performing iterative solution processing on the offline gain matrix by using an online policy model to obtain an optimized gain matrix; performing cumulative calculation processing on the optimized gain matrix and the system state of the augmented system of the first follower unmanned vehicle to obtain optimized control data; and performing formation control on the first follower unmanned vehicle based on the optimized control data.
2. The urban counter environment unmanned swarm system security reinforcement learning collaboration method according to claim 1, characterized in that, The method further comprises a construction method of the distributed resilient observer, and the construction method of the distributed resilient observer comprises the following steps: performing calculation processing on the attack strength of the denial-of-service attack of the network attack on the communication link between the first follower unmanned vehicle and the second follower unmanned vehicle by using a preset attack strength algorithm to obtain an observation control gain constraint value; performing calculation processing on the observation control gain constraint value by using a preset control gain algorithm to obtain a target observation control gain; constructing the distributed resilient observer based on the communication topology adjacency weight between the first follower unmanned vehicle and the second follower unmanned vehicle, the system state of the leader vehicle and the target observation control gain.
3. The urban confrontation environment-based unmanned swarm system security reinforcement learning collaboration method according to claim 2, characterized in that, The calculation formula of the step of calculating and processing attack intensity of a denial-of-service attack of a communication link of the first follower unmanned vehicle and the second follower unmanned vehicle based on a preset attack intensity algorithm is as follows: ; ; ; wherein, is the observed control gain constraint value; is a set of communication links of the leader vehicle and all follower unmanned vehicles; denotes the i-th follower unmanned vehicle in the unmanned swarm system, i = 1, 2,..., Nf; denotes the first follower unmanned vehicle in the unmanned swarm system; denotes the i-th follower unmanned vehicle in the unmanned swarm system, i = 1, 2,..., Nf; denotes the second follower unmanned vehicle in the unmanned swarm system; is the equivalent attenuation rate under the condition that a denial-of-service attack occurs on the communication link between the first follower unmanned vehicle and the second follower unmanned vehicle; is the attack intensity of the denial-of-service attack on the communication link between the first follower unmanned vehicle and the second follower unmanned vehicle; is the equivalent attenuation rate under the condition that no denial-of-service attack occurs on the communication link between the first follower unmanned vehicle and the second follower unmanned vehicle; is a set of communication links under the denial-of-service attack in The symbol represents the difference set operator of the set.
4. The urban confrontation environment-based safety reinforcement learning collaboration method for unmanned swarm systems according to claim 2, wherein, The calculation formula of the step of calculating and processing the observation control gain constraint value based on a preset control gain algorithm is as follows: ; ; wherein is the target observation control gain; is an identity matrix; is a first constant matrix; is a transpose of is a positive gain constant; ; is a minimum eigenvalue function; is an information delivery matrix after the unmanned swarm system is subjected to a denial-of-service attack; is a transpose of ; is an operator for constructing a diagonal matrix; is an inverse of a first diagonal element in the diagonal matrix; is an inverse of an th diagonal element in the diagonal matrix; is a positive definite parameter matrix related to the distributed resilient observer control gain.
5. The urban confrontation environment-based unmanned swarm system security reinforcement learning collaboration method according to claim 4, characterized in that, The calculation formula of the step of constructing the distributed resilient observer based on a communication topology adjacency weight between the first follower unmanned vehicle and the second follower unmanned vehicle, the system state of the leader vehicle, and the target observation control gain is as follows: ; wherein, is the first observation state; is the observation; is a target observation control gain; is a neighbor set of the first follower unmanned vehicle; is a denial of service attack model; is a communication topology adjacency weight between the first follower unmanned vehicle and the second follower unmanned vehicle; is the second observation state; is an expected platoon error between the first follower unmanned vehicle and the second follower unmanned vehicle; is a connection weight between the leader vehicle and the first follower unmanned vehicle; is a system state of the leader vehicle; is an expected platoon error between the first follower unmanned vehicle and the leader vehicle.
6. The urban confrontation environment-based unmanned swarm system security reinforcement learning collaboration method according to claim 5, characterized in that, The calculation formula of the denial-of-service attack model is as follows: ; ; wherein, is a set of communication links of the leader vehicle and all follower unmanned vehicles; is a set of communication links of the first follower unmanned vehicle and the second follower unmanned vehicle; is a set of communication links of the first follower unmanned vehicle and the second follower unmanned vehicle; is a union of time intervals in which the communication links of the first follower unmanned vehicle and the second follower unmanned vehicle are subject to denial-of-service attacks; is an initial time instant at which the unmanned swarm system operates; is a current time instant; is a fixed attack effect by which the communication links of the first follower unmanned vehicle and the second follower unmanned vehicle are allowed to be subject to denial-of-service attacks; is an attack intensity by which the communication links of the first follower unmanned vehicle and the second follower unmanned vehicle are subject to denial-of-service attacks.
7. The urban counter environment unmanned swarm system security reinforcement learning collaboration method according to claim 1, characterized in that, The step of performing safety estimation processing on the state of the leader vehicle based on the observation result to obtain a safety estimation result includes: In a case where the observation result is a derivative of the first observation state, it is determined that the state of the leader vehicle belongs to a safe state, and the safety estimation result is safe; otherwise, the state of the leader vehicle does not belong to a safe state, and the safety estimation result is unsafe.
8. The urban counter environment unmanned swarm system security reinforcement learning collaboration method according to claim 1, characterized in that, The calculation formula of the step of performing cumulative calculation processing based on the optimization gain matrix and the system state of the augmented system of the first follower unmanned vehicle to obtain optimization control data is as follows: ; wherein, is the optimized control data; is the optimized gain matrix.
Citation Information
Patent Citations
Distributed security formation control method for cross-domain cluster system in urban interference environment
CN120686895A