Multi-UAV cooperative interception method based on graph information consistency and deep learning
Through the graph information consistency algorithm and deep learning method, the unmanned boat position information extraction and global position information sharing are designed, which solves the problems of large computational complexity and low efficiency in large-scale unmanned boat systems and realizes efficient multi-unmanned boat collaborative interception.
Patent Information
- Application Number
- CN202510711588.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-05-29
AI Technical Summary
The existing multi-unmanned boat collaborative interception method has a large amount of computation in large-scale systems, which is difficult to meet the computing power and resource conditions. In addition, the existing method requires global information and has high computational complexity.
Through the graph information consistency algorithm and deep learning method, the unmanned boat position information extraction and global position information sharing are designed, and the multi-layer perceptron and strategy network are used for regulation. The state feedback control law and artificial potential field method are combined to control the unmanned boat motion, reducing the amount of calculation and improving the operation efficiency.
It has achieved a significant reduction in computing power in large-scale unmanned boat systems, high operating efficiency, and the ability to effectively coordinate and intercept invading unmanned boats, reducing information processing dimensions and improving the accuracy of global position information sharing.
Smart Images

Figure CN120233781B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of control and information technology, and particularly relates to a multi-unmanned boat collaborative interception method based on graph information consistency and deep learning. Background Art
[0002] In the past few years, deep learning-based multi-agent collaborative control technology has made significant progress, enabling more and more complex tasks to be solved and optimized through multi-agent systems, such as multi-target interception, area defense tasks, and intrusion detection.
[0003] In the coordinated interception problem of multiple unmanned aerial vehicles (UAVs), defending UAVs must coordinate to effectively intercept as many intruders as possible, while intruders must employ intelligent strategies to avoid interception and reach their target area. It is worth noting that existing methods often require global information, requiring real-time access to the states of all UAVs (position, velocity, remaining energy) and intruder trajectory predictions. This imposes a significant computational burden when dealing with large-scale UAV systems. Furthermore, these computational requirements and resources are difficult to meet in the coordinated interception problem of large-scale multi-UAV systems.
[0004] Against this backdrop, we propose a collaborative interception method for multiple unmanned aerial vehicles (UAVs) based on graph information consistency and deep learning. This method first represents the UAVs' positions as a map, then uses a graph information consistency algorithm to obtain global position information. Finally, deep learning is used to train the network to obtain control information.
[0005] The difficulty in designing a multi-UAV collaborative interception method based on graph information consistency and deep learning lies in:
[0006] First: How to design a technical framework to extract the location information of multiple unmanned boats and apply subsequent algorithms;
[0007] Second: How to design an algorithm to achieve global location information sharing with low computational cost?
[0008] Third: How to design a simple loss function to train the network and output the optimal control information to guide the unmanned boat to perform position changes. Summary of the Invention
[0009] The purpose of this invention is to propose a multi-unmanned boat collaborative interception method based on graph information consistency and deep learning. Through the graph information consistency algorithm and deep learning method, the computational complexity is greatly reduced and the operation efficiency is high.
[0010] To achieve the above objectives, the technical solution of the present invention is: a multi-unmanned boat collaborative interception method based on graph information consistency and deep learning, which specifically includes the following steps:
[0011] Step 1: Build a mission environment, which includes multiple defense unmanned boats and multiple intrusion unmanned boats in a set area; the goal of the defense unmanned boats is to collaboratively intercept the intrusion unmanned boats and prevent them from reaching the target area; the goal of the intrusion unmanned boats is to avoid the nearest defense unmanned boats while reaching the target area at the fastest speed; all defense unmanned boats and intrusion unmanned boats are set to have second-order dynamic models; the maximum speed of all unmanned boats is equal; execute steps 2 to 5 once at each sampling moment of the mission time until the mission time ends; execute steps 2 to 5 once at each sampling moment of the mission time until the mission time ends;
[0012] Step 2: Each defense drone i Build map information from your own perspective X i , the map information X i Set in the form of a graph matrix;
[0013] Step 3: Each defense unmanned boat communicates with its neighbor defense unmanned boats according to the communication constraints and runs the graph information consistency algorithm to make the map information constructed by each defense unmanned boat consistent. X i Converging to the true global graph information X ;
[0014] Step 4: Based on the multi-layer perceptron, i Designing a policy network W i , the policy network W i Input for defense unmanned boat i Constructed map information X i , the output is the control information Y i ;
[0015] Step 5: Based on regulatory information Y i , for map information X i Interior defense unmanned boat i Operate the location to get new map information ; while defending against unmanned boats i According to its own map location and control information Y i Select the next destination T i , correct its own speed direction and run towards the target position.
[0016] Preferably, each unmanned boat (defense or invasion) is designed based on the idea of feedback control and the state feedback control law, that is, all defense unmanned boats and invasion unmanned boats realize the control of speed and position through the second-order dynamic model and the state feedback control law; the mission environment simulation setting adopts a fixed sampling time interval , the specific dynamic model is:
[0017]
[0018] in, and It is a defense unmanned boat i The position, velocity and acceleration at time t, It is a defense unmanned boat i exist Position and speed at a given moment; and It's an invading unmanned boat. j Position, velocity, and acceleration at time t; It's an invading unmanned boat. j exist Position and speed at a given moment.
[0019] Preferably, the invading unmanned boat is controlled by artificial potential field method: target area P T For each invading unmanned boat j There is attraction and invasion of unmanned boats j Treat all defensive drones as obstacles and invade drones j The combined force F t Target attraction F a and the repulsive force of all defense drones F r The synthesis of F t , the position and speed of the invading unmanned boat at each sampling time interval Then update iteratively.
[0020] Preferably, target attractiveness F a Expressed as:
[0021]
[0022] in, k a is the coefficient of target attractiveness, P j It's an invading unmanned boat. j Current location, P T is the target area;
[0023] repulsive force F r Expressed as:
[0024]
[0025] in, k r is the coefficient of repulsive force, is a small constant to avoid division by zero errors, P i It is a defense unmanned boat i location, S is the number of defensive unmanned boats;
[0026] Combined Force F t Expressed as:
[0027]
[0028] The position and speed of the invading unmanned boat at each sampling time interval The update iteration is expressed as:
[0029]
[0030] Where, Indicates intrusion of unmanned boats j exist Position and speed at a given moment.
[0031] Optimum, defense unmanned boat i Constructed map information X i for The matrix of is the resolution of the map, corresponding to the set area of the mission environment; each matrix element Represents the matrix m Rank n The information corresponding to the column position is set as:
[0032]
[0033] in, m and n Used to traverse each element in the matrix, .
[0034] Preferably, the step 3 is specifically as follows:
[0035] Each defense drone i Initialize map information The map information built for yourself is set according to the position of the invading unmanned boat and the defending unmanned boat within your observation range. , if the first pixel of the map m Rank n If there is an intruding unmanned boat at the following location , if there is no unmanned boat , if there is a defense unmanned boat ;
[0036] Set the communication radius to r, and consider all defense UAVs with a distance less than r to be neighbor defense UAVs;
[0037] The graph information consistency algorithm includes multiple rounds of update iterations. In the kth round of update iteration, each defense unmanned boat i Broadcast its own map information in the current iteration Provide defense information for neighbors against unmanned boats and receive map information for neighboring defense against unmanned boats , where the set Defense Unmanned Boat i The collection of all neighbor defense unmanned boats; after receiving the information of neighbor defense unmanned boats, the defense unmanned boat i Your own map information and every neighbor I Map information Perform Boolean logic operations to update map information:
[0038]
[0039] in, It is a Boolean logic operation used for information fusion;
[0040] Obtaining Defense Drones i In the k +1 round of updated map information ; Continue to repeat the update iteration process until the map information no longer changes.
[0041] Preferably, the step 4 is specifically as follows:
[0042] Strategy Network W i Enter the Defense Unmanned Boat i Map information X i , where map information X i Flatten into a vector to fit the neural network input format:
[0043]
[0044] Output defense unmanned boati Regulatory information Y i , the value range is , corresponding to the defense unmanned boat i Choose between stationary, north, northwest, west, southwest, south, southeast, east, and northeast movement;
[0045] The strategy network is composed of a multi-layer perceptron. The output layer has 9 neurons, each corresponding to The Softmax activation function is used to convert the output into a probability distribution, and the value corresponding to the neuron with the highest probability is selected as the output value.
[0046] Preferably, the step 5 is specifically as follows:
[0047] Defense unmanned boat i According to regulatory information Y i Modify current map information X i , the operation is as follows:
[0048] when Keep the current location unchanged, map information X i constant;
[0049] when Adjust the defender's position and update the map information at the same time X i ; For each defense drone i , new map information Expressed as:
[0050]
[0051] in Based on regulatory information Y i Operation function to update the map;
[0052] Defense unmanned boat i Based on current location and control information Y i Select the target location at the current moment T i And correct the speed direction; for each defense unmanned boat, the target position is determined by the minimum Euclidean distance:
[0053]
[0054] in It is a defense unmanned boat iCurrent location, They are the position components of the current position in two directions respectively; It is a defense unmanned boat i The target location, are the position components of the target position in two directions respectively;
[0055] To correct the defense of unmanned boats i The velocity direction of the update formula is expressed as:
[0056]
[0057] in, v i It is a defense unmanned boat i Velocity vector before correction, is the maximum acceleration of the defense unmanned boat, is the sampling time interval, is a unit vector pointing to the target position, indicating the defense unmanned boat i Velocity direction:
[0058]
[0059] Defense unmanned boat i The actual position information update formula is:
[0060]
[0061] in It is a defense unmanned boat i Updated location information, is the corrected speed.
[0062] Preferably, the strategy network is trained based on the PPO algorithm w i , the training goal is to minimize the map information X i The Frobenius norm of the matrix obtained after the convolution and summation operation.
[0063] Preferably, the PPO algorithm is used to train the strategy network W i The specific steps are as follows:
[0064] (1) Setting up a strategy network for each defense drone W i Input map information built for itself X i , flattened into a vector , to adapt to the input format of the neural network and output control information , corresponding to the action choices of being still or moving in eight directions;
[0065] (2) During each round of training, the defense unmanned boat interacts with the simulation environment and collects data from the state ,action Y i ,award , action probability and the next state Interaction data composed of constructing training samples; Refers to the map information flattened into vectors, actions Y i It refers to regulatory information;
[0066] (3) Map information X i Convolution operation is performed to evaluate the cooperative interception effect of the defense unmanned boat in the current state; convolution summation operation is performed on the map information X i Set the convolution kernel to 3*3, for position The convolution operation is expressed as:
[0067]
[0068] in Is the position after convolution The result at indicates the position and the sum of its neighborhood;
[0069] (4) Calculate the Frobenius norm based on the convolution result to measure the aggregation and effectiveness of the current map state. The Frobenius norm calculation formula is:
[0070]
[0071] (5) Use the clipping strategy in the PPO algorithm to optimize the objective function:
[0072]
[0073] in, Represents the time step t expectations on represents the advantage function; represents the probability ratio of the new and old strategies, , represents the new policy action probability, represents the old policy action probability; Restricting the probability ratio to the interval To prevent the policy update from being too large, is the clipping threshold;
[0074] (6) Construct a joint optimization objective function, combining the PPO objective with the Frobenius norm term to form the following training objective:
[0075]
[0076] in, represents the joint objective function, The weight coefficient for regulating the influence of the Frobenius norm;
[0077] (7) Use stochastic gradient descent optimization algorithm to optimize the policy network parameters Perform iterative updates to maximize the joint objective function , realize the policy network W i training.
[0078] Compared with the prior art, the present invention has the following beneficial effects:
[0079] (1) By extracting the information features of the intruder and defender at different locations on the map as map information X i Effectively reduces the dimension of information processing;
[0080] (2) Through the graph information consistency algorithm, each defender can obtain global intrusion and defender location information;
[0081] (3) Training strategy network based on PPO algorithm W i , map information X i The Frobenius norm of the matrix obtained after the convolution and summation operation is used as the LOSS function for training, which makes the training objectives clear and the physical meaning clear. BRIEF DESCRIPTION OF THE DRAWINGS
[0082] Figure 1 Flow chart of the method of the present invention. DETAILED DESCRIPTION
[0083] The following is combined with Figure 1 , the technical solution of the present invention is described in detail.
[0084] The present invention proposes a multi-unmanned boat collaborative interception method based on graph information consistency and deep learning, which specifically includes the following steps:
[0085] Step 1: Build a mission environment, which includes multiple defense unmanned boats and multiple intrusion unmanned boats in a set area; the goal of the defense unmanned boats is to collaboratively intercept the intrusion unmanned boats and prevent them from reaching the target area; the goal of the intrusion unmanned boats is to avoid the nearest defense unmanned boats while reaching the target area at the fastest speed; all defense unmanned boats and intrusion unmanned boats are set to have second-order dynamic models; the maximum speed of all unmanned boats is equal; execute steps 2 to 5 once at each sampling moment of the mission time until the mission time ends; execute steps 2 to 5 once at each sampling moment of the mission time until the mission time ends;
[0086] Step 2: Each defense drone i Build map information from your own perspective X i , the map information X i Set in the form of a graph matrix;
[0087] Step 3: Each defense unmanned boat communicates with its neighbor defense unmanned boats according to the communication constraints and runs the graph information consistency algorithm to make the map information constructed by each defense unmanned boat consistent. X i Converging to the true global graph information X ;
[0088] Step 4: Based on the multi-layer perceptron, i Designing a policy network W i , the policy network W i Input for defense unmanned boat i Constructed map information X i , the output is the control information Y i ;
[0089] Step 5: Based on regulatory information Y i , for map information X i Interior defense unmanned boat i Operate the location to get new map information ; while defending against unmanned boats i According to its own map location and control information Y i Select the next destination T i , correct its own speed direction and run towards the target position.
[0090] In this embodiment, each UAV (defense or invasion) is designed with a state feedback control law based on the feedback control concept. That is, all defense UAVs and invasion UAVs use a second-order dynamic model and state feedback control law to control speed and position. The mission environment simulation setting uses a fixed sampling time interval. , the specific dynamic model is:
[0091]
[0092] in, and It is a defense unmanned boat i The position, velocity and acceleration at time t, It is a defense unmanned boat i exist Position and speed at a given moment; and It's an invading unmanned boat. j Position, velocity, and acceleration at time t; It's an invading unmanned boat. j exist Position and speed at a given moment.
[0093] In this embodiment, the invading unmanned boat is controlled by the artificial potential field method: the target area P T For each invading unmanned boat j There is attraction and invading unmanned boats j Treat all defensive drones as obstacles and invade drones j The combined force F t Target attraction F a and the repulsive force of all defense drones F r The synthesis of F t , the position and speed of the invading unmanned boat at each sampling time interval Then update iteratively.
[0094] In this embodiment, the target attractiveness F a Expressed as:
[0095]
[0096] in, k a is the coefficient of target attractiveness, P j It's an invading unmanned boat. j Current location, P Tis the target area;
[0097] repulsive force F r Expressed as:
[0098]
[0099] in, k r is the coefficient of repulsive force, is a small constant to avoid division by zero errors, P i It is a defense unmanned boat i location, S is the number of defensive unmanned boats;
[0100] Combined Force F t Expressed as:
[0101]
[0102] The position and speed of the invading unmanned boat at each sampling time interval The update iteration is expressed as:
[0103]
[0104] Where, Invasion of unmanned boats j exist Position and speed at a given moment.
[0105] Defense unmanned boat i The goal is to coordinate with other defense UAVs to intercept the invading UAVs by adjusting their positions. The motion control of the defense UAV depends on the dynamic information of its relative position and the invading UAV. The maximum speed of all UAVs in the environment is set to ,Right now:
[0106] .
[0107] In this embodiment, the defense unmanned boat i Constructed map information X i for The matrix of is the resolution of the map, corresponding to the set area of the mission environment; each matrix element Represents the matrix m Rank n The information corresponding to the column position is set as:
[0108]
[0109] in, m and n Used to traverse each element in the matrix, .
[0110] In this embodiment, step 3 is specifically as follows:
[0111] Each defense drone i Initialize map information The map information built for yourself is set according to the position of the invading unmanned boat and the defending unmanned boat within your observation range. , if the first pixel of the map m Rank n If there is an intruding unmanned boat at the following location , if there is no unmanned boat , if there is a defense unmanned boat ;
[0112] Set the communication radius to r, and consider all defense UAVs with a distance less than r to be neighbor defense UAVs;
[0113] The graph information consistency algorithm includes multiple rounds of update iterations. In the kth round of update iteration, each defense unmanned boat i Broadcast its own map information in the current iteration Provide defense information for neighbors against unmanned boats and receive map information for neighboring defense against unmanned boats , where the set Defense Unmanned Boat i The collection of all neighbor defense unmanned boats; after receiving the information of neighbor defense unmanned boats, the defense unmanned boat i Your own map information and every neighbor I Map information Perform Boolean logic operations to update map information:
[0114]
[0115] in, Indicates from neighbors I The collection of all map information received, It is a Boolean logic operation used for information fusion; according to the fusion rules, the defense unmanned boat i In the k The map information after +1 round of updates is:
[0116]
[0117] Continue to repeat the update iteration process until the map information no longer changes.
[0118] In this embodiment, step 4 is specifically as follows:
[0119] Strategy Network W i Enter the Defense Unmanned Boat i Map information X i , where map information X i Flatten into a vector to fit the neural network input format:
[0120]
[0121] Output defense unmanned boat i Regulatory information Y i , the value range is , corresponding to the defense unmanned boat i Choose between stationary, north, northwest, west, southwest, south, southeast, east, and northeast movement;
[0122] The strategy network is composed of a multi-layer perceptron. The output layer has 9 neurons, each corresponding to The Softmax activation function is used to convert the output into a probability distribution, and the value corresponding to the neuron with the highest probability is selected as the output value.
[0123] In this embodiment, the step 5 is specifically as follows:
[0124] Defense unmanned boat i According to regulatory information Y i Modify current map information X i , the operation is as follows:
[0125] when Keep the current location unchanged, map information X i constant;
[0126] when Adjust the defender's position and update the map information X i ; For each defense drone i , new map information Expressed as:
[0127]
[0128] in Based on regulatory information Y i Operation function to update the map;
[0129] Defense unmanned boat i According to the current location and control information and control information Y i Select the target location at the current moment T i And correct the speed direction; for each defense unmanned boat, the target position is determined by the minimum Euclidean distance:
[0130]
[0131] in It is a defense unmanned boat i Current location, They are the position components of the current position in two directions respectively; It is a defense unmanned boat i The target location, are the position components of the target position in two directions respectively;
[0132] To correct the defense of unmanned boats i The velocity direction of the update formula is expressed as:
[0133]
[0134] in, v i It is a defense unmanned boat i Velocity vector before correction, is the maximum acceleration of the defense unmanned boat, is the sampling time interval, is a unit vector pointing to the target position, indicating the defense unmanned boat i Velocity direction:
[0135]
[0136] Defense unmanned boat i The actual position information update formula is:
[0137]
[0138] in It is a defense unmanned boat i Updated location information, is the corrected speed.
[0139] In this embodiment, the strategy network is trained based on the PPO algorithm. w i , the training goal is to minimize the map information X iThe Frobenius norm of the matrix obtained after the convolution and summation operation.
[0140] In this embodiment, the strategy network is trained based on the PPO algorithm. W i The specific steps are as follows:
[0141] (1) Setting up a strategy network for each defense drone W i Input map information built for itself X i , flattened into a vector , to adapt to the input format of the neural network and output control information , corresponding to the action choices of being still or moving in eight directions;
[0142] (2) During each round of training, the defense unmanned boat interacts with the simulation environment and collects data from the state ,action Y i ,award , action probability and the next state Interaction data composed of constructing training samples; Refers to the map information flattened into vectors, actions Y i It refers to regulatory information;
[0143] (3) Map information X i Convolution operation is performed to evaluate the cooperative interception effect of the defense unmanned boat in the current state; convolution summation operation is performed on the map information X i Set the convolution kernel to 3*3, for position The convolution operation is expressed as:
[0144]
[0145] in Is the position after convolution The result at indicates the position and the sum of its neighborhood;
[0146] (4) Calculate the Frobenius norm based on the convolution result to measure the aggregation and effectiveness of the current map state. The Frobenius norm calculation formula is:
[0147]
[0148] (5) Use the clipping strategy in the PPO algorithm to optimize the objective function:
[0149]
[0150] in, Represents the time step t expectations on represents the advantage function; represents the probability ratio of the new and old strategies, , represents the new policy action probability, represents the old policy action probability; Restricting the probability ratio to the interval To prevent the policy update from being too large, is the clipping threshold;
[0151] (6) Construct a joint optimization objective function, combining the PPO objective with the Frobenius norm term to form the following training objective:
[0152]
[0153] in, represents the joint objective function, The weight coefficient for regulating the influence of the Frobenius norm;
[0154] (7) Use stochastic gradient descent optimization algorithm to optimize the policy network parameters Perform iterative updates to maximize the joint objective function , realize the policy network W i training.
[0155] At this point, all steps are completed. The present invention studies how to design a multi-unmanned boat collaborative interception method based on graph information consistency and deep learning. This method can be used to train a multi-unmanned boat system to collaboratively intercept multiple target invading unmanned boats. For the problem of multi-target interception, the mainstream existing algorithms are matching-based methods or differential game-based methods. However, these methods often require global information, and the amount of calculation is huge when dealing with large-scale unmanned boat systems. The designed new method greatly reduces the amount of calculation and has high operating efficiency through the graph information consistency algorithm and deep learning method.
[0156] The above are preferred embodiments of the present invention. Any changes made according to the technical solution of the present invention, as long as the resulting functions and effects do not exceed the scope of the technical solution of the present invention, shall fall within the scope of protection of the present invention.
Claims
1. A multi-unmanned boat collaborative interception method based on graph information consistency and deep learning, characterized by: The specific steps include: Step 1: Build a mission environment, which includes multiple defense UAVs and multiple intrusion UAVs in a set area. The goal of the defense UAVs is to collaboratively intercept the intrusion UAVs and prevent them from reaching the target area. The goal of the intrusion UAVs is to reach the target area at the fastest speed while avoiding the nearest defense UAVs. All defense UAVs and intrusion UAVs are assumed to have second-order dynamic models. The maximum speed of all UAVs is equal. Steps 2 to 5 are executed once at each sampling time of the mission time until the mission time ends. Step 2: Each defense unmanned boat i constructs map information X from its own perspective i , the map information X i Set in the form of a graph matrix; Step 3: Each defense unmanned boat communicates with its neighboring defense unmanned boats according to the communication constraints and runs the graph information consistency algorithm so that the map information X constructed by each defense unmanned boat is consistent with the original map information. i Converge to the true global graph information X; Step 4: Design a strategy network W for each defense unmanned boat i based on the multi-layer perceptron i , the policy network W i Input the map information X constructed for defending against unmanned boat i i , the output is the control information Y i ; Step 5: Based on the control information Y i , for map information X i Operate the position of the internal defense unmanned boat i to obtain new map information At the same time, the defense unmanned boat i is based on its own map position and control information Y i Select the next target position T i , correct its own speed direction and run towards the target position; The step 3 is specifically as follows: Each defense unmanned boat i initializes map information The map information built for yourself is set according to the position of the invading unmanned boat and the defending unmanned boat within your observation range. If there is an invading unmanned boat at the position of the mth row and nth column of the map pixel, then No unmanned boats If there is a defensive unmanned boat Set the communication radius to r, and consider all defense UAVs with a distance less than r to be neighbor defense UAVs; The graph information consistency algorithm includes multiple rounds of update iterations. In the kth round of update iteration, each defense unmanned boat i broadcasts its own map information in the current round of iteration. Provide defense information for neighbors against unmanned boats and receive map information for neighboring defense against unmanned boats The set N(i) represents the set of all neighboring defense unmanned boats of the defense unmanned boat i. After receiving the information of the neighboring defense unmanned boat, the defense unmanned boat i will send its own map information to the defense unmanned boat i. and map information of each neighbor I Perform Boolean logic operations to update map information: Where f(·) is a Boolean logic operation used for information fusion; Get the map information of the defense unmanned boat i in the k+1th round of update iteration Continue to repeat the update iteration process until the map information no longer changes.
2. The multi-unmanned boat cooperative interception method based on graph information consistency and deep learning according to claim 1 is characterized in that: Each unmanned boat is designed with a state feedback control law based on the idea of feedback control. That is, all defense unmanned boats and intrusion unmanned boats use the second-order dynamic model and state feedback control law to control speed and position. The mission environment simulation setting uses a fixed sampling time interval ∈. The specific dynamic model is: P i (t+∈)=P i (t)+∈v i (t) v i (t+∈)=v i (t)+∈a i (t) P j (t+e)=P j (t)+∈v j (t) v j (t+∈)=v j (t)+∈a j (t) Among them, P i (t), v i (t) and a i (t) is the position, velocity and acceleration of the defense unmanned boat i at time t, P i (t+∈),v i (t+∈) is the position and speed of the defense unmanned boat i at time t+∈; P j (t), v j (t) and a j (t) is the position, velocity and acceleration of the invading unmanned boat j at time t; P j (t+∈),v j (t+∈) is the position and velocity of the invading unmanned boat j at time t+∈.
3. The multi-unmanned boat cooperative interception method based on graph information consistency and deep learning according to claim 1 is characterized in that: The invading unmanned boat is controlled by artificial potential field method: the target area P T There is an attraction for each invading unmanned boat j, and the invading unmanned boat j regards all the defending unmanned boats as obstacles. The combined force F of the invading unmanned boat j is t is the target attraction F a and the repulsive force F of all defense unmanned boats r The synthesis of the force F t , the position and velocity of the invading unmanned boat are updated iteratively after each sampling time interval ∈.
4. The multi-unmanned boat cooperative interception method based on graph information consistency and deep learning according to claim 3 is characterized in that: Target attractiveness F a Expressed as: F a =-k a (P j -P T ) Among them, k a is the coefficient of target attractiveness, P j is the current position of the invading UAV j, P T is the target area; Repulsive force F r Expressed as: Among them, k r is the coefficient of the repulsive force, ε is a small constant to avoid division by zero errors, P i is the position of the defense unmanned boat i, S is the number of defense unmanned boats; Combined Force F t Expressed as: F t =F a +F r The position and velocity of the invading unmanned boat are updated iteratively after each sampling time interval ∈ as follows: P j (t+∈)=P j (t)+∈v j (t) v j (t+∈)=v j (t)+∈F t Where, P j (t+∈) and v j (t+∈) represents the position and speed of the invading unmanned boat j at time t+∈.
5. The multi-unmanned boat cooperative interception method based on graph information consistency and deep learning according to claim 1 is characterized in that: Map information constructed by defense unmanned boat i i It is an M×N matrix, where M×N is the resolution of the map, corresponding to the set area of the mission environment; each matrix element [X i ] m,n Represents the information corresponding to the position of the mth row and nth column of the matrix, which is set as: Here, m and n are used to traverse each element in the matrix, m=1, 2, ..., M, n=1, 2, ..., N.
6. The multi-unmanned boat cooperative interception method based on graph information consistency and deep learning according to claim 1 is characterized in that: The step 4 is specifically as follows: Strategy Network W i Enter the map information of the defense unmanned boat i X i , where map information X i Flatten into a vector to fit the neural network input format: Output the control information Y of the defense unmanned boat i i , the value range is {0, 1, 2, ..., 8}, corresponding to the operation of selecting the defense unmanned boat i to stay still, move north, northwest, west, southwest, south, southeast, east, and northeast; The policy network is composed of a multi-layer perceptron. The output layer has 9 neurons, each of which corresponds to a value from 0, 1, 2, ..., 8. The Softmax activation function is used to convert the output into a probability distribution, and the value corresponding to the neuron with the highest probability is selected as the output value.
7. The multi-unmanned boat cooperative interception method based on graph information consistency and deep learning according to claim 2 is characterized in that: The step 5 is specifically as follows: Defense unmanned boat i according to control information Y i Modify current map information X i , the operation is as follows: When Y i =0 when the current position remains unchanged, map information X i constant; When Y i ={1, 2, 3..., 8}, adjust the defender's position and update the map information X i ; For each defense drone i, new map information Expressed as: Where g is based on the control information Y i Operation function to update the map; Defense unmanned boat i based on current location and control information Y i Select the target position T at the current moment i And correct the speed direction; for each defense unmanned boat, the target position is determined by the minimum Euclidean distance: Among them, P i =(P i,x , P i,y ) is the current position of the defense unmanned boat i, P i,x 、P i,y are the position components of the current position in two directions respectively; T i =(T i,x , T i,y ) is the target position of the defense unmanned boat i, T i,x 、T i,y are the position components of the target position in two directions respectively; To correct the speed direction of the defense unmanned boat i, the update formula is expressed as: Among them, v i is the velocity vector of the defense unmanned boat i before correction, a max is the maximum acceleration of the defense unmanned boat, ∈ is the sampling time interval, is a unit vector pointing to the target position, indicating the velocity direction of the defense unmanned boat i: The actual position information update formula of the defense unmanned boat i is: in It is the updated position information of the defense unmanned boat i. is the corrected speed.
8. The multi-unmanned boat cooperative interception method based on graph information consistency and deep learning according to claim 7 is characterized in that: Training policy network w based on PPO algorithm i , the training goal is to minimize the map information X i The Frobenius norm of the matrix obtained after the convolution and summation operation.
9. The multi-unmanned boat cooperative interception method based on graph information consistency and deep learning according to claim 8 is characterized in that: The PPO algorithm is used to train the policy network W i The specific steps are as follows: (1) Set the strategy network W for each defense UAV i Input the map information X built for itself i , flattened into a vector Output control information Y in an input format that adapts to the neural network i ∈{0, 1, ..., 8}, corresponding to the action selection of being still or moving in eight directions; (2) During each round of training, the defense unmanned boat interacts with the simulation environment and collects data from the state Action Y i , reward r t , action probability and the next state X′ i Interaction data composed of constructing training samples; Refers to the map information flattened into vectors, action Y i It refers to regulating information; (3) Map information X i Perform convolution operation to evaluate the cooperative interception effect of the defense unmanned boat in the current state; perform convolution summation operation on the map information X i Set the convolution kernel to 3*3, and the convolution operation at position (m,n) is expressed as: where Z i,m,n It is the result at position (m,n) after convolution, which represents the sum of the position (m,n) and its neighborhood; (4) Calculate the Frobenius norm based on the convolution result to measure the aggregation and effectiveness of the current map state. The Frobenius norm calculation formula is: (5) Use the clipping strategy in the PPO algorithm to optimize the objective function: L ppo (θ)=E t [min(A t r t (θ), clip(r t (θ), 1-Δ, 1+α)A t )] Among them, E t [·] represents the expectation at time step t; A t represents the advantage function; r t (θ) represents the probability ratio of the new and old strategies, represents the new policy action probability, represents the old strategy action probability; clip(r t (θ), 1-Δ, 1+Δ) means limiting the probability ratio to the interval [1-Δ, 1+Δ] to prevent the policy update from being too large, and Δ is the clipping threshold; (6) Construct a joint optimization objective function, combining the PPO objective with the Frobenius norm term to form the following training objective: L i (θ)=L ppo (θ)-α||X i || F Among them, L i (θ) represents the joint objective function, and α>0 is the weight coefficient for regulating the influence of the Frobenius norm; (7) Use the stochastic gradient descent optimization algorithm to iteratively update the policy network parameters θ to maximize the joint objective function L i (θ), realize the policy network W i training.
Citation Information
Patent Citations
Unmanned ship cluster collaborative search method and system integrating global and local decisions
CN118605524A
A control method for aircraft / ship cooperative triggered communication path tracking based on grid search
JP7576373B1