Multi-unmanned-boat Area Patrol Method and System in an Adversarial Environment Based on Situation Information
By constructing situation information and using MADDPG algorithm, the coordination problems and security threats in unmanned boat patrols have been solved, and efficient regional patrols and safety improvements of multiple unmanned boats have been achieved.
Patent Information
- Application Number
- CN202310379144.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-10
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2043-04-10
AI Technical Summary
In a complex confrontational environment, the patrol range of a single unmanned boat is limited, making it difficult to achieve coordinated patrol between multiple unmanned boats, and early algorithms are easily exploited by the enemy, resulting in security threats.
By constructing global situation information and local situation information, combining trajectory prediction and MADDPG algorithm, action instructions are determined to realize regional patrol tasks of multiple unmanned boats.
It improves the robustness of the algorithm, realizes coordinated patrol between multiple unmanned boats, and enhances patrol safety in the confrontational environment.
Smart Images

Figure CN116382285B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of unmanned boat patrol, and particularly relates to a method and system for multi-unmanned-boat regional patrol in an adversarial environment based on situation information. Background Art
[0002] An unmanned surface vehicle (USV), as a surface mission platform, has advantages such as small size, strong concealment, and low cost compared with traditional large ships. In a complex adversarial environment, researching the use of an unmanned boat cluster to replace manned ships to perform patrol tasks in dangerous waters has significant strategic value that cannot be ignored.
[0003] Patrol refers to a continuous movement process for the purpose of safety in an environment. In an adversarial environment, using unmanned boats to perform patrol tasks around our mother boat, the goal of the task is to detect other boats invading our patrol task area as early as possible to prevent enemy boats from entering the safety distance of our mother boat. In actual tasks, our patrol boats also need to use technical means such as taking pictures of the discovered non-friendly boats to identify the attributes of the boats, so as to distinguish enemy boats and neutral boats such as fishing boats. This requires us to keep the discovered other boats within our detection range for a period of time. In addition, due to the large patrol scene and the limited detection range of a single patrol boat, how to enable multiple unmanned boats to cooperate and complete coordinated patrol within the task area is also a difficult point.
[0004] The research on patrol tasks has been carried out for a long time. However, early algorithms are mostly limited to a single intelligent agent and are difficult to apply when there is a task requirement for multi-boat coordination. Second, a large number of rule-based algorithms pursue the minimum revisit time for each location in the area, which is easily exploited by enemy boats to find loopholes in the patrol rules and thus successfully invade the safety distance of our mother boat, posing a threat to the safety of our mother boat. Summary of the Invention
[0005] The purpose of the present invention is to provide a method and system for multi-unmanned-boat regional patrol in an adversarial environment based on situation information, so as to make effective decisions in a complex adversarial environment and complete the task requirements of regional coordinated patrol.
[0006] To achieve the above purpose, the present invention provides the following solutions:
[0007] A method for multi-unmanned-boat regional patrol in an adversarial environment based on situation information includes:
[0008] Constructing global situation information of the battlefield according to an information set; the information set includes: patrol task information, nautical chart information, position information of our own boats, and position information of the discovered enemy boats;
[0009] Construct a local nautical chart using the detection information of the sensors of the current vessel itself;
[0010] Predict the future movement trajectories of all vessels based on the movement trajectories of the vessels around the current vessel and the movement trajectory of the current vessel, to obtain trajectory prediction information;
[0011] Combine the trajectory prediction information with the local nautical chart to form local situation information;
[0012] Determine a feature vector based on the global situation information, the local situation information, and the vessel movement attribute vector;
[0013] Train the MADDPG algorithm using the feature vector and the reward function; the MADDPG algorithm includes an Actor network and a Critic network;
[0014] Determine an action instruction through the trained Actor network; perform a regional patrol mission for multiple unmanned vessels according to the action instruction.
[0015] Optionally, predicting the future movement trajectories of all vessels based on the movement trajectories of the vessels around the current vessel and the movement trajectory of the current vessel specifically includes:
[0016] Perform kinematic modeling on the vessels according to the CTRA model to determine the state variables of the vessels at the current moment;
[0017] Calculate the change rates of the state variables at each of the current moments;
[0018] Determine the state variables at the next moment according to the change rates;
[0019] Determine the future movement trajectories of all vessels according to the state variables at the next moment.
[0020] Optionally, determining the feature vector based on the global situation information, the local situation information, and the vessel movement attribute vector specifically includes:
[0021] Input the global situation information and the local situation information into the trained CNN network to obtain an initial feature vector;
[0022] Concatenate the initial feature vector with the vessel movement attribute vector to obtain a feature vector.
[0023] Optionally, training the MADDPG algorithm using the feature vector and the reward function specifically includes:
[0024] Input the feature vector into the Actor network to obtain an action instruction at the next moment;
[0025] Patrol according to the action instruction at the next moment, and evaluate the action instruction at the next moment through the reward function and the Critic network;
[0026] Adjust the parameters of the Actor network according to the evaluation result.
[0027] Optionally, the expression of the reward function R is as follows:
[0028] R = r unknown + r catch + r escape + r intrusion + r collision
[0029] where r unknown is the global unknown degree reward of the patrol area, r catch is the enemy ship capture reward, r escape is the enemy ship escape reward, r intrusion is the enemy ship intrusion reward, r collision is the collision reward.
[0030] Optionally, the expression of the global unknown degree reward r unknown of the patrol area is as follows:
[0031] r unknown = r 0 × K unknown
[0032]
[0033] where r 0 represents an adjustable parameter, K unknown represents the unknown degree of the global situation information, n unknown represents the value of the unknown degree of the global situation information, N omniscient represents the global data of the global situation information.
[0034] Optionally, the expression of the enemy ship capture reward r catch is as follows:
[0035]
[0036] where r catchdist represents the distance between the position of the enemy ship when it is calibrated and our mother boat, r globol represents the radius of the patrol area, r 1 represents an adjustable parameter.
[0037] The present invention also provides a multi-unmanned-boat area patrol system in an adversarial environment based on situation information, including:
[0038] The global situation information construction module is used to construct the global situation information of the battlefield according to the information set; the information set includes: patrol task information, nautical chart information, the position information of friendly ships, and the position information of detected enemy ships;
[0039] The local nautical chart construction module is used to construct a local nautical chart by using the sensor detection information of the current ship itself;
[0040] The trajectory prediction module is used to predict the future movement trajectories of all ships based on the movement trajectories of the ships around the current ship and the movement trajectory of the current ship, and obtain trajectory prediction information;
[0041] The local situation information determination module is used to combine the trajectory prediction information with the local nautical chart to form local situation information;
[0042] The feature vector determination module is used to determine a feature vector based on the global situation information, the local situation information, and the ship movement attribute vector;
[0043] The training module is used to train the MADDPG algorithm through the feature vector and the reward function; the MADDPG algorithm includes an Actor network and a Critic network;
[0044] The action instruction determination module is used to determine an action instruction through the trained Actor network; and perform the area patrol task of multiple unmanned boats according to the action instruction.
[0045] Optionally, the feature vector determination module specifically includes:
[0046] The initial feature vector determination unit is used to input the global situation information and the local situation information into the trained CNN network to obtain an initial feature vector;
[0047] The splicing unit is used to splice the initial feature vector with the ship movement attribute vector to obtain a feature vector.
[0048] Optionally, the training module specifically includes:
[0049] The input unit is used to input the feature vector into the Actor network to obtain the action instruction for the next moment;
[0050] The evaluation unit is used to perform patrol according to the action instruction for the next moment, and evaluate the action instruction for the next moment through the reward function and the Critic network;
[0051] The parameter adjustment unit is used to adjust the parameters of the Actor network according to the evaluation result.
[0052] According to the specific embodiments provided by the present invention, the following technical effects are disclosed by the present invention:
[0053] The present invention uses global situation information to represent the global patrol form, and by means of the degree of uncertainty, represents the past patrol situation of the entire patrol scene; uses local situation information to represent high-precision local environmental information, and combines the means of Kalman filter + CTRV model to predict the future trajectories of the vessels appearing around the mother vessel, so as to strive to provide more information in the local situation information. The present invention avoids the influence of factors such as the number of vessels and the environmental complexity on the model decision-making, greatly improves the robustness of the entire algorithm, and finally realizes the requirements of the multi-unmanned vessel area patrol task. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0055] Figure 1 It is a flowchart of a multi-unmanned vessel area patrol method based on situation information provided in Embodiment 1 of the present invention;
[0056] Figure 2 It is a schematic diagram of mission chart information;
[0057] Figure 3 It is an initial rasterized global situation information map constructed according to the chart information;
[0058] Figure 4 It is a global situation information map reflecting the position information of the vessels on both sides of the enemy and us;
[0059] Figure 5 It is a global situation information map after expressing the patrol historical information;
[0060] Figure 6 It is an initial local situation information map rasterized according to the sensor data;
[0061] Figure 7 It is a future trajectory area probability map predicted according to the vessel historical data;
[0062] Figure 8 It is a local situation information map after expressing the future trajectory area probability map on the initial local situation information map;
[0063] Figure 9 It is a CNN network structure diagram for processing situation information;
[0064] Figure 10 It is the training structure diagram of MADDPG. Specific implementation manners
[0065] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0066] The object of the present invention is to provide a method and system for multi-unmanned boat area patrol in an adversarial environment based on situation information. The global situation information of the battlefield is constructed by using the task information of area patrol, the chart information, the position information of friendly boats obtained through communication and the position information of detected enemy boats; a single unmanned boat constructs a local high-precision chart by using its own sensor detection information, and combines the trajectory prediction information of all other boats appearing in the local area to form local situation information; the two pieces of information are understood as a two-channel situation information map, which is processed by a CNN network to become high-dimensional information; finally, combined with the physical information of the navigation of its own boat, a complete state vector is formed; trained by means of the MADDPG algorithm, and finally the action instructions required for driving the unmanned boat are output, so as to complete the area patrol task of multi-unmanned boats.
[0067] To make the above objects, features and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0068] Embodiment 1
[0069] Embodiment 1 of the present invention provides a method for multi-unmanned boat area patrol in an adversarial environment based on situation information. As Figure 1 shown, the method includes the following steps:
[0070] S1: Construct the global situation information of the battlefield according to the information set; the information set includes: patrol task information, chart information, position information of friendly boats, and position information of detected enemy boats.
[0071] In practical applications, obtain task information from the task center, including the chart information of the area to be patrolled (as Figure 2 shown) and the information of the mother boat that needs to be guarded against intrusion in this patrol task, including the position of the mother boat, the safe distance that the mother boat cannot be broken through by enemy boats, and the physical characteristics of friendly boats, etc.
[0072] After receiving the received chart information and the position information of the mother ship, first establish a global situation information map. Determine the position of the mother ship that needs to be patrolled and guarded as the exact center of the global situation. Based on the size of the patrol area in the issued patrol task and combined with the size M of the previously set M*M matrix, determine the distance represented by each grid in the global situation information map. Then, according to the position information, size, and shape of the island reef relative to the mother ship, set the grid points occupied by the island reef on the global grid map representing the situation information to the value 255, and the points occupied by the mother ship to 225. Establish a mapping relationship from Figure 2 to Figure 3 To this point, the task information has been well expressed on the situation information map. Distribute this initial situation map to each patrol boat participating in the patrol task.
[0073] After a single patrol boat receives this initial global situation map, based on its own position information and the position information of friendly boats obtained through communication, set the grid where it is located on the grid map to 175, and the other friendly boats to 200, as Figure 4 shown.
[0074] Then, as the patrol task progresses, according to the time difference △t from the time T when each sea area in the task area was last patrolled to the present, through the calculation of the uncertainty function F(t), map it to a value in the range of 0 - 100, and reflect it on the global grid map. 0 represents completely unknown, and 100 represents completely known. The current detectable range value of our boat is 100. For the grids outside the detection range, calculate the difference from the current time based on the time stamp of the last patrol to the corresponding position reserved, and then solve for the uncertainty of the grid. When each boat updates its global situation information, in addition to updating the uncertainty around itself, it will also obtain the position information and perception information of its own unmanned boats through communication within its own cluster, and thus iterate them on the global situation information map. The final global situation information obtained by each boat is as Figure 5 shown.
[0075] S2: Use the sensor detection information of the current boat itself to construct a local chart.
[0076] In practical applications, first, based on the sensor detection range of the current boat and combined with the size M of the previously set M*M matrix, determine the distance represented by each grid in the local situation information. Then, with the center origin of the grid occupied by the current boat, with a value of 175, map it to the local situation map according to the island reef information sensed by the sensor. For the grids covered by the island reef, we define the value as 255. Then, for the boat information that appears around the current boat, the grid occupied by the mother ship is 225, the grids covered by other friendly boats of our side have a value of 200, and the enemy boats are 150. For the remaining unoccupied grids, set them to 0, and establish a local high-precision grid map as Figure 6 shown.
[0077] S3: Based on the motion trajectories of the ships around the current ship and the motion trajectory of the current ship, the future motion trajectories of all ships are predicted to obtain trajectory prediction information. Specifically, it includes: kinematic modeling of the ship according to the CTRA model to determine the state variables of the ship at the current moment; calculating the change rate of each state variable at the current moment; determining the state variable at the next moment according to the change rate; and determining the future motion trajectories of all ships according to the state variables at the next moment.
[0078] First, the kinematic modeling of the ship is carried out according to the CTRA model, with state variables:
[0079]
[0080] Where (x, y) represents the horizontal and vertical coordinates in the current coordinate system, with the position of our mother ship as the origin. v represents the current speed of the ship, θ represents the current yaw angle of the ship, a represents the current linear acceleration, and ω represents the current yaw angular velocity of the ship.
[0081] In the CTRA model, assuming that the ship has a constant acceleration and a constant yaw rate, that is, the rate of change of both is 0, the rate of change of each state quantity is as follows:
[0082]
[0083] By integrating the above formula, we can get the relationship between the state variable X and time:
[0084]
[0085] Combining equations (1) and (2), when ω k ≠0:
[0086]
[0087] And when ω k =0, we have:
[0088]
[0089] In the above formula, v k Represents t k The linear speed of the ship at time a k Represents t k The ship's acceleration at time θ k Represents t k The ship's yaw angle at time, ω k Represents t k The ship’s yaw rate at time t, Δt represents k Time to t k+1 The time difference between the instants.
[0090] Since the motion model error caused by the CTRA model assumption must exist, it is necessary to perform filtering analysis by combining the predicted trajectory derived from the CTRA model with the measurement information of the radar sensor. The present invention uses UKF for:
[0091] First, considering the existence of noise, write the state equation with Gaussian white noise:
[0092] x k =f(x k-1 )+ω k
[0093] z k =h(x k-1 )+v k
[0094] In the above formula, f(·) is the motion equation; h(·) is the observation equation; ω k is Gaussian noise with covariance Q, which is the system noise; v k is Gaussian noise with covariance R, which is the observation noise.
[0095] Then use UKF to calculate the estimated value at the next moment, and then substitute it into equation (3) for iteration to obtain the required trajectory prediction content.
[0096] For the N-dimensional variable X, generate 2N + 1 Sigma sampling points. Here, assume that the mean of X at a certain moment is and the variance is P x , then the sampling points are:
[0097]
[0098]
[0099]
[0100] The weights of each point are:
[0101]
[0102]
[0103] In the above formula, λ = α 2 (n + κ) - n, α is the scaling factor, which can control the range of the Sigma point set, β reflects the high-order characteristics of the state information, and κ is the scaling factor.
[0104] After determining the sampling points, perform state transfer on each sampling point to obtain the new state
[0105] Then, the mean and variance of the next state prediction can be obtained:
[0106]
[0107]
[0108] Then, measurement update is performed. First, write out the state transition of the measurement:
[0109]
[0110] Then, the one-step prediction mean, variance, and covariance of the measurement value can be written as:
[0111]
[0112]
[0113]
[0114] Thereby, the state estimate value and the estimated variance at time K+1 can be obtained:
[0115]
[0116] K k+1 = P xz,k+1|k (P zz,k+1|k ) -1
[0117]
[0118] After obtaining the state estimate value, through continuous iteration, all subsequent trajectory prediction information can be obtained.
[0119] S4: Combine the trajectory prediction information with the local chart to form local situation information.
[0120] Due to the limitations of the grid map, in order to better express the prediction information, the uncertainty of the previously obtained trajectory prediction information is expressed on the grid map, that is, a trajectory probability region is expressed on the grid map. Thus, we obtain the prediction content of the ship trajectory and also obtain the local situation information.
[0121] Using the past trajectory information of each ship, reasonably describe its future behavior and give a reasonable probability distribution as Figure 7 shown. And map the probability to the local high-precision grid map as Figure 8 shown to form a local situation information map.
[0122] S5: Determine the feature vector based on the global situation information, the local situation information, and the ship motion attribute vector.
[0123] The processing of situation information is as follows Figure 9 As shown, the two-channel image is convolved through a CNN convolutional network. Finally, through a flattening operation, it is flattened into a one-dimensional vector, and then concatenated with the vessel motion attribute feature vector to form a feature vector.
[0124] The vessel motion attribute feature vector is as follows:
[0125] x = [v, a, θ, ω, v t-1 , θ t-1 ,
[0126] Here, v represents the current speed of the vessel, a represents the current linear acceleration, θ represents the current yaw angle of the vessel, ω represents the current yaw angular velocity of the vessel, v t-1 represents the desired speed in the output action at the previous decision, and θ t-1 represents the desired speed in the output action at the previous decision.
[0127] S6: Train the MADDPG algorithm through the feature vector and the reward function; the MADDPG algorithm includes an Actor network and a Critic network.
[0128] With the help of the MADDPG algorithm, each vessel constructs an Actor network and a Critic network, as Figure 10 shown. The Actor network outputs actions (desired speed, desired heading), which are then passed to the underlying motion controller of the unmanned boat for operation, and the action output is evaluated according to the previous reward function, so as to train a better cooperative patrol strategy.
[0129] The reward function consists of a global unknown degree reward for the patrol area, a captured enemy vessel reward, an escaped enemy vessel reward, an invaded enemy vessel reward, and a collision reward.
[0130] R = r unknown + r catch + r escape + r intrusion + r collision
[0131] The global unknown degree reward r unknown is related to the unknown degree information of all navigable areas on the global situation map, and the calculation method is as follows:
[0132] First, assume that all information in the world is known, that is, the unknown degree value of all navigable grids is 100, and a value N omniscient is obtained. Then, sum the unknown degree information of all the grids in the currently navigable areas in the global situation information to obtain the current global data n unknown, then compare the two to obtain the current degree of uncertainty The value varies between (0, 1). The higher the value, the better the mastered situation. The lower the value, the worse the current patrol situation. Set r unknown = r 0 ×K unknown .
[0133] Here, r 0 is an adjustable parameter, which is set according to the requirement level of the time interval for the ship to traverse the area in the task.
[0134] The reward for capturing enemy ships is related to the position where the enemy ships are captured by our ships. The closer the distance to our mother ship when captured, the smaller the reward. Define the radius of the entire patrol area as r globol , the distance between the position of the enemy ship when it is calibrated and our mother ship is r catchdist , require r catchdist > r safedist , that is, the capture must be carried out outside the safe distance r safedist initially defined in the patrol task.
[0135] Define the capture reward Here, r 1 is an adjustable parameter, which is defined according to the requirement of the task for the position of capturing enemy ships.
[0136] The reward for enemy ships escaping is due to the needs of objective reality. To capture an enemy ship, it is necessary to keep the enemy ship within the detection range of our sensors for a period of time to accurately complete the identification and positioning of the target. For this requirement, if an enemy ship appears within our detection range but escapes from our sensor detection range without being continuously tracked by our ships for a certain period of time, a penalty r escape , which is negative.
[0137] The reward for enemy ship intrusion r intrusion is given when the enemy ship successfully approaches our mother ship and the distance between the two is less than the safe distance r safedist initially defined in the patrol task, and it is negative.
[0138] The collision reward r collision is given when the distance between our ship and any other ship is too close and a collision occurs. After giving this reward, the current round of training ends prematurely.
[0139] The training process of the Actor network and the Critic network is as follows:
[0140] Construct an Actor network and a Critic network for each ship, centralized training, distributed execution.
[0141] It is assumed that there are N unmanned boats in the set training scenario. The parameter set of the policy network (Actor) of the N unmanned boats is θ = {θ 1 , θ 2 ,..., θ n}, the policy set is u = {u 1 , u 2 ,..., u n}, the observation set is O = {o 1 , o 2 ,..., o n}, o n represents the feature vector of the nth boat, the action set is a = {a 1 , a 2 ,..., a n}, and the reward function set is r = {R 1 , R 2 ,..., R n}.
[0142] Then the policy gradient of the unmanned boat i is:
[0143]
[0144] In the above formula, represents the action-value function, indicating the value of the unmanned boat i choosing a certain action in a certain state. D is the experience replay pool, which stores quadruples. The quadruple is composed of (x, a, x,, r), where x, represents the observation set at the next moment after executing the action set a. Each training will randomly select a part from the experience replay pool for training.
[0145] According to the principle of maximizing the state-action function, the update of the policy network is:
[0146]
[0147] β in the above formula is the learning rate.
[0148] The value network (critic) optimizes the network parameters by minimizing the time difference error, and its objective function is:
[0149]
[0150] In the above formula, μ' is the policy set of the target network. The update of the target network is carried out according to the soft update strategy: θ i ′ ← (1 - τ)θ′ i + τθ i , τ << 1, which restricts the update amplitude; θ i and θ iThey are the parameters of the value network and the target network respectively.
[0151] S7: Determine the action instruction through the trained Actor network; perform the area patrol task of multiple unmanned boats according to the action instruction.
[0152] Embodiment 2
[0153] In order to execute the method corresponding to the above Embodiment 1 to achieve the corresponding functions and technical effects, a multi-unmanned-boat area patrol system in an adversarial environment based on situation information is provided below.
[0154] The system includes:
[0155] A global situation information construction module, configured to construct global situation information of the battlefield according to the information set; the information set includes: patrol task information, nautical chart information, the position information of its own ships, and the position information of the detected enemy ships;
[0156] A local nautical chart construction module, configured to construct a local nautical chart by using the sensor detection information of the current ship itself;
[0157] A trajectory prediction module, configured to predict the future movement trajectories of all ships according to the movement trajectories of the ships around the current ship and the movement trajectory of the current ship, so as to obtain trajectory prediction information;
[0158] A local situation information determination module, configured to combine the trajectory prediction information with the local nautical chart to form local situation information;
[0159] A feature vector determination module, configured to determine a feature vector based on the global situation information, the local situation information, and the ship movement attribute vector;
[0160] A training module, configured to train the MADDPG algorithm through the feature vector and the reward function; the MADDPG algorithm includes an Actor network and a Critic network;
[0161] An action instruction determination module, configured to determine an action instruction through the trained Actor network; perform the area patrol task of multiple unmanned boats according to the action instruction.
[0162] Wherein, the feature vector determination module specifically includes:
[0163] An initial feature vector determination unit, configured to input the global situation information and the local situation information into the trained CNN network to obtain an initial feature vector;
[0164] A splicing unit, configured to splice the initial feature vector with the ship movement attribute vector to obtain a feature vector.
[0165] Among them, the training module specifically includes:
[0166] An input unit for inputting the feature vector into the Actor network to obtain an action instruction at the next moment;
[0167] An evaluation unit for patrolling according to the action instruction at the next moment, and evaluating the action instruction at the next moment through the reward function and the Critic network;
[0168] A parameter adjustment unit for adjusting the parameters of the Actor network according to the evaluation result.
[0169] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the system disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0170] In this article, specific examples are used to elaborate on the principles and implementation manners of the present invention. The descriptions of the above embodiments are only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, based on the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation on the present invention.
Claims
1. A multi-unmanned boat regional patrol method in an adversarial environment based on situation information, characterized in that, it includes: Constructing the global situation information of the battlefield according to the information set; The information set includes: patrol task information, nautical chart information, the position information of friendly ships, and the position information of detected enemy ships; Using the sensor detection information of the current boat itself to construct a local nautical chart; Predicting the future movement trajectories of all boats according to the movement trajectories of the boats around the current boat and the movement trajectory of the current boat to obtain trajectory prediction information; Combining the trajectory prediction information with the local nautical chart to form local situation information; Determining a feature vector based on the global situation information, the local situation information, and the boat movement attribute vector; Training the MADDPG algorithm through the feature vector and the reward function; the MADDPG algorithm includes an Actor network and a Critic network; Determining an action instruction through the trained Actor network; performing a multi-unmanned boat regional patrol task according to the action instruction.
2. The multi-unmanned boat regional patrol method in an adversarial environment based on situation information according to claim 1, characterized in that, Predicting the future movement trajectories of all boats according to the movement trajectories of the boats around the current boat and the movement trajectory of the current boat, specifically including: Performing kinematic modeling on the boats according to the CTRA model to determine the state variables of the boats at the current moment; Calculating the change rates of the state variables at each current moment; Determining the state variables at the next moment according to the change rates; Determining the future movement trajectories of all boats according to the state variables at the next moment.
3. The multi-unmanned boat regional patrol method in an adversarial environment based on situation information according to claim 1, characterized in that, Determining a feature vector based on the global situation information, the local situation information, and the boat movement attribute vector, specifically including: Inputting the global situation information and the local situation information into the trained CNN network to obtain an initial feature vector; Concatenating the initial feature vector with the boat movement attribute vector to obtain a feature vector.
4. The multi-unmanned boat regional patrol method in an adversarial environment based on situation information according to claim 1, characterized in that, Training the MADDPG algorithm through the feature vector and the reward function, specifically including: Inputting the feature vector into the Actor network to obtain the action instruction at the next moment; Performing patrol according to the action instruction at the next moment, and evaluating the action instruction at the next moment through the reward function and the Critic network; Adjusting the parameters of the Actor network according to the evaluation result.
5. The multi-unmanned boat regional patrol method in an adversarial environment based on situation information according to claim 4, characterized in that, The expression of the reward function R is as follows: R=r unknown +r catch +r escape +r intrusion +r collision Among them, r unknown is the global unknown degree reward for the patrol area, r catch is the enemy ship capture reward, r escape is the enemy ship escape reward, r intrusion is the enemy ship intrusion reward, r collision is the collision reward.
6. The multi-unmanned boat regional patrol method in an adversarial environment based on situation information according to claim 5, characterized in that, The global unknown degree reward r of the patrol area unknown has the following expression: r unknown =r 0 ×K unknown Among them, r 0 represents an adjustable parameter, K unknown represents the degree of uncertainty of the global situation information, n unknown represents the numerical value of the degree of uncertainty of the global situation information, N omniscient represents the global data of the global situation information.
7. The multi-unmanned boat regional patrol method in an adversarial environment based on situation information according to claim 5, characterized in that, Enemy ship capture reward r catch The expression is as follows: where r catchdist represents the distance between the position of the enemy ship when it is calibrated and our mother submarine, r globol represents the radius of the patrol area, r 1 represents an adjustable parameter.
8. A multi-unmanned boat area patrol system in an adversarial environment based on situation information, characterized in that, it includes: A global situation information construction module for constructing global situation information of the battlefield according to the information set; The information set includes: patrol task information, nautical chart information, the position information of its own vessels, and the position information of the detected enemy vessels; A local nautical chart construction module for constructing a local nautical chart by using the detection information of the sensors of the current vessel itself; A trajectory prediction module for predicting the future movement trajectories of all vessels according to the movement trajectories of the vessels around the current vessel and the movement trajectory of the current vessel to obtain trajectory prediction information; A local situation information determination module for combining the trajectory prediction information with the local nautical chart to form local situation information; A feature vector determination module for determining a feature vector based on the global situation information, the local situation information, and the vessel movement attribute vector; A training module for training the MADDPG algorithm through the feature vector and the reward function; the MADDPG algorithm includes an Actor network and a Critic network; An action instruction determination module for determining an action instruction through the trained Actor network; performing a multi-unmanned boat area patrol task according to the action instruction.
9. The multi-unmanned boat area patrol system in an adversarial environment based on situation information according to claim 8, characterized in that, The feature vector determination module specifically includes: An initial feature vector determination unit for inputting the global situation information and the local situation information into the trained CNN network to obtain an initial feature vector; A splicing unit for splicing the initial feature vector with the vessel movement attribute vector to obtain a feature vector.
10. The multi-unmanned boat area patrol system in an adversarial environment based on situation information according to claim 8, characterized in that, The training module specifically includes: An input unit for inputting the feature vector into the Actor network to obtain an action instruction for the next moment; An evaluation unit for patrolling according to the action instruction for the next moment, and evaluating the action instruction for the next moment through the reward function and the Critic network; A parameter adjustment unit for adjusting the parameters of the Actor network according to the evaluation result.
Citation Information
Patent Citations
Multi-unmanned aerial vehicle cooperative confrontation method and system, and storage medium
CN114721424A
Unmanned ship trajectory tracking control method based on Actor-Credit-Advantage network
CN115793455A