Unmanned ship cluster multi-target task allocation method based on improved DDQN algorithm

By improving the DDQN algorithm and combining it with the Boltzmann exploration strategy and priority experience replay, a task allocation model for unmanned surface vessel (USV) swarms is constructed. This solves the problems of low efficiency and poor stability of traditional DQN algorithm in complex environments, and achieves efficient and stable multi-objective allocation for USV swarms.

CN121998362APending Publication Date: 2026-05-08HARBIN ENG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HARBIN ENG UNIV
Filing Date
2026-01-29
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Traditional DQN algorithms suffer from low task allocation efficiency, poor stability, overestimation of Q-values, and low sample utilization in unmanned surface vessel (USV) swarm task allocation, making them difficult to adapt to complex and ever-changing marine environments.

Method used

An improved DDQN algorithm is adopted, combined with the Boltzmann exploration strategy and the priority experience replay mechanism, to construct a capability-constrained distributed task allocation model. Through temperature-controlled action selection and reward function optimization, training efficiency and sample utilization are improved.

Benefits of technology

It achieves efficient and stable multi-objective task allocation for unmanned surface vessel swarms in complex environments, avoids Q-value overestimation, and improves the robustness and convergence speed of the allocation strategy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121998362A_ABST
    Figure CN121998362A_ABST
Patent Text Reader

Abstract

The invention discloses an unmanned ship cluster multi-target task allocation method based on an improved DDQN algorithm, and belongs to the technical field of unmanned ship task planning. According to the method, the problems of low task allocation efficiency, poor task allocation stability, Q value over-estimation and low sample utilization rate of a traditional DQN algorithm are solved. According to the method, a distributed task allocation model containing capability constraint is constructed, the exploration capability of the DDQN architecture is enhanced in combination with a temperature-controllable Boltzmann strategy, the training efficiency is improved by using priority experience playback, and the sample utilization rate is ensured. According to the method, the unmanned ship cluster can autonomously learn the optimal allocation strategy through interaction with the environment without priori knowledge, so that the task allocation efficiency and robustness are effectively improved, the task allocation stability is ensured, and efficient and collaborative allocation of the unmanned ship cluster to multiple targets is realized. And meanwhile, action selection and value evaluation are decoupled by utilizing the DDQN architecture, so that the over-estimation deviation of the Q value is eliminated. The method can be applied to unmanned ship target task allocation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of unmanned surface vehicle (USV) mission planning technology, specifically relating to a multi-objective mission allocation method for USV swarms based on an improved DDQN algorithm. Background Technology

[0002] Unmanned surface vessels (USVs) play a vital role in marine resource exploration and other fields due to their small size, high level of intelligence, and good stealth capabilities. However, in the face of complex and ever-changing marine environments, a single USV often cannot meet mission requirements; therefore, collaborative operations among multiple USV swarms have become an inevitable trend. Among these, multi-target task allocation is a key aspect of swarm collaboration, directly determining the overall effectiveness of the system.

[0003] Traditional task allocation methods, including contract net protocols and auction algorithms, often rely on precise mathematical models or extensive prior knowledge. These methods suffer from high computational complexity and poor real-time performance when dealing with high-dimensional state spaces and dynamic environments. While reinforcement learning methods such as Deep Q-Networks (DQNs) offer new solutions to these problems, traditional DQN algorithms suffer from issues such as Q-value overestimation, low sample utilization, and susceptibility to local optima during action selection. These limitations make it difficult to guarantee the efficiency and convergence stability of task allocation in complex environments for unmanned surface vessel (USV) swarms.

[0004] In summary, the traditional DQN algorithm still suffers from problems such as low task allocation efficiency, poor task allocation stability, overestimation of Q-values, and low sample utilization. Therefore, researching a task allocation method that can adapt to dynamic environments, has fast convergence speed, and optimizes the allocation strategy has significant practical value. Summary of the Invention

[0005] This invention addresses the problems of low task allocation efficiency, poor task allocation stability, Q-value overestimation, and low sample utilization in the traditional DQN algorithm by proposing a multi-target task allocation method for unmanned surface vessel swarms based on an improved DDQN algorithm.

[0006] The technical solution adopted by this invention to solve the above-mentioned technical problems is: a multi-target task allocation method for unmanned surface vessel swarms based on an improved DDQN algorithm, the method specifically including the following steps:

[0007] Step 1: Construct the input state vector and action space of the improved DDQN algorithm, wherein the improved DDQN algorithm is obtained by replacing the action selection strategy with the Boltzmann exploration strategy;

[0008] The input state vector includes the coordinates of all unmanned surface vessels (USVs), the communication capability values, detection capability values, motion capability values ​​of all USVs, and the current mission status of all target USVs; the action space is the assigned number of the target USVs.

[0009] Step 2: Improve the DDQN algorithm by assigning target unmanned surface vessels (USVs) to each USV based on the input state vector.

[0010] Furthermore, the current mission state of the target unmanned surface vessel is defined as follows: for any target unmanned surface vessel, the mission state of the target unmanned surface vessel is a three-dimensional vector;

[0011] The first element in the three-dimensional vector represents the normalized remaining demand ratio of the target unmanned surface vessel's current communication capabilities;

[0012] The second element in the three-dimensional vector represents the normalized remaining demand ratio of the target unmanned surface vessel's current detection capability;

[0013] The third element in the three-dimensional vector represents whether the motion capability of the capture unmanned surface vessel assigned to the target meets the requirements. If the motion capability meets the requirements, the value of the third element in the three-dimensional vector is 1; otherwise, it is 0.

[0014] Furthermore, the method for calculating the normalized residual demand ratio of the communication capacity is as follows:

[0015]

[0016] in, This represents the normalized remaining demand ratio for communication capacity.

[0017] Furthermore, the method for calculating the normalized residual demand ratio of the detection capability is as follows:

[0018]

[0019] in, This represents the normalized remaining demand ratio for detection capabilities.

[0020] Furthermore, the reward function used in the training process of the improved DDQN algorithm is:

[0021]

[0022] in, This represents the total reward function value during the training process of the improved DDQN algorithm. Indicates the first The reward function value for capturing unmanned surface vessels. This indicates the total number of unmanned boats captured. Indicates the encirclement and capture of the first The reward function value of a team of unmanned surface vessels (USVs) targeting a specific USV. Indicates the first The reward function value of a target unmanned surface vessel. This indicates the total number of target unmanned surface vessels.

[0023] Furthermore, the first The reward function value for each unmanned surface vessel (USV) being hunted. for:

[0024]

[0025] in, Indicates the first The reward for each unmanned surface vessel used for the manhunt is based on the allocation results. It is a fixed value (which can be set according to the actual situation); This indicates a skill-matching reward, which will be available soon. The three capability values ​​of each pursuit unmanned surface vessel (USV) and the assigned target USV are respectively calculated as differences, and the sum of the three differences is the result. ; Indicates the first The inverse of the distance between the encirclement unmanned surface vessel and the assigned target unmanned surface vessel. , and All are weighting coefficients.

[0026] Furthermore, the encirclement of the first The reward function value of a team of unmanned surface vessels (USVs) targeting a specific objective. for:

[0027]

[0028] in, Indicates allocation to the first The total reward allocated to all unmanned surface vessels (USVs) within a team that targets a specific USV is the sum of the number of USVs captured within the team and the total reward allocated to all USVs captured. The product; Indicates allocation to the first The entire team surrounding the target unmanned surface vessel was tasked with capturing it. The negative number of the total distance to the target unmanned surface vessel. and All are weighting coefficients.

[0029] Furthermore, the first The reward function value of the target unmanned surface vessel for:

[0030]

[0031] in, Indicates the first The reward for each target unmanned surface vessel is allocated; if an unmanned surface vessel is captured, it will be allocated to the target unmanned surface vessel. If there is a target unmanned surface vessel, then The value is a set value (which can be set according to the actual situation). If no unmanned surface vessel is captured, it will be assigned to the first... If there is a target unmanned surface vessel, then The value of is 0; Indicates the first The ability-matching reward for each target unmanned surface vessel and All are weighting coefficients.

[0032] Furthermore, during the training process of the improved DDQN algorithm, the first... The probability of a sample being selected for:

[0033]

[0034] in, Indicates the first The priority of each sample , For positive integers, For hyperparameters, Indicates the first TD error for an empirical sample Indicates the first Priority of each empirical sample This represents the total number of experience samples in the experience replay pool.

[0035] Furthermore, when the improved DDQN algorithm outputs an action, the probability of each action being selected in the action space is:

[0036]

[0037] in, For temperature parameters, Indicates the selection action The value gained Indicates the selection action The value gained The base of the natural logarithm. Represents a set of actions;

[0038] The temperature parameter The decays linearly with each training round:

[0039]

[0040] in, Indicates the initial temperature parameter. Indicates the decay rate of the temperature parameter. Indicates the training round.

[0041] The beneficial effects of this invention are:

[0042] This invention constructs a distributed task allocation model with capability constraints, enhances the exploration capability of the DDQN architecture by combining a temperature-controlled Boltzmann strategy, improves training efficiency by utilizing priority experience replay, and ensures sample utilization. This method enables unmanned surface vessel (USV) swarms to autonomously learn the optimal allocation strategy through interaction with the environment without prior knowledge, effectively improving the efficiency and robustness of task allocation, ensuring the stability of task allocation, and achieving efficient and collaborative allocation of multiple targets by the USV swarm. Simultaneously, the DDQN architecture decouples action selection and value assessment, eliminating the overestimation bias of Q-values ​​to avoid affecting the allocation results. Attached Figure Description

[0043] Figure 1 This is a flowchart of a multi-objective task allocation method for unmanned surface vessel swarms based on an improved DDQN algorithm, according to the present invention.

[0044] Figure 2 A schematic diagram illustrating a scenario for assigning multiple objectives to a USV;

[0045] Figure 3 The task allocation results of the original DDQN algorithm are shown in the following diagram: in a scenario with 9 USVs being captured and 3 USVs being targeted.

[0046] Figure 4 The task allocation results are shown in the following diagram: In the scenario of 9 USVs being surrounded and 3 USVs being targeted, and the action selection strategy is replaced with the Boltzmann strategy.

[0047] Figure 5 The task allocation results of the method of the present invention are shown in the following diagram for a scenario with 9 USVs being captured and 3 target USVs.

[0048] Figure 6 A comparison chart of task allocation rewards for three methods in a scenario with 9 USVs being captured and 3 USVs being targeted;

[0049] Figure 7 A diagram showing the task allocation results in a scenario with 9 USVs being captured and 4 USVs being targeted;

[0050] Figure 8 A task allocation reward diagram for a scenario with 9 USVs being captured and 4 target USVs. Detailed Implementation

[0051] Establish a mathematical model for multi-objective task allocation in unmanned surface vessel (USV) swarms:

[0052] (1) Define the system The group consisted of several unmanned vessels used for encirclement and capture. , The collection of target unmanned surface vessels is The task assignment mapping is defined as a mapping from the set of unmanned surface vessels to the set of targets. .

[0053] (2) The communication, detection, and motion capabilities of each capture unmanned surface vessel (USV) and the target USV are quantified separately. In this invention, communication bandwidth is used as the communication capability, effective detection radius is used as the detection capability, and maximum speed is used as the motion capability. Furthermore, constraints for task allocation are set, including capability matching constraints and allocation conflict constraints.

[0054] Each target unmanned surface vessel (USV) must meet capability matching constraints, specifically:

[0055] For any target unmanned surface vessel (USV), the sum of the communication capabilities of all USVs in the USV team hunting down that target USV is denoted as: , It should be greater than or equal to the communication capability of the target unmanned surface vessel. ;

[0056] The sum of the detection capabilities of all unmanned surface vessels (USVs) within the USV team tasked with capturing the target is denoted as: , It should be greater than or equal to the detection capability of the unmanned surface vessel targeting the target. ;

[0057] The maneuverability of each unmanned surface vessel (USV) in the swarm to capture the target USV should be greater than or equal to the maneuverability of the target USV. In this invention, it is assumed that the mobility of each capture unmanned surface vessel is at least greater than that of a target unmanned surface vessel.

[0058] Each unmanned surface vessel (USV) used for the capture operation must satisfy the allocation conflict constraint, specifically that each USV can only be assigned to one target.

[0059] Based on this, the method of the present invention will be further described in detail with reference to the accompanying drawings.

[0060] Specific implementation method one: Combining Figure 1 This embodiment describes a method for multi-target task allocation in an unmanned surface vessel (USV) swarm based on an improved DDQN algorithm. The method specifically includes the following steps:

[0061] Step 1: Construct the input state vector and action space of the improved DDQN algorithm. The improved DDQN algorithm is obtained by replacing the action selection strategy with a temperature-controlled Boltzmann exploration strategy. That is, based on the traditional DDQN algorithm, only the action selection strategy of the DDQN algorithm is improved.

[0062] The input state vector includes the coordinates of all unmanned surface vessels (USVs), the communication capability values, detection capability values, motion capability values ​​of all USVs, and the current mission status of all target USVs.

[0063] The action space is the assigned target unmanned surface vessel number;

[0064] The current mission state of a target unmanned surface vessel is defined as follows: for any target unmanned surface vessel, the mission state of the target unmanned surface vessel is a three-dimensional vector;

[0065] The first element in the three-dimensional vector represents the normalized remaining demand ratio of the target unmanned surface vessel's current communication capability, and its value is the calculated normalized remaining demand ratio value (ranging from 0 to 1).

[0066] The method for calculating the normalized residual demand ratio of communication capacity is as follows:

[0067]

[0068] The second element in the three-dimensional vector represents the normalized remaining demand ratio of the target unmanned surface vessel's current detection capability. Its value is the calculated normalized remaining demand ratio (ranging from 0 to 1).

[0069] The method for calculating the normalized residual demand ratio of detection capability is as follows:

[0070]

[0071] The third element in the three-dimensional vector represents whether the motion capability of the capture unmanned surface vessel assigned to the target unmanned surface vessel meets the requirements. If the motion capability meets the requirements, the value of the third element in the three-dimensional vector is 1; otherwise, it is 0 (that is, the motion capability of the capture unmanned surface vessel assigned to the target unmanned surface vessel is greater than the motion capability of the target unmanned surface vessel. Since the present invention imposes constraints on the motion capability, the value of the third element must be 1).

[0072] The input state vector is a one-dimensional row vector. In the input state vector, firstly, there are the coordinates and three capability values ​​of the first capture unmanned surface vessel, then the coordinates and three capability values ​​of the second capture unmanned surface vessel, then the coordinates and three capability values ​​of the first target unmanned surface vessel, and finally the current mission state of each target unmanned surface vessel.

[0073] Step 2: Improve the DDQN algorithm by assigning target unmanned surface vessels (USVs) to each USV based on the input state vector.

[0074] It should be noted that in the process of allocation using the improved DDQN algorithm, each capture unmanned surface vessel is allocated separately. After the allocation of the first capture unmanned surface vessel is completed, the task state of the target unmanned surface vessel to which the first capture unmanned surface vessel is allocated will be changed, which will affect the input state vector when allocating the second capture unmanned surface vessel. That is, the allocation result of the previous capture unmanned surface vessels will affect the subsequent allocation.

[0075] The following is a detailed explanation of the relevant factors in the training process of the improved DDQN algorithm in this invention:

[0076] 1. Reward Function

[0077]

[0078] in, This represents the total reward function value during the training process of the improved DDQN algorithm. Indicates the first The reward function value for capturing unmanned surface vessels. This indicates the total number of unmanned boats captured. Indicates the encirclement and capture of the first The reward function value of a team of unmanned surface vessels (USVs) targeting a specific USV. Indicates the first The reward function value of a target unmanned surface vessel. Indicates the total number of target unmanned surface vessels;

[0079]

[0080] in, Indicates the first The reward for each unmanned surface vessel used for the manhunt is based on the allocation results. It is a fixed value; This indicates a skill-matching reward, which will be available soon. The three capability values ​​of each pursuit unmanned surface vessel (USV) and the assigned target USV are respectively calculated as differences, and the sum of the three differences is the result. ; Indicates the first The inverse of the distance between the capture unmanned surface vessel and the assigned target unmanned surface vessel;

[0081]

[0082] in, Indicates allocation to the first The total reward allocated to all unmanned surface vessels (USVs) within a team that targets a specific USV is the sum of the number of USVs captured within the team and the total reward allocated to all USVs captured. The product; Indicates allocation to the first The entire team surrounding the target unmanned surface vessel was tasked with capturing it. The negative of the total distance to each target unmanned surface vessel;

[0083]

[0084] in, Indicates the first The reward for each target unmanned surface vessel is allocated; if an unmanned surface vessel is captured, it will be allocated to the target unmanned surface vessel. If there is a target unmanned surface vessel, then The value is a set value; if no unmanned surface vessel is captured, it will be assigned to the first... If there is a target unmanned surface vessel, then The value of is 0. Initially, the first is... The first target unmanned surface vessel may not have been assigned to a capture unmanned surface vessel, but as the training process continues, the second... Each target unmanned surface vessel will be assigned to a capture unmanned surface vessel; Indicates the first The ability matching reward for each target unmanned surface vessel;

[0085] It should be noted that: the calculation is assigned to the first... The total communication and detection capabilities of all the target unmanned surface vessels (USVs) within the team will be compared with the total communication capabilities of the team. The communication capabilities of each target unmanned surface vessel are compared, and the team's total detection capability is compared with that of the first... The detection capabilities of each unmanned surface vessel (USV) are subtracted, and a three-dimensional vector is constructed using the two differences and 0. The 2-norm value of this three-dimensional vector is then calculated. If the obtained 2-norm value is greater than the set expected value, then... The value is 0, otherwise, It is a fixed reward value.

[0086] 2. Action Selection

[0087] When selecting actions, each available action is calculated separately. Probability of being selected :

[0088]

[0089] in, For temperature parameters, Indicates the selection action The value gained Indicates the selection action The value gained The base of the natural logarithm. This represents a set of actions, where the number of elements in the set is the same as the number of target unmanned surface vessels.

[0090] Temperature parameters in the early stages of training When the value is large, the action selection tends to be evenly distributed to encourage exploration; as the training rounds increase... Increase, The value decays linearly according to the formula:

[0091]

[0092] in, Indicates the initial temperature parameter. Indicates the decay rate of the temperature parameter. Indicates the training round;

[0093] Action selection probabilities tend to favor high Q-value actions in order to leverage existing experience.

[0094] 3. Sample selection in the experience replay pool

[0095] When storing empirical samples, calculate their TD error. .

[0096] During sampling, the first The probability of an empirical sample being selected Its priority Proportional, that is:

[0097]

[0098] in, Indicates the first Priority of each empirical sample , It is a positive constant (a small positive number used to ensure that all experiences have a non-zero sampling probability). This is a hyperparameter (used to control the influence of priority on sampling probability, with a value range of 0~1; in this invention, ...) The value is 0.6). Indicates the first TD error for an empirical sample Indicates the first Priority of each empirical sample This represents the total number of empirical samples.

[0099] To eliminate the bias caused by non-uniform sampling, importance sampling weights are introduced when calculating the loss function. , , This indicates the size of the experience pool (fixed at 2000 in this invention). This represents the weighted hyperparameter that corrects for sampling bias (it starts at 0.4 and increases linearly to 1, with a growth rate of 0.0003).

[0100] The upcoming Squared TD error of an empirical sample and Multiply to get the first The loss is calculated for each empirical sample. In this invention, the batch size during training is N, and N is 64. Weights are introduced. The error introduced by sampling can be amplified during gradient updates. The weighted loss function is as follows:

[0101]

[0102] Experimental Section

[0103] like Figure 2 As shown, the experiment sets up a scenario where multiple USVs surround multiple targets. In this scenario, multiple surrounding USVs move freely within a designated area to explore targets, and subsequently, several target USVs enter the detection range of the surrounding USVs. Once the surrounding USVs detect a target, they immediately assign the target according to the steps of the improved DDQN algorithm. In the experiment, a two-dimensional plane is defined... A square enclosure water environment was set up, with 9 USVs randomly placed in the environment, and 3 target USVs and 4 target USVs were allocated in the water area respectively.

[0104] like Figure 3 , Figure 4 and Figure 5 As shown, the task allocation results of three algorithms are represented respectively: the traditional DDQN algorithm, the DDQN+Boltzmann algorithm, and the DDQN+Boltzmann+PER (i.e., DDQN + Boltzmann policy + priority experience replay).

[0105] Figure 3 This is the task allocation result obtained using the original DDQN algorithm, from... Figure 3 It can be concluded that, under the guidance of the DDQN algorithm, the unmanned surface vessel (USV) completed the task allocation of target points, but the allocation scheme was not optimal. The USV team clearly had problems with target point allocation being too far away and low efficiency, resulting in unnecessary energy consumption losses.

[0106] Figure 4 This is the task allocation result after replacing the greedy action selection strategy in the original DDQN algorithm with the Boltzmann strategy. Figure 4The allocation results show that, compared to the original DDQN algorithm, the allocation scheme of the algorithm after replacing the action selection strategy is improved and does not cause excessive energy consumption. However, this allocation scheme still has the problems of excessively long paths and low efficiency.

[0107] Figure 5 This is the allocation result of the improved DDQN algorithm proposed in this invention, from... Figure 5 As can be seen, the allocation result not only achieves near-global optimum but also reduces energy consumption loss, such as... Figure 6 As shown, the experimental comparison verifies the effectiveness of the algorithm proposed in the multi-target task allocation scenario of multiple unmanned surface vessels.

[0108] Figure 6 The reward graphs for the three algorithms are shown in the scenarios of 9 encirclement unmanned surface vessels and 3 target unmanned surface vessels. From... Figure 6 As can be seen, the reward value of the original DDQN algorithm begins to converge around 250 rounds, but during training, the reward value oscillates around 0. After improving the greedy action selection strategy to the Boltzmann action selection strategy, the reward value begins to converge around 200 rounds, which is relatively higher than the original DDQN algorithm, but the stability is still lacking. Based on this, the method of prioritizing experience replay (i.e., the method of this invention) is added. As can be seen from the comparison in the figure, the reward value of the algorithm of this invention begins to converge around 250 rounds, and the reward value is significantly higher than the previous two algorithms. Moreover, it does not oscillate around 0 or become a negative reward value. The effectiveness of the improved algorithm of this invention can be seen from the comparison of reward values.

[0109] like Figure 7 and Figure 8 As shown, the task allocation results for 9 USVs to be captured and 4 target USVs are displayed, along with the reward convergence curve. From... Figure 7 The allocation results show that even with a larger number of targets (4), the algorithm can still effectively allocate tasks. All targets are assigned to the corresponding UAV teams, and the composition of each UAV team is reasonable, with no resource conflicts. Figure 8 The reward curves show that although the increased number of objectives leads to higher task complexity, the reward value still converges to a high positive range after training. In the early stages of training (approximately the first 200 rounds), the reward value is low and fluctuates significantly, indicating that the agent is exploring the environment and learning strategies. As training progresses, the reward value rises rapidly and stabilizes after about 250 rounds, demonstrating the robustness and adaptability of the improved DDQN algorithm in handling multi-objective task allocation, and its ability to cope with the challenges brought by changes in the number of objectives.

[0110] Thus, the multi-target task allocation method for unmanned surface vessel (USV) swarms based on the improved DDQN algorithm effectively solves the task allocation problem when multiple USVs face multiple targets by introducing a temperature-controlled Boltzmann exploration mechanism and a priority experience replay mechanism. It achieves an efficient and conflict-free allocation scheme and significantly improves the convergence speed and stability of the algorithm.

[0111] The above examples of the present invention are merely illustrative of the computational model and process of the present invention, and are not intended to limit the implementation of the present invention. Those skilled in the art will recognize that other variations or modifications can be made based on the above description. It is impossible to exhaustively list all possible implementations here. Any obvious variations or modifications derived from the technical solutions of the present invention are still within the scope of protection of the present invention.

Claims

1. A method for multi-objective task allocation in unmanned surface vessel (USV) swarms based on an improved DDQN algorithm, characterized in that, The method specifically includes the following steps: Step 1: Construct the input state vector and action space of the improved DDQN algorithm, wherein the improved DDQN algorithm is obtained by replacing the action selection strategy with the Boltzmann exploration strategy; The input state vector includes the coordinates of all unmanned surface vessels (USVs), the communication capability values, detection capability values, motion capability values ​​of all USVs, and the current mission status of all target USVs; the action space is the assigned number of the target USVs. Step 2: Improve the DDQN algorithm by assigning target unmanned surface vessels (USVs) to each USV based on the input state vector.

2. The method for multi-objective task allocation in unmanned surface vessel swarms based on the improved DDQN algorithm according to claim 1, characterized in that, The current mission state of the target unmanned surface vessel is defined as follows: for any target unmanned surface vessel, the mission state of the target unmanned surface vessel is a three-dimensional vector; The first element in the three-dimensional vector represents the normalized remaining demand ratio of the target unmanned surface vessel's current communication capabilities; The second element in the three-dimensional vector represents the normalized remaining demand ratio of the target unmanned surface vessel's current detection capability; The third element in the three-dimensional vector represents whether the motion capability of the capture unmanned surface vessel assigned to the target meets the requirements. If the motion capability meets the requirements, the value of the third element in the three-dimensional vector is 1; otherwise, it is 0.

3. The method for multi-objective task allocation in unmanned surface vessel swarms based on the improved DDQN algorithm according to claim 2, characterized in that, The method for calculating the normalized residual demand ratio of the communication capacity is as follows: in, This represents the normalized ratio of remaining demand for communication capacity.

4. The method for multi-objective task allocation in unmanned surface vessel swarms based on the improved DDQN algorithm according to claim 3, characterized in that, The method for calculating the normalized remaining demand ratio of the detection capability is as follows: in, This represents the normalized remaining demand ratio for detection capabilities.

5. A method for multi-objective task allocation in unmanned surface vessel swarms based on an improved DDQN algorithm according to claim 4, characterized in that, The reward function used in the training process of the improved DDQN algorithm is: in, This represents the total reward function value during the training process of the improved DDQN algorithm. Indicates the first The reward function value for capturing unmanned surface vessels. This indicates the total number of unmanned boats captured. Indicates the encirclement and capture of the first The reward function value of a team of unmanned surface vessels (USVs) targeting a specific USV. Indicates the first The reward function value of a target unmanned surface vessel. This indicates the total number of target unmanned surface vessels.

6. A method for multi-objective task allocation in unmanned surface vessel swarms based on an improved DDQN algorithm according to claim 5, characterized in that, The first The reward function value for each unmanned surface vessel (USV) being hunted. for: in, Indicates the first The reward for each unmanned surface vessel used for the manhunt is based on the allocation results. It is a fixed value; This indicates a skill-matching reward, which will be available soon. The three capability values ​​of each pursuit unmanned surface vessel (USV) and the assigned target USV are respectively calculated as differences, and the sum of the three differences is the result. ; Indicates the first The inverse of the distance between the encirclement unmanned surface vessel and the assigned target unmanned surface vessel. , and All are weighting coefficients.

7. A method for multi-objective task allocation in unmanned surface vessel swarms based on an improved DDQN algorithm as described in claim 6, characterized in that, The encirclement and capture The reward function value of a team of unmanned surface vessels (USVs) targeting a specific objective. for: in, Indicates allocation to the first The total reward allocated to all unmanned surface vessels (USVs) within a team that targets a specific USV is the sum of the number of USVs captured within the team and the total reward allocated to all USVs captured. The product; Indicates allocation to the first The entire team surrounding the target unmanned surface vessel was tasked with capturing it. The negative number of the total distance to the target unmanned surface vessel. and All are weighting coefficients.

8. A method for multi-objective task allocation in unmanned surface vessel swarms based on an improved DDQN algorithm according to claim 7, characterized in that, The first The reward function value of the target unmanned surface vessel for: in, Indicates the first The reward for each target unmanned surface vessel is allocated; if an unmanned surface vessel is captured, it will be allocated to the target unmanned surface vessel. If there is a target unmanned surface vessel, then The value is a set value; if no unmanned surface vessel is captured, it will be assigned to the first... If there is a target unmanned surface vessel, then The value of is 0; Indicates the first The ability-matching reward for each target unmanned surface vessel and All are weighting coefficients.

9. A method for multi-objective task allocation in unmanned surface vessel swarms based on an improved DDQN algorithm as described in claim 8, characterized in that, During the training process of the improved DDQN algorithm, the first... The probability of a sample being selected for: in, Indicates the first The priority of each sample , For positive integers, For hyperparameters, Indicates the first TD error for an empirical sample Indicates the first Priority of each empirical sample This represents the total number of experience samples in the experience replay pool.

10. A method for multi-objective task allocation in unmanned surface vessel swarms based on an improved DDQN algorithm according to claim 9, characterized in that, When the improved DDQN algorithm outputs an action, the probability of each action in the action space being selected is: in, For temperature parameters, Indicates the selection action The value gained Indicates the selection action The value gained The base of the natural logarithm. Represents a set of actions; The temperature parameter The decays linearly with each training round: in, Indicates the initial temperature parameter. Indicates the decay rate of the temperature parameter. Indicates the training round.