Multi-UUV dynamic target searching method based on strategy adaptive fusion
By employing a strategy-adaptive fusion method driven by weighted cumulative detection probability, combined with adaptive particle swarm optimization and regional serpentine traversal, the problem of low search efficiency and low success rate of UUVs in dynamic marine environments is solved, achieving efficient and complete dynamic target search.
Patent Information
- Application Number
- CN202511314012.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-15
- Publication Date
- 2025-11-14
AI Technical Summary
Existing UUV search algorithms are inefficient in dynamic marine environments, struggle to balance extensive exploration and precise convergence, and lack adaptability, resulting in low search success rates and insufficient robustness.
An adaptive fusion method based on weighted cumulative detection probability is adopted. Through adaptive particle swarm optimization, regional serpentine traversal, and coverage map optimization, the search strategy is dynamically adjusted to balance global exploration and fine convergence. Combined with the nonlinear adjustment of inertial weights and learning factors, multi-UUV collaborative search is achieved.
It significantly improves search efficiency and success rate, avoids path duplication and partial stagnation, and ensures efficient and complete search of dynamic targets.
Smart Images

Figure CN120947645A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to marine engineering equipment for multiple unmanned underwater vehicles (UUVs), particularly to the field of cooperative dynamic target search for UUVs, and especially to a strategy-adaptive fusion-based method for efficiently searching for maneuvering targets in unknown dynamic marine environments. Background Technology
[0002] In the field of modern marine exploration and safety, especially in the underwater rescue equipment manufacturing industry, the application of unmanned underwater vehicles (UUVs) is becoming increasingly crucial. These advanced devices typically serve as highly mobile hydrological observation platforms, playing an irreplaceable role in underwater search and rescue, environmental monitoring, and other missions. A typical and highly challenging application scenario is the rapid search and location of lost lifeboats and rafts, especially small, dynamic targets like inflatable life rafts whose trajectories are unpredictable due to ocean currents, following maritime accidents.
[0003] However, to efficiently accomplish such tasks, UUV platforms must rely on an advanced AI-optimized operating system for path planning and collaborative decision-making. Currently, existing search algorithms used in such operating systems have revealed numerous shortcomings in practical applications. Because the parameters of traditional algorithms are typically preset or linearly variable, they struggle to adapt to dynamically changing search environments. This often leads to algorithms missing targets due to premature convergence or repeatedly searching known areas later due to insufficient exploration capabilities, resulting in low search efficiency.
[0004] Existing search algorithms, such as traditional particle swarm optimization, genetic algorithms, or traversal search methods, often exhibit the following shortcomings when dealing with this type of problem:
[0005] (1) Low search efficiency: The parameters of traditional particle swarm optimization algorithms are usually preset or linearly variable, making it difficult to adapt to dynamically changing search environments. In the early stages of the search, it may get stuck in a local optimum due to premature convergence; in the later stages of the search, it may be unable to effectively clean up the remaining high-value regions due to insufficient exploration capabilities, resulting in long search times and low efficiency.
[0006] (2) Low search success rate: For highly mobile dynamic targets, their probability distribution changes rapidly over time. Traditional methods lack targeted strategy fusion mechanisms and often use a single strategy throughout, making it difficult to achieve an effective balance between extensive exploration and precise convergence, thus affecting the final search success rate.
[0007] (3) Lack of robustness: When the target probability distribution is multi-peaked, dispersed or sparse, traditional algorithms are prone to problems such as path repetition and omission of high-value areas, resulting in insufficient robustness and adaptability.
[0008] Therefore, there is an urgent need for a multi-UUV collaborative search method that can intelligently and dynamically adjust the search strategy based on real-time search results and integrate the advantages of multiple search modes to solve the problems of low efficiency and low success rate in existing technologies.
[0009] Meanwhile, compared to the present invention, the following methods differ as follows:
[0010] Currently, scholars both domestically and internationally have conducted extensive research on the problem of multi-agent cooperative search for dynamic targets. Most research focuses on improving various swarm intelligence algorithms to adapt to search tasks.
[0011] The paper "An Improved Particle Swarm Optimization Algorithm for Multi-Robot Search of Dynamic Targets" combines Newton's method with the traditional particle swarm optimization algorithm and switches the update mode with a certain probability to avoid stagnation in the later stages of the algorithm. However, this method does not fully consider the strategy requirements of different stages of the search task, and the probability switching mechanism is relatively fixed, lacking real-time feedback on the search effect.
[0012] The paper "Research on Dynamic Target Search for Multi-UAV Based on Cooperative Coevolution Motion-Encoded Particle Swarm Optimization" proposes a cooperative co-evolution particle swarm optimization algorithm. By decomposing the multi-UAV search problem into multiple sub-problems and optimizing them in parallel, it improves population diversity. However, its parameter settings are fixed, and its adaptability to environmental changes is limited.
[0013] Some studies have introduced target probability maps to guide the search. The paper "Cooperative Search Optimization of Unknown Dynamic Targets Based on Improved Target Probability Maps" improves the target probability map to fit the trajectory of dynamic targets and combines ant colony optimization and artificial potential field methods for path planning, effectively utilizing the prior information of the target. However, in the later stages of the search, when the target probability map becomes sparse or flat, this method is prone to reduced search efficiency due to excessive accumulation of pheromones or local extrema of the potential field.
[0014] In summary, to address the problems of existing technologies such as single search strategies, adaptive mechanisms limited to internal algorithm states, and low efficiency in the later stages of the search, this invention proposes a multi-UUV dynamic target search method with adaptive fusion of strategies to improve search efficiency and success rate, and to achieve adaptive parameter adjustment. Summary of the Invention
[0015] This invention aims to address the shortcomings in search efficiency, success rate, and robustness of multiple UUVs searching for dynamically uncertain targets in complex sea areas. To this end, a strategy-adaptive fusion method based on weighted cumulative detection probability is proposed: dynamically balancing global exploration and fine convergence, as well as the probability of acquiring residual targets, at different stages, maximizing the discovery probability and shortening the completion time under finite time constraints.
[0016] Specifically, the following steps are included:
[0017] Step 1:
[0018] Set up a two-dimensional map model of the simulated sea area, and define the area size MapSize and the grid size gridSize.
[0019] Step 2:
[0020] Initialize the overall parameters of the simulation environment and set the number of UUVs. num Set the time parameter Δt and the total simulation time T. total Let Np be the number of time steps. The formula for calculating the number of time steps is as follows:
[0021]
[0022] Step 3:
[0023] Set the motion performance parameters of the UUV, including the speed range, heading angular velocity range, and acceleration range. Set the motion parameters of the dynamic target and generate the trajectory. num A dynamic target trajectory, with the probability of target existence in each trajectory. p (i) Equal distribution:
[0024]
[0025] Step 4:
[0026] An initial probability map (POC) is obtained based on the current location of the target on each trajectory. The cumulative detection probability (CDP) is obtained by combining the instantaneous detection probabilities of all UUVs to the grid, and then weighted to obtain the cumulative detection probability (WCDP) as an indicator of task completion, as shown in the following formula:
[0027]
[0028] In the formula, WCDP previous It is the WCDP of the previous moment, Trajectory num It refers to the number of target trajectories. p (i) represents the probability of the target existing in the i-th trajectory, and CDP(i) represents the cumulative detection probability of the i-th trajectory at the current time.
[0029] Step 5:
[0030] Set the mode change flag (Mode) and define the growth rate (growth) to monitor the transition from the first-stage particle swarm optimization to the second-stage traversal search mode. When the second stage ends and WCDP fails to reach 1, the system proceeds to the third stage, which combines the harvest value map with the particle swarm algorithm.
[0031] Step 6:
[0032] The first stage involves optimizing the search path using an adaptive particle swarm optimization algorithm. Particle swarm parameters are initialized, and the fitness function is set to the increase of WCDP at each time step. The fitness of each particle is calculated using the fitness function, and the optimal fitness of individual particles and the global optimal particle fitness are updated.
[0033] Step 7:
[0034] Update the inertia weight w. The inertia weight w is updated in conjunction with the progress of WCDP. Set the partition threshold WCDP_threshold for w, and set the parameter b to control the update of w, as shown in the following formula:
[0035]
[0036] In the formula, w max The maximum inertia weight; w min This represents the minimum inertia weight.
[0037] Step 8:
[0038] The individual learning factor c1 and the social learning factor c2 are updated. To improve search efficiency, the individual learning factor c1 and the social learning factor c2 in the algorithm are nonlinearly adaptively adjusted according to WCDP. Their relationship satisfies the total conservation. The core calculation formula is as follows:
[0039] c1 = C sum -c2 (14)
[0040]
[0041] In the formula, c 2_max c is the maximum social learning factor; 2_min It is the minimum social learning factor.
[0042] Step 9:
[0043] Update the velocity and position of each particle.
[0044] Step 10:
[0045] Once the maximum number of iterations is reached, the global optimum is output, and the UUV movement is guided.
[0046] Step 11:
[0047] When the particle swarm optimization algorithm stalls or its search speed decreases, the second stage of traversal search is initiated. The Hungarian algorithm is used to select the nearest sub-region for each UUV as the traversal task area. Based on the UUV's position at the time of mode change, the nearest corner is selected as the traversal starting point, generating a serpentine traversal path.
[0048] Step 12:
[0049] UUV performs a traversal search according to the traversal path.
[0050] Step 13:
[0051] When the traversal phase ends, the UUV has passed through all traversal path points, and the WCDP has not reached 1, so it enters the third phase. Upon entering the third phase for the first time, a coverage map (Coverage_map) is drawn based on the historical UUV paths. A harvest value map (harvest_map) is then constructed using the Coverage_map, as shown in the following formula:
[0052] harvest_map(i,j)=POC(i,j)·(1-Coverage_map(i,j)) (16)
[0053] Step 14:
[0054] The third stage performs the same particle swarm search as the first stage, but the fitness function is changed to a weighted form that combines the harvest value map (harvest_map) and the growth rate of WCDP.
[0055] Step 15:
[0056] The output of the optimal global value guides the movement of the UUV until all targets have been searched.
[0057] The present invention has the following beneficial effects:
[0058] (1) The method described in this invention can quickly lock in an effective search direction in the early stage of the search and steadily improve the detection efficiency. This invention adopts an adaptive fusion transition search strategy triggered by the growth rate of the weighted cumulative detection probability. First, the particle swarm quickly expands the whole domain for exploration. When the gain slows down, it automatically transitions to regional serpentine traversal to fill blind spots. Then, it guides secondary optimization based on the historical coverage map and the harvest value map of the fusion prior probability, so as to take into account both the overall progress and the residual probability cleanup, thereby maintaining a stable upward progress curve in the middle and later stages.
[0059] (2) The method described in this invention can plan smoother and more efficient UUV navigation paths, avoiding ineffective maneuvers and local stagnation. Its key lies in overcoming the limitations of a single algorithm. When the efficiency of the particle swarm optimization algorithm, which relies on stochastic optimization, decreases, the system decisively switches to a deterministic partitioned serpentine traversal strategy that ensures complete coverage. This structured search mode guarantees no omissions in the search of designated areas, avoiding the path repetition and stagnation problems caused by getting trapped in local optima in traditional methods.
[0060] (3) This invention significantly improves the search success rate and overall efficiency through a strategy adaptive fusion mechanism. Figure 4 It can be seen that the weighted cumulative detection probability of the method described in this invention reaches 1 after 10 hours, while the weighted cumulative detection probability of the traditional particle swarm algorithm only reaches 0.68 after 10 hours. Based on data observation at 10 hours, the method described in this invention improves upon the traditional particle swarm algorithm by 32%. Regarding the success rate of target search, this invention significantly improves the search efficiency for all targets. Furthermore, the traditional particle swarm algorithm's weighted cumulative detection probability progresses slowly over a long period, failing to find subsequent targets, while the method described in this invention can perform a complete search for all targets, demonstrating its superior target search capability. Through a strategy adaptive fusion mechanism, the system can adopt the optimal strategy at different stages, thereby significantly improving the search success rate and overall efficiency. This invention can evaluate the effectiveness of the current strategy in real time, intelligently fusing and transitioning between three modes: adaptive particle swarm, partitioned serpentine traversal, and secondary search combined with a harvest value map. This design balances coverage breadth and local accuracy, solving the bottleneck of inconsistent efficiency of traditional single algorithms at different search stages, ensuring that the task can be executed efficiently and completely. Attached Figure Description
[0061] Figure 1 This is a flowchart of a multi-UUV dynamic target search method with strategy adaptive fusion as described in this invention;
[0062] Figure 2 This is a simulation diagram of the search path for 100 dynamic targets using 4 UUVs according to the method of the present invention;
[0063] Figure 3 Simulation diagram of the search path for 100 dynamic targets using 4 UUVs in a traditional particle swarm search;
[0064] Figure 4 The graph shows the weighted cumulative detection probability over time as described in the invention.
[0065] Figure 5 The graph shows the weighted cumulative detection probability over time for the traditional particle swarm optimization algorithm. Detailed Implementation
[0066] Figure 1 This is a flowchart of a multi-UUV dynamic target search method with policy adaptive fusion according to the present invention, including the following steps:
[0067] Step 1:
[0068] The simulated sea area is set as a two-dimensional raster map, with each row and column having MapSize=6 grids, and each grid has a size of gridSize=8, in kilometers.
[0069] Step 2:
[0070] Initialize the overall parameters of the simulation system and set the number of UUVs to UUV. num =4, set the time step to Δt = 0.1h, and the total simulation time to T. total =15h. The total number of time steps, Np, represents the total number of time steps in the simulation, including the initial probe and the number of probes performed in each subsequent time step. The formula for the total number of time steps is as follows:
[0071]
[0072] In the formula, T total Δt represents the total simulation time; Δt represents the time step.
[0073] Step 3:
[0074] The parameters for the UUV and the moving target are set, with the UUV's maximum speed at 10 knots and a heading angle range of 72 degrees; the trajectory is randomly generated in the simulation. num = 100 dynamic target trajectories, with the starting point of each trajectory randomly generated within the map area, the speed ranging from 1 to 6 knots, and the heading angle randomly varying between -180 degrees and 180 degrees.
[0075] Step 4:
[0076] Initialize the positions of the UUVs. In this simulation, four UUVs are deployed in the four corners of the map and given initial headings and velocities.
[0077] Based on all possible target trajectories generated, an initial target existence probability map (POC) is calculated. Initially, each trajectory has an equal initial probability, as shown in the following formula:
[0078]
[0079] Trajectory p (i) represents the probability of the target existing in the i-th path, Trajectory num This represents the number of trajectories.
[0080] The initial position of each trajectory is mapped to the corresponding raster index, and its existence probability is accumulated on that raster to form an initial POC map.
[0081] Step 5:
[0082] At the start of each simulation time step, based on the location of the i-th UUV and the location of the j-th target trajectory, the instantaneous detection probability IDP(j,i) is calculated. IDP(j,i) is based on the distance between the UUV and the target, as well as a fixed loss due to ocean wave noise. The maximum detection threshold is R = 3.6 km. The comprehensive instantaneous detection probability CIDP(i) of all UUVs on the same trajectory i is calculated, which is the probability that at least one UUV detects the target on that trajectory, as shown in the following formula:
[0083]
[0084] Step 6:
[0085] Using the comprehensive instantaneous detection probabilities CIDP(i) from the start of the simulation to the current time, the cumulative detection probability CDP(i) of the i-th trajectory is calculated. The intermediate matrix H is then calculated, where H(i,k) represents the minimum undetected probability from time i to k.
[0086]
[0087] The cumulative detection probability cdp at time i is calculated using the following iterative formula. H (i):
[0088] cdp H (i) = 1 - (1 - a) i-1 H(1,i)-cdp H (q-1)H(q,i)(1-a) i-q (5)
[0089] In the formula, 'a' is a constant related to the time step Δt. The final cdp is obtained. H The last element of the vector is the CDP at the current moment.
[0090] Then, the weighted cumulative detection probability (WCDP) at the current moment is calculated as an indicator of overall search progress, using the following formula:
[0091]
[0092] In the formula, WCDP previous It is the WCDP of the previous moment, Trajectory num It is the number of paths, Trajectory p(i) represents the probability of the i-th trajectory at the previous time step, and CDP(i) represents the cumulative detection probability of the i-th trajectory at the current time step.
[0093] Step 7:
[0094] Based on the new detection results, the existence probability of all target trajectories (Trajectory) is updated using Bayes' theorem. p This updates the entire POC map. The update formula is as follows:
[0095] Trajectory p_new (i) = Trajectory p_old (i)·(1-CDP(i)) (7)
[0096] In the formula, Trajectory p_new (i) represents the probability of the existence of the i-th trajectory at the current moment. p_old (i) represents the existence probability of the i-th trajectory at the previous moment, and CDP(i) represents the current cumulative detection probability of the i-th trajectory.
[0097] Step 8:
[0098] This invention sets a mode change flag, Mode, with an initial value of 0. When Mode = 0, the system executes an adaptive particle swarm optimization algorithm to optimize the search path. The particle swarm optimization algorithm can quickly and efficiently search the task area in the early stages. The system continuously monitors the WCDP growth rate, calculated using the following formula:
[0099] growth=(WCDP(n+1)-WCDP(n-min_window+1)) / min_window (8)
[0100] In the formula, min_window = 10 is a time window; WCDP(n) represents the WCDP at time n.
[0101] When Mode is 0, and WCDP > 0.6 and growth < 0.01, indicating that the UUV search efficiency is consistently low, this search mode has encountered difficulties and requires a different approach. In this case, the Mode flag is set to 1. When Mode = 1, the system performs a regional traversal search. The traversal mode UUV performs a serpentine traversal search of the target region. The UUV uses the Hungarian algorithm to allocate the nearest task sub-region, and then performs a serpentine search covering the task region according to the serpentine trajectory. When the traversal ends and WCDP has not reached 1, Mode is set to 2. When Mode = 2, the UUV reverts to particle swarm search, but a coverage map update is performed before the search. The UUV uses the coverage map and maximizes WCDP as the basis for updating the particle swarm fitness, quickly harvesting the remaining trajectory probabilities in the POC map.
[0102] During the entire search process, when WCDP reaches 1, the simulation will end early, indicating that the search is complete.
[0103] Step 9:
[0104] When Mode = 0, the system runs the particle swarm optimization algorithm. The particle swarm parameters are initialized; each particle represents a set of UUV motion information for the next time step, including the UUV's position, heading, and velocity. The number of iterations, iter_max = 50, and the number of particles, particle_num = 50. An inertia weight w is set. The inertia weight w is an important parameter balancing the algorithm's exploration and development capabilities, directly affecting the algorithm's convergence ability and optimization performance. The adaptive inertia weight implemented in this invention helps the algorithm find the optimal solution more quickly. An inertia weight update threshold b is set, as shown in the following formula:
[0105]
[0106] In the formula, WCDP_threshold controls the form of change in inertia weight; b is an intermediate variable that controls the change of w.
[0107] When t < 1, w gradually decreases as WCDP increases, and the algorithm shifts from exploration to fine-grained search as w decreases. When t ≥ 1, w quickly reverts to its maximum value and decreases exponentially as WCDP increases, ensuring that the system can still balance exploration and fine-grained search even when WCDP is at a high level, thus possessing better global exploration and fine-grained development capabilities. The update formula for w is shown below:
[0108]
[0109] In the formula, w max =0.9 is the maximum value of w; w min=0.6 is the minimum value of w.
[0110] Step 10:
[0111] The fitness of each particle is updated. After the fitness update, the individual best fitness and individual best position of each particle are updated, and the globally best particle is also updated. The fitness function fitness_1 of the particle swarm optimization algorithm is shown below:
[0112] fitness_1 = WCDP - WCDP previous (11)WCDP previous This represents the WCDP at the previous time step.
[0113] Step 11:
[0114] To enable the exploration and convergence strategies of the particle swarm optimization algorithm to be dynamically and interrelatedly adjusted based on search performance, this invention proposes a nonlinear adaptive adjustment mechanism that conserves the total amount of individual learning factors c1 and social learning factors c2. To maintain a dynamic balance throughout the process, we set a total constant C... sum =4. The core of this mechanism is that the changes in c1 and c2 are driven by the current weighted cumulative probe probability (WCDP). The formula for calculating the adaptive social learning factor c2 is designed as follows:
[0115]
[0116] In the formula, c 2_max =2.5 is the maximum value of c2; c 2_min =1 is the minimum value of c2.
[0117] After calculating c2, the adaptive individual learning factor c1 will be directly derived according to the law of conservation of total quantity:
[0118] c1 = C sum -c2 (13)
[0119] Equations 12 and 13 ensure the correlation between c1 and c2. When c2 decreases, c1 increases accordingly, making the particles more confident in their own historical experience, which is beneficial for exploration; when c2 increases, c1 decreases, prompting the particles to learn more from collective intelligence, which is beneficial for convergence.
[0120] Step 12:
[0121] Update the velocity and position of each particle, velocity V p Update the velocity vector of the p-th particle, used to update the particle's position in each iteration. The formula is as follows:
[0122] V p_new =w·V p_old+c1·r1·(PBest p -X p )+c2·r2·(GBest-X p (14)
[0123] In the formula, w is the inertia weight; c1 is the individual learning factor; c2 is the social learning factor; r1 and r2 are random numbers between 0 and 1; PBest p The optimal position for an individual particle; GBest is the globally optimal particle position; X p This represents the particle's position.
[0124] The position update formula is as follows:
[0125] X p_new =X p_old +V p (15)
[0126] In the formula, X p_new This is the updated particle position; X p_old It represents the particle position before the update; V p The velocity of the particle.
[0127] When determining particle positions, boundary constraints are applied to the particle positions. When the number of iterations reaches the maximum value iter_max, the position of the globally optimal particle is output as the motion information of the UUV at the current moment, including position, velocity, and heading.
[0128] Step 13:
[0129] When Mode=1, the system performs a region-based traversal search. Upon first entering this stage, the system assigns an optimal dedicated search region to each UUV. First, a cost matrix C is constructed, whose elements are the Euclidean distances from each UUV to the center of each pre-divided region:
[0130] C(i,j)=||UUV pos (i)-Region center (j)|| (16)
[0131] In the formula, C(i,j) is the element in the i-th and j-th column of matrix C; UUV pos (i) represents the coordinate position of the i-th UUV; Region center (j) represents the coordinates of the center of the j-th region.
[0132] The Hungarian algorithm is used for region allocation. The Hungarian algorithm is a combinatorial optimization algorithm for solving assignment problems. It can find an optimal matching scheme that minimizes the total travel distance of all UUVs in the allocation of UUVs to corresponding sub-regions.
[0133] Step 14:
[0134] After allocating the area, when Mode=1, the system generates an efficient and comprehensive coverage path for each UUV within its designated area and directs the UUV to navigate along that path. First, to generate the optimal traversal path, the system needs to determine the starting point. This starting point should be the center of the grid corresponding to the nearest corner of the area to the UUV's current position. This ensures that the UUV can enter its assigned area with the shortest possible journey to begin its search. Second, the system determines the main scanning direction of the traversal. By comparing the width and height of the allocated area, it decides whether to perform a row-first or column-first serpentine scan to reduce the number of turns and improve efficiency. Then, based on the determined starting corner and main scanning direction, the system generates an ordered sequence of waypoints. This sequence plans the path row by row or column by column, using the grid center as the unit, and reverses direction at the end of each row or column, forming a serpentine trajectory. This ensures that every grid within the allocated area is covered without generating duplicate paths. The UUV executes the traversal path, and at each time step, the UUV heads towards its current target point in the path sequence.
[0135] Step 15:
[0136] When the traversal ends and WCDP fails to reach 1, the system enters Mode=2 mode, and the system runs the particle swarm optimization algorithm again for path optimization. A harvest value map is constructed, and the map's value is determined by two parts: the current POC map and the historical UUV coverage map (Coverage_map). The formula for calculating coverage is as follows:
[0137]
[0138] In the formula, center(i,j) represents the coordinates of the center point of the grid in Coverage_map; N history It is the total number of all historical UUV trajectory points; UUV pos This indicates the location of the UUV; σ = 3 is a standard deviation parameter used to control the influence range of a single detection point.
[0139] The entire Coverage_map is normalized so that its values range from [0, 1], where 1 represents the most fully covered area and 0 represents a completely uncovered area. The system combines the latest Probability of Target (POC) map and the newly calculated coverage map to create the final Harvest_map. The value of this map is determined by the following formula:
[0140] harvest_map(i,j)=POC(i,j)·(1-Coverage_map(i,j)) (18)
[0141] In the formula, POC(i,j) represents the target probability map value; Coverage_map(i,j) represents the coverage map value.
[0142] Step 16:
[0143] When Mode = 2 and the harvest value map has been obtained, the system reverts to the same adaptive particle swarm optimization phase as when Mode = 0. The fitness function of the particle swarm algorithm is modified, and the fitness function fitness_2 formula is shown below:
[0144] fitness_2=α·ΔWCDP+(1-α)·harvest_map(i,j) (19)
[0145] In the formula, α = 0.3 is the weight; harvest_map represents the harvest value map value of the grid where the UUV is located.
[0146] Step 17:
[0147] When Mode=2, the system will continue to optimize the search path using the particle swarm optimization algorithm combined with the harvest value map until WCDP reaches 1 or the number of simulation iterations reaches the total number of simulation steps.
[0148] The method of this invention was simulated using MATLAB software. Figure 2 This is a search route map for searching 100 dynamic targets using 4 UUVs according to the method described in this invention; Figure 3 A search roadmap for 4 UUVs searching 100 dynamic targets in traditional particle swarm optimization. Figure 4 This is a graph showing the weighted cumulative detection probability over time as described in this invention. Figure 5 The graph shows the weighted cumulative detection probability over time for the traditional particle swarm optimization algorithm. Figure 2 It can be intuitively seen that the UUV search path in the method described in this invention does not experience a long-term search stagnation, and the movement of the UUV in each stage is smooth and covers a wide range. Figure 3 In conventional particle swarm optimization, the search path for UUVs is concentrated in the center of the region, resulting in a long-term stagnation and failure to effectively search for the target; from Figure 4 The method described in this invention successfully searched for all dynamic targets in just 10 hours, while... Figure 5Traditional particle swarm optimization (PSO) for dynamic targets took 15 hours and failed to search all targets. Furthermore, after 10 hours, the weighted cumulative detection probability reached only 0.68, and the curve had already stagnated. This demonstrates that the search efficiency of this invention is significantly higher and faster than the traditional PSO algorithm. Based solely on 10 hours of data, the method described in this invention achieves a 32% improvement over traditional PSO and is able to search all targets. Figure 4 It can be observed that the weighted cumulative detection probability curve of the method described in this invention rises significantly without stagnation. It effectively addresses the situation where the target is not found during the transitions at each stage, and can improve the current problem through other strategies when the weighted cumulative detection probability progresses slowly. Figure 5 Traditional particle swarm optimization methods often fail to escape the state of prolonged target stagnation when stuck in the search, resulting in local convergence and prolonged search stagnation. The above data demonstrates that the method described in this invention has advantages in search efficiency, search success rate, and the rationality of UUV movement paths, providing a new solution for dynamic target search path planning.
Claims
1. A multi-UUV dynamic target search method with adaptive strategy fusion, characterized in that, Includes the following steps: Step 1: A two-dimensional map model of the mission area is set up, and the overall parameters of the simulation environment are initialized, including setting the number of UUVs. num Simulation time step Δt, total simulation time T total And obtain the total number of time steps Np; Step 2: Initialize the positions of multiple UUVs in the mission area and generate the trajectory. num A dynamic target trajectory is generated, and based on all possible target trajectories, an initial target existence probability map (POC) is calculated. Initially, each trajectory has an equal probability of existence. The initial position of each trajectory is mapped to the corresponding grid, and its existence probability is accumulated on that grid to form an initial POC map. Step 3: At each time step, the comprehensive instantaneous detection probability CIDP is calculated based on the current UUV position and the positions of all possible target trajectories, and the cumulative detection probability CDP(i) of each trajectory is updated. The weighted cumulative detection probability (WCDP), which measures the overall progress of the mission, is updated according to the following formula: In the formula, WCDP previous It is the WCDP of the previous moment; Trajectory p (i) represents the probability of the target existing in the i-th trajectory; CDP(i) represents the cumulative detection probability of the i-th trajectory at the current time. Step 4: The UUV search path is optimized in the first stage by the particle swarm optimization algorithm. The particle swarm parameters are initialized, and the fitness function `fitness_1` aims to maximize the WCDP at each time step. The fitness of each particle is calculated, and the individual optimal position and fitness of each particle are updated. The globally optimal particle and its fitness are also updated. The inertia factor `w` in the particle swarm optimization algorithm is adjusted as the WCDP increases, as shown in the following formula: In the formula, w max The maximum inertia weight; w min is the minimum inertia weight; WCDP_threshold is the partitioning threshold of w; b is the intermediate variable set for w partitioning; Step 5: In the particle swarm optimization algorithm, the individual learning factor c1 and the social learning factor c2 are nonlinearly adaptively adjusted based on the current WCDP, and their relationship satisfies the total conservation. The calculation formula is as follows: c1=C sum -c2 (4) In the formula, c 2_max c is the maximum value of the social learning factor. 2_min C represents the minimum value of the social learning factor. sum This represents the total amount of individual learning factor c1 and social learning factor c2. Step 6: Update the velocity and position of each particle, and constrain the particle position to the correct region; after the iteration, output the globally optimal particle position and update the position of the UUV; Step 7: At each time step, the growth rate of WCDP is monitored in real time. If the growth of WCDP does not exceed the threshold within the time window of min_window steps, the system will enter the second stage of UUV regional traversal search mode. When entering this stage for the first time, a sub-region is assigned to each UUV. This assignment is completed by constructing a cost matrix and calling the Hungarian algorithm to ensure that the path distance from all UUVs to each sub-region is the shortest. Step 8: After the area allocation is completed, a serpentine traversal path is generated for each UUV within its exclusive area. The starting point of the path is determined based on the principle of minimum distance between the UUV and the four corners of the area. The UUV will travel along the generated path that can completely cover its area of responsibility. Step 9: After all UUVs have completed area traversal, if not all targets have been searched (i.e., WCDP does not reach 1), the process automatically proceeds to the third stage; a harvest value map (harvest_map) is constructed. This map's construction relies on a coverage map (Coverage_map) calculated based on the historical trajectories of all UUVs. The calculation formula is as follows: harvest_map(i,j)=POC(i,j)·(1-Coverage_map(i,j)) (7) In the formula, center(i,j) represents the coordinates of the center point of the grid in Coverage_map; N history It is the total number of historical trajectory points of all UUVs; UUV pos Indicates the coordinate position of the UUV; σ is a standard deviation parameter used to control the influence range of a single detection point; Step 10: The particle swarm optimization (PSO) algorithm is restarted to guide multiple UUVs to move towards the region with the highest value in the harvest value map. In this stage, the fitness function `fitness_2` of the PSO algorithm is a composite value, weighted by the estimated WCDP gain and the value of the harvest value map `harvest_map` obtained on the harvest value map. Its calculation formula is as follows: fitness_2=α·ΔWCDP+(1-α)·harvest_map(i,j) (8) In the formula, α is the fitness weight; this process continues until the search task ends.
Citation Information
Cited By
Underwater vehicle optimization cooperative detection array position method
CN122155043A