Unmanned ship intelligent search and rescue method and system
Through the multi-agent deep deterministic policy gradient algorithm and the improved PV-RCNN target detection model, the path planning and target detection problems of the unmanned boat search and rescue system in complex water environments were solved, and efficient collaborative search and precise target positioning of the unmanned boat cluster were achieved, thereby improving the search and rescue efficiency and safety.
Patent Information
- Application Number
- CN202511240904.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-02
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-09-02
AI Technical Summary
The existing unmanned boat search and rescue system has problems in complex water environments, such as insufficient adaptability of path planning to dynamic environments, lack of real-time feedback of collaborative strategies, and low target detection accuracy, resulting in low search and rescue efficiency and safety risks.
A multi-agent deep deterministic policy gradient algorithm is combined with an improved PV-RCNN target detection model. Dynamic environmental parameters are collected through distributed perception nodes to generate collaborative search paths. The motion control of the unmanned boat cluster is adjusted in real time. Point cloud and image information are integrated for target detection, forming a closed-loop linkage between detection and motion control.
It has achieved efficient collaborative search by unmanned boat clusters in complex water environments, improved the accuracy of target detection and search and rescue efficiency, and ensured the safety and real-time nature of search and rescue operations.
Smart Images

Figure CN120722907A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of unmanned boat search and rescue, and in particular to an unmanned boat intelligent search and rescue method and system. Background Art
[0002] In the field of water search and rescue, traditional search and rescue methods are limited by factors such as complex hydrological environments, adverse weather conditions, and delayed manual operation responses, making it difficult to quickly locate targets and efficiently rescue them. With the development of artificial intelligence and unmanned system technologies, unmanned boats have gradually become important equipment for water search and rescue due to their maneuverability, flexibility, and strong adaptability. However, the perception range of a single unmanned boat is limited, and it is prone to missed detections in large search and rescue areas. In addition, the autonomous decision-making ability is insufficient when faced with environmental interference such as dynamically changing water currents and wind. Therefore, unmanned boat swarm search and rescue systems based on multi-agent collaborative mechanisms and advanced target detection technologies have become a research hotspot. They aim to improve search and rescue coverage and target detection efficiency through multi-boat collaborative operations to meet emergency rescue needs in complex water environments.
[0003] Existing technologies have two significant shortcomings in the application of unmanned boat search and rescue. On the one hand, the collaborative path planning of multiple unmanned boats lacks real-time adaptability to dynamic environments. Traditional path planning methods mostly generate fixed paths based on preset environmental models. When environmental parameters such as water flow speed and wind force level suddenly change, it is difficult to quickly adjust the motion trajectory of each boat, resulting in the risk of collision between boats or blind spots in the search and rescue area. In addition, the collaborative strategy between boats does not fully consider the feedback of target detection results, and cannot dynamically optimize the search resource allocation according to the distribution of targets. On the other hand, the target detection model is not adaptable enough to complex water scenes. When dealing with water surface reflections, wave interference, and partial occlusion of the target, traditional target detection methods have difficulty in accurately extracting the three-dimensional features of the target, resulting in low target recognition accuracy. In addition, the detection results do not form a closed-loop linkage with the unmanned boat motion control, and cannot provide accurate target position information support for the boat group to adjust the search and rescue strategy in real time. Summary of the Invention
[0004] In order to overcome the shortcomings and deficiencies of the prior art, the present invention provides an unmanned boat intelligent search and rescue method and system.
[0005] The technical solution adopted by the present invention is an unmanned boat intelligent search and rescue method, comprising the following steps: Step S1: Deploy a swarm of unmanned boats driven by a multi-agent deep deterministic policy gradient algorithm to collect dynamic environmental parameters within the search and rescue area through distributed sensing nodes. The dynamic environmental parameters include water velocity vector, wind force level, three-dimensional coordinates of water obstacles, and initial probability distribution area of the target object; Step S2: Based on the dynamic environment parameters, the Actor-Critic framework in the multi-agent deep deterministic policy gradient algorithm is used to generate the initial collaborative search path of the unmanned boat cluster. The initial collaborative search path must meet the communication distance constraints and collision avoidance safety thresholds between the unmanned boats. Step S3: During the initial collaborative path search performed by the unmanned boat cluster, the improved PV-RCNN target detection model is activated, and water image data is collected in real time through the multispectral imaging equipment carried by the unmanned boat cluster. When the improved PV-RCNN target detection model extracts features from the water image data, it fuses the point cloud data with the image pixel information to construct a 3D candidate frame; Step S4: The improved PV-RCNN object detection model classifies and regresses the 3D candidate boxes, outputs the 3D spatial coordinates and confidence values of the target objects, and transmits the 3D spatial coordinates and confidence values to the state space of the multi-agent deep deterministic policy gradient algorithm; Step S5: The multi-agent deep deterministic policy gradient algorithm evaluates the value function of the current policy through the Critic network based on the updated information in the state space, and adjusts the motion control parameters of each unmanned boat in combination with the Actor network; Step S6: Repeat steps S3 to S5 until the target object confidence value output by the improved PV-RCNN target detection model reaches a preset threshold. At this time, the multi-agent deep deterministic policy gradient algorithm generates a target encirclement trajectory for the unmanned boat cluster, and the target encirclement trajectory adapts to the motion state parameters of the target object and real-time environmental interference factors.
[0006] Furthermore, in step S2, when the multi-agent deep deterministic policy gradient algorithm generates the initial collaborative search path, a policy optimization function based on the topological structure of the search and rescue area is introduced:
[0007] in, Indicates that the status The agent selects an action The probability distribution of is the activation function, are weight coefficients, represents the Euclidean distance between the i-th and j-th unmanned boats, Indicates that the search and rescue area The water flow gradient vector at the coordinate, represents the repulsive potential field function of the k-th obstacle; in step S5, the motion control parameter adjustment process must satisfy:
[0008] in, is the propeller output power of the i-th unmanned boat at time t, is the power attenuation coefficient, is the policy update rate, Represents the gradient of the action-value function with respect to the action.
[0009] Furthermore, in step S3, when the improved PV-RCNN target detection model fuses point cloud data with image pixel information, a feature enhancement module is used:
[0010] in, is the fused feature matrix, is the feature fusion coefficient, is the point cloud feature matrix, is the image feature matrix, represents the matrix Hadamard product, represents the matrix concatenation operation, represents the convolution operation; in step S4, the regression loss function of the three-dimensional candidate box is defined as:
[0011] in, is the translation error loss, is the size error loss, is the rotation angle error loss, are the weight coefficients of each loss.
[0012] Furthermore, in step S5, when the critic network of the multi-agent deep deterministic policy gradient algorithm evaluates the value function, a global reward distribution mechanism is introduced:
[0013] in, is the global reward value, is the number of unmanned boats, is the local reward coefficient of the i-th unmanned boat, is the local reward of the i-th unmanned boat, is the local reward of the j-th unmanned boat, is the synergy reward factor, is the behavior similarity function between the i-th and j-th unmanned boats; in step S2, the path planning cost function of the initial collaborative search path is:
[0014] in, is the total path length, is the sailing time of the i-th unmanned boat, is the number of obstacles, is the avoidance cost of the k-th obstacle, are the cost weight coefficients respectively.
[0015] Furthermore, in step S3, a dynamic exposure compensation model is used when preprocessing the water area image data collected by the multispectral imaging device:
[0016] in, is the pixel value of the image after compensation, is the original image pixel value, is the light attenuation coefficient, for Depth of water at coordinates wavelength Corresponding light absorption coefficient; in step S4, the confidence value correction formula output by the improved PV-RCNN target detection model is:
[0017] in, is the corrected confidence level, is the original confidence, is the signal-to-noise ratio correction factor, is the signal-to-noise ratio of the compensated image.
[0018] Furthermore, in step S5, when the Actor network adjusts the steering angle of the unmanned boat servo, an environment adaptive correction model is introduced:
[0019] in, for The steering angle of the servo at the moment, is the steering angle of the servo at time t, is the angle adjustment coefficient, is the wind impact weight, is the wind speed at time t, is the water flow influence weight, is the water flow velocity value at time t; in step S2, the path smoothness index of the initial collaborative search path is calculated as:
[0020] in, is the path smoothness, is the number of path segments of the i-th unmanned boat, is the heading angle of the kth path of the i-th unmanned boat.
[0021] Furthermore, in step S6, during the target encirclement trajectory generation process, the multi-agent deep deterministic policy gradient algorithm adopts a dynamic encirclement radius calculation model:
[0022] in, is the enclosed radius at time t, is the initial enclosure radius, is the shrinkage rate coefficient, is the minimum enclosing radius, is the target object's moving speed, is the average speed of the unmanned boat cluster; in step S3, the point cloud feature extraction of the improved PV-RCNN target detection model adopts the voxel enhancement method:
[0023] in, Voxel The eigenvalues of is the point cloud set included in the voxel, is the weight of point p, is the reflection intensity coefficient, is the reflection intensity value of point p, is the original eigenvector of point p.
[0024] Furthermore, in step S4, the improved PV-RCNN target detection model adopts a multi-level confidence screening mechanism when classifying the three-dimensional candidate boxes:
[0025] in, is the final confidence level, Output confidence for the classifier, is the correction factor, is the intersection-over-union ratio of the candidate box B and the target real box G, is the noise level of image I; in step S5, the experience replay pool of the multi-agent deep deterministic policy gradient algorithm adopts a priority sampling strategy:
[0026] in, is the sampling probability of the jth experience, is the TD error of the j-th experience, is a small constant, is the priority index, The capacity of the experience replay pool.
[0027] Furthermore, in step S1, when the distributed sensing nodes collect dynamic environmental parameters, a spatiotemporal fusion filtering model is used:
[0028] in, is the environmental parameter value after fusion at time t, is the current data weight, is the real-time data collected at time t, is the size of the historical data window, is the weight of the k-th historical data, for The historical data at the moment; in step S2, when the multi-agent deep deterministic policy gradient algorithm generates the initial collaborative search path, the energy constraint condition is satisfied:
[0029] in, is the total time of path planning, is the power consumption of the i-th unmanned boat at time t, is the total energy budget of the UAV swarm.
[0030] An unmanned boat intelligent search and rescue system, comprising: A distributed dynamic environmental parameter acquisition unit, which acquires water flow, wind force, and obstacle-related data through a sensor array deployed in the search and rescue area, and establishes a two-way data transmission link with the multi-agent decision-making control unit; A multi-agent decision control unit receives data transmitted by a distributed dynamic environment parameter acquisition unit, runs a multi-agent deep deterministic policy gradient algorithm to generate path planning instructions and motion control signals, and sends them to the unmanned boat cluster execution unit and the target detection coordination unit respectively; The unmanned boat swarm execution unit includes multiple unmanned boat individuals, receives the motion control signal output by the multi-agent decision-making control unit, adjusts the position and attitude through the propeller and steering gear adjustment module, and feeds back its own state data to the multi-agent decision-making control unit; A multispectral image acquisition unit, which is mounted on each unmanned boat of the unmanned boat cluster execution unit, captures water area image data in real time and transmits it to the improved three-dimensional target detection unit via a high-speed data interface; An improved 3D target detection unit, which runs an improved PV-RCNN target detection model, processes the image data transmitted by the multispectral image acquisition unit, and outputs the target's 3D coordinates and confidence information to the multi-agent decision control unit. The cluster collaborative communication unit establishes a wireless communication network between the individual unmanned boats in the unmanned boat cluster execution unit, conducts real-time interaction of status data and detection information, and feeds back the network topology status to the multi-agent decision-making control unit.
[0031] Beneficial Effects: The present invention proposes an intelligent search and rescue method and system for unmanned boats. By integrating a multi-agent deep deterministic policy gradient algorithm with an improved PV-RCNN target detection model, the method effectively overcomes the shortcomings of existing technologies and brings significant beneficial effects. In terms of multi-boat collaboration, with the help of the Actor-Critic framework of the multi-agent deep deterministic policy gradient algorithm, the unmanned boat swarm can dynamically adjust the collaborative search path based on real-time collected environmental parameters such as water flow and wind force. The safety of the movement is ensured through communication and collision avoidance constraints between each boat. At the same time, the motion control parameters are continuously optimized based on the target detection results to achieve dynamic allocation of search and rescue resources, solving the problems of traditional path planning's lack of adaptability to dynamic environments and the lack of feedback in collaborative strategies. In terms of target detection, the improved PV-RCNN target detection model fuses point cloud and image information to construct a three-dimensional candidate frame, improving the accuracy of target three-dimensional feature extraction in complex water environments. The detection results are directly connected to the state space of the multi-agent algorithm, forming a closed-loop linkage between detection and motion control, providing the boat swarm with accurate target location information. This overcomes the shortcomings of traditional models' poor adaptability to complex scenes and the disconnection between detection and control, and overall improves search and rescue efficiency and target detection capabilities. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 is a flow chart of the method steps of the present invention; Figure 2 It is a diagram of the system unit composition of the present invention. DETAILED DESCRIPTION
[0033] It should be noted that, unless there is a conflict, the embodiments in this application and the features described in the embodiments can be combined with each other. The application is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0034] like Figure 1 As shown, an unmanned boat intelligent search and rescue method includes the following steps: Step S1: Deploy a swarm of unmanned boats driven by a multi-agent deep deterministic policy gradient algorithm to collect dynamic environmental parameters within the search and rescue area through distributed sensing nodes. The dynamic environmental parameters include water velocity vector, wind force level, three-dimensional coordinates of water obstacles, and initial probability distribution area of the target object; Specifically, step S1 comprehensively and accurately collects dynamic environmental parameters of the search and rescue area, providing a reliable decision-making basis for subsequent path planning, target detection, and collaborative control. The collected dynamic environmental parameters cover multiple key dimensions. Among them, the water velocity vector includes not only the speed of the water flow, but also the direction of the water flow. This parameter is directly related to the calculation of the energy consumption of the unmanned boat during navigation and the deviation correction between the actual navigation trajectory and the preset trajectory. The wind force level reflects the degree of influence of the wind on the navigation stability of the unmanned boat. Different wind force levels will cause the unmanned boat to be subjected to different lateral forces, which in turn affects its course-keeping ability. The three-dimensional coordinates of water obstacles can provide accurate spatial location information for the collision avoidance algorithm of the unmanned boat cluster, ensuring that the unmanned boat will not collide with obstacles such as reefs and sunken ships during navigation. The initial probability distribution area of the target object is delineated based on the initial information of the search and rescue mission. It can guide the unmanned boat cluster to focus its main search efforts on the area where the target object is most likely to appear in the early stage of the search and rescue, thereby improving the efficiency of the search and rescue and avoiding the waste of resources. The accuracy and real-time performance of these parameters directly determine whether the subsequent multi-agent deep deterministic policy gradient algorithm and the improved PV-RCNN target detection model can work normally and efficiently, and are the prerequisite for the smooth operation of the entire search and rescue system.
[0035] During the specific implementation process, it is first necessary to complete the deployment and debugging of the multi-agent deep deterministic policy gradient algorithm in the unmanned boat swarm control system. The technicians will implant the core code of the algorithm into the main control unit of each unmanned boat and conduct compatibility tests between the algorithm and the unmanned boat hardware equipment to ensure that the algorithm can normally call the sensors, thrusters, servos and other equipment of the unmanned boat. Subsequently, the distributed perception node is started to collect dynamic environmental parameters. The flow velocity sensor on the unmanned boat will sample the water velocity at a frequency of 5 times per second, and the sampling cycle lasts for 10 minutes. By statistically analyzing these sampled data, the average value and fluctuation range of the water velocity vector are calculated. The measurement range of the flow velocity is 0-5 meters / second, and the direction measurement accuracy is ±5 degrees; the anemometer will monitor the wind force level in the search and rescue area in real time. Its measurement range is 0-12 levels, with a measurement accuracy of ±0.5 levels, and the wind speed value corresponding to the wind force level (such as level 0 corresponds to 0-0.2 meters / second, level 1 corresponds to 0.3-1.5 meters / second, etc.) is recorded; the three-dimensional lidar will measure the surrounding water area. Obstacles are scanned 10 times per second across a 360-degree horizontal and -15 to +90-degree vertical range. Point cloud data processing technology converts the scanned obstacle information into three-dimensional coordinates with an accuracy of ±0.1 meter. Furthermore, combining alarm information from the search and rescue mission, the geographic characteristics of the incident site, and historical search and rescue data, professional geographic information system software is used to delineate an initial probability distribution area for the target. This area, centered at the incident center, has a radius of 1-5 kilometers. The probability of target occurrence at different locations within the area is weighted based on factors such as the topography and currents of the water area. All collected dynamic environmental parameters are aggregated via the local area network between the UAVs, with a transmission rate maintained at above 10Mbps. The aggregated data is verified by a data verification module to remove outliers and erroneous data. Once verified, it is stored in the UAV cluster's shared data center, which utilizes a distributed storage architecture with a total storage capacity of at least 1000GB to ensure it can accommodate long-term environmental parameter data.
[0036] Step S2: Based on the dynamic environment parameters obtained in step S1, the Actor-Critic framework in the multi-agent deep deterministic policy gradient algorithm is used to generate an initial collaborative search path for the unmanned boat cluster. The initial collaborative search path must meet the communication distance constraints and collision avoidance safety thresholds between the unmanned boats. Specifically, step S2 generates a scientific and reasonable initial collaborative search path for the UAV swarm based on the dynamic environmental parameters collected in step S1, thereby ensuring that the UAV swarm can conduct search operations efficiently and orderly in the early stages of the search and rescue operation. The Actor-Critic framework in the multi-agent deep deterministic policy gradient algorithm plays a core role in this step. The Actor network is responsible for generating multiple possible search path solutions for each UAV, while the Critic network is responsible for evaluating and optimizing these path solutions. This approach can achieve an organic combination of the local optimal path of each individual UAV and the global optimal path of the UAV swarm, avoiding the problem of global inefficiency caused by individual optimization. The communication distance constraint is a crucial condition for ensuring effective collaboration within a swarm of unmanned vehicles. It specifies the minimum and maximum communication distances that must be maintained between adjacent unmanned vehicles, ensuring stable data transmission and information exchange between them, thereby achieving coordinated coverage of the search and rescue area. The collision avoidance safety threshold is a safety line set for the swarm, specifying the minimum safe distance that must be maintained between unmanned vehicles and between unmanned vehicles and obstacles. This threshold setting fundamentally prevents collisions between unmanned vehicles during navigation, ensuring the safety of the entire search and rescue operation. These three aspects together constitute the core constraints of the initial collaborative search path, each of which is essential. Their proper setting directly impacts the quality of the initial path and the smooth progress of subsequent search and rescue operations.
[0037] During implementation, the Actor Network of the multi-agent deep deterministic policy gradient algorithm first reads the dynamic environmental parameters stored in the shared data center in step S1, including the water velocity vector, wind speed level, three-dimensional coordinates of water obstacles, and the initial probability distribution area of the target. Based on these parameters, it then generates at least 20 potential search paths for each unmanned vehicle. These paths cover different subregions of the initial probability distribution area of the target, and each path includes the coordinates of multiple waypoints and the estimated arrival time at each waypoint. Next, the Critic Network conducts a comprehensive and detailed evaluation of the potential search paths. The evaluation metrics primarily include: the target probability area covered by the path (i.e., the proportion of the high-probability subregion of the initial probability distribution area covered by the path); the overall energy consumption of the swarm (i.e., the total energy consumed by the swarm when navigating the path); and the collision safety factor (i.e., the ratio of the distance between the path and the obstacle, as well as the distance between adjacent unmanned vehicles, to the collision safety threshold; a higher ratio indicates a higher score). Based on these evaluation metrics, the critic network assigns a comprehensive score to each potential path. Then, through multiple rounds of iterative optimization, the path combination with the highest comprehensive score is selected from all potential paths to serve as the initial collaborative search path. During this selection process, strict communication distance constraints and collision avoidance safety thresholds must be met. The communication distance is set to 500-800 meters, ensuring stable communication signals between UAVs while avoiding overlapping search areas due to close proximity. The collision avoidance safety threshold is generally set at 1.5-2 times the length of the UAV. For a 6-meter UAV, the collision avoidance safety threshold is 9-12 meters. Once the initial collaborative search path is determined, the algorithm transmits the path parameters, including the specific waypoint coordinates of each UAV (accurate to 3 decimal places), the estimated arrival time at each waypoint (accurate to seconds), and the dwell time at each waypoint (10-30 seconds), to the corresponding UAV control system via an encrypted transmission protocol. The UAV control system then interprets these parameters and converts them into navigation instructions for its own navigation, preparing for subsequent search operations.
[0038] Step S3: During the initial collaborative path search performed by the unmanned boat cluster, the improved PV-RCNN target detection model is activated, and water image data is collected in real time through the multispectral imaging equipment carried by the unmanned boat cluster. When the improved PV-RCNN target detection model extracts features from the water image data, it fuses the point cloud data with the image pixel information to construct a 3D candidate frame; Specifically, step S3 organically combines the improved PV-RCNN object detection model with multispectral imaging equipment, overcoming the limitations of traditional visual detection technology in complex aquatic environments and achieving real-time, accurate perception of aquatic targets. Compared to traditional single-band imaging equipment, multispectral imaging equipment can capture image information from multiple wavelengths in aquatic environments, including visible light and near-infrared bands. This enables it to effectively enhance the contrast between the target and the background in complex conditions such as water surface reflections and wave interference, thereby increasing the likelihood of target identification. The improved PV-RCNN object detection model is an optimization of the traditional PV-RCNN model. It can achieve a deep fusion of point cloud data and image pixel information. Through this fusion, the model can more comprehensively capture the spatial and appearance characteristics of the target, thereby improving its three-dimensional form perception. The construction of a 3D candidate box is a key output of this step. It can initially define the spatial range where the target may be located, laying a solid foundation for accurate classification and localization of the target in subsequent steps, ensuring that the target detection process is carried out in an orderly and efficient manner.
[0039] During implementation, once the UAV swarm begins navigating along the initial collaborative search path generated in step S2, the system automatically activates the multispectral imaging equipment onboard each UAV. The equipment's operating parameters are set as follows: an image acquisition frequency of 15-20 frames per second ensures real-time monitoring of the water environment without missing any information about the target's motion; spectral band coverage ranges from 400 to 1000 nanometers, including visible light from 400 to 760 nanometers and near-infrared light from 760 to 1000 nanometers. By imaging in these different bands, it is possible to effectively distinguish between the target and the background; and the image resolution is set to 1920 × 1080 pixels, ensuring sufficient image detail for subsequent feature extraction. The collected multispectral image data is transmitted in real time to the improved PV-RCNN target detection model via the UAV's internal high-speed data bus, maintaining a transmission rate above 100 Mbps to prevent data transmission delays from affecting the real-time detection. After receiving image data, the improved PV-RCNN object detection model first performs image preprocessing operations, including image denoising and contrast enhancement. Denoising uses an adaptive median filter algorithm to effectively remove salt-and-pepper and Gaussian noise from the image. The filter window size automatically adjusts between 3×3 and 7×7 according to noise intensity. Contrast enhancement uses a histogram equalization algorithm to improve the grayscale difference between the target and the background. The model then extracts pixel features from the image, including color and texture. This extraction process utilizes a multi-layer convolutional neural network, with each convolution kernel size being 3×3 and a stride of 1. By stacking multiple convolutional layers, the model gradually extracts deeper features of the image. Simultaneously, the point cloud sensor onboard the unmanned vehicle simultaneously collects 3D point cloud data of the surrounding environment. The point cloud sensor scans 20 times per second, covering 360 degrees horizontally and -30 to +30 degrees vertically. The point cloud data density is no less than 50 points per square meter. After point cloud data is collected, it will first be filtered to remove invalid point clouds caused by factors such as water surface fluctuations and sensor noise. The filtering adopts a statistical filtering algorithm, sets the number of neighborhood points to 50, and the standard deviation multiple to 1.5. Point clouds outside this range will be judged as noise points and eliminated. The processed point cloud data will be fused with the previously extracted image pixel features. The fusion process adopts a feature mapping algorithm. By matching and associating the coordinate information of the two-dimensional image pixels with the coordinate information of the three-dimensional point cloud, a one-to-one correspondence between the pixels and the point cloud is established, and finally a three-dimensional candidate box containing the possible location of the target object is constructed. The size range of the three-dimensional candidate box is pre-set according to the size characteristics of common search and rescue targets, with a length range of 0.5-3 meters, a width range of 0.3-2 meters, and a height range of 0.2-1.5 meters to ensure that it can cover the vast majority of possible targets.
[0040] Step S4: The improved PV-RCNN object detection model classifies and regresses the 3D candidate boxes, outputs the 3D spatial coordinates and confidence values of the target objects, and transmits the 3D spatial coordinates and confidence values to the state space of the multi-agent deep deterministic policy gradient algorithm; Specifically, step S4 performs precise classification and regression on the 3D candidate boxes constructed in step S3, transforming the raw perception data into target location information with practical decision-making value, providing a key basis for subsequent path adjustment and coordinated control of the UAV swarm. The classification process primarily determines the presence and specific type of target within each 3D candidate box. This process allows for the selection of candidate boxes that truly contain the target from a large number of candidate boxes, while eliminating those false candidates caused by factors such as surface debris and light and shadow variations. The regression process, based on classification, further accurately calculates the target's coordinate position in 3D space, enabling the UAVs to accurately determine the target's specific location. The confidence value quantitatively assesses the reliability of the classification and regression results. It reflects the model's confidence in the detection results and provides an important reference for the subsequent algorithm to determine whether to adjust the search strategy. The results of these three aspects together constitute the key input parameters of the state space of the multi-agent deep deterministic policy gradient algorithm, achieving a qualitative leap from raw data perception to effective information extraction. It serves as a critical bridge between target detection and collaborative decision-making, directly affecting the response speed and processing accuracy of the entire search and rescue system to targets.
[0041] In practice, the improved PV-RCNN object detection model first refines the features of the 3D candidate boxes generated in step S3. For each candidate box, the model extracts multiple features within it, including texture, shape, and depth. Texture features are extracted using a gray-level co-occurrence matrix algorithm, which calculates the spatial distribution relationship between pixels of varying grayscale values within the candidate box to determine texture parameters such as energy, entropy, and contrast. Shape features are extracted using an edge detection algorithm to obtain the outline of the object within the candidate box, then calculate parameters such as the contour's perimeter, area, and bounding rectangle. Depth features are calculated based on point cloud data and reflect the vertical dimensions and morphological characteristics of the object. After feature extraction, the model uses a classifier to identify and judge these features. This classifier utilizes a multi-class classification model based on deep learning. This model has been trained on a large number of water object samples and can recognize a variety of common search and rescue targets, including human bodies, life jackets, life rafts, and shipwrecks. Based on the degree of feature matching, the classifier determines the presence and type of an object within each candidate box, with a classification accuracy of at least 90%. For candidate boxes identified as containing an object, the model's regressor further calculates the object's 3D coordinates. This coordinate system is a 3D rectangular coordinate system established with the current UAV's position as the origin, with the X-axis pointing directly in front of the UAV, the Y-axis pointing to the right, and the Z-axis pointing perpendicular to the water surface and upward. The coordinates are measured with an accuracy of ±0.1 meters. The regressor also outputs a confidence score between 0 and 1 based on the degree of feature matching. Values closer to 1 indicate higher reliability, and confidence scores greater than 0.7 are considered highly reliable. All detection results, including object type, 3D coordinates, and confidence scores, undergo a format conversion process to a unified data format. This data is then transmitted in real time to the state space database of the multi-agent deep deterministic policy gradient algorithm via the UAV's onboard communication module. This communication module utilizes wireless communication technology with a transmission rate of 5 Mbps and a transmission latency of no more than 0.5 seconds to ensure real-time information delivery. After receiving this information, the state space database will immediately update the original environmental state information and integrate the new target object information into the database, providing the latest and most accurate data support for the algorithm decision in the subsequent step S5.
[0042] Step S5: The multi-agent deep deterministic policy gradient algorithm evaluates the value function of the current policy through the Critic network based on the updated information in the state space, and adjusts the motion control parameters of each unmanned boat in combination with the Actor network. The motion control parameters include the propeller output power, the steering angle of the servo, and the search path correction coefficient. Specifically, step S5, through real-time computation using a multi-agent deep deterministic policy gradient algorithm, converts target detection information into coordinated action commands for the UAV swarm, thereby achieving adaptive optimization of the search and rescue strategy. In this step, the critic network's assessment of the current strategy provides a quantitative benchmark for decision-making. Its evaluation results comprehensively consider the target's location, environmental disturbances, and the swarm's state, directly determining whether motion parameters need to be adjusted. The actor network, based on this evaluation, generates specific control commands, ensuring that each UAV's actions meet global coordination requirements while also responsive to local environmental changes. Adjustments to motion control parameters include thruster power, servo angles, and path correction coefficients. The dynamic changes in these parameters enable the UAVs to maintain swarm coordination while adapting to disturbances such as sudden changes in water flow and wind fluctuations, significantly improving the robustness of the search and rescue process. Furthermore, this step establishes a closed-loop "detection-assessment-decision-execution" system, enabling the swarm to update its search trajectory in real time based on the target's movement, avoiding the efficiency losses associated with traditional fixed-path methods and laying the foundation for rapid target acquisition.
[0043] In practice, the critic network first reads real-time data from a state-space database via a high-speed data interface (transmission rate no less than 20 Mbps). This data includes the 3D coordinates of targets (accuracy ±0.2 meters), confidence values (range 0-1), and other state parameters, such as the GPS positioning information (update frequency 1 Hz), speed (0-8 knots), and remaining battery power (0-100%) of each unmanned vehicle, output by the improved PV-RCNN model. Environmental data such as the water velocity vector (sampling frequency 10 Hz, measurement range 0-3 m / s) and wind speed (0-6, sampling interval 5 seconds) are also included. The value function is calculated using a three-layer fully connected neural network. The input layer contains 28 feature dimensions, and the hidden layers have 128 and 64 nodes, respectively. The output layer outputs a value score from 0 to 100. A score below a preset threshold of 55 triggers a policy adjustment. The Actor network generates control parameter adjustment instructions based on the value score deviation (deviation = threshold - current score): thruster output power adjustment is based on the product of the target distance (measurement range 0-1000 meters) and the water resistance coefficient (0.8-1.5), with an adjustment step of 5% of the rated power and a range limited to 30%-90% to ensure smooth power changes; the servo steering angle is calculated based on the target azimuth (-180° to 180°), with a steering rate not exceeding 5° / s and a maximum single adjustment angle of 10° to avoid drastic changes in the boat's posture; the path correction coefficient is positively correlated with the target confidence. For every 0.1 increase in confidence, the correction coefficient increases by 0.05, ranging from 0.8 to 1.2, to guide the boat group to converge towards high-confidence areas. Control commands are sent to each UAV execution unit via an encrypted wireless link (delay ≤ 0.5 seconds). The execution unit's MCU (main frequency 1GHz) completes command parsing within 10ms, drives the electronic speed control module (response time ≤ 20ms) to adjust the propeller speed, and the servo controller (accuracy ±0.5°) to perform steering actions. At the same time, it updates the waypoint parameters of the local path planning module (update frequency 5Hz).
[0044] Step S6: Repeat steps S3 to S5 until the target object confidence value output by the improved PV-RCNN target detection model reaches a preset threshold. At this time, the multi-agent deep deterministic policy gradient algorithm generates a target encirclement trajectory for the unmanned boat cluster, and the target encirclement trajectory adapts to the motion state parameters of the target object and real-time environmental interference factors.
[0045] Specifically, step S6 achieves precise lock-on and secure encirclement of the target through continuous iterative optimization, ensuring a smooth transition from the search phase to the rescue phase of the search and rescue operation. By repeatedly executing steps S3 through S5, the improved PV-RCNN model continuously tracks the target's position changes as it moves. The confidence threshold (0.8-0.9) provides a clear trigger for the encirclement action, avoiding invalid operations due to false detections. The encirclement trajectory generated by the multi-agent deep deterministic policy gradient algorithm must meet multiple constraints, including uniform distances between each unmanned vehicle (deviation ≤ 5 meters), a safe distance from the target (3-10 meters), and adaptability to environmental interference. Dynamic trajectory adjustment ensures the encirclement pattern remains stable despite changes in water currents and wind speeds. Furthermore, this step incorporates the target's motion parameters (speed 0-2 m / s, acceleration 0-0.5 m / s²) into trajectory planning, enabling the encirclement action to predict the target's movement trends, significantly reducing the risk of escape and creating stable conditions for subsequent deployment of rescue equipment or transfer of personnel.
[0046] During implementation, the system first sets an iteration termination condition: when the improved PV-RCNN model outputs a target confidence score ≥ 0.85 for five consecutive times (at 1-second intervals), the encirclement trajectory generation process is triggered. Prior to this, steps S3 through S5 are repeated every 2 seconds. In S3, the multispectral imaging device maintains a capture rate of 15 frames per second. The generated 3D candidate boxes are processed in S4 and output as real-time target coordinates (updated at 5Hz). In S5, the Critic network performs a value assessment every 0.5 seconds. Based on the assessment results, the Actor network adjusts thruster power (by ±3% of rated power) and servo angles (by ±2°), gradually shrinking the swarm toward the target area. To generate the encirclement trajectory, the multi-agent algorithm first fits the target's motion vector based on the last 10 seconds of motion data (with a sampling interval of 0.5 seconds). This includes velocity (measurement error ≤ 0.1 m / s) and direction (error ≤ 3°). The environmental interference factor (0.1-0.3) is calculated based on the current water velocity (0-2 m / s) and wind speed (0-10 m / s). Trajectory planning uses a circular encirclement method, with an initial radius set at 1.5 times the target's maximum dimensions (length × width × height) (range 5-15 meters). The UAVs are evenly distributed around the circumference (angular deviation ≤ 5°), with spacing between adjacent UAVs at 1.2 times the radius (minimum 8 meters). Once generated, the trajectory is transmitted to each UAV via a collaborative control protocol (UDP, port 5005). The onboard control system decomposes the trajectory into continuous waypoints (one every 10 meters) with a coordinate accuracy of ±0.5 meters and an arrival time error of ≤2 seconds. During execution, the inertial navigation system (positioning accuracy ±0.3 meters) verifies the deviation between the actual position and trajectory every 0.5 seconds. When the deviation exceeds 1 meter, the Actor network immediately generates correction instructions to adjust the thruster power (±5%) and the servo angle (±3°) to ensure that the deviation is controlled within 0.5 meters within three control cycles (1.5 seconds) until a stable encirclement situation is established.
[0047] Preferably, in step S2, when the multi-agent deep deterministic policy gradient algorithm generates the initial collaborative search path, a policy optimization function based on the topological structure of the search and rescue area is introduced:
[0048] in, Indicates that the status The agent selects an action The probability distribution of is the activation function, are weight coefficients, represents the Euclidean distance between the i-th and j-th unmanned boats, Indicates that the search and rescue area The water flow gradient vector at the coordinate, represents the repulsive potential field function of the k-th obstacle; in step S5, the motion control parameter adjustment process must satisfy:
[0049] in, is the propeller output power of the i-th unmanned boat at time t, is the power attenuation coefficient, is the policy update rate, Represents the gradient of the action-value function with respect to the action.
[0050] Specifically, the process of generating the initial collaborative search path using a multi-agent deep deterministic policy gradient algorithm was optimized, and the adjustment of motion control parameters was refined. The introduced policy optimization function comprehensively considers the distance between unmanned boats, water flow gradients, and obstacle repulsion. The integration of these parameters enables the initial path to better fit the actual conditions of the aquatic environment while satisfying communication and collision avoidance constraints, thereby reducing path deviations caused by environmental interference. In terms of motion control parameter adjustment, the power attenuation coefficient and policy update rate are set to make the changes in the thruster output power smoother, avoiding the loss of the unmanned boat power system caused by power mutations, while ensuring that policy adjustments can respond to changes in target detection results in a timely manner. During implementation, the value range of each weight coefficient is first determined according to the actual environment of the search and rescue area, where the distance weight is between 0.3-0.5, the water gradient weight is between 0.2-0.4, and the obstacle rejection weight is between 0.1-0.3. Then, the optimal value is determined through multiple simulation tests; the power attenuation coefficient is generally set to 0.95-0.98, and the strategy update rate is set to 0.01-0.05 according to the frequency of target detection, so as to achieve efficient coordination of path planning and motion control.
[0051] Preferably, in step S3, when the improved PV-RCNN target detection model fuses point cloud data with image pixel information, a feature enhancement module is used:
[0052] in, is the fused feature matrix, is the feature fusion coefficient, is the point cloud feature matrix, is the image feature matrix, represents the matrix Hadamard product, represents the matrix concatenation operation, represents the convolution operation; in step S4, the regression loss function of the three-dimensional candidate box is defined as:
[0053] in, is the translation error loss, is the size error loss, is the rotation angle error loss, are the weight coefficients of each loss.
[0054] Specifically, the PV-RCNN object detection model's feature fusion method and the regression loss calculation for 3D candidate bounding boxes are improved. The feature enhancement module uses a feature fusion coefficient to fuse point cloud features with image features in a reasonable proportion. This not only preserves the 3D spatial information of the point cloud data but also fully utilizes the texture details of the image data, enhancing the richness of the fused features and enabling the model to more accurately describe the characteristics of the target object. The regression loss function comprehensively considers translation, scale, and rotation errors. By setting different weighting coefficients, it can prioritize different error terms based on the actual needs of target detection. For example, in search and rescue, when high target position accuracy is required, the weight of the translation error loss can be increased. During implementation, the feature fusion coefficient is dynamically adjusted between 0.4 and 0.6 based on the lighting conditions in the aquatic environment. In good lighting, the weight of image features is increased, while in poor lighting, the weight of point cloud features is increased. The weighting coefficients for the translation, scale, and rotation error losses are initially set to 0.5, 0.3, and 0.2, respectively. These weights are then iteratively optimized based on the loss trends during model training until the model detection accuracy reaches the preset standard.
[0055] Preferably, in step S5, when the critic network of the multi-agent deep deterministic policy gradient algorithm evaluates the value function, a global reward distribution mechanism is introduced:
[0056] in, is the global reward value, is the number of unmanned boats, is the local reward coefficient of the i-th unmanned boat, is the local reward of the i-th unmanned boat, is the local reward of the j-th unmanned boat, is the synergy reward factor, is the behavior similarity function between the i-th and j-th unmanned boats; in step S2, the path planning cost function of the initial collaborative search path is:
[0057] in, is the total path length, is the sailing time of the i-th unmanned boat, is the number of obstacles, is the avoidance cost of the k-th obstacle, are the cost weight coefficients respectively.
[0058] Specifically, the global reward allocation mechanism and the cost function for the initial collaborative search path of the multi-agent deep deterministic policy gradient algorithm were improved. This global reward allocation mechanism, through local reward coefficients and collaborative reward factors, both recognizes the search contribution of individual UAVs and encourages collaborative cooperation among the swarm. This prevents some UAVs from excessively pursuing local rewards and undermining overall coordination, making reward allocation more consistent with the objectives of swarm search and rescue. The path planning cost function comprehensively considers total path length, navigation time, and obstacle avoidance costs, enabling assessment and control of overall search and rescue costs during the path planning phase, reducing unnecessary energy consumption and time waste. During implementation, the local reward coefficient was adjusted between 0.6 and 0.8 based on the UAV's search range and detection performance, and the collaborative reward factor was set between 0.2 and 0.4 to balance local and global interests. The weighting coefficients for total path length, navigation time, and obstacle avoidance costs were selected between 0.4 and 0.6, 0.2 and 0.3, and 0.1 and 0.2, respectively. The specific values were determined based on the size and complexity of the search and rescue area to ensure that the cost function truly reflects the actual costs of the search and rescue mission.
[0059] Preferably, in step S3, when preprocessing the water area image data collected by the multispectral imaging device, a dynamic exposure compensation model is used:
[0060] in, is the pixel value of the image after compensation, is the original image pixel value, is the light attenuation coefficient, for Depth of water at coordinates wavelength Corresponding light absorption coefficient; in step S4, the confidence value correction formula output by the improved PV-RCNN target detection model is:
[0061] in, is the corrected confidence level, is the original confidence, is the signal-to-noise ratio correction factor, is the signal-to-noise ratio of the compensated image.
[0062] Specifically, the preprocessing of image data collected by multispectral imaging equipment and the correction of target detection confidence have been optimized. The dynamic exposure compensation model adjusts the image pixel values according to the water depth and light absorption coefficient. This can effectively improve the uneven image brightness caused by changes in water depth and differences in light absorption, enhance image clarity and contrast, and provide higher-quality image data for subsequent feature extraction and target detection. The confidence correction formula combines the image's signal-to-noise ratio to correct the original confidence, so that the confidence value can more accurately reflect the reliability of the target detection results and reduce misjudgments due to image quality issues. During implementation, the light attenuation coefficient is selected between 0.1 and 0.3 based on the turbidity of the water area, and the light absorption coefficient corresponding to the wavelength is determined based on the common water quality parameters of the water area. The signal-to-noise ratio correction factor is set to 0.05-0.1. Through statistical analysis of the detection results of a large number of images of different quality, the optimal value of the correction factor is determined to ensure that the corrected confidence level can truly reflect the possibility of the target's existence.
[0063] Preferably, in step S5, when the Actor network adjusts the steering angle of the unmanned boat servo, an environment adaptive correction model is introduced:
[0064] in, for The steering angle of the servo at the moment, is the steering angle of the servo at time t, is the angle adjustment coefficient, is the wind impact weight, is the wind speed at time t, is the water flow influence weight, is the water flow velocity value at time t; in step S2, the path smoothness index of the initial collaborative search path is calculated as:
[0065] in, is the path smoothness, is the number of path segments of the i-th unmanned boat, is the heading angle of the kth path of the i-th unmanned boat.
[0066] Specifically, improvements have been made to the adjustment of the steering angle of the unmanned boat's rudder and the smoothness evaluation of the initial collaborative search path. The environmental adaptive correction model takes the influence of wind and water flow into consideration when adjusting the rudder angle. By setting the weights of wind and water flow, the rudder steering can actively offset the heading deviation caused by environmental interference, thereby improving the navigation stability and heading accuracy of the unmanned boat in complex water environments. The calculation of the path smoothness index is performed by averaging the cosine values of the heading angles of each path segment to quantitatively evaluate the smoothness of the path. The higher the smoothness, the lower the navigation energy consumption of the unmanned boat, and the time loss caused by frequent steering can be reduced. During implementation, the angle adjustment coefficient is determined between 0.02-0.05 based on the tonnage and power performance of the unmanned boat. The wind influence weight and water flow influence weight are set to 0.3-0.5 and 0.5-0.7 respectively according to the degree of their influence on navigation in the actual environment. In the calculation of path smoothness, the number of path segments is determined according to the size and complexity of the search and rescue area. The number of path segments for each unmanned boat is between 10-20. Through the feedback of the smoothness index, the initial path is iteratively optimized until the smoothness reaches above 0.8.
[0067] Preferably, in step S6, during the target encirclement trajectory generation process, the multi-agent deep deterministic policy gradient algorithm adopts a dynamic encirclement radius calculation model:
[0068] in, is the enclosed radius at time t, is the initial enclosure radius, is the shrinkage rate coefficient, is the minimum enclosing radius, is the target object's moving speed, is the average speed of the unmanned boat cluster; in step S3, the point cloud feature extraction of the improved PV-RCNN target detection model adopts the voxel enhancement method:
[0069] in, Voxel The eigenvalues of is the point cloud set included in the voxel, is the weight of point p, is the reflection intensity coefficient, is the reflection intensity value of point p, is the original eigenvector of point p.
[0070] Specifically, the generation of target encirclement trajectories and the point cloud feature extraction of the PV-RCNN target detection model are optimized. The dynamic encirclement radius calculation model takes into account the initial radius, contraction rate, minimum radius, and the velocity relationship between the target and the swarm. This allows the encirclement radius to shrink dynamically over time, ensuring the safety of the encirclement process while also adjusting the encirclement range in real time based on the target's motion state, improving the success rate of encirclement. The voxel enhancement method for point cloud feature extraction enhances the discriminability of point cloud features by introducing point weights and reflection intensity coefficients, enabling the model to more accurately identify the target's three-dimensional features. This can effectively improve target detection accuracy, especially when the target and the surrounding point cloud features are similar. During implementation, the initial enclosure radius is set to 10-20 meters based on the possible activity range of the target object, the shrinkage rate coefficient is selected between 0.01-0.03, and the minimum enclosure radius is determined to be 3-5 meters based on the size of the unmanned boat and the rescue operation space; the point weight is allocated between 0.5-1.0 based on the density and reliability of the point cloud, and the reflection intensity coefficient is set to 0.1-0.3. By training a large amount of point cloud data, the optimal parameter values are determined to improve the effect of feature extraction.
[0071] Preferably, in step S4, when the improved PV-RCNN target detection model classifies the three-dimensional candidate boxes, a multi-level confidence screening mechanism is adopted:
[0072] in, is the final confidence level, Output confidence for the classifier, is the correction factor, is the intersection-over-union ratio of the candidate box B and the target real box G, is the noise level of image I; in step S5, the experience replay pool of the multi-agent deep deterministic policy gradient algorithm adopts a priority sampling strategy:
[0073] in, is the sampling probability of the jth experience, is the TD error of the j-th experience, is a small constant, is the priority index, The capacity of the experience replay pool.
[0074] Specifically, improvements were made to the PV-RCNN object detection model's 3D candidate box classification mechanism and the experience replay pool sampling strategy for the multi-agent deep deterministic policy gradient algorithm. A multi-level confidence screening mechanism combines the intersection-over-union (IoU) of candidate boxes with the ground-truth box and the image noise level to correct the confidence of the classifier output. This effectively eliminates low-quality detection results caused by noise interference and inaccurate candidate box positioning, improving object classification accuracy. The experience replay pool's priority sampling strategy determines sampling probability based on the empirical TD error, enabling the algorithm to focus more on experiences that contribute significantly to policy improvement during training, accelerating convergence and enhancing policy performance. During implementation, the intersection-over-union ratio correction coefficient and the noise level correction coefficient are set between 0.1-0.3 and 0.05-0.15, respectively. The optimal combination of correction coefficients is determined by analyzing the detection results in different scenarios. The priority index is set to 0.5-0.7, and the capacity of the experience replay pool is determined to be 10,000-50,000 based on the complexity of the algorithm and the amount of training data to ensure the diversity and effectiveness of the sampling and promote the algorithm to quickly converge to the optimal strategy.
[0075] Preferably, in step S1, when the distributed sensing nodes collect dynamic environmental parameters, a spatiotemporal fusion filtering model is used:
[0076] in, is the environmental parameter value after fusion at time t, is the current data weight, is the real-time data collected at time t, is the size of the historical data window, is the weight of the k-th historical data, for The historical data at the moment; in step S2, when the multi-agent deep deterministic policy gradient algorithm generates the initial collaborative search path, the energy constraint condition is satisfied:
[0077] in, is the total time of path planning, is the power consumption of the i-th unmanned boat at time t, is the total energy budget of the UAV swarm.
[0078] Specifically, the distributed sensing nodes' process for collecting dynamic environmental parameters and the energy constraints of the initial collaborative search path are optimized. A spatiotemporal fusion filtering model reduces random errors in environmental parameter collection by weighted fusion of current and historical data, making the collected environmental parameters more stable and accurate, and providing reliable environmental data support for subsequent path planning and strategy adjustments. Energy constraints ensure that the initial collaborative search path planning is carried out within the total energy budget of the UAV swarm, avoiding interruptions to the search and rescue mission due to energy depletion and improving the sustainability of the mission. During implementation, the current data weight is set to 0.6-0.8, and the historical data window size is determined to be 5-10 sampling periods based on the rate of change of environmental parameters. The weights of historical data are distributed according to the time decay law, with recent data having a higher weight and long-term data having a lower weight. The total energy budget is determined based on the UAV's battery capacity and the expected duration of the search and rescue mission. During path planning, an energy consumption model is used to calculate the total energy consumption of each path in real time to ensure that it does not exceed 90% of the total energy budget, leaving a certain energy margin for emergencies.
[0079] like Figure 2 As shown, an unmanned boat intelligent search and rescue system includes: A distributed dynamic environmental parameter acquisition unit, which acquires water flow, wind force, and obstacle-related data through a sensor array deployed in the search and rescue area, and establishes a two-way data transmission link with the multi-agent decision-making control unit; A multi-agent decision control unit receives data transmitted by a distributed dynamic environment parameter acquisition unit, runs a multi-agent deep deterministic policy gradient algorithm to generate path planning instructions and motion control signals, and sends them to the unmanned boat cluster execution unit and the target detection coordination unit respectively; The unmanned boat swarm execution unit includes multiple unmanned boat individuals, receives the motion control signal output by the multi-agent decision-making control unit, adjusts the position and attitude through the propeller and steering gear adjustment module, and feeds back its own state data to the multi-agent decision-making control unit; A multispectral image acquisition unit, which is mounted on each unmanned boat of the unmanned boat cluster execution unit, captures water area image data in real time and transmits it to the improved three-dimensional target detection unit via a high-speed data interface; An improved 3D target detection unit, which runs an improved PV-RCNN target detection model, processes the image data transmitted by the multispectral image acquisition unit, and outputs the target's 3D coordinates and confidence information to the multi-agent decision control unit. The cluster collaborative communication unit establishes a wireless communication network between the individual unmanned boats in the unmanned boat cluster execution unit, conducts real-time interaction of status data and detection information, and feeds back the network topology status to the multi-agent decision-making control unit.
[0080] An intelligent search and rescue method and system for unmanned boats, which demonstrate significant advantages in multi-boat collaboration. Through the Actor-Critic framework of the multi-agent deep deterministic policy gradient algorithm, the unmanned boat cluster can dynamically adjust the collaborative search path based on real-time collected environmental parameters such as water flow and wind force. Each unmanned boat strictly adheres to the communication distance constraints and collision avoidance safety thresholds, effectively ensuring safety during movement. At the same time, based on the target detection results, the motion control parameters such as the thruster output power and the steering angle of the servo are continuously optimized to achieve dynamic allocation of search and rescue resources, successfully solving the problems of insufficient adaptability of traditional path planning to dynamic environments and lack of feedback on collaborative strategies.
[0081] The improved PV-RCNN object detection model plays a key role in target detection. When extracting features from water image data, this model fuses point cloud data with image pixel information to construct 3D candidate boxes, significantly improving the accuracy of 3D target feature extraction in complex water environments. Furthermore, the 3D spatial coordinates and confidence values of the target objects output by the model are transmitted in real time to the state space of a multi-agent deep deterministic policy gradient algorithm, forming a closed-loop linkage between detection and motion control. This provides accurate target position information for the unmanned vehicle swarm, overcoming the shortcomings of traditional models, such as poor adaptability to complex scenarios and the disconnect between detection and control.
[0082] Overall, this unmanned boat intelligent search and rescue method and system achieves complementary advantages through the close integration of a multi-agent deep deterministic policy gradient algorithm and an improved PV-RCNN target detection model. The dynamic adjustment capabilities of multi-boat collaboration and its precise target detection capabilities complement each other, expanding search and rescue coverage and improving target detection efficiency. This effectively addresses various challenges in complex water environments, comprehensively enhancing the overall effectiveness of unmanned boat search and rescue, and successfully addressing the shortcomings of existing technologies in dynamic adaptation and precise detection.
[0083] In the description of the present invention, it should be noted that, unless otherwise expressly specified or limited, the terms "disposed," "installed," "connected," "connected," and "fixed" should be understood in a broad sense. For example, they may refer to fixed connections, detachable connections, or integral connections; they may refer to mechanical connections or electrical connections; they may refer to direct connections or indirect connections through an intermediate medium; and they may refer to internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on specific circumstances.
[0084] While embodiments of the present invention have been shown and described, it will be understood by those skilled in the art that various equivalent changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. An unmanned boat intelligent search and rescue method, characterized in that: The following steps are involved: Step S1: Deploy a swarm of unmanned boats driven by a multi-agent deep deterministic policy gradient algorithm to collect dynamic environmental parameters within the search and rescue area through distributed sensing nodes. The dynamic environmental parameters include water velocity vector, wind force level, three-dimensional coordinates of water obstacles, and initial probability distribution area of the target object; Step S2: Based on the dynamic environment parameters, the Actor-Critic framework in the multi-agent deep deterministic policy gradient algorithm is used to generate the initial collaborative search path of the unmanned boat cluster. The initial collaborative search path must meet the communication distance constraints and collision avoidance safety thresholds between the unmanned boats. Step S3: During the initial collaborative path search performed by the unmanned boat cluster, the improved PV-RCNN target detection model is activated, and water image data is collected in real time through the multispectral imaging equipment carried by the unmanned boat cluster. When the improved PV-RCNN target detection model extracts features from the water image data, it fuses the point cloud data with the image pixel information to construct a 3D candidate frame; Step S4: The improved PV-RCNN object detection model classifies and regresses the 3D candidate boxes, outputs the 3D spatial coordinates and confidence values of the target objects, and transmits the 3D spatial coordinates and confidence values to the state space of the multi-agent deep deterministic policy gradient algorithm; Step S5: The multi-agent deep deterministic policy gradient algorithm evaluates the value function of the current policy through the Critic network based on the updated information in the state space, and adjusts the motion control parameters of each unmanned boat in combination with the Actor network; Step S6: Repeat steps S3 to S5 until the target object confidence value output by the improved PV-RCNN target detection model reaches a preset threshold. At this time, the multi-agent deep deterministic policy gradient algorithm generates a target encirclement trajectory for the unmanned boat cluster, and the target encirclement trajectory adapts to the motion state parameters of the target object and real-time environmental interference factors.
2. The unmanned boat intelligent search and rescue method according to claim 1, characterized in that: In step S2, when the multi-agent deep deterministic policy gradient algorithm generates the initial collaborative search path, a policy optimization function based on the topological structure of the search and rescue area is introduced: in, Indicates that the status The next agent chooses an action The probability distribution of is the activation function, are weight coefficients, represents the Euclidean distance between the i-th and j-th unmanned boats, Indicates that the search and rescue area The water flow gradient vector at the coordinate, represents the repulsive potential field function of the k-th obstacle; in step S5, the motion control parameter adjustment process must satisfy: in, is the propeller output power of the i-th unmanned boat at time t, is the power attenuation coefficient, is the policy update rate, Represents the gradient of the action-value function with respect to the action.
3. The unmanned boat intelligent search and rescue method according to claim 1, characterized in that: In step S3, when the improved PV-RCNN target detection model fuses point cloud data with image pixel information, a feature enhancement module is used: in, is the fused feature matrix, is the feature fusion coefficient, is the point cloud feature matrix, is the image feature matrix, represents the matrix Hadamard product, represents the matrix concatenation operation, represents the convolution operation; in step S4, the regression loss function of the three-dimensional candidate box is defined as: in, is the translation error loss, is the size error loss, is the rotation angle error loss, are the weight coefficients of each loss.
4. The unmanned boat intelligent search and rescue method according to claim 1, characterized in that: In step S5, when the critic network of the multi-agent deep deterministic policy gradient algorithm evaluates the value function, a global reward distribution mechanism is introduced: in, is the global reward value, is the number of unmanned boats, is the local reward coefficient of the i-th unmanned boat, is the local reward of the i-th unmanned boat, is the local reward of the j-th unmanned boat, is the synergy reward factor, is the behavior similarity function between the i-th and j-th unmanned boats; in step S2, the path planning cost function of the initial collaborative search path is: in, is the total path length, is the sailing time of the i-th unmanned boat, is the number of obstacles, is the avoidance cost of the k-th obstacle, are the cost weight coefficients respectively.
5. The unmanned boat intelligent search and rescue method according to claim 1, characterized in that: In step S3, a dynamic exposure compensation model is used when preprocessing the water area image data collected by the multispectral imaging device: in, is the pixel value of the image after compensation, is the original image pixel value, is the light attenuation coefficient, for Depth of water at coordinates wavelength Corresponding light absorption coefficient; in step S4, the confidence value correction formula output by the improved PV-RCNN target detection model is: in, is the corrected confidence level, is the original confidence, is the signal-to-noise ratio correction factor, is the signal-to-noise ratio of the compensated image.
6. The unmanned boat intelligent search and rescue method according to claim 1, characterized in that: In step S5, when the Actor network adjusts the steering angle of the unmanned boat servo, an environment adaptive correction model is introduced: in, for The steering angle of the servo at the moment, is the steering angle of the servo at time t, is the angle adjustment coefficient, is the wind impact weight, is the wind speed at time t, is the water flow influence weight, is the water flow velocity value at time t; in step S2, the path smoothness index of the initial collaborative search path is calculated as: in, is the path smoothness, is the number of path segments of the i-th unmanned boat, is the heading angle of the kth path of the i-th unmanned boat.
7. The unmanned boat intelligent search and rescue method according to claim 1, characterized in that: In step S6, during the target encirclement trajectory generation process, the multi-agent deep deterministic policy gradient algorithm adopts a dynamic encirclement radius calculation model: in, is the enclosed radius at time t, is the initial enclosure radius, is the shrinkage rate coefficient, is the minimum enclosing radius, is the target object's moving speed, is the average speed of the unmanned boat cluster; in step S3, the point cloud feature extraction of the improved PV-RCNN target detection model adopts the voxel enhancement method: in, Voxel The eigenvalues of is the point cloud set included in the voxel, is the weight of point p, is the reflection intensity coefficient, is the reflection intensity value of point p, is the original eigenvector of point p.
8. The unmanned boat intelligent search and rescue method according to claim 1, characterized in that: In step S4, the improved PV-RCNN target detection model adopts a multi-level confidence screening mechanism when classifying the three-dimensional candidate boxes: in, is the final confidence level, Output confidence for the classifier, is the correction factor, is the intersection-over-union ratio of the candidate box B and the target real box G, is the noise level of image I; in step S5, the experience replay pool of the multi-agent deep deterministic policy gradient algorithm adopts a priority sampling strategy: in, is the sampling probability of the jth experience, is the TD error of the j-th experience, is a small constant, is the priority index, The capacity of the experience replay pool.
9. The unmanned boat intelligent search and rescue method according to claim 1, characterized in that: In step S1, when the distributed sensing nodes collect dynamic environmental parameters, a spatiotemporal fusion filtering model is used: in, is the environmental parameter value after fusion at time t, is the current data weight, is the real-time data collected at time t, is the size of the historical data window, is the weight of the k-th historical data, for The historical data at the moment; in step S2, when the multi-agent deep deterministic policy gradient algorithm generates the initial collaborative search path, the energy constraint condition is satisfied: in, is the total time of path planning, is the power consumption of the i-th unmanned boat at time t, is the total energy budget of the UAV swarm.
10. An unmanned boat intelligent search and rescue system, characterized in that: include: A distributed dynamic environmental parameter acquisition unit, which acquires water flow, wind force, and obstacle-related data through a sensor array deployed in the search and rescue area, and establishes a two-way data transmission link with the multi-agent decision-making control unit; A multi-agent decision control unit receives data transmitted by a distributed dynamic environment parameter acquisition unit, runs a multi-agent deep deterministic policy gradient algorithm to generate path planning instructions and motion control signals, and sends them to the unmanned boat cluster execution unit and the target detection coordination unit respectively; The unmanned boat swarm execution unit includes multiple unmanned boat individuals, receives the motion control signal output by the multi-agent decision-making control unit, adjusts the position and attitude through the propeller and steering gear adjustment module, and feeds back its own state data to the multi-agent decision-making control unit; A multispectral image acquisition unit, which is mounted on each unmanned boat of the unmanned boat cluster execution unit, captures water area image data in real time and transmits it to the improved three-dimensional target detection unit via a high-speed data interface; An improved 3D target detection unit, which runs an improved PV-RCNN target detection model, processes the image data transmitted by the multispectral image acquisition unit, and outputs the target's 3D coordinates and confidence information to the multi-agent decision control unit. The cluster collaborative communication unit establishes a wireless communication network between the individual unmanned boats in the unmanned boat cluster execution unit, conducts real-time interaction of status data and detection information, and feeds back the network topology status to the multi-agent decision-making control unit.
Citation Information
Patent Citations
Unmanned aerial vehicle landing method and equipment based on MPC-NDQN, and medium
CN118192584A
Intelligent decision and response method and system based on data analysis
CN119091234A
Multi-unmanned ship cooperative hunting method based on channel attention mechanism deep reinforcement learning algorithm
CN119168601A
Route planning method and device, electronic equipment and storage medium
CN119618240A
3D point cloud target detection method based on state space model
CN119741697A
Cited By
Multi-modal sensor external parameter automatic calibration method and system
CN121114949A
Behavior simulation method of mobile agent applied to emergency rescue scene
CN121995791A