Water surface unmanned ship multi-source target detection method based on domestic platform and related equipment

Through the application of multi-source sensor data fusion and the application of the domestic deep learning framework MNN, the detection accuracy and platform dependence of surface unmanned boats in complex environments is solved, and high-precision target detection is achieved all-day and all-weather, and has independent intellectual property rights and efficient processing capabilities.

CN120259622APending Publication Date: 2025-07-04YICHANG TESTING TECHNIQUE RESEARCH INSTITUTE
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311485689.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-09
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing surface unmanned boat target detection system is insufficient in complex environments, and relies on foreign software and hardware platforms to achieve all-day and all-weather operations, and the domestic deep learning framework is not used in surface unmanned boats.

Method used

Multi-source sensor data fusion analysis is adopted, and information is collected using visible light cameras, infrared cameras and lidar, combined with the domestic lightweight deep learning front-end reasoning framework MNN, which is deployed on the domestic software and hardware platform for object detection, and the detection accuracy and efficiency are improved through data buffering and synchronization processing.

Benefits of technology

It realizes high-precision object detection all day and all day, improves the accuracy and confidence of the detection algorithm, has independent intellectual property rights, supports multi-threaded parallel processing, and has good scalability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259622A_ABST
    Figure CN120259622A_ABST
Patent Text Reader

Abstract

The invention discloses an unmanned surface vehicle multi-source target detection method and related equipment based on a domestic deep learning software and hardware platform. In the method, a data acquisition unit senses the environment by using a visible light camera, an infrared camera and a laser radar multi-source sensor; the single-source data target detection unit performs target detection on each data source, and comprises a target detection algorithm based on a domestic deep learning reasoning framework MNN and a traditional laser radar point cloud target clustering algorithm; the multi-source data fusion analysis unit performs data synchronization and fusion analysis on output results of all algorithms to obtain a final accurate target detection result; the system provided by the invention overcomes the defect of sensing capability of a single sensor, supports all-day and all-weather operation, and improves the accuracy and confidence of a detection algorithm; algorithm software is deployed in a domestic ruggedized computer provided with a Feiteng D2000 processor and a Jingjia micro JH920 display card, and is autonomous and controllable.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of computer vision and unmanned surface vehicle technology, and particularly relates to a multi-source target detection method and related equipment for an unmanned surface vehicle on a domestic platform. Background Art

[0002] Unmanned surface vehicles have the characteristics of small size, high speed, strong mobility, etc., and can be used to perform dangerous tasks and ensure personnel safety. They are widely used in civil and military fields such as water quality monitoring, water surface environment exploration, detection of enemy ships, offshore patrol, and open sea search and rescue. The demand for unmanned surface vehicles in our country is increasing day by day. In order to enable unmanned surface vehicles to have higher environmental perception and autonomous collision avoidance capabilities in complex environments, it is necessary to study high-precision and high-speed water surface target detection algorithms. As a lightweight, low-power and information-rich sensor, visible light cameras have become the mainstream information acquisition devices for unmanned surface vehicles.

[0003] In recent years, the development of deep learning algorithms has promoted a new round of transformation in the artificial intelligence industry, making breakthrough progress in technical fields such as computer vision, natural language processing, and autonomous driving. Deep learning algorithms utilize the advantages of data-driven, and can learn hierarchical and abstract data features, overcoming the defects of single-dimensional manual features, cumbersome design, and weak discriminative power. With the superior performance of Faster R-CNN and YOLO series models on general target datasets such as PASCAL VOC and MS COCO, many research works have also applied them to the unmanned surface vehicle water surface target detection task. For example, YANG et al. used the YOLOv3 model to achieve real-time detection of unmanned surface vehicles, and Borja Bovcon et al. comprehensively evaluated the performance of target detection algorithms such as SSD, FCOS, YOLOv4, and Mask R-CNN on the water surface target dataset.

[0004] However, there are still many problems and challenges in the actual application of current deep learning target detection algorithms for unmanned surface vehicles. First of all, due to the complex and changeable real water surface environment, affected by rainy and foggy weather, motion blur, light change and other situations, the quality of visible light images captured by unmanned surface vehicles will be damaged, greatly reducing the accuracy of the detection algorithm. Therefore, relying solely on visible light cameras as sensing devices is difficult to achieve stable operation of the algorithm throughout the day and in all weather conditions in actual tasks.

[0005] Secondly, most existing unmanned boat target detection systems are implemented using deep learning software frameworks such as Facebook's Pytorch and Google's TensorFlow, and are deployed and run on NVIDIA graphics cards. Both the software and hardware platforms are developed by foreign manufacturers. Therefore, such systems have no independent intellectual property rights and cannot be deployed and applied on domestic software and hardware platforms. Although China has independently developed deep learning frameworks such as Baidu PaddlePaddle and Tsinghua Jittor, and companies such as Horizon Robotics and Cambricon have successively launched domestic artificial intelligence chips, they have not been applied in the field of unmanned boat target detection on the water surface. Summary of the Invention

[0006] In view of this, the present invention provides a multi-source target detection method for unmanned boats on the water surface based on domestic deep learning software and hardware platforms. The data acquisition unit collects different modality information of the unmanned boat on the water surface through a variety of different sensors for environmental perception. The different modality information includes visible light images, infrared images, and 3D point cloud data;

[0007] The single-source data target detection unit uses the domestic lightweight deep learning front-end inference framework MNN to deploy and apply the deep learning target detection algorithm pre-trained on non-domestic platforms on domestic software and hardware platforms, and processes the visible light images, infrared images, and 3D point cloud data respectively;

[0008] The multi-source data fusion analysis unit receives the different detection results of the information collected by a variety of different sensors output by the single-source data target detection unit; reads the data timestamps of the different detection results, performs time series caching and alignment, and takes the union of the results of the variety of different sensors; constructs an output target set that includes all target categories detected by the variety of different sensors; merges overlapping detection frames, outputs the final target detection result after fusion processing, and sends the different detection results output by the single-source data target detection unit to the lower computer.

[0009] Specifically, the data acquisition unit specifically includes using a visible light camera, an infrared camera, and a lidar as sensors for the unmanned boat to perform environmental perception to complement different modality information on the water surface; designing a data buffer queue to cache the sensor data input by the data acquisition unit to prevent memory overflow caused by infinite stacking of data, and using a read-write lock for access to ensure the security of data processing, and scheduling the operation of the algorithm in a multi-threaded parallel manner to accelerate the overall operation efficiency of the software.

[0010] Specifically, in the single-source data target detection unit, the domestic lightweight deep learning front-end inference framework MNN is used to deploy and apply the deep learning target detection algorithm pre-trained on a non-domestic platform to a domestic software and hardware platform to process visible light images and infrared images. Specifically, it includes: based on the non-domestic deep learning framework PyTorch, respectively building target detection network models for processing visible light images and infrared images; using the open-source water surface target detection dataset MODS to train the target detection model for processing visible light images; selecting the infrared video annotation data in the open-source Singapore Maritime Dataset SMD, and using a non-domestic NVIDIA graphics card to train the target detection model to obtain a pre-trained model; using a water surface unmanned boat equipped with a visible light camera and an infrared camera to collect data on-site in the water area, and after manual annotation, respectively fine-tuning the pre-trained model obtained in the previous step; using the domestic deep learning framework MNN to convert the parameter files of the two target detection models fine-tuned in the previous step; deploying the two converted target detection models on a computer equipped with a first domestic central processing unit and a first domestic GPU processing unit, with the model inputs being visible light images and infrared images respectively, and the output being the target circumscribed rectangle containing the category results.

[0011] Specifically, the detection of the 3D point cloud data in the single-source data target detection unit specifically includes: downsampling the 3D point cloud data generated by the lidar to reduce the data computation amount without losing environmental information; clustering the disordered point cloud data to form separate target clusters; calculating the following characteristic values of the target clusters generated by clustering: target length, target width, average reflection intensity, reflection intensity variance, and distance; screening out reasonable obstacles according to the characteristic values of the target clusters, deleting the outliers generated by the shaking and fluctuation of the unmanned boat in actual applications, as well as the targets that do not conform to the actual situation.

[0012] Specifically, the multi-source data fusion and analysis unit specifically includes receiving the visible light camera target detection results, infrared camera target detection results, and lidar target detection results output by the single-source data target detection unit; designing a data synchronization module to cache the detection results of each sensor, reading the timestamps of each frame of data in the data buffer queue, and sorting and caching them in time sequence; taking the visible light image data generation time as the reference, matching the infrared light image detection results and lidar detection results that are closest to it in time and differ by no more than 200 milliseconds, and constructing an output target set by taking the union of the three results, that is, including ship, human, buoy, bird / airplane and other obstacle targets detected by the visible light image target detection algorithm model and the infrared image target detection algorithm model, as well as the near obstacle targets obtained by the lidar point cloud processing algorithm; merging overlapping targets and eliminating redundant prediction results with a high degree of overlap. If a certain target is detected by X sensors at the same time, then set the confidence of this result to X; transmit the final detection result after fusion analysis and the single-source sensor detection results to the lower computer through the network.

[0013] Specifically, the first domestic central processing unit is the Feiteng D2000 processor, and the first domestic GPU processor is the Jingjiawei JH920 graphics card.

[0014] The present invention also proposes a multi-source target detection system for a surface unmanned boat based on a domestic deep learning software and hardware platform. The system includes a data acquisition unit, a single-source data target detection unit, and a multi-source data fusion and analysis unit.

[0015] The data acquisition unit includes a variety of different sensors for environmental perception, and is used to collect different modality information of the surface unmanned boat. The different modality information includes visible light images, infrared images, and 3D point cloud data.

[0016] The single-source data target detection unit is used to deploy and apply the deep learning target detection algorithm pre-trained on a non-domestic platform on the domestic software and hardware platform by using the domestic lightweight deep learning front-end inference framework MNN, and respectively process the visible light images, infrared images, and 3D point cloud data.

[0017] The multi-source data fusion and analysis unit is used to receive the different detection results of the single-source data target detection unit for the information collected by a variety of different sensors; read the data timestamps of different detection results, perform time series caching and alignment, and take the union of the results of the variety of different sensors; construct an output target set including all target categories detected by the variety of different sensors; merge overlapping detection frames, output the final target detection result after fusion processing, and send the different detection results output by the single-source data target detection unit to the lower computer.

[0018] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the multi-source target detection method for a surface unmanned boat based on a domestic deep learning software and hardware platform is implemented. The processor includes a first domestic AI processor, a second domestic AI processor, and an NVIDIA processor.

[0019] Beneficial effects:

[0020] (1) The present invention can quickly and accurately detect targets and obstacles encountered by a surface unmanned boat during navigation. By using multi-source sensor perception data fusion analysis, it overcomes the defects of the single-sensor perception ability, supports all-day and all-weather operations, and improves the accuracy and confidence of the detection algorithm.

[0021] (2) The present invention uses the domestic lightweight deep learning front-end inference framework MNN as the software platform and the Feiteng D2000 processor and Jingjiawei JH920 graphics card as the domestic hardware platform. The core technology is not subject to foreign companies and has completely independent intellectual property rights.

[0022] (3) The present invention develops a multi-source target detection software for a surface unmanned boat, which supports multi-thread parallelism, speeds up the operation efficiency of the target detection algorithm, and uses data buffering and data synchronization to ensure the security of memory access and the consistency of detection results.

[0023] (4) The present invention realizes a multi-modal detection scheme based on multiple heterogeneous sensors. Visible light images provide color and details, infrared images conduct all-weather observations, and lidar provides distance and three-dimensional structure information, realizing the omnidirectional detection of surface targets.

[0024] (5) The present invention designs a configurable model construction, training, conversion, and deployment process, which can be flexibly adapted to different detection tasks, such as replacing the type of sensor, increasing the detection category, etc., and has good scalability. Description of the drawings

[0025] Figure 1 is a flow chart of the multi-source target detection method for a surface unmanned boat based on a domestic deep learning software and hardware platform proposed by the present invention;

[0026] Figure 2 is a schematic flow chart of the MNN target detection model used for prior generation in the present invention. Detailed implementation manners

[0027] The following are specific embodiments in conjunction with the drawings to describe the present invention in detail.

[0028] The present invention provides a multi-source target detection method for surface unmanned boats based on domestic deep learning software and hardware platforms. The overall technical route of its system design and implementation is as follows Figure 1 shown. The technical solution includes: a data acquisition unit, a single-source data target detection unit, and a multi-source data fusion analysis unit. In this example, the software system is implemented in the C++ programming language, relying on the domestic lightweight deep learning front-end inference framework MNN. The operating hardware environment is a domestic rugged computer configured with a Feiteng D2000 processor and a Jingjiawei JH920 graphics card.

[0029] The data acquisition unit is the main software process, responsible for receiving data from the information acquisition devices of the surface unmanned boat, including visible light cameras, infrared cameras, and lidar. To prevent abnormal algorithm loading and processing in the single-source data target detection unit, only when the data acquisition unit receives data sent by a certain sensor, the algorithm module thread corresponding to it is initialized. The data acquisition unit implements two ways of reading data, actively reading the data sampled from the sensor from files and memory, or passively receiving the data sent by the sensor by listening to a fixed port through the UDP / TCP network protocol. Since the sensor data sampling frequency usually exceeds 30Hz, the speed of data generation is likely to be greater than the processing speed of the subsequent MNN target detection model. To prevent memory overflow caused by infinite data stacking, a data buffer queue is designed, and a read-write lock is used to access it to ensure the security of data processing. The specific implementation is as follows:

[0030] Three std::vector queue variables with a length of 5 are established, corresponding to the data of the visible light camera, infrared camera, and lidar sensors respectively; each time sensor data is obtained actively or passively, a write lock is applied to the queue, and the push_back() function interface is called to add new data to the end of the queue; when the queue length is greater than 5, a write lock is applied to the queue, and the pop_front() function interface is called to delete the data at the head of the queue. When the algorithm processing of the corresponding single-source data target detection unit ends and enters the waiting period, a read lock is applied to the queue, and the first() function interface is called to read the data at the head of the queue as the algorithm input. The designed data buffer queue discards the unprocessed data farthest from the current time to maintain the real-time nature of data processing.

[0031] The single-source data target detection unit contains three independent threads, which respectively run algorithms for processing the three sensors. For both visible light images and infrared light images, an end-to-end target detection algorithm based on deep learning is used for inference, and then non-maximum suppression is used to eliminate redundant prediction boxes. As Figure 2 shown, before running this software system, the used MNN target detection model is generated in advance. The specific implementation process is as follows:

[0032] (1) Model construction: Based on the non-domestic deep learning framework PyTorch, the open-source object detection algorithm YOLOv5 is used to build object detection network models for processing visible light images and infrared images respectively. To ensure the real-time operation of the algorithm, the lightweight model with YOLOv5s configuration is selected. Modify the number of detection classes of the model to adapt to the training set used.

[0033] (2) Model pre-training: Use the open-source water surface object detection dataset MODS to train the object detection model for processing visible light images, which contains about 24,000 images, and the target categories are 3 types: ships, humans, and other obstacles. Then select the infrared video labeled data in the open-source Singapore Maritime Dataset SMD, which contains 30 infrared videos of the water surface taken on shore, and the target categories are 5 types: ships, humans, buoys, flying birds / airplanes, and other targets. The data preprocessing methods and data augmentation algorithms used all follow the default configuration of the open-source YOLOv5, and use the NVIDIA RTX 2080 graphics card to train the above two object detection models.

[0034] (3) Model fine-tuning: Since the open-source datasets used in the previous step are all collected in foreign waters, in order to make the algorithm more robust, use a surface unmanned boat equipped with visible light cameras and infrared cameras to collect data on the Zhanghe River water area, use the Labelme manual annotation software to annotate the data collected on site, and then use the newly added data to fine-tune the pre-trained models obtained in the previous step respectively;

[0035] (4) Model parameter file conversion: Use the MNNConvert tool of the domestic deep learning framework MNN to convert the parameter files of the two object detection models generated in the previous step to obtain parameter files with the file format suffix of "mnn".

[0036] (5) Model deployment: Write the corresponding model call code to deploy the two converted object detection models, relying on the MNN inference library and the OpenCL acceleration library, and supported by the underlying driver of the Jingjiawei JH920 graphics card. The model inputs are visible light images and infrared images respectively, and the output is the target bounding rectangle containing the category results.

[0037] In addition, the lidar point cloud data processing algorithm also occupies a separate thread. The input is the lidar 3D point cloud data, which includes the distance of the reflector from the lidar, the horizontal rotation angle of the beam, and the reflection intensity, and the output is the close-range obstacles around the unmanned boat. The algorithm is implemented based on the C++ PCL point cloud processing library and does not depend on any graphics card driver and deep learning library. The specific implementation steps are as follows:

[0038] (1) Downsampling filtering: Based on pcl::VoxelGrid <pointt>Data structures and function interfaces such as filter() are used to downsample the 3D point cloud data generated by the lidar, reducing the data computation volume without losing environmental information.

[0039] (2) Clustering: Since the point cloud data collected by the lidar is disorderly and there is no topological relationship between each point, in order to perform subsequent obstacle detection and obstacle state calculation, the point cloud must be clustered and segmented first. The purpose of clustering is to divide the data of tens of thousands of points in one frame into individual obstacles, and each obstacle contains a certain number of points. The DBSCAN algorithm with a representative base density is used. It defines a cluster as the largest set of density-connected points, can divide regions with sufficient high density into clusters, and can discover clusters of any shape in a spatial database. Since most of the laser from the lidar is absorbed by the water when it hits the water surface, and a small part is specularly reflected without echo, there is no point cloud data on the water surface, so the classical plane segmentation algorithm is not used here.

[0040] (3) Target feature calculation: Traverse each target cluster generated in the previous step and calculate the following feature values: target length, target width, average reflection intensity, reflection intensity variance, and distance. The target length and target width are calculated through the minimum bounding rectangle of the target's 2D plane projection. The average reflection intensity and reflection intensity variance are calculated from the reflection intensities of all points in each target cluster. The distance is the average value of the lengths of all points in the target cluster from the origin of the lidar coordinate system.

[0041] (4) Abnormal target elimination: Since the unmanned boat will shake and fluctuate in actual applications, some abnormal target clusters will be generated and usually disappear in the next frame, unable to form continuous targets. Abnormal values can be deleted through dynamic matching and screening. Therefore, it is necessary to associate multi-frame lidar data and calculate the target cluster features. Considering that there are few water surface targets and they are relatively scattered, the nearest neighbor data matching method is used to associate the obstacles in multi-frame lidar data. A difference function is constructed by fusing target feature information such as position, size, and laser reflection intensity, and the difference between the current target cluster and each target in the previous target cluster list is calculated respectively. The two targets with the smallest difference function values within the gate are matched together. Targets that are successfully dynamically matched for 3 consecutive frames are normal targets and are output, while unmatched target clusters are abnormal targets and are deleted. In addition, considering that extremely small floating objects on the water surface do not pose a threat, targets with too small target length and target width are deleted; the unmanned boat and its on-board equipment may generate point clouds with too close distances, which also need to be deleted.

[0042] The multi-source data fusion and analysis unit first receives the visible light camera target detection results, infrared camera target detection results, and lidar target detection results output by the single-source data target detection unit. Then, the data synchronization module caches the detection results of each sensor, reads the timestamps of each frame of data in the data buffer queue, and caches them in sequence. Taking the generation time of the visible light image data as the reference, it matches the infrared light image detection results and lidar detection results that are closest in time and within no more than 200 milliseconds of difference, and performs subsequent multi-source fusion analysis operations. Next, it constructs an output target set by taking the union of the three results, that is, it includes ships, humans, buoys, flying birds / airplanes, and other obstacle targets detected by the visible light image target detection algorithm model and the infrared image target detection algorithm model, as well as the nearby obstacle targets obtained by the lidar point cloud processing algorithm. Immediately afterwards, it merges overlapping targets and eliminates redundant prediction results with a high degree of overlap. If a certain target is detected by X (X = 1, 2, 3) sensors at the same time, the confidence level of this result is set to X. Finally, it transmits the final detection results after fusion analysis and the single-source sensor detection results to the lower computer through the network.

[0043] The present invention also proposes a multi-source target detection system for a surface unmanned boat based on a domestic deep learning software and hardware platform. The system includes a data acquisition unit, a single-source data target detection unit, and a multi-source data fusion and analysis unit. The data acquisition unit includes a variety of different sensors for environmental perception, and is used to collect different modality information of the surface unmanned boat. The different modality information includes visible light images, infrared images, and 3D point cloud data. The single-source data target detection unit is used to deploy and apply the deep learning target detection algorithms pre-trained on non-domestic platforms on the domestic software and hardware platform by using the domestic lightweight deep learning front-end inference framework MNN, and respectively process the visible light images, infrared images, and 3D point cloud data. The multi-source data fusion and analysis unit is used to receive the different detection results of the information collected by a variety of different sensors output by the single-source data target detection unit; read the data timestamps of different detection results, perform time series caching and alignment, and take the union of the results of the variety of different sensors; construct an output target set including all target categories detected by the variety of different sensors; merge overlapping detection frames, output the final target detection results after fusion processing, and send the different detection results output by the single-source data target detection unit to the lower computer. The features of this system and method embodiment correspond one by one and will not be elaborated here.

[0044] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the multi-source target detection method for the unmanned surface vehicle based on the domestic deep learning software and hardware platform is implemented. The processor includes a first domestic AI processor, a second domestic AI processor, and an NVIDIA processor.

[0045] In summary, the above are only the preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

[0046] For those skilled in the art, it is obvious that the embodiments of the present invention are not limited to the details of the above exemplary embodiments, and the present invention can be implemented in other specific forms without departing from the spirit or basic characteristics of the embodiments of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the embodiments of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the embodiments of the present invention. Any reference signs in the claims should not be regarded as limiting the claimed rights. In addition, it is obvious that the word "including" does not exclude other units or steps, and the singular does not exclude the plural. The multiple units, modules, or devices stated in the system, apparatus, or terminal claims can also be implemented by the same unit, module, or device through software or hardware. The terms first, second, etc. are used to denote names and do not denote any particular order.

[0047] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the embodiments of the present invention and not to limit them. Although the technical solutions of the embodiments of the present invention have been described in detail with reference to the above preferred embodiments, those of ordinary skill in the art should understand that any modification or equivalent replacement of the technical solutions of the embodiments of the present invention should not depart from the spirit and scope of the technical solutions of the embodiments of the present invention.< / pointt>

Claims

1. A multi-source target detection method for surface unmanned boats based on domestic deep learning software and hardware platforms, characterized in that, The data acquisition unit collects different modality information of the unmanned surface vehicle through a variety of different sensors for environmental perception. The different modality information includes visible light images, infrared images, and 3D point cloud data; The single-source data target detection unit uses the domestic lightweight deep learning front-end inference framework MNN to deploy and apply the pre-trained deep learning target detection algorithm on a non-domestic platform to a domestic software and hardware platform, and processes the visible light images, infrared images, and 3D point cloud data respectively; The multi-source data fusion analysis unit receives the different detection results of the information collected by multiple different sensors output by the single-source data target detection unit; reads the data timestamps of the different detection results, performs time series caching and alignment, and takes the union of the results of the multiple different sensors; Construct an output target set, including all target categories detected by the multiple different sensors; merge overlapping detection frames, output the final target detection result after fusion processing, and send the different detection results output by the single-source data target detection unit to the lower computer.

2. The multi-source target detection method for the unmanned surface vehicle based on the domestic deep learning software and hardware platform according to claim 1, wherein: The data acquisition unit specifically includes using a visible light camera, an infrared camera, and a lidar as sensors for the unmanned surface vehicle to perform environmental perception to complement different modality information on the water surface; designing a data buffer queue to cache the sensor data input by the data acquisition unit to prevent memory overflow caused by infinite data stacking, and using a read-write lock for access to ensure data processing security, and scheduling the operation of the algorithm in a multi-threaded parallel manner to accelerate the overall software operation efficiency.

3. The multi-source target detection method for surface unmanned boats based on domestic deep learning software and hardware platforms according to claim 1, characterized in that: In the single-source data target detection unit, the domestic lightweight deep learning front-end inference framework MNN is used to deploy and apply the pre-trained deep learning target detection algorithm on a non-domestic platform to a domestic software and hardware platform to process visible light images and infrared images. Specifically, it includes: based on the non-domestic deep learning framework PyTorch, respectively build target detection network models for processing visible light images and infrared images; use the open-source water surface target detection dataset MODS to train the target detection model for processing visible light images; select the infrared video annotation data in the open-source Singapore Marine Dataset SMD, and use a non-domestic NVIDIA graphics card to train the target detection model to obtain a pre-trained model; use the unmanned surface vehicle equipped with a visible light camera and an infrared camera to collect data on-site in the water area, and after manual annotation, fine-tune the pre-trained model obtained in the previous step respectively; use the domestic deep learning framework MNN to convert the parameter files of the two fine-tuned target detection models generated in the previous step; deploy the two converted target detection models on a computer equipped with a first domestic central processing unit and a first domestic GPU processing unit, with the model inputs being visible light images and infrared images respectively, and the output being the target circumscribed rectangle containing the category results.

4. The multi-source target detection method for a surface unmanned boat based on a domestic deep learning software and hardware platform according to claim 3, characterized in that: The detection of the 3D point cloud data in the single-source data target detection unit specifically includes: downsampling the 3D point cloud data generated by the lidar to reduce the data computation amount without losing environmental information; clustering the disordered point cloud data to form separate target clusters; calculating the following eigenvalue of the target clusters generated by clustering: target length, target width, average reflection intensity, reflection intensity variance, and distance; screening out reasonable obstacles according to the eigenvalue of the target clusters, deleting outliers caused by the shaking and fluctuation of the unmanned boat in practical applications, and targets that do not conform to the actual situation.

5. The multi-source target detection method for surface unmanned boats based on domestic deep learning software and hardware platforms according to claim 1, characterized in that: The multi-source data fusion and analysis unit specifically includes receiving the visible light camera target detection result, infrared camera target detection result, and lidar target detection result output by the single-source data target detection unit; designing a data synchronization module to cache the detection results of each sensor, reading the timestamp of each frame of data in the data buffer queue, and sorting and caching them in time sequence; Taking the generation time of the visible light image data as the reference, matching the infrared light image detection result and lidar detection result that are closest in time and within 200 milliseconds of it, and constructing an output target set by taking the union of the three results, that is, including ship, human, buoy, bird / airplane, and other obstacle targets detected by the visible light image target detection algorithm model and infrared image target detection algorithm model, as well as near obstacle targets obtained by the lidar point cloud processing algorithm; Merging overlapping targets, eliminating redundant prediction results with a high degree of overlap. If a certain target is detected by X sensors at the same time, the confidence of this result is set to X; transmitting the final detection result after fusion analysis and the single-source sensor detection result to the lower computer through the network.

6. The multi-source target detection method for unmanned surface vessels based on domestic deep learning software and hardware platforms according to claim 3, characterized in that: The first domestic central processing unit is the Feiteng D2000 processor, and the first domestic GPU processor is the Jingjiawei JH920 graphics card.

7. A multi-source target detection system for a surface unmanned boat based on a domestic deep learning software and hardware platform, the system includes a data acquisition unit, a single-source data target detection unit, and a multi-source data fusion and analysis unit, characterized in that The data acquisition unit includes a variety of different sensors for environmental perception, and is used to collect different modality information of the surface unmanned boat. The different modality information includes visible light images, infrared images, and 3D point cloud data; The single-source data target detection unit is used to deploy and apply the deep learning target detection algorithm pre-trained on a non-domestic platform on the domestic software and hardware platform by using the domestic lightweight deep learning front-end inference framework MNN, and process the visible light image, infrared image, and 3D point cloud data respectively; The multi-source data fusion and analysis unit is used to receive the different detection results of the information collected by a variety of different sensors output by the single-source data target detection unit; reading the data timestamps of different detection results, performing time series caching and alignment, and taking the union of the results of the variety of different sensors; Construct an output target set that includes all target categories detected by the multiple different sensors; merge overlapping detection boxes, output the final target detection result after fusion processing, and send the different detection results output by the single-source data target detection unit to the lower computer.

8. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, it implements the multi-source target detection method for a surface unmanned boat based on a domestic deep learning software and hardware platform as described in any one of claims 1 to 6. The processor includes a first domestic AI processor, a second domestic AI processor, and an NVIDIA processor.

Citation Information

Cited By

  • Marine ship intelligent detection and three-dimensional positioning method based on multi-modal data cooperation

    CN121259288A