Intelligent multi-mode sea object identification system and method
By integrating multiple sensor data and optimizing feature extraction and fusion using Marshall measurement learning, the problems of low accuracy and poor robustness of marine object recognition are solved, and efficient identification and safe navigation of intelligent multimodal marine object recognition system are realized.
Patent Information
- Application Number
- CN202411797833.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-09
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-09
AI Technical Summary
The prior art is difficult to effectively use multiple sensor data for sea objects recognition, resulting in low recognition accuracy and poor robustness, and it is difficult to meet the navigation safety needs of smart ships in complex scenarios.
Integrate water surface optical images, radar images and lidar point cloud data, optimize the feature extraction and fusion process through Marshall measurement learning, and build an intelligent multimodal maritime object recognition system to realize the fusion and feature extraction of multimodal data.
Through multimodal data fusion, active and accurate identification of maritime encounter targets is achieved, recognition accuracy and robustness are improved, system costs are reduced, and the penetration and practicality in the field of smart ships is improved.
Smart Images

Figure CN119942470A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent ships, and in particular to an intelligent multi-modal marine object recognition system and method. Background Art
[0002] Accurate identification of maritime targets can be used to determine the safety risk levels of different encounter objects during the navigation of intelligent ships, and to make more reasonable and safer navigation decisions in a timely manner for different types of encounter objects, which plays a vital role in preventing ship collisions and planning more reasonable local paths.
[0003] In order to achieve autonomous path planning and collision avoidance, ships usually use a variety of sensors to improve the ability to identify encounter objects. Optical sensors obtain image pixels containing rich details and semantic information; LiDAR is used to obtain point cloud data containing accurate three-dimensional position and contour information; navigation radar can obtain electromagnetic wave reflection data or radar image data of medium and long-distance targets, etc. It can be seen that these sensor data show typical multimodal characteristics, and it is difficult to achieve multimodal data fusion using existing multi-sensor data fusion methods. In addition, compared with vehicles traveling on land, the huge inertia and relatively small damping of ships sailing on the water significantly reduce the maneuverability of ships, and the reaction time when encountering safety risks is greatly shortened, so the accuracy of encounter object recognition during autonomous navigation is required to be higher. Therefore, how to use multimodal sensor data to improve the recognition accuracy of encounter objects has always been a challenging task.
[0004] In recent years, through the efforts of researchers in related fields, this field has made exciting progress. For the problem of multimodal maritime encounter object recognition, early research mainly realized target detection and tracking based on multi-sensor fusion, focusing on identifying the target's attribute information such as distance, direction, and speed. Although detecting and tracking the motion state of encounter objects is very necessary and critical for the path planning of intelligent ships, it is far from enough for implementing accurate and precise safety decisions. With the development of artificial intelligence, recognition methods based on optical images, a single data type, have made great progress in recent years. However, the quality of surface images is often affected by weather and light, which makes the recognition of maritime objects based solely on single sensor data insufficiently robust in real applications. Using multi-sensor data to further improve the perception ability of intelligent systems has a positive practical significance.
[0005] Methods for maritime target recognition: Existing early studies have extracted multiple structured features from synthetic aperture radar (SAR) images to achieve the recognition of oil tankers and cargo ships. Jeong et al. used the AdaBoost method to fuse the recognition results based on texture features and the recognition results based on discrete Fourier transform features to improve the accuracy of target recognition. Since the above methods require the use of SAR remote sensing technology and facilities, they are expensive and have the defects of identifying a single target object, making them difficult to be promoted and applied to the actual perception scenarios of intelligent ships. In terms of general target recognition, Li et al. proposed a data fusion method based on improved DS evidence theory to achieve the fusion of millimeter wave radar and infrared images and applied it to the automatic target recognition system. This method solves the paradox problem caused by inconsistent evidence and inequality by improving the combination rules, thereby improving the effectiveness and accuracy of target recognition. Eitel et al. used two independent convolutional neural networks (CNN) to fuse ordinary visible light images and depth images collected by depth cameras, thereby achieving better recognition results than single images. However, the test objects of this algorithm only include six structured targets. Kim et al. used a doubly weighted neural network to extract features from SAR and infrared images, and used linear and nonlinear fusion strategies to perform decision-level fusion processing on them. The experimental results showed that the nonlinear fusion method has better recognition effect.
[0006] The current limitations of multimodal marine object recognition are: target recognition methods based on optical remote sensing images and SAR images have the ability to recognize marine targets, but they require the use of expensive sensor equipment and are mainly used in the military field, making it difficult to extend to the field of intelligent ships. Target recognition technology based on computer vision and deep learning has made great progress in recent years. However, most of them are based on a single modality, images, and do not make reasonable use of data from multiple sensors, resulting in recognition accuracy that cannot meet the actual application requirements of ship intelligent navigation in complex scenarios. Therefore, it is urgent to propose a method that uses the multi-sensor system carried by intelligent ships to achieve active and accurate recognition of encountered targets, thereby ensuring the navigation safety of intelligent ships. Summary of the invention
[0007] The purpose of the present invention is to provide an intelligent multimodal marine object recognition system and method. The present invention integrates water surface optical images, radar images and lidar point cloud data, and uses Mahalanobis metric learning to optimize feature extraction and fusion processes, thereby solving the problems of traditional single sensor systems in complex marine environments, such as expensive equipment, low target recognition accuracy and poor robustness.
[0008] To achieve this purpose, the present invention designs an intelligent multimodal marine object recognition system, comprising: a data acquisition module for collecting data of three modalities, namely, optical images, laser point clouds and radar images during the navigation of a ship using a ship-borne optical camera, a navigation radar and a laser radar, to construct a multimodal data set;
[0009] The cross-modal dataset construction module is used to select corresponding samples from the water surface optical image, radar image and lidar point cloud data of the multi-modal dataset for multi-modal matching to form a cross-modal dataset;
[0010] The feature extraction module is used to construct a feature extraction model to extract features for different modal data in the cross-modal dataset;
[0011] The unimodal regularization module is used to construct the similarity measurement function of each unimodality through Mahalanobis metric learning, and train the feature vectors of each unimodality as the regularization loss function of the fully connected layer to obtain a unimodal feature extraction model with enhanced distinguishing ability, which is used to extract the feature vectors of each unimodality;
[0012] The feature fusion module is used to concatenate and fuse the single-modal feature vectors to obtain the concatenated and fused feature vectors;
[0013] The multimodal regularization module is used to construct a multimodal combined metric function through Mahalanobis metric learning, and train the concatenated and fused feature vectors as the regularization loss function of the fully connected layer to obtain a multimodal feature extraction model with enhanced distinguishing ability, which is used to extract multimodal feature vectors;
[0014] The recognition module model is used to map the multimodal feature vector to a specific marine object target category by adding a Softmax layer to obtain the marine object target category recognition probability.
[0015] The beneficial effects of the present invention are:
[0016] 1) By integrating data from multiple sensors, such as marine radar, lidar, and optical cameras, active and accurate recognition of targets encountered at sea can be achieved, making up for the shortcomings of single-modality recognition technology;
[0017] 2) By utilizing the multi-sensor system carried by smart ships, there is no need to purchase expensive additional sensor equipment, which reduces system costs and improves the popularity and practicality in the field of smart ships;
[0018] 3) The similarity measurement based on Mahalanobis distance is used to extract single-modal features and design loss functions for multi-modal feature fusion, which can effectively improve the feature representation ability of single modality and the common feature expression ability of multi-modality, which not only enhances the generalization ability of the model, but also improves the quality of fused features. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 A module workflow diagram of an intelligent multi-modal marine object recognition method of the present invention;
[0020] Figure 2 The present invention is a flow chart of an intelligent multi-modal marine object recognition method. DETAILED DESCRIPTION
[0021] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments:
[0022] Embodiment 1:
[0023] like Figures 1-2 An intelligent multi-modal marine object recognition system is shown, comprising:
[0024] The data acquisition module is used to collect data in three modes, namely, optical images, laser point clouds and radar images, during the ship's voyage using the ship's optical camera, navigation radar and lidar, and construct a multimodal data set;
[0025] The cross-modal dataset construction module is used to select corresponding samples from the water surface optical image, radar image and lidar point cloud data of the multi-modal dataset for multi-modal matching to form a cross-modal dataset;
[0026] The feature extraction module is used to construct a feature extraction model to extract features for different modal data in the cross-modal dataset;
[0027] The unimodal regularization module is used to construct the similarity measurement function of each unimodality through Mahalanobis metric learning, and use the similarity measurement function of each unimodality as the regularization loss function of the fully connected layer to train the feature vectors of each unimodality, so as to obtain a unimodal feature extraction model with enhanced distinguishing ability, which is used to extract the feature vectors of each unimodality;
[0028] The feature fusion module is used to concatenate and fuse the single-modal feature vectors to obtain the concatenated and fused feature vectors;
[0029] The multimodal regularization module is used to construct a multimodal combination metric function through Mahalanobis metric learning, and use the multimodal combination metric function as the regularization loss function of the fully connected layer to train the concatenated and fused feature vectors to obtain a multimodal feature extraction model with enhanced distinguishing ability, which is used to extract multimodal feature vectors;
[0030] The recognition module is used to map the multimodal feature vector to a specific marine object target category by adding a Softmax layer to obtain the marine object target category recognition probability.
[0031] In the above technical solution, in the data acquisition module, the specific process is:
[0032] The module includes a ship-borne optical camera, navigation radar and lidar, which are used to collect water surface optical images, radar images and lidar point cloud data during the ship's travel;
[0033] The three modal data are aligned using timestamps and spatial positions so that the three modalities describe the same scene or the same target in the time series. The aligned data constitute a multimodal dataset.
[0034] In the above technical solution, in the cross-modal dataset construction module, the specific construction process is as follows: in the multimodal dataset, select marine object samples, the marine object categories include container ships, bulk carriers, passenger ships, yachts and other marine objects, and use the LabelMe tool to annotate the marine objects in the optical image in the multimodal dataset with rectangular frames and classification labels, and the samples of marine objects of the three modalities constitute a cross-modal sample pair, and all cross-modal sample pairs are combined into a cross-modal dataset. For the modal data of navigation radar and lidar, there is no need to annotate them separately to reduce the annotation workload. The three modal data use the annotation information of the optical image for subsequent training and learning.
[0035] The optical image with visual annotation information is used as the label of other modalities, and weakly supervised metric learning is used to accurately identify the encountering objects, thereby obtaining the final target recognition result. On the one hand, it can overcome the problem that a single modality cannot utilize the useful associations between different modal data, and on the other hand, it can also reduce the time and labor cost of manually annotating various modal data.
[0036] In the above technical solution, the feature extraction module is used to construct a feature extraction model to extract features for different modal data in the cross-modal data set, including:
[0037] Optical image feature extraction: The backbone network ResNet50 is used as a feature extractor for optical image modal data to extract key features and obtain optical image modal features;
[0038] ResNet50 builds its architecture by stacking multiple residual blocks. Each residual unit contains two convolutional layers and a skip connection. The skip connection directly adds the input to the output to form residual learning. This structure enables ResNet50 to effectively solve the gradient vanishing problem in deep neural networks, allowing the network to become deeper, extract richer feature information, and improve the model's expressiveness and accuracy. Specifically, the input optical image first passes through an initial convolution layer, which uses a 7x7 convolution kernel with a step size of 2 and outputs 64 feature maps. It then passes through batch normalization and ReLU activation function. Then, the data passes through a 3x3 maximum pooling layer with a step size of 2 to reduce the spatial dimension. After that, the data enters four convolutions, each of which consists of a different number of residual blocks. The number of convolutional layers and filters in these residual blocks gradually increases, starting from 64 and doubling after each stage until it reaches 512. Finally, after the optical image modality data is extracted through ResNet50, the output of each image is a feature vector with a length of 1024.
[0039] Laser point cloud feature extraction: After cleaning and denoising, coordinate conversion, depth mapping and distortion correction of the laser point cloud modal data, the laser point cloud modality is projected onto an image and input into a convolutional neural network for feature extraction to obtain the laser point cloud modal features;
[0040] Among them, the specific operations of preprocessing and converting the LiDAR point cloud modal data into images are: cleaning, denoising and removing outliers on the LiDAR point cloud modal data; using the rotation matrix and translation vector to determine the transformation relationship between the optical camera and the LiDAR; then using the principle of projection and the transformation relationship to map the three-dimensional coordinate depth in the optical camera coordinate system to the normalized image coordinates in the two-dimensional image coordinate system; and performing distortion correction on the depth mapped image.
[0041] For the laser point cloud modality, in order to achieve the registration and alignment of the laser radar three-dimensional data with the optical image, it is first necessary to calibrate the camera and laser radar in the device. The calibration process includes determining the camera's internal parameters (such as focal length, optical center, etc.) and external parameters (the position and attitude of the camera relative to the reference coordinate system), as well as the position and attitude of the laser radar relative to the same reference coordinate system. Through calibration, the relative position relationship between the camera and the laser radar in the device is obtained, that is, the transformation relationship between the laser radar coordinate system and the camera coordinate system is determined. The transformation relationship is described by a rotation matrix (Rotation Matrix) and a translation vector (Translation Vector). First, the original point cloud data is transformed from the laser radar coordinate system to the camera coordinate system with the optical camera as the coordinate origin:
[0042] P camera =RP lidar +T;
[0043] Among them, P camera and P lidar They represent the position of a point in the camera coordinate system and the lidar coordinate system respectively. The rotation matrix R and translation vector T are obtained through calibration, which are a 3×3 matrix and a 3×1 vector respectively.
[0044] Subsequently, the original point cloud data is cleaned to remove noise points and outliers, and then depth mapping is performed. The three-dimensional point coordinates in the camera coordinate system are converted to normalized image coordinates in the two-dimensional image coordinate system through the principle of projection using the camera's intrinsic parameters and the coordinates of the three-dimensional points in the camera coordinate system:
[0045] Assume that the coordinates of a point in the three-dimensional space in the camera coordinate system are (x, y, z), and convert them to the normalized coordinates (u, v) in the image coordinate system:
[0046]
[0047]
[0048] Where f is the focal length of the camera, (c x ,c v ) are the principal point coordinates of the camera.
[0049] Subsequently, necessary distortion correction and further optimization processing are performed on the image after depth mapping to improve the quality of the two-dimensional image after the LiDAR data is projected. Finally, a convolutional neural network model is constructed to extract features from the normalized image and obtain a 1024-dimensional feature vector of the modal data.
[0050] Radar image feature extraction: De-noising is performed on the radar image modal data. After denoising, a convolutional neural network is constructed to extract features and obtain the radar image modal features.
[0051] For radar image modal data, the radar image is first denoised to remove speckle noise and background noise. In this system, the installation positions of the optical camera and the navigation radar are close, and the scene devices observed by the two overlap in geographic space. Therefore, the collected image data has a high consistency in spatial information, so as to fully retain the mutual reference and complementation capabilities of the two modal data. By adjusting the beam width of the radar and the focal length of the optical camera, the spatial resolution of the two images is matched, so that the radar image and the optical image are closer in details, which is convenient for pixel-level comparison and analysis. Radar images can also use optical image labels to reduce the workload of annotation. For radar image data, since the image is formed by the echo signal, the image data is less information-intensive than the optical image. Therefore, a convolutional neural network is constructed for feature extraction to reduce the calculation amount of the modal data and improve the overall calculation speed of the system; finally, a 1024-length feature vector of the modal data is obtained.
[0052] In the above technical solution, in the unimodal regularization module, in the process of feature extraction for three unimodal data, in order to constrain the consistency of the feature semantics of different modal outputs, a similarity measurement function based on Mahalanobis distance is constructed, which specifically includes:
[0053] The principal component analysis is used to reduce the dimension of the eigenvector of each modal data to 512 degrees, retaining the main components to reduce the computational complexity, and calculate the covariance matrix S and the corresponding inverse matrix S of the modal feature after dimensionality reduction. -1 , construct a unimodal feature extraction loss function based on Mahalanobis distance, and obtain each unimodal feature extraction model with enhanced distinguishing ability by calculating the similarity between sample features to extract each unimodal feature vector;
[0054] Taking the optical image modality as an example, a modal similarity metric function based on Mahalanobis distance is constructed:
[0055]
[0056] Among them, S -1 It is the inverse of the covariance matrix S that is unique to the optical image modality, and x′ and y′ represent the eigenvectors extracted from the modality data;
[0057] In order to improve the distinguishability of internal features in unimodal data, the similarity measurement function based on Mahalanobis distance is used as the loss function of the unimodal feature network to quantify the similarity between sample features. The more similar the unimodal samples are, the shorter the distance between the feature vectors extracted by the neural network is. The more dissimilar the samples are, the greater the distance between the feature vectors is. The similarity measurement function based on Mahalanobis distance improves the feature expression of unimodal data by a single neural network by considering the distribution characteristics of unimodal features. The loss function for unimodal feature extraction is:
[0058] L Mahalanobis (a,p,n)=log(1+exp(D M (a,p)-D M (a,n)+margin))+γ1·I;
[0059] Where a, p, and n represent a set of samples of single-modal data (such as optical image modal data, laser point cloud modal data, and radar image modal data), p and n represent positive samples that are closer to sample a and negative samples that are farther away, respectively. M (a,p),D M (a,n) represents the Mahalanobis distance between the unimodal sample a and the closer positive sample p and the farther negative sample n respectively; margin is a hyperparameter used to ensure the distance gap between positive and negative samples; the γ1·I regularization term is added to balance the loss function of the unimodal feature extraction network, I is the unit matrix, and γ1 is a smaller hyperparameter used to ensure the effectiveness of the loss function during network propagation.
[0060] The three unimodal feature extraction network models are trained respectively using the above-mentioned unimodal network loss function to obtain a unimodal feature extraction model with enhanced distinction ability. The model can be used to extract the optical image modal feature vector, radar image modal feature vector and laser point cloud modal feature vector of the input data.
[0061] In the above technical solution, in the feature fusion module, each single-modal feature vector is integrated into a unified framework and trained through a joint optimization strategy, specifically including:
[0062] Through the splicing operation, the three unimodal feature vectors are spliced, and the spliced common feature vector is nonlinearly transformed through the fusion layer to obtain the fused feature vector. The three unimodal feature extraction networks can extract feature vectors with discrimination, and then multimodal feature fusion and joint optimization are performed to integrate the feature vectors extracted by network models of different modes into a unified framework. First, through the splicing operation, the three unimodal vectors are spliced into a feature vector with a shape of (512, 3). Then, two fully connected layers are connected as a fusion layer to fuse the three unimodal feature vectors. The role of the fusion layer is to perform nonlinear transformation on the spliced feature vectors.
[0063] In the above technical solution, in the multimodal regularization module, a multimodal combination metric function based on the Mahalanobis distance is constructed, and it is used as a regularization loss function to fuse the concatenated feature vectors to obtain a multimodal feature extraction model with distinguishing ability. The loss function of the multimodal combination metric based on the Mahalanobis distance is specifically:
[0064]
[0065] Among them, L c / r is the loss of classification and regression tasks, λ and γ2 are weights used to count the differences between different modal features, f(x) is the mapping function of the fusion layer network to the input feature x, that is, the multimodal feature after fusion, D M (f(x i ),f(x j )) is represented as the feature vector f(x) after multimodal fusion i ) and f(x j )Mahalanobis distance.
[0066] The data acquisition module collects three modal data to construct a multimodal data set, uses the multimodal data set to select corresponding samples for matching, uses the optical image of the marine objects to annotate with rectangular boxes and classification labels, uses the annotation information of the optical image for the three modal data, and constructs the samples of the marine objects of the three modalities into cross-modal sample pairs. All cross-modal sample pairs are combined into a cross-modal data set, and the cross-modal data set is input into the feature extraction module for training. The unimodal attitude measurement function regularization module is used to enhance the distinguishing ability of unimodal features, and each unimodal feature vector is extracted, and each unimodal feature vector is spliced. The spliced unimodal feature vectors are input into the multimodal regularization module for fusion to enhance the features with multimodal expression capabilities, extract the multimodal feature vector, and convert the multimodal feature vector into a probability distribution form through the added Softmax layer to obtain the target recognition result. Through the above content, a trained multimodal attitude measurement learning model is obtained.
[0067] During the testing phase, different modal data in the same scenario are input into the trained multimodal attitude metric learning model to predict the recognition results of the encounter objects.
[0068] By comparing and analyzing the performance differences between single-mode marine object recognition and multi-modal recognition methods, under the same conditions, the multi-modal recognition method is superior to the single-mode recognition method in recognition accuracy. When affected by other factors and the marine radar and lidar data are not obtained, the device degenerates into a single-mode target detection algorithm based on optical images. Through experimental comparison, in the public data set Seaship7000, when only optical images are available, the target detection accuracy of the device is comparable to the results of the mainstream target detection algorithm, and the recognition accuracy is 95.7%. In the multi-modal data set collected by this device, the recognition accuracy of marine objects reaches 97.7%, and when only optical image target detection is performed, the recognition accuracy is 95.4%. The performance of the method of the present invention is better than the target detection algorithm using a single optical mode.
[0069] Embodiment 2:
[0070] An intelligent multi-modal marine object recognition method comprises the following steps:
[0071] The data acquisition module is used for data,construction of multimodal datasets;
[0072] Select corresponding samples from the water surface optical image, radar image and lidar point cloud data of the multimodal data set for multimodal matching to form a cross-modal data set;
[0073] The feature extraction module is used to construct a feature extraction model to extract features for different modal data in the cross-modal dataset;
[0074] Through Mahalanobis metric learning, a similarity metric function is constructed under each single modality, and the feature vector of each single modality is trained as the regularized loss function of the fully connected layer to obtain a single modality feature extraction model with enhanced distinguishing ability, which is used to extract each single modality feature vector;
[0075] Each single-mode feature vector is concatenated and fused to obtain a concatenated and fused feature vector;
[0076] Through Mahalanobis metric learning, a multimodal combined metric function is constructed and used as the regularized loss function of the fully connected layer to train the concatenated and fused feature vectors, thereby obtaining a multimodal feature extraction model with enhanced distinguishing ability, which is used to extract multimodal feature vectors.
[0077] By adding a Softmax layer, the multimodal feature vector is mapped to the specific marine object target category to obtain the marine object target category recognition probability.
[0078] This method collects multimodal data through three-modal data acquisition devices, and accurately identifies marine objects through single-modal feature extraction and multimodal feature fusion framework. Unlike the traditional method of simply concatenating or weighted summing feature vectors of different modalities, this method realizes adaptive selection and weighting of different modal features by designing a specific fusion layer, thereby more effectively extracting and utilizing complementary information in multimodal data. Secondly, in the design of single-modal and multi-modal metric functions, the similarity measurement based on Mahalanobis distance is used for single-modal feature extraction and multi-modal feature fusion loss function design, which can effectively improve the single-modal feature representation ability and multi-modal common feature expression ability. Not only does it enhance the generalization ability of the model, but it also improves the quality of fused features.
[0079] This intelligent multimodal marine object recognition method integrates water surface optical images, radar images and lidar point cloud data to construct a cross-modal data set, and uses a feature extraction model to extract features from multimodal data; through Mahalanobis metric learning, similarity measurement functions are constructed at the unimodal and multimodal levels, and loss functions are constructed using similarity measurement functions to train unimodal feature extraction models and multimodal feature extraction models to enhance the model's ability to distinguish different types of targets; combined with the Softmax classifier, accurate classification of marine objects is achieved. This method not only improves recognition accuracy, but also enhances robustness in complex environments.
[0080] By integrating multiple perception methods, it is possible to capture target characteristics more comprehensively and improve recognition accuracy; the feature representation is optimized using Mahalanobis metric learning, so that good performance can be maintained even in the case of large changes in lighting or bad weather conditions; it provides strong technical support for the fields of ocean monitoring and navigation safety, and helps to improve the intelligence level of related industries. Therefore, this method is of great significance for promoting the effective management and protection of marine resources.
[0081] Embodiment 3:
[0082] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0083] Embodiment 4:
[0084] A computer program product comprises a computer program, wherein the computer program implements the steps of the above method when executed by a processor.
[0085] The contents not described in detail in this specification belong to the prior art known to those skilled in the art. It should be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0086] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0087] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0088] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0089] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit its protection scope. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the field should understand that after reading the present invention, those skilled in the art can still make various changes, modifications or equivalent substitutions to the specific implementation methods of the invention, but these changes, modifications or equivalent substitutions are all within the protection scope of the pending claims of the invention.
[0090] The contents not described in detail in this specification belong to the prior art known to professional and technical personnel in this field.
Claims
1. An intelligent multimodal marine object recognition system, characterized in that: include: The data acquisition module is used to construct a multimodal data set using water surface optical images, radar images, and lidar point cloud data; The cross-modal dataset construction module is used to select corresponding samples from the water surface optical image, radar image and lidar point cloud data of the multi-modal dataset for multi-modal matching to form a cross-modal dataset; The feature extraction module is used to construct a feature extraction model to extract features for different modal data in the cross-modal dataset; The unimodal regularization module is used to construct the similarity measurement function of each unimodality through Mahalanobis metric learning, and train the feature vectors of each unimodality as the regularization loss function of the fully connected layer to obtain a unimodal feature extraction model with enhanced distinguishing ability, which is used to extract the feature vectors of each unimodality; The feature fusion module is used to concatenate the single-modal feature vectors and fuse them through the multi-modal regularization module to obtain the fused feature vector; The multimodal regularization module is used to construct a multimodal combined metric function through Mahalanobis metric learning, and train the concatenated and fused feature vectors as the regularization loss function of the fully connected layer to obtain a multimodal feature extraction model with enhanced distinguishing ability, which is used to extract multimodal feature vectors; The recognition module is used to map the multimodal feature vector to a specific marine object target category by adding a Softmax layer to obtain the marine object target category recognition probability.
2. The intelligent multimodal marine object recognition system according to claim 1, characterized in that: During data collection, the specific process is as follows: The system includes ship-borne optical cameras, navigation radars and lidars, which are used to collect water surface optical images, radar images and lidar point cloud data during the ship's travel; The three modal data are aligned using timestamps and spatial positions so that the three modalities describe the same scene or the same target in the time series. The aligned data constitute a multimodal dataset.
3. The intelligent multimodal marine object recognition system according to claim 1, characterized in that: In the cross-modal dataset construction module, the specific construction process is as follows: in the multimodal dataset, marine object samples are selected, and the marine objects in the optical images in the multimodal dataset are annotated with rectangular frames and classification labels. The samples of marine objects in the three modalities constitute a cross-modal sample pair, and all cross-modal sample pairs are combined into a cross-modal dataset.
4. The intelligent multimodal marine object recognition system according to claim 1, characterized in that: The feature extraction module is used to construct a feature extraction model to extract features for different modal data in the cross-modal dataset, including: Optical image feature extraction: The backbone network ResNet50 is used as a feature extractor for optical image modal data to extract key features and obtain optical image modal features; Laser point cloud feature extraction: After cleaning and denoising, coordinate conversion, depth mapping and distortion correction of the laser point cloud modal data, the laser point cloud modality is projected onto an image and input into a convolutional neural network for feature extraction to obtain the laser point cloud modal features; Among them, the specific operations of preprocessing and converting the LiDAR point cloud modal data into images are: cleaning, denoising and removing outliers of the LiDAR point cloud modal data; using the rotation matrix and translation vector to determine the transformation relationship between the optical camera and the LiDAR; then using the principle of projection and the transformation relationship, mapping the three-dimensional coordinate depth in the optical camera coordinate system to the normalized image coordinates in the two-dimensional image coordinate system; and performing distortion correction on the image after depth mapping; Radar image feature extraction: De-noising is performed on the radar image modal data. After denoising, a convolutional neural network is constructed to extract features and obtain the radar image modal features.
5. The intelligent multimodal marine object recognition system according to claim 1, characterized in that: In the single-mode regularization module, the eigenvector of each modal data is reduced in dimension, and the covariance matrix S and the corresponding inverse matrix S of the modal feature after dimensionality reduction are calculated. -1 , construct a unimodal similarity measurement function based on Mahalanobis distance as the unimodal feature extraction loss function, calculate the similarity between sample features, enhance the distinguishability of internal features under unimodality, and obtain each unimodal feature vector with enhanced distinguishability. The loss function of the unimodal feature extraction model is: L Mahalanobis (a,p,n)=log(1+exp(D M (a,p)-D M (a,n)+margin))+γ1·I; Among them, a, p, and n represent a set of samples of unimodal data, p and n represent positive samples that are close to sample a and negative samples that are far away, respectively. M (a,p),D M (a,n) represents the Mahalanobis distance between the unimodal sample a and the closer positive sample p and the farther negative sample n respectively; margin is a hyperparameter used to ensure the distance gap between positive and negative samples; the γ1·I regularization term is added to balance the loss function of the unimodal feature extraction network, I is the unit matrix, and γ1 is a smaller regularization parameter.
6. The intelligent multimodal marine object recognition system according to claim 1, characterized in that: In the feature fusion module, the single-modal feature vectors are integrated into a unified framework and trained through a joint optimization strategy, including: Through the splicing operation, the three single-modal feature vectors are spliced, and the spliced common feature vector is nonlinearly transformed through the fusion layer to obtain the fused feature vector.
7. The intelligent multimodal marine object recognition system according to claim 1, characterized in that: In the multimodal regularization module, a multimodal combined metric function based on the Mahalanobis distance is constructed and used as a regularized loss function to train the fused feature vector to obtain common features with distinguishing ability. The loss function of the multimodal combined metric based on the Mahalanobis distance is specifically: Among them, L c / r is the loss of classification and regression tasks, λ and γ2 are parameters used to count the differences between features of different modalities, f(x) is the mapping function of the fusion layer network to the input feature x, and D M (f(x i ),f(x j )) is represented as the feature vector f(x) after multimodal fusion i ) and f(x j ) is the Mahalanobis distance between .
8. An intelligent multi-modal marine object recognition method, characterized in that: The steps include: The data acquisition module is used for data,construction of multimodal datasets; Select corresponding samples from the water surface optical image, radar image and lidar point cloud data of the multimodal data set for multimodal matching to form a cross-modal data set; The feature extraction module is used to construct a feature extraction model to extract features for different modal data in the cross-modal dataset; Through Mahalanobis metric learning, a unimodal feature extraction loss function based on Mahalanobis distance is constructed, and the feature vectors of each unimodality are trained to obtain a unimodal feature extraction model with enhanced distinguishing ability, which is used to extract each unimodal feature vector; Each single-mode feature vector is concatenated and fused to obtain a concatenated and fused feature vector; Through Mahalanobis metric learning, a multimodal combined metric function is constructed and used as the regularized loss function of the fully connected layer to train the concatenated and fused feature vectors, thereby obtaining a multimodal feature extraction model with enhanced distinguishing ability, which is used to extract multimodal feature vectors. By adding a Softmax layer, the multimodal feature vector is mapped to the specific marine object target category to obtain the marine object target category recognition probability.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to claim 8 are implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method described in claim 8 are implemented.
Citation Information
Patent Citations
Target detection and identification method and system based on multi-source information fusion and storage medium
CN116091883A
Method and system for 3D object detection by using multi-modal expert knowledge
CN117475425A
Efficient robot vision system based on deep learning and multi-modal fusion
CN118865042A
VISION-LiDAR FUSION METHOD AND SYSTEM BASED ON DEEP CANONICAL CORRELATION ANALYSIS
US20220366681A1