Multi-modal fusion driven ship navigation target detection method and device

The ship navigation target detection method driven by multimodal fusion utilizes feature extraction and fusion of visual, marine radar and AIS data, and projects it onto the semantic space of bird's-eye view for decoding. This solves the problems of spatiotemporal heterogeneity and semantic differences in multimodal data fusion, and improves the reliability and accuracy of target detection.

CN122049335APending Publication Date: 2026-05-15NINGBO OCEAN SHIPPING CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NINGBO OCEAN SHIPPING CO LTD
Filing Date
2026-02-02
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing multimodal data fusion methods struggle to overcome the spatiotemporal heterogeneity and semantic differences between multimodal data in ship target detection, resulting in low consistency and robustness in target recognition under complex sea conditions.

Method used

By acquiring visual perception data, marine radar data, and AIS data of ships, a feature extraction network is used to extract visual feature vectors, marine radar feature vectors, and AIS feature vectors. The marine radar feature vector is used as the dominant feature vector and is supplemented and fused with the visual and AIS feature vectors. The data is then projected and mapped onto the bird's-eye view semantic space for decoding, and finally input into the detection head for target detection.

Benefits of technology

It improves the consistency and robustness of target detection under complex sea conditions, ensures the reliability and accuracy of detection results, and overcomes the problems of spatiotemporal heterogeneity and semantic differences in multimodal data fusion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122049335A_ABST
    Figure CN122049335A_ABST
Patent Text Reader

Abstract

The invention provides a multi-modal fusion-driven ship navigation target detection method and device, and the method comprises the steps: extracting the feature vectors of ship vision, marine radar, AIS and ship navigation data, taking the vector of the marine radar as a dominant vector, taking the other feature vectors as supplementary fusion, and achieving the modal complementation. And mapping the fusion feature vector to a bird's-eye view semantic space to obtain bird's-eye view features, decoding the bird's-eye view features, and inputting the decoded bird's-eye view features into a detection head for target detection to obtain a target detection result. In this way, through multi-modal fusion complementation and semantic space mapping, the problem that existing multi-modal data fusion only stays at a data layer for direct splicing or shallow feature level fusion can be avoided, and the space-time isomerism and semantic difference between multi-modal data are overcome. The consistency and robustness of target detection in the ship navigation process are improved under the complex sea condition, and the reliability and accuracy of the detection result are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of ship inspection technology, and in particular relates to a multimodal fusion-driven method and device for detecting ship navigation targets. Background Technology

[0002] In ensuring maritime navigation safety, situational awareness is a key technology for intelligent ships, and ship target detection is a core function of situational awareness. Target detection also provides crucial support for identifying surrounding vessels, obstacles, and other critical objects. The detection results are input into the decision-making system, supporting dynamic collision avoidance decisions during intelligent navigation. To improve accuracy, existing target detection solutions increasingly employ multimodal data, integrating data from marine radar, AIS, and vision to enhance perception precision and system fault tolerance. However, current multimodal data fusion methods often rely on direct data layer stitching or shallow feature-level fusion. This fails to overcome the spatiotemporal heterogeneity and semantic differences between multimodal data, resulting in low consistency and robustness in target recognition under complex sea conditions. Summary of the Invention

[0003] In view of the shortcomings of the prior art, the purpose of this invention is to provide a multimodal fusion-driven method and apparatus for detecting ship navigation targets, so as to solve the above problems.

[0004] This invention provides a multimodal fusion-driven method for ship navigation target detection, comprising: acquiring visual perception data, marine radar data, AIS data, and ship navigation data; extracting features from the visual perception data, marine radar data, AIS data, and ship navigation data using a feature extraction network matched with each modal data, obtaining visual feature vectors, marine radar feature vectors, AIS feature vectors, and ship navigation feature vectors; using the marine radar feature vector as the dominant feature vector, supplementing and fusing the dominant feature vector with the visual feature vector, the AIS feature vector, and the ship navigation feature vector, to obtain a fused feature vector; projecting and mapping the fused feature vector onto a bird's-eye view semantic space according to the polar coordinate imaging characteristics of marine radar and the mapping relationship characteristics of the visual coordinate system, to obtain bird's-eye view features; decoding the bird's-eye view features to obtain fused decoded features; and inputting the fused decoded features into a detection head for target detection to obtain target detection results.

[0005] In one embodiment of the present invention, the dominant feature vector is supplemented and fused using the visual feature vector, the AIS feature vector, and the ship's navigation feature vector to obtain a fused feature vector. This includes: sequentially using each of the visual feature vector, the AIS feature vector, and the ship's navigation feature vector as a target feature vector; and supplementing and fusing the dominant feature vector using a multi-head attention mechanism based on the target feature vector to obtain the fused feature vector.

[0006] In one embodiment of the present invention, the method further includes: constructing a query vector corresponding to each feature point in the dominant feature vector based on the dominant feature vector; using the target feature vector as a key-value pair; and supplementing and fusing the dominant feature vector based on the query vector and the key-value pair to obtain a fused feature vector.

[0007] In one embodiment of the present invention, the dominant feature vector is supplemented and fused according to the query vector and the key value to obtain the fused feature vector, including: determining the similarity between each query vector and each key vector; determining the vector weight of each key vector according to the similarity between each query vector and each key vector; and supplementing and fusing the dominant feature vector according to the vector weight of each key vector and the value vector corresponding to each key vector to obtain the fused feature vector.

[0008] In one embodiment of the present invention, before obtaining the fused feature vector, the method further includes: determining the relative bearing of the target to be detected and the ship, and determining the relative distance between the target to be detected and the ship; and embedding the relative bearing, relative bearing and modal timestamps of each data during the supplementary fusion process through a multi-head attention mechanism.

[0009] In one embodiment of the present invention, feature extraction is performed on the visual perception data, the marine radar data, the AIS data, and the ship's navigation data using a feature extraction network that matches each modal data, to obtain visual feature vectors, marine radar feature vectors, AIS feature vectors, and ship's navigation feature vectors. This includes: extracting features from the visual perception data and marine radar data using a visual backbone network to obtain a visual feature vector output by the visual backbone network, and obtaining a marine radar feature vector output by the visual backbone network; and extracting features from the AIS data and ship's navigation data using a multilayer perceptron to obtain an AIS feature vector output by the multilayer perceptron, and obtaining a ship's navigation feature vector output by the multilayer perceptron.

[0010] In one embodiment of the present invention, the detection head includes a convolutional network, a category encoding module, and a decoding module. The fused decoded features are input into the detection head for target detection to obtain a target detection result. This includes: performing convolution processing on the fused decoded features using the convolutional network to obtain decoded convolutional features; performing category encoding on the decoded convolutional features using the category encoding module to obtain category-encoded features; decoding the category-encoded features using the decoding module; and using the decoded results for target detection to obtain a target detection result.

[0011] In one embodiment of the present invention, the decoding module includes three sequentially connected decoding sub-networks, each decoding sub-network including a transformer decoding layer and a network prediction head connected in sequence.

[0012] This invention also provides a multimodal fusion-driven ship navigation target detection device, comprising: an acquisition module configured to acquire visual perception data, marine radar data, AIS data, and ship navigation data of a ship; an extraction module configured to extract features from the visual perception data, marine radar data, AIS data, and ship navigation data using a feature extraction network matched with each modal data, to obtain visual feature vectors, marine radar feature vectors, AIS feature vectors, and ship navigation feature vectors; a fusion module configured to use the marine radar feature vector as the dominant feature vector, and supplement and fuse the dominant feature vector using the visual feature vector, the AIS feature vector, and the ship navigation feature vector to obtain a fused feature vector; a projection module configured to project and map the fused feature vector onto a bird's-eye view semantic space based on the polar coordinate imaging characteristics of marine radar and the mapping relationship characteristics of the visual coordinate system, to obtain bird's-eye view features; and a detection module configured to decode the bird's-eye view features to obtain fused decoded features, input the fused decoded features to a detection head for target detection, and obtain target detection results.

[0013] The present invention also provides an electronic device, comprising: one or more processors; and a storage device for storing one or more programs, which, when executed by one or more processors, cause the electronic device to perform the steps of the above-described method.

[0014] The present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a computer's processor, causes the computer to perform the steps of implementing the above-described method.

[0015] The beneficial effects of this technical solution are as follows: This technical solution acquires visual perception data, marine radar data, AIS data, and ship navigation data. A feature extraction network is used to extract features from these data, resulting in visual feature vectors, marine radar feature vectors, AIS feature vectors, and ship navigation feature vectors. The marine radar feature vector is used as the dominant feature vector, and the visual feature vector, AIS feature vector, and ship navigation feature vector are used to supplement and fuse the dominant feature vector. This avoids the semantic fragmentation problem caused by traditional data layer splicing and reduces semantic differences between different data modalities. The fused feature vector is projected onto the bird's-eye view semantic space to obtain bird's-eye view features. Mapping the fused feature vector to the same semantic level avoids the spatiotemporal heterogeneity problem in multimodal data fusion. The bird's-eye view features are decoded to obtain fused decoded features, which are then input into the detection head for target detection to obtain the target detection results. This improves the consistency and robustness of target detection under complex sea conditions, ensuring the reliability and accuracy of the detection results.

[0016] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit the invention. Attached Figure Description

[0017] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention. It is obvious that the drawings described below are merely some embodiments of the invention, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort. In the drawings:

[0018] Figure 1 This is a flowchart illustrating a multimodal fusion-driven ship navigation target detection method according to an exemplary embodiment of the present invention; Figure 2 This is a schematic diagram of the structure of a multimodal fusion-driven ship navigation target detection device according to an exemplary embodiment of the present invention; Figure 3 A schematic diagram of a computer system suitable for implementing embodiments of the present invention is shown. Detailed Implementation

[0019] The embodiments of the present invention will be described below with reference to the accompanying drawings and preferred embodiments. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be understood that the preferred embodiments are only for illustrating the present invention and not for limiting the scope of protection of the present invention.

[0020] It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Therefore, the drawings only show the components related to the present invention and are not drawn according to the actual number, shape and size of the components. In actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.

[0021] In the following description, numerous details are explored to provide a more thorough explanation of embodiments of the invention. However, it will be apparent to those skilled in the art that embodiments of the invention may be practiced without these specific details. In other embodiments, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring embodiments of the invention.

[0022] Currently, common sensing methods in ship intelligent navigation perception systems mainly include marine radar, AIS (Automatic Identification System), and camera visual perception systems. However, various sensing devices have significant limitations in complex scenarios (such as poor visibility, port navigation, and collision avoidance in narrow waterways). For example, marine radar can obtain echo images in polar coordinates and has strong anti-interference capabilities, but lacks target category, shape, and precise contour; AIS can obtain static and dynamic navigation data of surrounding ships, but cannot detect targets with AIS turned off, and suffers from latency and false alarms; visual images acquired by cameras have rich appearance features such as contour, texture, and color, but are sensitive to changes in ambient light, and their performance degrades significantly under poor visibility conditions. Therefore, a single sensor is insufficient to meet the comprehensive requirements of accuracy and real-time performance for target detection in complex environments. Therefore, multi-sensor fusion has become a development trend in ship target detection technology, achieving complementary enhancement by integrating information from multiple sources such as radar, AIS, and vision. However, AIS data is unstructured messages, radar data is echo signals or images, and camera data is optical images. The data modalities and semantic dimensions differ significantly among various sensing devices. How to achieve semantic complementarity across modal data, overcome the limitations of single-modal scenarios, and overcome the spatiotemporal heterogeneity and semantic differences between multimodal data has become a current research challenge in ship target detection.

[0023] In view of this, this application proposes a multimodal fusion-driven method and apparatus for detecting ship navigation targets, in order to solve the above problems.

[0024] Please see Figure 1 , Figure 1 This is a flowchart illustrating a multimodal fusion-driven ship navigation target detection method according to an exemplary embodiment of the present invention. Figure 1 As shown, in an exemplary embodiment, the multimodal fusion-driven ship navigation target detection method includes steps S101 to S105, and each step is described in detail below.

[0025] S101, acquire the ship's visual perception data, marine radar data, AIS data, and the ship's navigation data; Specifically, visual perception data includes image or video information of the environment surrounding the ship, which can be acquired by visible light cameras installed on the ship.

[0026] It is understandable that after collecting relevant data through a visible light camera, the collected data can be preprocessed, such as denoising, contrast enhancement, and edge detection, to obtain preprocessed visual perception data.

[0027] It should be noted that marine radar, based on electromagnetic echo imaging, is a legally mandated navigational equipment for ships; while lidar relies on 3D point cloud reconstruction and is primarily used in high-precision, short-range detection scenarios such as autonomous driving. Currently, multimodal fusion of lidar and video is extensively studied in autonomous driving, but its deployment on intelligent ships is relatively low due to factors such as equipment cost, adaptability to the marine environment, and current regulatory requirements. Therefore, fusion algorithms relying on lidar point cloud information are difficult to directly apply to the situational awareness tasks of intelligent ships. This embodiment acquires marine radar data from the ship for subsequent data fusion to adapt to the situational awareness tasks of intelligent ships.

[0028] Understandably, marine radar data can be acquired through marine radar equipment on board a ship, and it can represent information such as the distance, speed, and bearing of an object. AIS data can be acquired through AIS equipment installed on board a ship, and it can provide static information (such as ship name, call sign, length, and ship type) and dynamic information (such as heading, speed, and position). Ship navigation data includes the ship's position data, heading, and speed over land. Among these, the ship's position data can be acquired through devices such as the Global Positioning System (GPS), the heading and heading data can be acquired from the compass, and the speed over land can be acquired from the speed log.

[0029] In addition, after acquiring relevant data through marine radar, AIS equipment, global positioning system, compass and log, the data can also be preprocessed to make the data more accurate and reliable.

[0030] S102, the visual perception data, the marine radar data, the AIS data and the ship navigation data are extracted by a feature extraction network that matches each modal data to obtain visual feature vectors, marine radar feature vectors, AIS feature vectors and ship navigation feature vectors. Specifically, the feature extraction network is a network that transforms raw perceptual data into structured feature vectors. It should be noted that since different perceptual data have different data structures, during the feature extraction process, a feature extraction network matched to each perceptual data can be used to extract features from each data set. This allows for better preservation of information from the original data and improves the accuracy and effectiveness of feature extraction.

[0031] S103, the marine radar feature vector is used as the dominant feature vector, and the visual feature vector, the AIS feature vector, and the ship's navigation feature vector are used to supplement and fuse them to obtain a fused feature vector; Specifically, the feature vector of marine radar can directly reflect core navigation-related information such as target position and distance, serving as the primary basis for anchoring core judgments. By supplementing and fusing visual feature vectors, AIS feature vectors, and the ship's own navigation feature vectors, visual features can supplement target appearance details, AIS features can provide semantic information such as target identity and planned route, and the ship's own navigation features can reflect relative motion relationships. This multi-dimensional supplementation makes the features more comprehensive. Simultaneously, different feature vectors have different error sources; after fusion, they can be cross-checked, reducing the influence of noise and errors from a single sensor or data source, and improving feature stability. This avoids the problem of data overwhelming key semantic information in traditional multimodal fusion, improves the semantic representation between multimodal data, and enables the fused feature vector to more comprehensively and accurately characterize target characteristics.

[0032] S104. Based on the polar coordinate imaging characteristics of marine radar and the mapping relationship characteristics of the visual coordinate system, the fused feature vector is projected and mapped to the bird's-eye view semantic space to obtain the bird's-eye view features. Specifically, marine radar feature vectors are based on polar coordinates, while visual feature vectors are based on pixel coordinates. Direct use of these can lead to spatial misalignment. Projecting them into a bird's-eye view semantic space (a top-down two-dimensional plane coordinate system) achieves coordinate unification. This semantic space can intuitively present the relative position, distance, and distribution of targets on the sea surface, aligning with the spatial judgment logic of crew members or automated navigation systems, thus reducing the understanding cost for subsequent tasks. In other words, the fused feature vectors are transformed from the original data space to the bird's-eye view semantic space, enabling the resulting bird's-eye view features to achieve spatial alignment and modeling of multi-source data. This preserves the key information of the fused feature vectors while conforming to the spatial logic and semantic definition of a bird's-eye view perspective.

[0033] S105, decode the bird's-eye view features to obtain fused decoded features, input the fused decoded features into the detection head for target detection, and obtain the target detection result.

[0034] According to the technical solution provided in this application, visual perception data, marine radar data, AIS data, and ship navigation data of a vessel are acquired. Feature extraction networks are used to extract features from these data to obtain visual feature vectors, marine radar feature vectors, AIS feature vectors, and ship navigation feature vectors. The marine radar feature vector is used as the dominant feature vector, and the visual feature vector, AIS feature vector, and ship navigation feature vector are used to supplement and fuse the dominant feature vector. This avoids the semantic fragmentation problem caused by traditional data layer splicing and reduces semantic differences between different data modalities. The fused feature vector is projected onto the bird's-eye view semantic space to obtain bird's-eye view features. Mapping the fused feature vector to the same semantic level avoids the spatiotemporal heterogeneity problem in multimodal data fusion. The bird's-eye view features are decoded to obtain fused decoded features, which are then input into a detection head for target detection to obtain target detection results. This improves the consistency and robustness of target detection under complex sea conditions, ensuring the reliability and accuracy of the detection results.

[0035] In some embodiments, the dominant feature vector is supplemented and fused using the visual feature vector, the AIS feature vector, and the ship's navigation feature vector to obtain a fused feature vector, including: taking each of the visual feature vector, the AIS feature vector, and the ship's navigation feature vector as a target feature vector in sequence; and supplementing and fusing the dominant feature vector using a multi-head attention mechanism based on the target feature vector to obtain the fused feature vector.

[0036] Specifically, in the process of supplementing and fusing the dominant feature vector with the visual feature vector, the AIS feature vector, and the ship's navigation feature vector, one of the visual feature vector, the AIS feature vector, and the ship's navigation feature vector can be used as the target feature vector, and the target feature vector can be used to supplement and fuse the dominant feature vector.

[0037] Understandably, multi-head attention mechanisms can map input sequences to multiple different subspaces for attention computation, thereby capturing richer feature information. In this embodiment, through a multi-head attention mechanism, the target feature vector can supplement and fuse the dominant feature vector from different perspectives, resulting in a fused feature vector containing more comprehensive and accurate feature information. Specifically, the multi-head attention mechanism generates multiple different attention weights, which are applied to the dominant feature vector, effectively integrating key information from the target feature vector into the dominant feature vector, thus obtaining the fused feature vector.

[0038] Furthermore, after the supplementation and fusion of a certain target feature vector is completed, the remaining terms of the two remaining terms are used as target feature vectors, and the dominant feature vector is fused through a multi-head attention mechanism. This process is repeated until all vectors are fused into the dominant feature vector to obtain the final fused feature vector.

[0039] In some embodiments, the method further includes: Based on the dominant feature vector, construct the query vector corresponding to each feature point in the dominant feature vector; use the target feature vector as a key-value pair; and supplement and fuse the dominant feature vector through a multi-head attention mechanism based on the query vector and the key-value pair to obtain the fused feature vector.

[0040] Specifically, after using the visual feature vector as the dominant feature vector, the query vector corresponding to each feature point in the dominant feature vector can be determined based on the salient features of the target to be detected in two-dimensional space, such as edges, textures, colors, and local structures included in the visual feature vector.

[0041] Furthermore, the target feature vector is used as a key-value pair to participate in cross-modal attention computation. In this process, cross-modal attention computation can be implemented based on the multi-head attention mechanism in the transformer.

[0042] In some embodiments, the dominant feature vector is supplemented and fused based on the query vector and key value to obtain a fused feature vector, including: determining the similarity between each query vector and each key vector, and determining the vector weight of each key vector based on the similarity between each query vector and each key vector; and supplementing and fusion the dominant feature vector based on the vector weight of each key vector and the value vector corresponding to each key vector to obtain a fused feature vector.

[0043] Specifically, for each feature point in the dominant feature vector, the dot product similarity between its corresponding query vector and all key vectors in the target feature vector is calculated. Then, the weight distribution is obtained through the softmax function. The weight is then multiplied by its corresponding value vector, and finally, a weighted sum is performed to obtain the supplementary representation of the target feature vector to the dominant feature vector.

[0044] It is understandable that the above process is the fusion process of a certain target feature vector. The fusion process of other target feature vectors can refer to the above process, and will not be elaborated on here.

[0045] Thus, when supplementing and fusing the dominant feature vector using a multi-head attention mechanism based on the query vector and key-value pairs, the multi-head attention mechanism calculates the similarity between the query vector and each key-value pair separately, and generates different vector weights based on the similarity. These vector weights are applied to the dominant feature vector, allowing key information from the target feature vector to be integrated into the dominant feature vector according to different weights, thereby obtaining the fused feature vector. In this way, the fused feature vector can more comprehensively retain the information of the original data, improving the accuracy and effectiveness of feature extraction.

[0046] It should be noted that in step S105, the decoding process of the bird's-eye view features can utilize a target query mechanism based on a Transformer decoding structure to extract and identify target instances from the bird's-eye view features. Each target query vector interacts with the bird's-eye view features during the training phase through a self-attention mechanism, gradually converging into a semantic representation of the potential target in the scene. Furthermore, the obtained features are decoded by a decoder using a SECOND network structure to obtain fused decoded features. Finally, the fused decoded features are fed into the detection head for inference to obtain the final target detection result.

[0047] In some embodiments, before obtaining the fused feature vector, the method further includes: The relative bearing and relative distance between the target and the ship are determined; during the supplementary fusion process using a multi-head attention mechanism, the relative bearing and the modal timestamps of each data point are embedded.

[0048] Specifically, when determining the relative bearing of a target and a vessel, the relative bearing angle can be calculated using mathematical methods such as trigonometric functions, by combining position information from marine radar data or AIS data with the vessel's own navigation data. Relative distance can also be determined using marine radar data, which directly provides information on the distance between the target and the vessel.

[0049] It is understandable that after obtaining the relative orientation and relative distance, these are embedded into the process of supplementing and fusing data through a multi-head attention mechanism. The modal timestamps of each data point record the time information of different sensory data acquisitions; embedding these together allows the fusion process to take into account the temporal characteristics of the data.

[0050] Furthermore, when embedding relative bearing, relative distance, and modal timestamps of each data, the relative bearing, relative distance, and modal timestamps can be encoded and transformed into vector forms suitable for multi-head attention mechanisms. These vectors then participate in attention calculations together with the dominant feature vector and the target feature vector. This allows the fused feature vector to not only contain the feature information of the original perceived data but also the spatial positional relationship between the target and the ship, as well as the temporal information of the data. This further enriches the connotation of the fused feature vector and improves the accuracy and reliability of target detection.

[0051] In some embodiments, feature extraction is performed on the visual perception data, the marine radar data, the AIS data, and the ship's navigation data using a feature extraction network that matches each modal data, to obtain visual feature vectors, marine radar feature vectors, AIS feature vectors, and ship's navigation feature vectors, including: Feature extraction is performed on visual perception data and marine radar data using a visual backbone network to obtain the visual feature vector output by the visual backbone network, and the marine radar feature vector output by the visual backbone network is also obtained. Feature extraction is performed on AIS data and ship navigation data using a multilayer perceptron to obtain the AIS feature vector output by the multilayer perceptron, and the ship navigation feature vector output by the multilayer perceptron is also obtained.

[0052] Specifically, both visual perception data and marine radar data are spatially structured data, containing spatial information such as pixel arrangement, edge contours, and regional relationships. When extracting features from visual perception data and marine radar data, a visual backbone network (such as YOLOv8, convolutional neural network (CNN), Transformer, etc.) is used to gradually extract target edge details and abstract semantic features from visual perception data and marine radar data, respectively, to obtain visual feature vectors and marine radar feature vectors.

[0053] Furthermore, both AIS data and ship navigation data are one-dimensional numerical vectors, lacking the spatial structure of visual perception data and marine radar data. Only numerical conversion and dimensional mapping are required. Therefore, when extracting features from visual perception data and marine radar data, a multilayer perceptron is used to extract AIS feature vectors and ship navigation feature vectors.

[0054] In this way, the characteristics of different types of data can be fully utilized, and appropriate feature extraction methods can be adopted. This results in high-quality feature vectors, improving the accuracy and reliability of target detection.

[0055] In some embodiments, the detection head includes a convolutional network, a class encoding module, and a decoding module. The fused decoded features are input into the detection head for target detection to obtain the target detection result, including: The fused decoded features are processed by convolutional network to obtain decoded convolutional features; the decoded convolutional features are then categorically encoded using a class encoding module to obtain categorically encoded features; the categorically encoded features are then decoded using a decoding module, and the decoded results are used for target detection to obtain the target detection results.

[0056] Specifically, firstly, a convolutional network is used to perform convolutional processing on the fused decoded features to extract spatial hierarchical information. Then, the decoded convolutional features are input into the category encoding module for semantic mapping to generate category encoded features. Finally, based on the category encoded features, the target geometry and attribute information are restored through the decoding module, thus forming a progressively refined feature processing flow.

[0057] It is understandable that convolution processing can enhance the spatial structure representation capability of multimodal fusion features, category encoding can accurately map spatial features to the semantic space of ship target categories, and the decoding stage can perform targeted restoration of category semantic features. The three are connected sequentially and progress in one direction, which can overcome the shortcomings of traditional detection heads in decoding multi-source fusion features.

[0058] In some embodiments, the decoding module includes three sequentially connected decoding subnetworks, each decoding subnetwork including a transformer decoding layer and a network prediction head connected in sequence.

[0059] Specifically, the 3-layer decoding sub-network is a cascaded three-layer feature processing unit, which aims to achieve staged decoding to handle semantic features of different complexities; the transformer decoding layer is a feature processing layer based on a self-attention mechanism, which aims to dynamically capture long-distance dependencies between features; the network prediction head is an output layer that can map features to detection parameters, that is, transform features into specific detection results.

[0060] In this way, through the three-layer cascaded design of the decoding module, the category-encoded features are gradually refined and extracted during the layer-by-layer transmission process. Each decoding sub-network receives the output of the previous stage, dynamically calculates attention weights through the transformer decoding layer to strengthen key features and handle semantic differences, and then generates intermediate detection parameters through the network prediction head, finally outputting high-precision target detection results.

[0061] It should be noted that the final output of the target detection result includes the target location prediction vector and the target category prediction probability distribution. The target location prediction vector includes the target's center point coordinates, the target's bounding box width and height, and the target's rotation angle. Furthermore, the target detection result is processed by non-maximum suppression to remove redundant bounding boxes before outputting the target set.

[0062] All of the above-mentioned optional technical solutions can be combined in any way to form the optional embodiments of this application, and will not be described in detail here.

[0063] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the process of the embodiments of this application.

[0064] The following are embodiments of the apparatus described in this application, which can be used to execute the embodiments of the method described in this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the method described in this application.

[0065] Figure 2 This is a schematic diagram illustrating the structure of a multimodal fusion-driven ship navigation target detection device, as shown in an exemplary embodiment of the present invention. Figure 2 As shown, the exemplary device includes: The acquisition module 201 is configured to acquire the ship's visual perception data, marine radar data, AIS data, and the ship's navigation data; The extraction module 202 is configured to extract features from the visual perception data, the marine radar data, the AIS data, and the ship navigation data through a feature extraction network that matches each modal data, to obtain visual feature vectors, marine radar feature vectors, AIS feature vectors, and ship navigation feature vectors. The fusion module 203 is configured to use the marine radar feature vector as the dominant feature vector, and supplement and fuse the dominant feature vector with the visual feature vector, the AIS feature vector and the ship navigation feature vector to obtain a fused feature vector. The projection module 204 is configured to project and map the fused feature vector onto the bird's-eye view semantic space based on the polar coordinate imaging characteristics of the marine radar and the mapping relationship characteristics of the visual coordinate system, so as to obtain the bird's-eye view features. The detection module 205 is configured to decode the features of the bird's-eye view to obtain fused decoded features, input the fused decoded features into the detection head for target detection, and obtain the target detection result.

[0066] In some embodiments, the fusion module 203 is further configured to sequentially use each of the visual feature vector, the AIS feature vector, and the ship navigation feature vector as a target feature vector; and to supplement and fuse the dominant feature vector according to the target feature vector through a multi-head attention mechanism to obtain the fused feature vector.

[0067] In some embodiments, the fusion module 203 is further configured to construct a query vector corresponding to each feature point in the dominant feature vector based on the dominant feature vector; use the target feature vector as a key-value pair; and perform supplementary fusion on the dominant feature vector based on the query vector and the key-value pair to obtain a fused feature vector.

[0068] In some embodiments, the fusion module 203 is further configured to determine the similarity between each query vector and each key vector; determine the vector weight of each key vector based on the similarity between each query vector and each key vector; and supplement and fuse the dominant feature vector based on the vector weight of each key vector and the value vector corresponding to each key vector to obtain a fused feature vector.

[0069] In some embodiments, the fusion module 203 is further configured to determine the relative bearing of the target to be detected and the ship, and to determine the relative distance between the target to be detected and the ship; and to embed the relative bearing, relative bearing and modal timestamps of each data during the supplementary fusion process through a multi-head attention mechanism.

[0070] In some embodiments, the fusion module 203 is further configured to extract features from visual perception data and marine radar data through a visual backbone network to obtain a visual feature vector output by the visual backbone network and a marine radar feature vector output by the visual backbone network; and to extract features from AIS data and ship navigation data through a multilayer perceptron to obtain an AIS feature vector output by the multilayer perceptron and a ship navigation feature vector output by the multilayer perceptron.

[0071] In some embodiments, the detection head includes a convolutional network, a category encoding module, and a decoding module. The detection module 205 is further configured to perform convolution processing on the fused decoded features using the convolutional network to obtain decoded convolutional features; perform category encoding on the decoded convolutional features using the category encoding module to obtain category encoded features; decode the category encoded features using the decoding module; and perform target detection using the decoded results to obtain target detection results.

[0072] In some embodiments, the decoding module includes three sequentially connected decoding subnetworks, each decoding subnetwork including a transformer decoding layer and a network prediction head connected in sequence.

[0073] According to the apparatus provided in this application embodiment, visual perception data, marine radar data, AIS data, and ship navigation data of a vessel are acquired. Feature extraction networks are used to extract features from these data to obtain visual feature vectors, marine radar feature vectors, AIS feature vectors, and ship navigation feature vectors. The marine radar feature vector is used as the dominant feature vector, and the visual feature vector, AIS feature vector, and ship navigation feature vector are used to supplement and fuse the dominant feature vector. This avoids the semantic fragmentation problem caused by traditional data layer splicing and reduces semantic differences between different data modalities. The fused feature vector is projected onto the bird's-eye view semantic space to obtain bird's-eye view features. Mapping the fused feature vector to the same semantic level avoids the spatiotemporal heterogeneity problem in multimodal data fusion. The bird's-eye view features are decoded to obtain fused decoded features, which are then input into a detection head for target detection to obtain target detection results. This improves the consistency and robustness of target detection under complex sea conditions, ensuring the reliability and accuracy of the detection results.

[0074] It should be noted that the multimodal fusion-driven ship navigation target detection device and the multimodal fusion-driven ship navigation target detection method provided in the above embodiments belong to the same concept. The specific operation methods of each module and unit have been described in detail in the method embodiments and will not be repeated here. In practical applications, the target detection device provided in the above embodiments can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. This is not a limitation here.

[0075] Embodiments of the present invention also provide an electronic device, including: one or more processors; and a storage device for storing one or more programs, which, when executed by one or more processors, cause the electronic device to implement the methods provided in the above embodiments.

[0076] Figure 3 A schematic diagram of a computer system suitable for implementing embodiments of the present invention is shown. It should be noted that... Figure 3 The computer system 300 of the electronic device shown is merely an example and should not impose any limitation on the functionality and scope of the embodiments of the present invention.

[0077] like Figure 3As shown, the computer system 300 includes a Central Processing Unit (CPU) 301, which can perform various appropriate actions and processes, such as executing the methods described in the above embodiments, based on programs stored in Read-Only Memory (ROM) 302 or programs loaded from storage portion 308 into Random Access Memory (RAM) 303. The RAM 303 also stores various programs and data required for system operation. The CPU 301, ROM 302, and RAM 303 are interconnected via a bus 304. An Input / Output (I / O) interface 305 is also connected to the bus 304.

[0078] The following components are connected to I / O interface 305: an input section 306 including a keyboard, mouse, etc.; an output section 307 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 308 including a hard disk, etc.; and a communication section 309 including a network interface card such as a LAN (Local Area Network) card, modem, etc. The communication section 309 performs communication processing via a network such as the Internet. A drive 310 is also connected to I / O interface 305 as needed. Removable media 311, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 310 as needed so that computer programs read from them can be installed into storage section 308 as needed.

[0079] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing computer programs for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 309, and / or installed from removable medium 311. When the computer program is executed by central processing unit (CPU) 301, it performs various functions defined in the system of the present invention.

[0080] It should be noted that the computer-readable medium shown in the embodiments of the present invention can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. A computer-readable storage medium can be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In the present invention, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, wherein a computer-readable computer program is carried. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media can also be any computer-readable medium other than computer-readable storage media, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The computer program contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wired, etc., or any suitable combination thereof.

[0081] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. Each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0082] The units described in the embodiments of the present invention can be implemented in software or hardware, and the described units can also be located in a processor. The names of these units do not necessarily limit the specific unit itself.

[0083] Another aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a computer's processor, causes the computer to perform the method as described above. This computer-readable storage medium may be included in the electronic device described in the above embodiments, or it may exist independently and not assembled into the electronic device.

[0084] Another aspect of the present invention provides a computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods described in the various embodiments above.

[0085] The above embodiments are merely illustrative of the principles and effects of the present invention and are not intended to limit the present invention. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in the present invention should still be covered by the present invention.

Claims

1. A multimodal fusion-driven method for detecting ship navigation targets, characterized in that, include: Acquire visual perception data, marine radar data, AIS data, and the ship's own navigation data; The visual perception data, the marine radar data, the AIS data, and the ship's navigation data are extracted using a feature extraction network that matches each modal data, resulting in visual feature vectors, marine radar feature vectors, AIS feature vectors, and ship navigation feature vectors. The marine radar feature vector is used as the dominant feature vector, and the visual feature vector, the AIS feature vector, and the ship's navigation feature vector are used to supplement and fuse the dominant feature vector to obtain a fused feature vector. Based on the polar coordinate imaging characteristics of marine radar and the mapping relationship characteristics of the visual coordinate system, the fused feature vector is projected and mapped onto the bird's-eye view semantic space to obtain the bird's-eye view features. The bird's-eye view features are decoded to obtain fused decoded features, which are then input into the detection head for target detection to obtain the target detection result.

2. The method according to claim 1, characterized in that, The process of supplementing and fusing the dominant feature vector with the visual feature vector, the AIS feature vector, and the ship's navigation feature vector to obtain a fused feature vector includes: Each of the visual feature vector, the AIS feature vector, and the ship's navigation feature vector is taken as the target feature vector in sequence. Based on the target feature vector, the dominant feature vector is supplemented and fused using a multi-head attention mechanism to obtain the fused feature vector.

3. The method according to claim 2, characterized in that, The method further includes: Based on the dominant feature vector, construct a query vector corresponding to each feature point in the dominant feature vector; Use the target feature vector as a key-value pair; The dominant feature vector is supplemented and fused based on the query vector and the key value to obtain the fused feature vector.

4. The method according to claim 3, characterized in that, The step of supplementing and fusing the dominant feature vector based on the query vector and the key value to obtain the fused feature vector includes: Determine the similarity between each query vector and each key vector; The vector weight of each key vector is determined based on the similarity between each query vector and each key vector. The dominant feature vector is supplemented and fused based on the vector weight of each key vector and the value vector corresponding to each key vector to obtain the fused feature vector.

5. The method according to any one of claims 2-4, characterized in that, Before obtaining the fused feature vector, the process also includes: Determine the relative bearing of the target to be detected and the ship, and determine the relative distance between the target to be detected and the ship; The relative orientation, the relative orientation, and the modal timestamps of each data are embedded during the supplementary fusion process using a multi-head attention mechanism.

6. The method according to claim 1, characterized in that, The step involves using a feature extraction network that matches the modal data to extract features from the visual perception data, the marine radar data, the AIS data, and the ship's navigation data, resulting in visual feature vectors, marine radar feature vectors, AIS feature vectors, and ship navigation feature vectors, including: The visual perception data and the marine radar data are subjected to feature extraction by a visual backbone network to obtain the visual feature vector output by the visual backbone network, and the marine radar feature vector output by the visual backbone network is also obtained. The AIS data and the ship's navigation data are subjected to feature extraction by a multilayer perceptron to obtain the AIS feature vector output by the multilayer perceptron and the ship's navigation feature vector output by the multilayer perceptron.

7. The method according to claim 1, characterized in that, The detection head includes a convolutional network, a category encoding module, and a decoding module. The step of inputting the fused decoded features into the detection head for target detection to obtain the target detection result includes: The fused decoded features are processed by the convolutional network to obtain decoded convolutional features; The category encoding module is used to perform category encoding on the decoded convolutional features to obtain category-encoded features; The category-encoded features are decoded using the decoding module, and the decoded results are used for target detection to obtain the target detection result.

8. The method according to claim 7, characterized in that, The decoding module comprises three sequentially connected decoding sub-networks, each of which contains a transformer decoding layer and a network prediction head connected in sequence.

9. A multimodal fusion-driven ship navigation target detection device, characterized in that, include: The acquisition module is configured to acquire the ship's visual perception data, marine radar data, AIS data, and the ship's navigation data; The extraction module is configured to extract features from the visual perception data, the marine radar data, the AIS data, and the ship navigation data through a feature extraction network that matches each modal data, to obtain visual feature vectors, marine radar feature vectors, AIS feature vectors, and ship navigation feature vectors. The fusion module is configured to use the marine radar feature vector as the dominant feature vector, and supplement and fuse the dominant feature vector with the visual feature vector, the AIS feature vector, and the ship's navigation feature vector to obtain a fused feature vector. The projection module is configured to project and map the fused feature vector onto the bird's-eye view semantic space based on the polar coordinate imaging characteristics of the marine radar and the mapping relationship characteristics of the visual coordinate system, so as to obtain the bird's-eye view features. The detection module is configured to decode the bird's-eye view features to obtain fused decoded features, input the fused decoded features into the detection head for target detection, and obtain the target detection result.

10. An electronic device, characterized in that, include: One or more processors and a memory, wherein a computer program is stored in the memory, and when the one or more processors execute the computer program, the device performs the method as described in any one of claims 1 to 8.