Sea surface vessel detection method and system based on tri-modal attention fusion
By using the trimodal attention fusion method, Swin Transformer and DH-DINO algorithms, the accuracy and robustness issues of surface ship detection in complex environments are solved, and efficient detection effects are achieved in all weather and multi-scale.
Patent Information
- Application Number
- CN202510815179.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-09-12
AI Technical Summary
Existing technologies find it difficult to effectively fuse visible light, short-wave infrared, and medium-wave infrared images in complex sea surface environments, resulting in insufficient accuracy and robustness in sea surface ship detection, especially in all-weather conditions where it is difficult to meet detection needs.
A three-channel network based on Swin Transformer and the DH-DINO algorithm are used to realize dynamic alignment and feature fusion of trimodal images through feature extraction and attention cross-fusion. A dual-path deformable attention fusion module is designed to handle the problem of weak alignment of different scales and modalities.
It achieves all-weather, multi-scale, and efficient sea surface ship detection, improves the accuracy and robustness of detection, and enhances the adaptability of the model in complex environments.
Smart Images

Figure CN120635392A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image processing technology, and in particular relates to a method and system for detecting sea surface ships based on trimodal attention fusion. Background Art
[0002] With the development of my country's marine science and technology, intelligent water transportation technology has gradually become popular, forming a high-quality development system for the shipping industry. Surface vessel detection is a key technology in fields such as ocean monitoring and maritime traffic management. It can accurately and in real time detect and identify surface vessel targets, which is of great significance for preventing maritime accidents, maintaining maritime safety, and planning maritime traffic.
[0003] Currently, detecting ships on the sea surface faces numerous challenges. For example, sea scenes are susceptible to weather, lighting, and wave conditions. For example, fog, rain, snow, and low-light conditions at night can significantly reduce the clarity of visible light images, making target identification difficult. Furthermore, the complex sea surface background is susceptible to false detections or missed detections due to interference from waves and reflected light. Furthermore, the scale of ship targets in images varies significantly. A distant ship may occupy only a few pixels, while a close-up one may occupy the majority of the image area. Traditional visible light-based target detection performs poorly in complex environments and struggles to meet practical needs.
[0004] Short-wave infrared (SWIR) has a wavelength range of 0.75μm to 2.5μm and primarily senses near-infrared light reflected from targets. Its wavelength is close to visible light, resulting in higher imaging resolution, the ability to capture more detail, and strong penetration of scattering media such as haze and smoke, making it suitable for low-visibility conditions. Medium-wave infrared (MWIR) has a wavelength range of 3μm to 8μm and primarily senses the thermal radiation energy of targets. It is suitable for detecting medium-temperature targets. Unlike short-wave infrared, it is based on the thermal radiation characteristics of the target, so it does not require an external light source and is suitable for nighttime and low-light conditions, allowing detection of targets in almost complete darkness. Visible light images, on the other hand, provide high-resolution texture and color information and are suitable for scenes with ample light.
[0005] Cross-modal detection improves target detection performance by integrating data from different sensors and fully utilizing the complementary information of each modality. This is of great significance for the detection of surface ships. First, it has sufficient all-weather detection capabilities. Visible light is suitable for scenes with abundant daylight, short-wave infrared is suitable for low-visibility conditions such as haze, and medium-wave infrared is suitable for nighttime detection. The fusion of three modalities can cover all-weather scenarios and improve the robustness of the system in different environments. Second, visible light can provide high-resolution texture information, short-wave infrared can enhance the target outline, and medium-wave infrared provides thermal radiation information. The combination of these three can comprehensively characterize the target characteristics, improve detection accuracy, and effectively suppress noise interference in a single modality.
[0006] In addition, cross-modal detection faces some technical challenges. Different modalities may have significant differences in imaging principles and feature distribution. Direct fusion may lead to information redundancy or conflict. How to effectively fuse multimodal information while fully utilizing the complementarity of each modality is the core issue of multimodal algorithms.
[0007] To address the above problems, there is an urgent need for a cross-modal attention sea surface ship detection method and system that can effectively integrate visible light, short-wave infrared and medium-wave infrared. Summary of the Invention
[0008] The purpose of this invention is to address the shortcomings of current surface ship detection methods and propose a cross-modal attention fusion detection method and system based on visible light, short-wave infrared and medium-wave infrared images to comprehensively characterize the information of surface ships and achieve all-weather, multi-scale and high-efficiency surface ship detection.
[0009] The present invention mainly solves the problem of surface ship detection through two approaches: on the one hand, the present invention utilizes the complementarity between trimodal images and adopts a three-channel network based on Swin Transformer to dynamically align different modalities through feature extraction and attention-based feature cross-fusion, thereby realizing multimodal information interaction and reducing information conflicts caused by differences between modalities; on the other hand, the present invention proposes the DH-DINO algorithm, which integrates a dual-path deformable attention fusion module on the basis of the DINO detection algorithm to handle targets of different scales and weak alignment of modalities, thereby improving the robustness and accuracy of target detection.
[0010] The object of the present invention is achieved through the following technical solutions:
[0011] According to a first aspect of this specification, a method for detecting sea surface ships based on trimodal attention fusion is provided, the method comprising:
[0012] Step 1: Three-band image acquisition: Use continuously zoomed and registered visible light, shortwave infrared, and medium wave infrared cameras to simultaneously acquire visible light images, shortwave infrared images, and medium wave infrared images;
[0013] Step 2, coarse alignment of the three-band images, includes the following sub-steps:
[0014] 2.1 Use video editing software to align the three video streams. Based on the video timestamps and the ship's movement trajectory, align the multimodal video streams on the timeline to ensure that images at the same time point correspond to the same scene.
[0015] 2.2 By adjusting the translation, rotation and scaling parameters of each frame of the image, the position of the target in the image is basically consistent, completing the rough alignment of the three-band images;
[0016] 2.3 Export the roughly aligned video frames as image sequences and select high-quality images for subsequent feature extraction and object detection;
[0017] Step 3, ship target detection based on coarsely aligned three-band images, includes the following sub-steps:
[0018] 3.1 Repeat steps 1 and 2 to obtain a set of three-band images, manually annotate the ships in the images, and create a three-band visible light-infrared image registration dataset;
[0019] 3.2 The filtered dataset is enhanced by adding noise, translation, scaling, rotation and other operations. Adversarial training is used to generate adversarial samples, and the training data is expanded to obtain a three-band visible light-infrared image fusion dataset.
[0020] 3.3 The improved DH-DINO algorithm is trained on the three-band visible light-infrared image fusion dataset to obtain the sea surface ship detection model;
[0021] 3.4 Input the three-band image into the trained sea surface ship detection model to obtain the sea surface ship detection results.
[0022] Furthermore, in the step 1, the three-band camera used is specifically composed of a short-wave infrared camera, a medium-wave infrared camera and a visible light camera on the same optical axis, the camera lenses are in the same plane, and the baseline distance is 4 to 8 cm; the visible light camera uses a continuous zoom visible light camera with a focal length of 500 mm and a resolution of 3840*2160, which can capture high-resolution texture and color information; the short-wave infrared camera is an ultra-wide spectrum continuous zoom camera with a focal length of 420 mm, a spectral range of 0.4 to 1.7 μm, an InGaAs focal plane detector, and a camera resolution of 640*512; the medium-wave infrared camera is an ultra-low light night vision camera with a minimum illumination of 10 -4 LUX, with a focal length of 100mm and a resolution of 610*512. Before acquisition, the three cameras are spatially aligned and calibrated. A gimbal is used to ensure consistent imaging perspectives, reducing the complexity of subsequent alignment. The captured images are stored as separate video streams for easy processing.
[0023] Furthermore, in step 2.2, manual adjustments are made based on the content of each frame of the video. These adjustments primarily involve translation, rotation, and scaling. The image position is translated to maintain a consistent position across the three bands; the image rotation angle is adjusted to eliminate perspective differences caused by environmental influences and inherent field of view differences during the capture process; and the image size is adjusted to eliminate differences caused by varying resolutions across the three devices. By overlaying the aligned images, the primary target is checked for alignment, and further fine-tuning is performed.
[0024] Furthermore, in step 2.3, the roughly aligned video frames are exported as image sequences at fixed time intervals, and after data sorting and cleaning, they are stored as visible light, short-wave infrared, and medium-wave infrared image pairs for subsequent feature extraction and target detection.
[0025] Furthermore, in step 3.1, a three-band visible-light-infrared image registration dataset is prepared, specifically as follows: images of ships entering, leaving, and mooring at different distances and scenes are collected at different ports, where the ship types include cargo ships, passenger ships, fishing boats, speedboats, etc.; three-band cameras with the same optical axis are mounted on drones, shores, and ships, and the drones are kept at a suitable height from the sea surface. All-weather sea surface ship images are collected at a suitable distance from the ships, including images of images at all time periods throughout the day, especially during time periods with large changes in illumination and seawater temperature. Data collection is repeated, and data is collected and sorted under different weather conditions such as sunny days, rainy days, and foggy days according to different weather conditions. An all-weather three-band visible-light-infrared ship dataset for complex sea conditions and rich scenes is constructed. Then, the shortwave infrared images, mediumwave infrared images, and visible light images are registered through the alignment and registration operations in step 2 to obtain a three-band visible-light-infrared image registration dataset.
[0026] Furthermore, in step 3.2, data enhancement is performed on the obtained dataset to expand the diversity of the training data, and adversarial samples are generated through adversarial training to improve the model's resistance to noise and interference. The specific process is as follows:
[0027] a) Data augmentation: By performing a series of transformations on the original image, diverse training samples are generated to improve the model's generalization capabilities. Specific operations include noise addition, translation, scaling, and rotation. The purpose of noise addition is to simulate noise interference in real scenes, such as sensor noise or environmental noise, and improve the model's robustness to noise. Operations such as translation and rotation can simulate changes in the position and angle of the target in the image. Image transformations such as brightness and contrast adjustments can also be performed to improve the model's adaptability to weakly aligned targets.
[0028] b) Adversarial Training: Generate adversarial examples so that the model learns to resist noise and interference during training, thereby improving robustness. The goal of adversarial example generation is to produce adversarial examples that can deceive the model. First, a generative adversarial network (GAN) is used to generate adversarial examples. Then, the adversarial examples and the original examples are mixed as training data. An adversarial loss term is added to the loss function to optimize the total loss function. Through adversarial training, the model learns to resist noise and interference, improving its stability in practical applications, providing high-quality training data for the subsequent detection task in step 3.3, and improving the model's generalization ability.
[0029] Furthermore, in step 3.3, the surface ship detection model is trained on the three-band visible light-infrared image fusion dataset constructed in step 3.2. The backbone network is used to extract the feature maps of the three modalities, which are then input into the two-way deformable attention fusion module for feature interaction and fusion. Finally, the DINO detection head is used to output the location and category information of the ship target. The specific process is as follows:
[0030] a) Feature Extraction: We use the Swin-Large network, a large-scale variant of the Swin Transformer pre-trained on public datasets, to build the backbone network. We use a layered design and a windowed attention mechanism to efficiently extract multi-scale features. It consists of four stages, reducing the resolution and increasing the number of channels at each stage, and ultimately outputs a five-layer feature map, which is the low-level features of the input image and the high-level features of the four stages. Assume that the input image is Where H and W are the height and width of the image respectively. First, the input image is divided into non-overlapping patches of size 4×4, and then each patch is flattened into a vector and mapped to the feature space with a dimension of 192 through linear transformation, that is, The next four stages gradually extract multi-scale features through Swin Transformer Block and Patch Merging. Swin Transformer Block extracts features by using windowed multi-head self-attention (W-MSA) and shifted windowed multi-head self-attention (SW-MSA) in pairs, performing multi-head self-attention calculations within the window, thereby significantly reducing the amount of computation and improving computational efficiency. Patch Merging reduces the resolution by 2 times and increases the number of channels by 2 times by grouping, splicing, and linearly transforming feature maps, thereby extracting higher-dimensional feature representations.
[0031] Feature extraction for trimodal images involves simultaneously inputting a corresponding set of three-modal images into the backbone network. That is, each layer of feature maps extracted has three corresponding feature maps, which then enter the subsequent dual-path deformable attention fusion module for feature fusion.
[0032] b) Dual-Path Deformable Attention Fusion Module: Two main convolution-based cross-attention modules are designed to fuse features from the three modalities. Two offset prediction modules are designed to dynamically adjust the medium-wave infrared and short-wave infrared feature maps. An adaptive weight allocation network is designed to assign different weights to the feature maps of different modalities. The offset prediction module generates an adaptive offset sampling field through a lightweight convolutional network, modeling the weak alignment difference between the visible light and two infrared images, thereby guiding feature alignment between the modalities. The cross-attention module implements positional interaction and fusion between the modalities through a multi-scale convolutional attention mechanism and position encoding, focusing on extracting correlation information between the modalities. The adaptive weight allocation network performs weighted fusion of the obtained features and uses the fused features for subsequent target detection. Considering the poor quality of short-wave infrared data at night, a medium-wave-dominated degradation fusion strategy is designed to ensure that only the interactive features of visible light and medium-wave infrared are used in night mode, avoiding the impact of low-quality data on model performance.
[0033] For each five-layer feature map extracted by the backbone network, it is divided into two paths: visible light-shortwave infrared and visible light-medium wave infrared, and then enters two offset prediction modules. First, the visible light feature map and shortwave infrared signatures Splicing is performed in the channel dimension to obtain the joint feature F cat , then use 3*3 convolution, after BN layer and ReLU activation function, to F cat Perform preliminary integration and dimensionality reduction. Then, two standard residual blocks and bottleneck blocks are used to enhance the feature extraction capability. Then, two prediction heads are used: the offset prediction head uses a 3*3 convolution layer to predict the two-dimensional offset of the feature point in 9 directions, each offset includes a longitudinal and lateral offset; the attention prediction head uses a 3*3 convolution layer to output the attention weight of each direction to control the contribution of the offsets in 9 different directions. Finally, it is normalized by the tanh activation function, and the absolute value of the offset is controlled by the scaling factor. The shortwave infrared feature map is guided and sampled based on the offset to obtain the offset shortwave infrared feature map. Similarly, the visible light characteristic map and mid-wave infrared signatures The same steps are followed to obtain the offset mid-wave infrared characteristic map
[0034] The cross-attention module first calculates cross-attention based on the visible light feature map and the offset short-wave infrared feature map, and generates Query (Q), Key (K), and Value (V) matrices through convolutional layers with different kernel sizes, which increases the model's ability to extract multi-level correlations between features from each modality. The specific definition of cross-attention is as follows:
[0035]
[0036] Among them, i is the number of layers of feature maps, k is the size of the convolution kernel that generates the Q / K / V matrix, and d k represents the dimension of the channel, and Respectively represent the convolutional layers that generate the Query, Key, and Value matrices, and They are the visible light feature map and the offset shortwave infrared feature map. The dot product of the visible light query and the shortwave infrared key is calculated, and the attention weight matrix is generated through the softmax function to reflect the correlation between the visible light feature and the shortwave infrared feature. Finally, the attention weight matrix is used to perform weighted summation on the shortwave infrared value to obtain the fusion feature, which enhances the model's adaptability to complex environments.
[0037] The calculation of cross attention based on the visible light feature map and the offset medium-wave infrared feature map is similar to the short-wave processing. The specific definition of cross attention is as follows:
[0038]
[0039] in, and They are the visible light feature map and the offset medium-wave infrared feature map. The fusion features it generates reflect the correlation between the visible light features and the thermal radiation information, making up for the deficiency of visible light being ineffective at night.
[0040] The sizes of the convolution kernels of the two channels are 1*1 (k1), 5*5 (k2) and 7*7 (k3) respectively. After obtaining the fusion feature matrices of different levels, the matrices obtained by k2 and k3 are upsampled, the feature map size is unified, and the concatenation is performed along the channel dimension to obtain the visible light-shortwave infrared fusion features. and visible-mid-wave infrared fusion features Will and These five feature maps are input into the adaptive weight allocation network. First, the channels are spliced and mapped through 1*1 convolution to obtain the attention weight matrix, which represents the attention weight of each spatial position to the five features. It is normalized by the softmax function and is defined as follows:
[0041]
[0042] in, is the attention weight matrix calculated from the feature map of the i-th layer, and B is the batch size batchsize.
[0043] Finally and These five feature maps are multiplied by the corresponding attention weights and then summed to obtain the final fusion feature Input into the DINO detection head.
[0044] c) DINO detection framework: Based on the original transformer structure in the DINO detection head, the encoder encodes input features to generate high-level semantic features. A two-stage mechanism is adopted. The encoder generates preliminary candidate boxes, which serve as input to the decoder to further optimize the prediction results. The decoder uses self-attention and cross-attention mechanisms to generate target queries and predict target categories and bounding boxes. Finally, the category prediction head and bounding box prediction head predict the target category probability and bounding box coordinates. Using the optimized training scheme used in DN-DETR, noise is added to the target labels and bounding boxes during the training phase to generate noise samples, which are then used for training to improve the model's resistance to noise.
[0045] According to a second aspect of the present specification, a surface ship detection system based on trimodal attention fusion is provided, the system comprising:
[0046] Three-band visible-infrared image acquisition module: This module uses a three-band camera to simultaneously acquire visible light, shortwave infrared, and medium-wave infrared images. The three-band camera consists of a shortwave infrared camera, a medium-wave infrared camera, and a visible light camera aligned on the same optical axis, with the camera lenses located in the same plane.
[0047] Data augmentation module: Generates diverse training samples by adding noise, translating, scaling, and rotating the original image. Generates adversarial samples using a generative adversarial network, mixing adversarial samples with the original sample as training data, and adding an adversarial loss term to the loss function.
[0048] Surface ship detection module: The data enhancement module is used to obtain an enhanced data set, and a surface ship detection model is obtained by training on the data set; the three-band image is input into the trained surface ship detection model to obtain the surface ship detection results.
[0049] According to a third aspect of this specification, an electronic device is provided, comprising a memory and a processor, wherein the memory is coupled to the processor; wherein the memory is used to store program data, and the processor is used to execute the program data to implement the surface ship detection method based on trimodal attention fusion as described in the first aspect.
[0050] According to a fourth aspect of the present specification, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the method for detecting sea surface ships based on trimodal attention fusion as described in the first aspect is implemented.
[0051] According to a fifth aspect of the present specification, a computer program product is provided, comprising a computer program / instruction, which, when executed by a processor, implements the surface ship detection method based on trimodal attention fusion as described in the first aspect.
[0052] The beneficial effects of the present invention are as follows:
[0053] 1. All-weather multi-modal fusion detection capability: Through the coordinated acquisition and fusion of visible light, short-wave infrared, and medium-wave infrared tri-modal images, it can effectively cope with ship inspection tasks in complex sea conditions and all-weather environments.
[0054] 2. Efficient cross-modal feature fusion and dynamic adjustment: A dual-path deformable attention fusion module is designed to achieve feature interaction and fusion between modalities through a multi-scale convolutional attention mechanism and position encoding, significantly improving the adaptability to complex scenarios and enhancing detection accuracy and generalization performance.
[0055] 3. Optimized training and detection framework: The improved DINO detection framework is used to further improve model performance and has broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0057] Figure 1 Flowchart of a method for detecting sea vessels based on trimodal attention fusion provided by an embodiment of the present invention;
[0058] Figure 2 A block diagram of an implementation model for a surface ship detection provided by an embodiment of the present invention;
[0059] Figure 3 Schematic diagram of a dual-path deformable attention fusion module used in an embodiment of the present invention;
[0060] Figure 4 A schematic diagram of an offset prediction module used in an embodiment of the present invention;
[0061] Figure 5A schematic diagram of an implementation of a sea surface ship detection system based on trimodal attention fusion provided by an embodiment of the present invention;
[0062] Figure 6 The figure is a schematic structural diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION
[0063] In order to better understand the technical solution of the present application, the embodiments of the present application are described in detail below with reference to the accompanying drawings.
[0064] It should be clear that the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0065] The terms used in the embodiments of the present application are for the purpose of describing specific embodiments only and are not intended to limit the present application. The singular forms "a", "an", "the" and "the" used in the embodiments of the present application and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise.
[0066] like Figure 1 As shown, this embodiment provides a method for detecting sea surface ships based on trimodal attention fusion, the steps are as follows:
[0067] 1. Use a three-band visible-infrared camera to collect shortwave infrared images, visible light images, and medium-wave infrared images in all weather and multiple scenes. Align the time axis and image content of the collected video stream data, including transformations such as translation, rotation, and scaling to ensure target alignment. Then, perform annotation, preprocessing, and image enhancement to construct a three-band visible-infrared image fusion dataset.
[0068] 2. Build a surface ship detection model for trimodal target detection, and use the model pre-trained on public datasets such as the COCO dataset to train the model on the constructed dataset.
[0069] Reference for the specific implementation block diagram of the sea surface ship detection model Figure 2 , using the backbone network to extract trimodal feature maps, then inputting them into the dual-path deformable attention fusion module for feature interaction and fusion, and finally using the DINO detection head to output the ship target position and category information; specifically including the following steps:
[0070] 2.1 The backbone network is constructed using Swin-Large, a large-scale variant of the pre-trained Swin Transformer. This network efficiently extracts multi-scale features through a layered design and a windowed attention mechanism. It consists of four stages, each of which reduces resolution and increases the number of channels. The final output is a five-layer feature map consisting of the low-level features of the input image and the high-level features from the four stages. Feature extraction for trimodal images involves simultaneously inputting a set of three modal images into the backbone network. This means that each extracted feature map layer has three corresponding feature maps, which then enter the subsequent dual-path deformable attention fusion module for feature fusion.
[0071] 2.2 Construct a two-way deformable attention fusion module to fuse the features of the three modalities. The structure of this module is as follows: Figure 3 As shown. Two main convolution-based cross-attention modules are designed to fuse the features of the three modalities. Two offset prediction modules are designed to dynamically adjust the medium-wave infrared and short-wave infrared feature maps. An adaptive weight allocation network is designed to assign different weights to the feature maps of different modalities. The offset prediction module generates an adaptive offset sampling field through a lightweight convolutional network to model the weak alignment difference between visible light and two infrared images, thereby guiding the feature alignment between the modalities. The cross-attention module realizes position interaction and fusion between modalities through a multi-scale convolutional attention mechanism and position encoding, focusing on extracting the correlation information between the modalities. The adaptive weight allocation network performs weighted fusion on the obtained features and uses the fused features for subsequent target detection. At the same time, considering the poor quality of short-wave infrared data at night, a medium-wave-dominated degradation fusion strategy is designed to ensure that only the interactive features of visible light and medium-wave infrared are used in night mode to avoid the impact of low-quality data on model performance.
[0072] Each layer of feature maps extracted by the backbone network is divided into two paths: visible light-shortwave infrared and visible light-medium wave infrared, and then enters two offset prediction modules; Figure 4 As shown, the visible light feature map and the shortwave infrared feature map are first concatenated in the channel dimension to obtain a joint feature. A 3x3 convolution is then used, followed by a batch normalization layer and a ReLU activation function, to integrate and reduce the dimensionality of the joint feature. Feature extraction is then enhanced using two standard residual blocks and a bottleneck block. The offset prediction head uses a 3x3 convolution layer to predict the two-dimensional offsets of feature points in nine directions, and the attention prediction head uses a 3x3 convolution layer to output the attention weights for each direction. Finally, the tanh activation function is used for normalization, and a scaling factor is used to control the absolute value of the offset. The shortwave infrared feature map is then guided and sampled based on the offset to obtain the offset shortwave infrared feature map. Similarly, the visible light feature map and the medium-wave infrared feature map undergo the same steps to obtain the offset medium-wave infrared feature map.
[0073] The cross-attention module first obtains the convolution layers of three different convolution kernels used by the Query, Key and Value matrices, generates the attention weight matrix through the softmax function, and performs weighted summation to obtain the visible light-shortwave infrared fusion features and the visible light-medium wave infrared fusion features. The visible light feature map, the offset shortwave infrared feature map, the offset medium wave infrared feature map, the visible light-shortwave infrared fusion features and the visible light-medium wave infrared fusion features are input into the adaptive weight distribution network. First, the channels are spliced and mapped through 1*1 convolution to obtain the attention weight matrix, which is normalized by the softmax function. Finally, the visible light feature map, the offset shortwave infrared feature map, the offset medium wave infrared feature map, the visible light-shortwave infrared fusion features and the visible light-medium wave infrared fusion features are multiplied by the corresponding normalized attention weights and then summed to obtain the final fusion features, which are input into the DINO detection head.
[0074] 2.3DINO detection framework: Based on the original transformer structure in the DINO detection head, the encoder encodes the input features to generate high-level semantic features. A two-stage mechanism is adopted, that is, the encoder stage generates preliminary candidate boxes, which serve as the input of the decoder to further optimize the prediction results; the decoder uses self-attention and cross-attention mechanisms to generate target queries and predict the target category and bounding box. Finally, the category prediction head and bounding box prediction head predict the target category probability and bounding box coordinates.
[0075] The present invention also discloses a surface ship detection system based on trimodal attention fusion, such as Figure 5 As shown, the system includes the following modules:
[0076] Three-band visible-infrared image acquisition module: This module uses a three-band camera to simultaneously acquire visible light, shortwave infrared, and medium-wave infrared images. The three-band camera consists of a shortwave infrared camera, a medium-wave infrared camera, and a visible light camera aligned on the same optical axis, with the camera lenses located in the same plane. For module implementation, refer to step 1.
[0077] Data augmentation module: Generates diverse training samples by adding noise, translating, scaling, rotating, and other operations on the original image. Generates adversarial samples through a generative adversarial network, mixes adversarial samples with the original samples as training data, and adds an adversarial loss term to the loss function to optimize the total loss function.
[0078] Surface ship detection module: Use the data enhancement module to obtain an enhanced dataset, and use the improved DH-DINO algorithm to train on the dataset to obtain a surface ship detection model. Input the three-band image into the trained surface ship detection model to obtain the surface ship detection results. The implementation of this module can refer to step 2 above.
[0079] The present invention is aimed at the task of detecting ships on the sea surface under complex sea conditions. By collecting three-band images and designing an improved three-band target detection algorithm, it solves the problem that traditional detection methods are not effective under complex sea conditions and weather conditions, and solves the problem that cross-modal detection is difficult to fully integrate modal information, thus achieving ship detection with high practicality.
[0080] Accordingly, the present application also provides an electronic device, comprising: one or more processors; a memory for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the above-mentioned method for detecting sea surface ships based on trimodal attention fusion. Figure 6 As shown in FIG, a hardware structure diagram of any device with data processing capability in the method for detecting sea surface ships based on trimodal attention fusion provided by an embodiment of the present invention is shown, except Figure 6 In addition to the processor, memory, and network interface shown, any device with data processing capabilities in which the apparatus in the embodiment is located may also include other hardware, generally based on the actual functions of the device with data processing capabilities, which will not be described in detail.
[0081] Accordingly, the present application also provides a computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, implement the above-mentioned method for detecting sea surface ships based on trimodal attention fusion. The computer-readable storage medium can be an internal storage unit of any device with data processing capabilities described in any of the aforementioned embodiments, such as a hard disk or memory. The computer-readable storage medium can also be an external storage device, such as a plug-in hard disk, a smart media card (SMC), an SD card, a flash card, etc. equipped on the device. Furthermore, the computer-readable storage medium can also include both an internal storage unit and an external storage device of any device with data processing capabilities. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capabilities, and can also be used to temporarily store data that has been output or is to be output.
[0082] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the contents disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only.
[0083] It will be understood that the present application is not limited to the exact construction that has been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof.
[0084] The above description is only a preferred embodiment of the present invention. Although the present invention has been disclosed as a preferred embodiment, it is not intended to limit the present invention. Any person skilled in the art can use the above disclosed methods and technical contents to make many possible changes and modifications to the technical solution of the present invention without departing from the scope of the technical solution of the present invention, or modify it into an equivalent embodiment with equivalent changes. Therefore, any simple modification, equivalent change and modification made to the above embodiment based on the technical essence of the present invention without departing from the content of the technical solution of the present invention still falls within the scope of protection of the technical solution of the present invention.
Claims
1. A method for detecting sea vessels based on trimodal attention fusion, characterized in that: include: S1, three-band image acquisition: using continuously zoomed and registered visible, shortwave, and medium-wave cameras to simultaneously acquire visible light, shortwave infrared, and medium-wave infrared images; S2, three-band image coarse alignment: align the three video streams on the time axis according to the ship's movement trajectory. By adjusting the translation, rotation, and scaling parameters of each frame, the position of the target in the image is basically consistent. This completes the three-band image coarse alignment and exports the coarsely aligned video frames as an image sequence. S3, ship target detection based on coarsely aligned three-band images, includes the following sub-steps: S3.1, repeat S1 and S2 to obtain a set of three-band images, annotate the ships in the images, and create a three-band visible light-infrared image registration dataset; S3.2, perform data augmentation on the three-band visible-infrared image registration dataset, use adversarial training to generate adversarial samples, expand the training data, and obtain a three-band visible-infrared image fusion dataset; S3.
3. Train on a three-band visible-infrared image fusion dataset to obtain a surface ship detection model. The surface ship detection model includes a backbone network, a dual-path deformable attention fusion module, and a DINO detection head. The backbone network is used to extract trimodal feature maps, which are then input into the dual-path deformable attention fusion module for feature interaction and fusion. Finally, the DINO detection head is used to output ship target location and category information. The dual-path deformable attention fusion module includes an offset prediction module, a convolution-based cross-attention module, and an adaptive weight allocation network. The offset prediction module generates an adaptive offset sampling field through a lightweight convolutional network, modeling the weak alignment difference between visible light and two infrared images, thereby guiding feature alignment between the modalities. The cross attention module realizes position interaction and fusion between modalities through multi-scale convolutional attention mechanism and position encoding, and extracts correlation information between modalities; the adaptive weight distribution network performs weighted fusion on multimodal features to obtain fused features; S3.4, input the three-band image into the trained sea surface ship detection model to obtain the sea surface ship detection result.
2. The method for detecting sea vessels based on trimodal attention fusion according to claim 1, characterized in that: The three-band cameras used are a short-wave infrared camera, a medium-wave infrared camera, and a visible light camera with the same optical axis, and the camera lenses are in the same plane; the visible light camera uses a continuous zoom visible light camera with a focal length of 500mm, the short-wave infrared camera is an ultra-wide spectrum continuous zoom camera with a focal length of 420mm, and the medium-wave infrared camera is an ultra-low light night vision camera with a focal length of 100mm.
3. The method for detecting sea vessels based on trimodal attention fusion according to claim 1, characterized in that: The three-band visible-light-infrared image registration dataset is prepared specifically as follows: images of ships entering, leaving, and mooring at different distances and scenes at different ports are collected, where the ship types include cargo ships, passenger ships, fishing boats, and speedboats; three-band cameras with the same optical axis are mounted on drones, shores, and ships to collect all-weather images of ships at sea, and the images are collected separately under different weather conditions to construct an all-weather three-band visible-light-infrared ship dataset for complex sea conditions and rich scenes; and then alignment and registration operations are performed to register shortwave infrared images, mediumwave infrared images, and visible light images to obtain a three-band visible-light-infrared image registration dataset.
4. The method for detecting sea vessels based on trimodal attention fusion according to claim 1, characterized in that: The backbone network uses a variant of the pre-trained Swin Transformer, Swin-Large, to extract multi-scale features through a layered design and a windowed attention mechanism. It consists of four stages, reducing the resolution and increasing the number of channels at each stage, and finally outputting a five-layer feature map, namely the low-level features of the input image and the high-level features of the four stages; A set of three-modal images are simultaneously input into the backbone network, and three feature maps corresponding to each layer of feature maps are extracted, which are then input into the dual-path deformable attention fusion module for feature fusion.
5. The method for detecting sea vessels based on trimodal attention fusion according to claim 1, characterized in that: Each layer of feature maps extracted by the backbone network enters two offset prediction modules in two ways: visible light-shortwave infrared and visible light-medium wave infrared. The offset prediction module first splices the visible light feature map and a certain infrared feature map in the channel dimension to obtain a joint feature. Then, a 3*3 convolution is used to integrate and reduce the dimension of the joint feature through a BN layer and a ReLU activation function. The feature extraction capability is then enhanced through two standard residual blocks and a bottleneck block. The offset prediction head uses a 3*3 convolution layer to predict the two-dimensional offset of the feature point in 9 directions, and the attention prediction head uses a 3*3 convolution layer to output the attention weight of each direction. Finally, the tanh activation function is used for normalization, the scaling factor is used to control the absolute value of the offset, and the infrared feature map is guided and sampled based on the offset to obtain the offset infrared feature map.
6. The method for detecting sea vessels based on trimodal attention fusion according to claim 1, characterized in that: The implementation of the cross-attention module is specifically as follows: obtaining the convolution layers of three different convolution kernels used by the Query, Key and Value matrices, generating the attention weight matrix through the softmax function, weighted summing to obtain the visible light-shortwave infrared fusion features and the visible light-medium wave infrared fusion features, and inputting the visible light feature map, the offset shortwave infrared feature map, the offset medium wave infrared feature map, the visible light-shortwave infrared fusion features and the visible light-medium wave infrared fusion features into the adaptive weight distribution network. First, the channels are spliced and mapped through 1*1 convolution to obtain the attention weight matrix, which is normalized by the softmax function. Finally, the visible light feature map, the offset shortwave infrared feature map, the offset medium wave infrared feature map, the visible light-shortwave infrared fusion features and the visible light-medium wave infrared fusion features are multiplied by the corresponding normalized attention weights and then summed to obtain the final fusion features, which are input into the DINO detection head.
7. A system for implementing the method according to any one of claims 1 to 6, characterized in that: The system includes: visible light-infrared image acquisition module, data enhancement module and sea surface ship detection module; The three-band visible light-infrared image acquisition module uses a three-band camera to simultaneously acquire visible light images, short-wave infrared images, and medium-wave infrared images. The three-band camera is composed of a short-wave infrared camera, a medium-wave infrared camera, and a visible light camera that are aligned on the same optical axis, and the camera lenses are in the same plane; The data enhancement module generates diverse training samples by adding noise, translating, scaling, and rotating the original image. The adversarial samples are generated through a generative adversarial network, and the adversarial samples and original samples are mixed as training data. An adversarial loss term is added to the loss function. The sea surface ship detection module uses the data enhancement module to obtain an enhanced data set, and performs training on the data set to obtain a sea surface ship detection model; the three-band image is input into the trained sea surface ship detection model to obtain a sea surface ship detection result.
8. An electronic device comprising a memory and a processor, characterized in that: The memory is coupled to the processor; wherein, the memory is used to store program data, and the processor is used to execute the program data to implement the surface ship detection method based on trimodal attention fusion as described in any one of claims 1-6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method for detecting sea surface ships based on trimodal attention fusion as described in any one of claims 1 to 6 is implemented.
10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, the method for detecting sea surface ships based on trimodal attention fusion as described in any one of claims 1 to 6 is implemented.
Citation Information
Cited By
Channel buoy detection method based on fusion of multi-mode pulse neural network and visual Transform
CN121616952A