Multi-modal image ship target detection method and system
By constructing a lightweight feature extraction network, a cross-scale attention module and a multi-scale feature fusion module, the detection accuracy and efficiency problems in multimodal image ship target detection are solved, and efficient ship target recognition is achieved.
Patent Information
- Application Number
- CN202510988394.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-07-14
- Filing Date
- 2025-07-17
- Publication Date
- 2025-09-05
AI Technical Summary
Existing multimodal image ship target detection methods have problems such as the contradiction between classification and detection tasks, difficulty in cross-domain migration parameter adaptation, and fixed limitations of network architecture, resulting in low detection accuracy and efficiency.
A ship target detection network is constructed by adopting a lightweight feature extraction network, a cross-scale attention module, a dynamic position encoding Transformer and a multi-scale feature fusion module, combined with depthwise separable convolution, batch normalization, activation function and channel shuffling, and the network performance is optimized by an improved loss function.
It improves the accuracy and efficiency of ship target detection, enhances position perception and feature expression capabilities, and is suitable for low-computation scenarios.
Smart Images

Figure CN120599231A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of target detection, and in particular to a multimodal image ship target detection method and system. Background Art
[0002] Ship target detection technology has important applications in maritime traffic monitoring, marine resource management, and search and rescue missions. Synthetic aperture radar (SAR), with its all-weather imaging capabilities, has become a key technology for ship target detection in adverse weather conditions. However, SAR images suffer from complex scattering mechanisms, significant speckle noise, and the significant impact of observation angle on target detection, significantly different from natural optical visible light images. While visible light images offer rich texture detail and clear color information, they are dependent on lighting conditions, resulting in poor image quality at night and on rainy days. Multimodal fusion of SAR and visible light images can complement each other's strengths. SAR provides all-weather monitoring capabilities, compensating for the limitations of visible light due to weather and lighting conditions. The rich detail of visible light images enables more accurate identification of vessel type, appearance, and other information. Therefore, designing ship target detection methods based on a multimodal approach combining SAR and visible light images is crucial for improving the accuracy of maritime ship target detection.
[0003] Most current approaches focus on multimodal image detection, relying primarily on convolutional neural networks (CNNs) pre-trained on ImageNet as the detector backbone. However, this approach has inherent limitations. First, the objectives of classification and detection are inherently conflicting: classification strives for translation invariance, while detection requires precise object location. Consequently, the parameter optimization direction of ImageNet-pretrained networks deviates from the requirements of the detection task. Second, the design of ImageNet-pretrained networks favors large receptive field feature extraction, which inhibits the localization of small objects and the capture of detailed features. More critically, SAR images differ significantly from natural images in imaging mechanisms, noise distribution, and object representation, making it difficult for ResNet-based cross-domain transfer parameters to adapt to the target domain characteristics. Furthermore, the rigid structure of existing frameworks limits the scope for network optimization tailored to SAR image characteristics (such as speckle reduction and multi-scale object representation) and object detection tasks, preventing researchers from fully exploiting domain prior knowledge by adjusting the network architecture.
[0004] Therefore, there is an urgent need to provide a new multimodal image ship target detection method to improve the accuracy and efficiency of ship target detection. Summary of the Invention
[0005] The purpose of this application is to provide a multimodal image ship detection method and system that can improve the accuracy and efficiency of ship target detection.
[0006] To achieve the above objectives, this application provides the following solutions:
[0007] In a first aspect, the present application provides a multimodal image ship target detection method, the multimodal image ship target detection method comprising:
[0008] Acquire multimodal images of ship targets; multimodal images include SAR images and visible light images;
[0009] Preprocess the multimodal images and construct a ship target dataset. The preprocessing process includes image resizing and pixel value normalization. The ship target dataset includes a training set and a test set.
[0010] A ship target detection network is constructed and trained using a training set. The ship target detection network includes a lightweight feature extraction module, a cross-scale attention module, a dynamic position encoding transformer, and a multi-scale feature fusion module. The lightweight feature extraction module includes multiple lightweight feature extraction networks and is used to determine the first-scale feature map, the second-scale feature map, and the third-scale feature map based on the multimodal image. The lightweight feature extraction network includes depthwise separable convolution, batch normalization, activation function, pointwise convolution, and channel shuffling. The scales of the first-scale feature map and the second-scale feature map are both larger than the third-scale feature map. The cross-scale attention module includes three parallel dilated convolution branches, a depthwise separable convolution branch, and a grouped convolution branch, and is used to perform attention enhancement on the third-scale feature map to determine an attention-enhanced third-scale feature map. The dynamic position encoding transformer is used to perform position and geometry perception enhancement on the first-scale feature map, the second-scale feature map, and the attention-enhanced third-scale feature map to determine the corresponding position-sensitive features. The multi-scale feature fusion module is used to fuse the position-sensitive features to determine a multi-scale feature pyramid with three levels.
[0011] The trained ship target detection network is used to perform ship target detection on the multimodal images in the test set.
[0012] Optionally, the processing process of the cross-scale attention module is:
[0013] Based on the dilated convolution branch, the depth-wise separable convolution branch, and the grouped convolution branch, the first feature, the second feature, and the third feature corresponding to the third-scale feature map are determined respectively;
[0014] Performing weighted fusion on the first feature, the second feature, and the third feature based on a spatial attention weight;
[0015] According to the weighted fusion results, the attention-enhanced third-scale feature map is determined.
[0016] Optionally, the weighted fusion of the first feature, the second feature, and the third feature based on the spatial attention weight specifically includes:
[0017] Using the formula Determine the spatial attention weight α; where softmax() is the activation function, W2 is the second weight matrix for weight distribution, ReLU is the activation function, W1 is the first weight matrix for feature compression, GAP is the global average pooling, F1 is the first feature, F2 is the second feature, and F3 is the third feature. is element-wise addition;
[0018] Using the formula Perform weighted fusion on the first feature, the second feature and the third feature; wherein, F out is the third-scale feature map of attention enhancement, α i is the spatial attention weight, i is the feature number, i=1,2,3.
[0019] Optionally, the processing process of the dynamic position encoding Transformer is:
[0020] Mapping the first-scale feature map, the second-scale feature map, and the attention-enhanced third-scale feature map to corresponding position codes respectively;
[0021] According to the position code, the corresponding feature map containing the position information is determined respectively;
[0022] According to the feature map containing position information, the corresponding position-sensitive feature is determined.
[0023] Optionally, the processing process of the multi-scale feature fusion module is:
[0024] Using the formula and Determine a multi-scale feature pyramid with 3 levels;
[0025] Among them, P1" is the first level multi-scale feature pyramid, Conv() is the convolution operation, Conv 1×1 () is a 1×1 convolution operation, P 2' is the second up-sampled scale feature map, F1' is the first position sensitive feature, P 2” is the second-level multi-scale feature pyramid, P 3” is the third level multi-scale feature pyramid, F3' is the third position sensitive feature, ↑ is upsampling, ↓ is downsampling, It is element-wise addition.
[0026] Optionally, the loss function of the ship target detection network is determined as follows:
[0027] Using the formula Determine the improved classification loss function L cls ; Where N is the batch size in each training, t is the sample index, α t is the dynamic weighting coefficient, according to the formula OK, p t The classification branch of the ship target detection network predicts the probability that the t-th input sample belongs to the true category, γ is the adjustment factor, and the formula is: OK, II(p G =c) is the indicator function, when the predicted label p G When the true label is c, the indicator function is 1, otherwise it is 0, N pos is the number of positive samples of all categories in the current batch, ε is the smoothing term, N total is the number of samples in the current batch, and C is the number of categories;
[0028] Using the formula L 1oc =0.8*(1-GIoU)+0.2*L1 to determine the positioning loss function L 1oc ; Among them, GIoU is the intersection-over-union function, L1 is the regularization function, and the formula L1=|x pred -x gt |+|y pred -y gt | OK, (x gt ,y gt ) is the center point coordinate of the real frame, (x pred ,y pred ) is the center point coordinate of the prediction box;
[0029] Using the formula Determine the consistency regularization loss function L consist ; Among them, f() is the feature of the third level feature pyramid, is a multimodal image obtained by preprocessing the multimodal image. is the SAR image; is a visible light image;
[0030] Using the formula L total =1.0*L cls +0.7*L loc +0.3*L consist Determine the comprehensive loss value L of the loss function of the ship target detection network total .
[0031] In a second aspect, the present application provides a multimodal image ship target detection system, the multimodal image ship target detection system comprising:
[0032] A multimodal image acquisition module is used to acquire multimodal images of ship targets; multimodal images include SAR images and visible light images;
[0033] The ship target dataset construction module is used to preprocess multimodal images and construct a ship target dataset. The preprocessing process includes image resizing and pixel value normalization. The ship target dataset includes a training set and a test set.
[0034] The ship target detection network construction module is used to construct the ship target detection network and train the ship target detection network using the training set; the ship target detection network includes: a lightweight feature extraction module, a cross-scale attention module, a dynamic position encoding Transformer and a multi-scale feature fusion module; the lightweight feature extraction module includes multiple lightweight feature extraction networks and is used to determine the first scale feature map, the second scale feature map and the third scale feature map according to the multimodal image; the lightweight feature extraction network includes: depthwise separable convolution, batch normalization, activation function, point-by-point convolution and channel shuffling; the first scale feature map and The scale of the second-scale feature map is larger than that of the third-scale feature map; the cross-scale attention module includes: 3 parallel void convolution branches, depth-separable convolution branches and grouped convolution branches, and is used to enhance the attention of the third-scale feature map and determine the attention-enhanced third-scale feature map; the dynamic position encoding Transformer is used to perform position and geometric perception enhancement on the first-scale feature map, the second-scale feature map and the attention-enhanced third-scale feature map respectively, and determine the corresponding position-sensitive features; the multi-scale feature fusion module is used to fuse the position-sensitive features and determine a multi-scale feature pyramid with 3 levels;
[0035] The ship target detection module is used to use the trained ship target detection network to perform ship target detection on the multimodal images in the test set.
[0036] According to the specific embodiments provided in this application, this application has the following technical effects:
[0037] The present application provides a multimodal image ship target detection method and system. Through the depth-wise separable convolution in the lightweight feature extraction network, it can reduce the computational complexity and accelerate the inference speed while ensuring spatial details; through the cross-scale attention module, the attention of the third-scale feature map is enhanced, which can improve the expression ability of the third-scale feature map and the overall performance of the ship target detection network; through the dynamic position encoding Transformer, the geometric perception enhancement of the feature map can be performed, which can enhance the position perception ability and detection accuracy of the ship target detection network; through the multi-scale feature fusion module, the position sensitive features are fused, which further improves the detection performance of the ship target detection network. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0039] Figure 1 Schematic diagram of a flow chart of a multimodal image ship target detection method in one embodiment of the present application;
[0040] Figure 2 Schematic diagram of the process of processing multimodal images by a ship target detection network in one embodiment of the present application;
[0041] Figure 3 This is a schematic diagram of the structure of a lightweight feature extraction network in one embodiment of the present application;
[0042] Figure 4 This is a schematic diagram of the structure of the cross-scale attention module in one embodiment of the present application;
[0043] Figure 5 Schematic diagram of the structure of the dynamic position encoding Transformer in one embodiment of the present application. DETAILED DESCRIPTION
[0044] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0045] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0046] In an exemplary embodiment, Figure 1 As shown, a multimodal image ship target detection method is provided, and the multimodal image ship target detection method includes the following S1 to S5.
[0047] S1: Acquire a multimodal image of the ship target, which includes a SAR image and a visible light image.
[0048] S2: Preprocess the multimodal images and construct the ship target dataset.
[0049] When preprocessing the multimodal image, the size of the multimodal image is first adjusted, that is, the multimodal image of the ship target is adjusted to 300×300×3. Then the pixel value normalization processing is performed on the adjusted multimodal image. The pixel value normalization processing is specifically as follows: the pixel value of each channel of the adjusted multimodal image is standardized. Specifically, the pixel value of the original range of 0-255 is first divided by 127.5 and then subtracted by 1, so that the pixel value range is mapped to the interval [-1,1] to obtain the normalized pixel value; controllable Gaussian noise is superimposed on the basis of the standardized pixel value to enhance the anti-interference ability of the ship target detection network. The noise intensity is determined by a n Coefficient adjustment, the noise source obeys the mean of 0 and the variance of σ 2 The expression of the pixel value superimposed with controllable Gaussian noise is: noise = normalized pixel value + a n × noise source, the preprocessing process is recorded as Aug.
[0050] The preprocessing process simulates the noise interference in real scenes. By adding noise to the image, the ship target detection network is forced to learn noise-invariant features, reducing the ship target detection network's dependence on the training set.
[0051] Considering that visible light images are easily obscured by clouds and fog, resulting in missing ship target information, we perform oblique box annotation on the preprocessed SAR images to obtain the true box of the ship target in the SAR image, which corresponds to the output of the entire detection task. We also annotate the type of ship target to obtain the true label. In this application, we assume that there are C types of labels. We construct a ship target dataset, which includes the preprocessed multimodal images of ship targets, the true box, and the true label. The ship target dataset is divided into a training set and a test set.
[0052] S3: Build a ship target detection network and train it using the training set.
[0053] The ship target detection network includes: lightweight feature extraction module, cross-scale attention module (Cross-Scale Attention Module, CSAM), dynamic position encoding transformer (Dynamic Position Embedding Transformer, DPE-Tfmer), multi-scale feature fusion module, detection head module and loss function. The ship target detection network processes multimodal images as follows: Figure 2 shown.
[0054] S3 specifically includes:
[0055] S31: Establish a lightweight feature extraction module.
[0056] The lightweight feature extraction module includes multiple lightweight feature extraction networks and is used to determine a first-scale feature map, a second-scale feature map, and a third-scale feature map according to the multimodal image.
[0057] In this application, the lightweight feature extraction network (LiteBlock) includes: depthwise separable convolution (DepthwiseConv), batch normalization (BatchNorm), activation function (HardSwish), pointwise convolution (PointwiseConv) and channel shuffle (ChannelShuffle), a total of five layers, such as Figure 3 As shown in the figure, the depthwise separable convolution is composed of 3×3 spatial convolutions, which are used to independently process each input channel, significantly reducing the number of parameters while keeping the number of channels unchanged. Pointwise convolution achieves channel dimension transformation through 1×1 convolution to control the expressiveness of features. The number of channels is doubled during this process. Channel shuffling groups and rearranges channels to enhance cross-channel information interaction. This application uses a combined structure of depthwise separable convolution and channel shuffling to reduce the amount of computation while ensuring spatial details, achieving efficient multi-scale feature extraction.
[0058] The lightweight feature extraction network is further used to construct units with different functions in the lightweight feature extraction module. The lightweight feature extraction module specifically includes a high-resolution feature extraction unit, a first multi-scale feature generation unit, and a second multi-scale feature generation unit. This application adopts progressive downsampling, gradually abstracting semantics through a three-level cascade structure to generate and retain multi-scale feature information.
[0059] In the high-resolution feature extraction unit, a lightweight feature extraction network is set up and targeted modifications are made; the multimodal image (300×300×3) is spliced and fused after a convolution layer to obtain a 300×300×16 feature map, which is then input into the lightweight feature extraction network. After a 3×3 depthwise separable convolution (stride=2, output 32 channels), it is batch normalized and activated by HardSwish, and then expanded to 64 channels through a 1×1 point-by-point convolution. Finally, a channel shuffling operation is performed and the number of groups is set to 4 to break the fixed connection mode between channels. The final output is a feature map with a scale of 150×150×64, i.e., the first-scale feature map.
[0060] Although the feature map (300×300×16) obtained from the multimodal image has rich spatial details, the number of channels is small (16 channels). At the beginning of the calculation, the resolution needs to be greatly reduced (300→150), and a stronger nonlinear transformation is required to expand from 16 channels to 64 channels. Therefore, the step size is set to 2.
[0061] In the first multi-scale feature generation unit, three lightweight feature extraction networks are set, the convolution step is set to 1, the output channel is set to 128, the first-scale feature map is downsampled to 75×75×128, and the output scale is 75×75×128, that is, the second-scale feature map.
[0062] In the second multi-scale feature generation unit, three lightweight feature extraction networks are set, the convolution step is set to 2, the output channel is set to 256, the second-scale feature map is downsampled to 38×38×256, and the output scale is 38×38×256 feature map, that is, the third-scale feature map.
[0063] The lightweight feature extraction network reduces computational complexity while maintaining spatial details and speeds up inference by setting up depthwise separable convolutions, making it suitable for low-computation scenarios such as edge devices. In addition, the modular design of the lightweight feature extraction network has dual advantages: on the one hand, sharing the same basic units at each stage can reduce the complexity of engineering implementation and improve structural consistency; on the other hand, the network depth can be flexibly adjusted by simply increasing the number of lightweight feature extraction networks, achieving convenient scalability. In addition, the lightweight feature extraction network enhances cross-channel information interaction through channel shuffling, breaking the fixed connection pattern between channels, making the extracted feature maps richer and more diverse, and helping the ship target detection network learn more effective feature representations. Finally, through progressive downsampling, a cascade structure is constructed to gradually abstract semantics, generate and retain multi-scale feature information, and provide rich semantic information for subsequent feature fusion and ship target detection.
[0064] S32: Building a cross-scale attention module.
[0065] like Figure 4 As shown in the figure, the cross-scale attention module includes: 3 parallel dilated convolution branches (DilConv), depthwise separable convolution branch (DwConv) and grouped convolution branch (GrConv), and is used to enhance the attention of the third-scale feature map and determine the third-scale feature map for attention enhancement.
[0066] In a specific embodiment, the third-scale feature map (38×38×256) contains more semantic information but has lower sensitivity to details. Therefore, a cross-scale attention module is designed to enhance cross-scale context perception and achieve adaptive fusion of multi-receptive field features.
[0067] The input of the cross-scale attention module is the third-scale feature map, and the output is the attention-enhanced third-scale feature map of the same size. To address the limitations of traditional single-scale convolution and the problem of cross-scale feature alignment, this application proposes a multi-branch heterogeneous convolution structure and adopts a multi-branch dilated convolution combined with adaptive weight fusion.
[0068] S32 specifically includes:
[0069] S321: Based on the dilated convolution branch, the depthwise separable convolution branch, and the grouped convolution branch, respectively determine the first feature, the second feature, and the third feature corresponding to the third-scale feature map.
[0070] Specifically, the third-scale feature map is simultaneously fed into three parallel branches, namely, the 3×3 dilated convolution branch DilConv(3,1,1), the 5×5 depth-wise separable convolution DwConv(5,1,1), and the 7×7 grouped convolution branch GrConv(7,1,1); the dilated convolution branch (dilation=2) captures a wide range of context, the depth-wise separable convolution branch extracts medium-scale features, and the grouped convolution branch obtains local details. The first feature F1, second feature F2, and third feature F3 corresponding to the third-scale feature map are determined through the above three branches respectively.
[0071] S322: Perform weighted fusion on the first feature, the second feature, and the third feature based on the spatial attention weight.
[0072] Specifically, the first, second and third features output by the three parallel branches are subjected to global average pooling (GAP), and then normalized by a two-layer fully connected network and Softmax to generate spatial attention weights, and finally the multi-scale features are weightedly fused.
[0073] The spatial attention weight α can be dynamically generated using the following formula:
[0074]
[0075] Among them, softmax() is the activation function, W2 is the second weight matrix, W2∈R 3h×3 , used to implement weight distribution, ReLU is the activation function, W1 is the first weight matrix, W1∈R h×3h , used to achieve feature compression, GAP is global average pooling, F1 is the first feature, F2 is the second feature, F3 is the third feature, is element-by-element addition, and h is the number of channels of the third-scale feature map.
[0076] S323: Determine an attention-enhanced third-scale feature map based on the weighted fusion result.
[0077] The weighting is performed according to the following formula:
[0078]
[0079] Among them, F out is the third feature map of attention enhancement, α i is the spatial attention weight corresponding to the i-th feature, i is the feature number, i = 1, 2, 3.
[0080] The cross-scale attention module constructs a multi-branch heterogeneous convolution consisting of three branches. Through the dilated convolution branch, the depthwise separable convolution branch, and the grouped convolution branch, it captures large-scale context, medium-scale features, and local details, respectively, addressing the limitations of traditional single-scale convolution. Adaptive weight fusion is then performed, generating spatial attention weights through global average pooling and softmax activation function normalization. This weighted fusion of multi-scale features enables the ship target detection network to adaptively focus on features at different scales, improving both the feature representation capability and the detection performance of the ship target detection network.
[0081] S33: Build a dynamic position encoding Transformer.
[0082] The dynamic position encoding transformer is used to perform position and geometric perception enhancement on the first scale feature map, the second scale feature map and the attention-enhanced third scale feature map, and determine the corresponding position-sensitive features. Its structure is as follows Figure 5 shown.
[0083] The dynamic position encoding transformer takes as input a first-scale feature map, a second-scale feature map, and an attention-enhanced third-scale feature map, and outputs position-sensitive features of the same scale. To overcome the problem of missing position information in traditional transformers for object detection, this application proposes a learnable dynamic position encoding transformer.
[0084] S33 specifically includes:
[0085] S331: Map the first-scale feature map, the second-scale feature map, and the attention-enhanced third-scale feature map to corresponding position codes respectively.
[0086] Specifically, first map the position (x, y) of the feature map to the corresponding position code PE (x,y) , the mapping formula is:
[0087]
[0088] Where x is the horizontal coordinate (column index) of a position in the feature map, y is the vertical coordinate (row index) of a position in the feature map, parameter F is the number of frequency components in the Transformer, and f is the sequence number of the frequency component, f = 0, ..., F-1. f is the corresponding first learnable parameter, μ f is the corresponding second learnable parameter, both of which participate in the position encoding calculation and are optimized in the loss propagation. Used to adjust the frequency component. In this application, F is set to 64 and the position encoding dimension D is set to 128, that is, the position information is dynamically encoded in a 128-dimensional space through a combination of 64 sine / cosine waves of different frequencies.
[0089] S332: Determine corresponding feature maps containing position information according to the position codes.
[0090] The first scale feature map, the second scale feature map and the attention-enhanced third scale feature map are flattened respectively, and then added to the corresponding position code to obtain the corresponding feature map containing position information
[0091] S333: Determine corresponding position-sensitive features based on the feature map containing position information.
[0092] The feature map containing position information is input into the improved multi-head attention layer, and the calculation formula is as follows:
[0093]
[0094] Among them, IoU prioris the IoU statistic between the preset anchor box and the real box, Attention is the output of multi-head attention calculation, Q is the query matrix, and the formula is: Determine, K is the key matrix, by the formula Determine, V is the value matrix, by the formula OK, d k is the scaling parameter for scaling dot product attention, W Q 、W K 、W V are parameters optimized by back propagation during model training, where W Q is the query matrix weight, W K is the bond matrix weight, W V is the value matrix weight, is a feature map containing position information, where b is the number of the feature map containing position information, b=1, 2, 3.
[0095] The feature map containing position information passes through a two-layer dynamic position encoding Transformer to obtain the corresponding position-sensitive features, namely the first position-sensitive feature F1' with a scale of 150×150×64, the second position-sensitive feature F2' with a scale of 75×75×128, and the third position-sensitive feature F3' with a scale of 38×38×256.
[0096] The dynamic position encoding Transformer encodes the position information of the first scale feature map, the second scale feature map and the attention-enhanced third scale feature map into position-sensitive features, overcoming the problem of missing position information in target detection by the traditional Transformer and enhancing the position perception ability of the ship target detection network. At the same time, in the improved multi-head attention layer, IoU is introduced in the standard attention calculation. prior The bias term uses the IoU statistics of the preset anchor boxes to guide the attention distribution, so that the ship target detection network can better focus on the features related to the true label, thereby improving the accuracy and robustness of the ship target detection network.
[0097] S34: Establish a multi-scale feature fusion module.
[0098] The multi-scale feature fusion module is used to fuse position-sensitive features to determine a three-level multi-scale feature pyramid. The input of the multi-scale feature fusion module is the first position-sensitive feature, the second position-sensitive feature, and the third position-sensitive feature, and the output is the fused multi-scale feature pyramid.
[0099] This application proposes an upsampling-downsampling bidirectional interactive architecture: in the upsampling path, high-level features are amplified by bilinear interpolation and then added element-by-element to adjacent low-level features; in the downsampling path, low-level features are downsampled by 3×3 convolution (step 2) and then horizontally fused with high-level features. Specifically, the third position-sensitive feature is first upsampled to 75×75 resolution and fused with the second position-sensitive feature to generate the second upsampled scale feature map P 2' The second upsampled scale feature map is further upsampled to 150×150 and horizontally fused with the first position sensitive feature to obtain the first-level multi-scale feature pyramid P1". The second upsampled scale feature map is fused with the first-level multi-scale feature pyramid downsampled to 75×75 resolution to obtain the second-level multi-scale feature pyramid P2". The third position sensitive feature is fused with the second-level multi-scale feature pyramid downsampled to 38×38 to obtain the third-level multi-scale feature pyramid P3".
[0100] The fusion process includes:
[0101] P 3' =Conv 1×1 (F3')↑.
[0102]
[0103] The bidirectional feature pyramid in the multi-scale feature fusion module uses an upsampling-downsampling bidirectional interactive architecture. This structure optimizes cross-level feature fusion, enabling the ship target detection network to fully utilize feature information from different levels, improving the richness of the feature map and the detection performance of the ship target detection network. Finally, horizontal fusion is performed through element-by-element addition, preserving the detailed information of features at different levels and further enhancing the expressiveness of features.
[0104] S35: Establish a detection head module.
[0105] The detection head module inputs a multi-scale feature pyramid with three different levels. For each level of features, a regression branch is constructed, using a 3×3 convolutional layer connected to a 1×1 convolutional layer with 5 output channels. The predicted box and its confidence are determined. Simultaneously, a classification branch is constructed, using a 3×3 convolutional layer connected to a 1×1 convolutional layer. The feature map output by the 1×1 convolutional layer is expanded into a vector. Finally, a fully connected layer containing C nodes is connected to predict C types of ship targets. A softmax activation function is then applied to predict the target presence probability and determine the predicted label.
[0106] S36: Establish loss function.
[0107] The loss function is used to achieve balanced optimization of classification and positioning. The input of the loss function is the predicted label and the true label, and the output is the comprehensive loss value of the ship target detection network.
[0108] S36 specifically includes:
[0109] S361: Determine the classification loss function.
[0110] To address the issues of category imbalance and positioning accuracy, this application proposes a dynamic weighted loss architecture. The classification loss uses an improved FocalLoss, and its basic form is:
[0111]
[0112] Among them, L cls is the classification loss function, which measures the difference between the predicted label and the true label of the ship target detection network, N is the batch size in each training, t is the sample index, t=1,…,N, α t is the dynamic weighting coefficient, p t The classification branch of the ship target detection network predicts the probability that the t-th input sample belongs to the true category, and γ is the adjustment factor.
[0113] Considering the different numbers of samples of different categories in the training set, this class imbalance problem will affect the accuracy of classification. Therefore, a dynamic weighting coefficient is designed to be dynamically adjusted according to the frequency of occurrence of the category. It is used to balance the impact of samples of different categories on the performance of the algorithm in the class imbalance problem, thereby ensuring that the classification accuracy of samples of different categories with different numbers is as consistent as possible. The calculation formula is as follows:
[0114]
[0115] Among them, II(p G =c) is the indicator function, i.e. the predicted label p G Is it the real label c? If so, the indicator function is 1, otherwise 0, N pos is the number of positive samples of all categories in the current batch, that is, the number of predicted labels whose intersection-over-union ratio between the predicted box and the true box exceeds the set threshold (0.5), and ε is a smoothing term, which is 0.01 here.
[0116] The adjustment factor controls the contribution of difficult and easy samples to the loss, alleviating the class imbalance problem. In this application, the adjustment factor changes adaptively with the proportion of positive samples as follows:
[0117]
[0118] Among them, N total is the number of samples in the current batch, corresponding to the number of real boxes.
[0119] S362: Determine the positioning loss function.
[0120] The positioning loss function combines the GIoU metric and L1 regularization to balance the position sensitivity and scale invariance of the predicted box relative to the real box. The calculation formula is as follows:
[0121] L 1oc =0.8*(1-GIoU)+0.2*L1.
[0122] Among them, L 1oc As the positioning loss function, GIoU (Generalized Intersection over Union) is used to evaluate the improvement index of the overlap between the predicted box and the real box. It is an existing technology. In this application, L1 regularization uses the center point coordinate position difference, (x gt ,y gt ) is the center point coordinate of the real frame, (x pred ,y pred ) is the center point coordinate of the prediction box, then the calculation formula of the L1 loss function is:
[0123] L1=|x pred -x gt |+|y pred -y gt |.
[0124] S363: Determine the consistency regularization loss.
[0125] The consistency regularization loss forces the feature distance between the original image and the enhanced image to be minimized, and the formula is as follows:
[0126]
[0127] Among them, L consist is the consistency regularization loss function, f() is the feature of the third level feature pyramid, is a multimodal image obtained by preprocessing the multimodal image. is the SAR image; For visible light images.
[0128] S364: Determine the comprehensive loss value.
[0129] The weighted sum of the three constitutes the comprehensive loss value, and the calculation formula is as follows:
[0130] L total =1.0*L cls +0.7*L loc +0.3*L consist .
[0131] The loss function improves the classification loss function (FocalLoss). By dynamically adjusting the dynamic weighting coefficient and adjustment factor, the loss weight is adaptively adjusted according to the frequency of class occurrence and the proportion of positive samples. This solves the problems of class imbalance and localization accuracy, allowing the ship target detection network to better focus on difficult-to-classify samples, improving classification accuracy and the robustness of the ship target detection network. Secondly, the localization loss function combines the GIoU metric with L1 regularization to balance the position sensitivity and scale invariance of the predicted box relative to the ground-truth box, enabling the ship target detection network to more accurately predict the position and size of the target detection box, thereby improving localization accuracy. Finally, the consistency regularization loss forces the feature distance between the original image and the enhanced image to be minimized, further enhancing the robustness and generalization ability of the ship target detection network, allowing the ship target detection network to maintain stable performance under different data augmentation conditions.
[0132] The Adam optimizer was used, with an initial learning rate of 0.0002. Cosine annealing was used to gradually reduce the learning rate to accelerate convergence and prevent overfitting. The batch size was set to 64, meaning 64 images were input for each training run. The training epochs were set to 500, with the parameters of the ship detection network updated once per iteration to obtain the final trained ship detection network.
[0133] S4: Use the trained ship target detection network to perform ship target detection on the multimodal images in the test set.
[0134] The multimodal images to be tested in the test set are input into the trained ship target detection network to obtain the target detection frame of the multimodal images to be tested in the test set.
[0135] Based on the same inventive concept, embodiments of the present application also provide a multimodal image ship target detection system for implementing the aforementioned method. The solution provided by this system is similar to the solution described in the aforementioned method. Therefore, the specific limitations of one or more embodiments of the multimodal image ship target detection system provided below can be found in the above-mentioned limitations on multimodal image ship target detection and will not be further elaborated here.
[0136] In an exemplary embodiment, a multimodal image ship target detection system is provided, wherein the multimodal image ship target detection system includes:
[0137] A multimodal image acquisition module is used to acquire multimodal images of ship targets; multimodal images include SAR images and visible light images;
[0138] The ship target dataset construction module is used to preprocess multimodal images and construct a ship target dataset. The preprocessing process includes image resizing and pixel value normalization. The ship target dataset includes a training set and a test set.
[0139] The ship target detection network construction module is used to construct the ship target detection network and train the ship target detection network using the training set; the ship target detection network includes: a lightweight feature extraction module, a cross-scale attention module, a dynamic position encoding Transformer and a multi-scale feature fusion module; the lightweight feature extraction module includes multiple lightweight feature extraction networks and is used to determine the first scale feature map, the second scale feature map and the third scale feature map according to the multimodal image; the lightweight feature extraction network includes: depthwise separable convolution, batch normalization, activation function, point-by-point convolution and channel shuffling; the first scale feature map and The scale of the second-scale feature map is larger than that of the third-scale feature map; the cross-scale attention module includes: 3 parallel void convolution branches, depth-separable convolution branches and grouped convolution branches, and is used to enhance the attention of the third-scale feature map and determine the attention-enhanced third-scale feature map; the dynamic position encoding Transformer is used to perform position and geometric perception enhancement on the first-scale feature map, the second-scale feature map and the attention-enhanced third-scale feature map respectively, and determine the corresponding position-sensitive features; the multi-scale feature fusion module is used to fuse the position-sensitive features and determine a multi-scale feature pyramid with 3 levels;
[0140] The ship target detection module is used to use the trained ship target detection network to perform ship target detection on the multimodal images in the test set.
[0141] This document uses specific examples to illustrate the principles and implementation methods of this application. The description of the above examples is only intended to help understand the method and core concept of this application. At the same time, for those skilled in the art, based on the concept of this application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.
Claims
1. A multimodal image ship target detection method, characterized in that: The multimodal image ship target detection method comprises: Acquire multimodal images of ship targets; multimodal images include SAR images and visible light images; Preprocess the multimodal images and construct a ship target dataset. The preprocessing process includes image resizing and pixel value normalization. The ship target dataset includes a training set and a test set. A ship target detection network is constructed and trained using a training set. The ship target detection network includes a lightweight feature extraction module, a cross-scale attention module, a dynamic position encoding transformer, and a multi-scale feature fusion module. The lightweight feature extraction module includes multiple lightweight feature extraction networks and is used to determine the first-scale feature map, the second-scale feature map, and the third-scale feature map based on the multimodal image. The lightweight feature extraction network includes depthwise separable convolution, batch normalization, activation function, pointwise convolution, and channel shuffling. The scales of the first-scale feature map and the second-scale feature map are both larger than the third-scale feature map. The cross-scale attention module includes three parallel dilated convolution branches, a depthwise separable convolution branch, and a grouped convolution branch, and is used to perform attention enhancement on the third-scale feature map to determine an attention-enhanced third-scale feature map. The dynamic position encoding transformer is used to perform position and geometry perception enhancement on the first-scale feature map, the second-scale feature map, and the attention-enhanced third-scale feature map to determine the corresponding position-sensitive features. The multi-scale feature fusion module is used to fuse the position-sensitive features to determine a multi-scale feature pyramid with three levels. The trained ship target detection network is used to perform ship target detection on the multimodal images in the test set.
2. The multimodal image ship target detection method according to claim 1, characterized in that: The processing process of the cross-scale attention module is: Based on the dilated convolution branch, the depth-wise separable convolution branch, and the grouped convolution branch, the first feature, the second feature, and the third feature corresponding to the third-scale feature map are determined respectively; Performing weighted fusion on the first feature, the second feature, and the third feature based on a spatial attention weight; According to the weighted fusion results, the third-scale feature map for attention enhancement is determined.
3. The multimodal image ship target detection method according to claim 2, characterized in that: The weighted fusion of the first feature, the second feature, and the third feature based on the spatial attention weight specifically includes: Using the formula Determine the spatial attention weight α; where softmax() is the activation function, W2 is the second weight matrix for weight distribution, ReLU is the activation function, W1 is the first weight matrix for feature compression, GAP is the global average pooling, F1 is the first feature, F2 is the second feature, and F3 is the third feature. is element-wise addition; Using the formula Perform weighted fusion on the first feature, the second feature and the third feature; wherein, F out is the third-scale feature map of attention enhancement, α i is the spatial attention weight, i is the feature number, i=1,2,3.
4. The multimodal image ship target detection method according to claim 1, characterized in that: The processing process of the dynamic position encoding Transformer is as follows: Mapping the first-scale feature map, the second-scale feature map, and the attention-enhanced third-scale feature map to corresponding position codes respectively; According to the position code, the corresponding feature map containing the position information is determined respectively; According to the feature map containing position information, the corresponding position-sensitive feature is determined.
5. The multimodal image ship target detection method according to claim 1, characterized in that: The processing process of the multi-scale feature fusion module is as follows: Using the formula and Determine a multi-scale feature pyramid with 3 levels; Among them, P1" is the first level multi-scale feature pyramid, Conv() is the convolution operation, Conv 1×1 () is a 1×1 convolution operation, P 2' is the second up-sampled scale feature map, F1' is the first position sensitive feature, P 2” is the second-level multi-scale feature pyramid, P 3” is the third level multi-scale feature pyramid, F3' is the third position sensitive feature, ↑ is upsampling, ↓ is downsampling, It is element-wise addition.
6. The multimodal image ship target detection method according to claim 1, characterized in that: The loss function of the ship target detection network is determined as follows: Using the formula Determine the improved classification loss function L cls ; Where N is the batch size in each training, t is the sample index, α t is the dynamic weighting coefficient, according to the formula OK, p t The classification branch of the ship target detection network predicts the probability that the t-th input sample belongs to the true category, γ is the adjustment factor, and the formula is: Determine, II(p G =c) is the indicator function, when the predicted label p G When the true label is c, the indicator function is 1, otherwise it is 0, N pos is the number of positive samples of all categories in the current batch, ε is the smoothing term, N total is the number of samples in the current batch, and C is the number of categories; Using the formula L 1oc =0.8*(1-GIoU)+0.2*L1 to determine the positioning loss function L 1oc ; Among them, GIoU is the intersection-over-union function, L1 is the regularization function, and the formula L1=|x pred -x gt |+|y pred -y gt | OK, (x gt ,y gt ) is the center point coordinate of the real frame, (x pred ,y pred ) is the center point coordinate of the prediction box; Using the formula Determine the consistency regularization loss function L consist ; Among them, f() is the feature of the third level feature pyramid, is a multimodal image obtained by preprocessing the multimodal image. is the SAR image; is a visible light image; Using the formula L total =1.0*L cls +0.7*L loc +0.3*L consist Determine the comprehensive loss value L of the loss function of the ship target detection network total .
7. A multimodal image ship target detection system, characterized in that: The multimodal image ship target detection system includes: A multimodal image acquisition module is used to acquire multimodal images of ship targets; multimodal images include SAR images and visible light images; The ship target dataset construction module is used to preprocess multimodal images and construct a ship target dataset. The preprocessing process includes image resizing and pixel value normalization. The ship target dataset includes a training set and a test set. The ship target detection network construction module is used to construct the ship target detection network and train the ship target detection network using the training set; the ship target detection network includes: a lightweight feature extraction module, a cross-scale attention module, a dynamic position encoding Transformer and a multi-scale feature fusion module; the lightweight feature extraction module includes multiple lightweight feature extraction networks and is used to determine the first scale feature map, the second scale feature map and the third scale feature map according to the multimodal image; the lightweight feature extraction network includes: depthwise separable convolution, batch normalization, activation function, point-by-point convolution and channel shuffling; the first scale feature map and The scale of the second-scale feature map is larger than that of the third-scale feature map; the cross-scale attention module includes: 3 parallel void convolution branches, depth-separable convolution branches and grouped convolution branches, and is used to enhance the attention of the third-scale feature map and determine the attention-enhanced third-scale feature map; the dynamic position encoding Transformer is used to perform position and geometric perception enhancement on the first-scale feature map, the second-scale feature map and the attention-enhanced third-scale feature map respectively, and determine the corresponding position-sensitive features; the multi-scale feature fusion module is used to fuse the position-sensitive features and determine a multi-scale feature pyramid with 3 levels; The ship target detection module is used to use the trained ship target detection network to perform ship target detection on the multimodal images in the test set.
Citation Information
Cited By
Insulator ultraviolet corona discharge target detection method based on YOLO-SM
CN120912997A