SiDAT model SAR image ship wake detection method based on multi-modal fusion attention mechanism

Through the SiDAT model of multimodal fusion attention mechanism, the SAR image features are extracted and fused using twin networks and DWT-ATT modules, and the low contrast and fuzzy problems of ship wake detection are solved, achieving efficient and accurate detection effects.

CN120388302AActive Publication Date: 2025-07-29HARBIN ENG UNIV

Patent Information

Application Number
CN202510401722.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-29
Estimated Expiration
2045-04-01

AI Technical Summary

Technical Problem

The ship's wake detection in SAR images has low contrast and blur problems, and the characteristics are affected by the sea surface state and ship speed, resulting in low detection accuracy and insufficient learning ability.

Method used

The SiDAT model adopts a multimodal fusion attention mechanism, and features are extracted and fusion through twin network structure and DWT-ATT module, using the similarity and differences between the visual domain and the wavelet domain, combined with the cross attention mechanism and the self-attention mechanism, improve the accuracy of feature expression and detection.

Benefits of technology

It realizes efficient and accurate SAR image ship wake detection, enhances the fusion of global and local feature information, and improves the accuracy and robustness of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388302A_ABST
    Figure CN120388302A_ABST
Patent Text Reader

Abstract

The invention provides a SiDAT model SAR image ship wake detection method based on a multi-modal fusion attention mechanism. According to the method, similarity extraction of high-level semantic feature items of visual domain and wavelet domain two-mode images is realized through a twin network structure, then effective fusion of visual domain and wavelet domain features of different scales is realized through a DWT-ATT module, and effective fusion of global and local feature information of different scales is enhanced. And efficient and accurate SAR image ship wake detection is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target detection of synthetic aperture radar (SAR), and particularly to a method for detecting ship wakes in SAR images based on a SiDAT model with a multi-modal fusion attention mechanism. Background Art

[0002] In recent years, deep learning methods have made remarkable progress in target detection in SAR images. However, the detection of ship wakes in SAR images still faces many challenges. First, due to the imaging mechanism of SAR images, ship wakes usually appear in the images in the form of low contrast and blurriness. Traditional algorithms are difficult to effectively extract the detailed features of these weak reflection signals, resulting in low accuracy of wake detection. Second, the features of ship wakes are affected by sea surface conditions, ship speed, etc., showing strong time-variability and dependence. This makes the feature information provided by training samples limited in the absence of a large number of samples, thus affecting the learning ability of wake features. Therefore, how to accurately detect ship wakes in SAR images is a challenge that needs to be solved urgently at present. Summary of the Invention

[0003] The purpose of the present invention is to solve the problems in the prior art, and a method for detecting ship wakes in SAR images based on a SiDAT model with a multi-modal fusion attention mechanism is proposed.

[0004] The present invention is realized through the following technical solutions. The present invention proposes a method for detecting ship wakes in SAR images based on a SiDAT model with a multi-modal fusion attention mechanism, and the method includes the following steps:

[0005] Step 1: Input the SAR image into a siamese network for feature extraction. Select a multi-scale convolutional neural network as the model for feature extraction. After feature extraction by the visual branch and the wavelet branch, the input image is transformed into a feature vector and concatenated with the subsequent fusion features;

[0006] Step 2: Perform feature fusion on the visual features and the wavelet features. Input the features of each layer with different scales into the DWT-ATT module, and generate fusion features that fuse the visual domain and the wavelet domain through wavelet transform and the calculation of the attention mechanism;

[0007] Step 3: Initialize to generate learnable fusion features and introduce a cross-attention mechanism for updating in subsequent layers, which can extract global features at the initial stage of the network and gradually fuse high-level features;

[0008] Step 4: Concatenate the visual features, the wavelet features, and the fusion features to obtain a feature vector with rich ship wake features, and input it into a fully connected layer to obtain the final detection result.

[0009] Furthermore, a feature extraction network model is constructed based on the Siamese network structure, which consists of two multi-scale sub-networks sharing parameters; each sub-network extracts rich feature representations from different scales; let the network inputs be X1 and X2, and the outputs after network convolution be feature vectors f(X1) and f(X2), and the similarity between them is calculated through the Euclidean distance, and the similarity metric is expressed as,

[0010] d(X1,X2) = ||f(X1) - f(X2)||2

[0011] Since the two branches share the same parameters, when calculating the similarity loss, the gradients of the two branches will act on the shared parameters simultaneously.

[0012] Furthermore, the loss function is expressed as:

[0013]

[0014] where y is the annotation and m is the distance threshold; the gradient update of the shared parameters is expressed as:

[0015]

[0016] The two branches share the parameter W and act on the same set of weights together during gradient update.

[0017] Furthermore, in step 2, the DWT-ATT module calculates the attention mechanism for the visual features and wavelet features of the current layer, and the obtained fused features are then calculated by the self-attention mechanism to obtain high-weight attention to important features while suppressing irrelevant or redundant information.

[0018] Furthermore, in step 2, each DWT-ATT module takes the visual features and wavelet features of the corresponding scale layer as inputs, and after passing through a linear layer, further wavelet transform is performed by the Haar wavelet basis to obtain four wavelet subbands: X LL , X LH , X HL and X HH ; X LL represents the encoding of the low-frequency component, containing coarse-grained structural information; X LH , X HL , X HH represent the high-frequency components, containing fine-grained textures; the four subbands are concatenated along the channel dimension to obtain which is expressed as,

[0019]

[0020] Feature concatenation output is converted into the key K DWT matrix, value VDWT Matrix and Query Q DWT Matrix, forming Q with visual features vis , K vis , V vis Perform cross-attention mechanism calculation, denoted as

[0021]

[0022] where d k represents the scale factor and represent the features after cross-attention in the wavelet domain and the visual domain respectively; through concatenate the results obtained from the cross-attention calculation to obtain the fused features, and retain the features after each layer of wavelet transform as the input for the next layer of wavelet transform to achieve further wavelet extraction; perform self-attention calculation on the features output by the DWT-ATT module to obtain the re-weighting of the features, and perform cross-attention calculation with the fused features updated layer by layer to obtain the finally updated fused features of this layer

[0023] Furthermore, in step 3, use the feature fusion module to process the obtained fused features, ensure the consistency of the fused features of the current layer and the previous layer in scale by using a 1×1 convolution block, perform cross-attention calculation with the fused features updated from the previous layer, and generate the updated fused features for subsequent processing

[0024] Furthermore, in step 3, initialize a learnable fused feature at the first layer of the multi-scale convolutional neural network, and fuse it with the features extracted layer by layer through the feature fusion module The fusion process is denoted as

[0025]

[0026] where n represents the number of layers of the multi-scale convolutional neural network, and CrossAttention(·) is the cross-attention mechanism calculation

[0027] Furthermore, step 4 adopts a loss function that is the weighted sum of the object classification loss, the directional bounding box regression loss, the wake feature regression loss, and the fused feature optimization loss, denoted as

[0028]

[0029] where represents the object classification loss represents the directional bounding box regression loss represents the wake feature regression loss Denote the fusion feature optimization loss, and λ1, λ2, λ3, λ4 denote the weight hyperparameters.

[0030] The present invention also provides an electronic device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the method for detecting ship wakes in SAR images based on a multi-modal fusion attention mechanism of the SiDAT model are implemented.

[0031] The present invention also provides a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, the steps of the method for detecting ship wakes in SAR images based on a multi-modal fusion attention mechanism of the SiDAT model are implemented.

[0032] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0033] The present invention provides a SiDAT model based on a multi-modal fusion attention mechanism for a method for detecting ship wakes in SAR images. This method extracts the similarity of two-modal images in the visual domain and the wavelet domain at the high-level semantic feature items through a siamese network structure, and then effectively fuses the features of different scales in the visual domain and the wavelet domain through a DWT-ATT module, enhancing the effective fusion of global and local feature information at different scales. Efficient and accurate detection of ship wakes in SAR images is achieved. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0035] Figure 1 It is a flowchart of the method for detecting ship wakes in SAR images based on a multi-modal fusion attention mechanism of the SiDAT model described in the present invention.

[0036] Figure 2 It is a structural framework diagram of the network model.

[0037] Figure 3 It is a structural diagram of the DWT-ATT module.

[0038] Figure 4 It is a schematic diagram of the detection result of ship wakes in SAR in the embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0039] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0040] In view of the problems in the prior art, the present invention proposes a SiDAT model based on a multi-modal fusion attention mechanism for a SAR image ship wake detection method. In order to solve the problem that the ship wake features in the original SAR image are not clear enough, according to the idea of a multi-scale feature extraction network, a feature fusion module is added layer by layer in the network. Corresponding features after wavelet transformation are generated based on the feature maps of each layer scale, and visual features and wavelet features are fused to form higher-level semantic features, making the expression of ship wake features more abundant. At the same time, a siamese network structure is introduced. Through the way of sharing weights and contrastive learning, it automatically captures the similarity of the two-modal images in high-level semantic features, makes full use of the feature information of the two kinds of images, enables the network to obtain the subtle differences and internal connections between images in different domains, and thus improves the accuracy and robustness of the model.

[0041] Specifically, in combination with Figures 1 - 4 , the present invention proposes a SAR image ship wake detection method based on a SiDAT model with a multi-modal fusion attention mechanism. The method includes the following steps:

[0042] Step 1: Input the SAR image into the siamese network for feature extraction. Select a multi-scale convolutional neural network as the model for feature extraction. After feature extraction by the visual branch and the wavelet branch, the input image is transformed into a feature vector and concatenated with the subsequent fused features;

[0043] The dataset used in Step 1 is the OpenSARWake ship wake dataset. Considering the problem that the original SAR image does not comprehensively express the features of ship wakes, the siamese network structure is used to extract features in the visual domain and the wavelet domain respectively, and finally feature concatenation is performed.

[0044] Based on the siamese network structure, a feature extraction network model is constructed, which consists of two multi-scale sub-networks with shared parameters; each sub-network extracts rich feature representations from different scales; let the network inputs be X1 and X2, and the outputs after network convolution be feature vectors f(X1) and f(X2). The similarity between the two is calculated through the Euclidean distance, and the similarity metric is expressed as

[0045] d(X1,X2) = ||f(X1) - f(X2)||2

[0046] Since the two branches share the same parameters, when calculating the similarity loss, the gradients of the two branches will act on the shared parameters simultaneously.

[0047] The loss function is expressed as:

[0048]

[0049] where y is the annotation and m is the distance threshold; the gradient update of the shared parameters is expressed as:

[0050]

[0051] The two branches share the parameter W and act on the same set of weights together during gradient update.

[0052] Step 2: Perform feature fusion on the visual features and wavelet features. Input the features of different scales of each layer into the DWT-ATT module, and generate fused features that integrate the visual domain and the wavelet domain through wavelet transform and the calculation of the attention mechanism;

[0053] In Step 2, the DWT-ATT module calculates the attention mechanism for the visual features and wavelet features of the current layer. The obtained fused features are then calculated through the self-attention mechanism to obtain high-weight attention to important features, while suppressing irrelevant or redundant information.

[0054] In Step 2, each DWT-ATT module takes the visual features and wavelet features of the corresponding scale layer as inputs. After passing through a linear layer, further wavelet transform is performed by the Haar wavelet basis to obtain four wavelet subbands: X LL , X LH , X HL and X HH ; X LL represents the encoding of the low-frequency component, containing coarse-grained structural information; X LH , X HL , X HH represent high-frequency components, containing fine-grained textures; the four subbands are concatenated along the channel dimension to obtain which is expressed as,

[0055]

[0056] Feature concatenation output is converted into a key K DWT matrix, a value V DWT matrix, and a query Q DWT matrix through a convolutional layer, and cross-attention mechanism calculation is performed with the Q vis , K vis , V vis formed by the visual features, which is expressed as,

[0057]

[0058] Among them, d k represents the scale factor, and respectively represent the features after cross-attention in the wavelet domain and the visual domain; through the results obtained by cross-attention calculation are concatenated to obtain the fused feature, and the feature after each layer of wavelet transform is retained as the input of the next layer of wavelet transform to realize further wavelet extraction; the features output by the DWT-ATT module are subjected to self-attention calculation to obtain the reweighting of the features, and cross-attention calculation is performed with the fused feature updated layer by layer to obtain the finally updated fused feature of this layer.

[0059] Step 3: By initializing to generate learnable fused features and introducing a cross-attention mechanism for updating in subsequent layers, global features can be extracted at the initial stage of the network and high-level features can be gradually fused;

[0060] Step 3 uses the feature fusion module to process the obtained fused features, ensures the consistency of the fused features of the current layer and the previous layer in scale by using a 1×1 convolution block, performs cross-attention calculation with the fused feature updated from the previous layer, and generates the updated fused feature for subsequent processing.

[0061] In Step 3, a learnable fused feature is initialized in the first layer of the multi-scale convolutional neural network, and is fused with the features extracted layer by layer through the feature fusion module The fusion process is expressed as,

[0062]

[0063] where n represents the number of layers of the multi-scale convolutional neural network, and CrossAttention(·) is the cross-attention mechanism calculation.

[0064] Step 4: The visual features, wavelet features, and fused features are concatenated to obtain a feature vector with rich ship wake features, which is input into the fully connected layer to obtain the final detection result.

[0065] Step 4 adopts a loss function that is the weighted sum of the object classification loss, the directional bounding box regression loss, the wake feature regression loss, and the fused feature optimization loss, and jointly optimizes the classification ability, positioning accuracy, and fused expression of multi-modal features of the network, which is specifically expressed as,

[0066]

[0067] where, represents the object classification loss, Represents the directional bounding box regression loss, Represents the wake feature regression loss, Represents the fused feature optimization loss, where λ1, λ2, λ3, λ4 represent weight hyperparameters.

[0068] Embodiment

[0069] The objective of the present invention is to solve the problem of SAR image ship wake detection. Using a deep learning network to efficiently and accurately detect SAR ship wake targets and output the corresponding position information and wake direction. To achieve this goal, an embodiment of the present invention provides a SiDAT model SAR image ship wake detection method based on a multi-modal fusion attention mechanism, and the method includes the following steps:

[0070] Step 1: Input the SAR image into a siamese network for feature extraction. Select a multi-scale convolutional neural network as the feature extraction model. After feature extraction by the visual branch and the wavelet branch, the input image is transformed into a feature vector and concatenated with the subsequent fused features;

[0071] The dataset used in Step 1 is the OpenSARWake ship wake dataset, which contains 4,719 satellite data. The images cover three different SAR radar bands: X, C, and L. Divide the existing samples into a dataset, where the training set accounts for 80% of the total number of images, and the test set accounts for 20%. The training set and the test set are generated randomly. Randomly select a part of the training set as the validation set. The input images are uniformly adjusted to a fixed size of 768×768, with a batch size of 16. The training is carried out for 100 iterations.

[0072] Construct a feature extraction network model based on the siamese network structure, which consists of two multi-scale sub-networks with shared parameters; each sub-network extracts rich feature representations from different scales; let the network inputs be X1 and X2, and the outputs after network convolution are feature vectors f(X1) and f(X2). Calculate the similarity between the two through the Euclidean distance, and the similarity metric is expressed as,

[0073] d(X1,X2) = ||f(X1) - f(X2)||2

[0074] Since the two branches share the same parameters, when calculating the similarity loss, the gradients of the two branches will act on the shared parameters simultaneously.

[0075] The loss function is expressed as:

[0076]

[0077] Among them, y is the annotation, and m is the distance threshold; the gradient update of the shared parameters is expressed as:

[0078]

[0079] Two branches share the parameter W and jointly act on the same set of weights during gradient update. This mechanism can effectively improve the training efficiency while ensuring the symmetry of similarity learning.

[0080] Step 2: Perform feature fusion on the visual features and wavelet features. Input the features of different scales in each layer into the DWT-ATT module, and generate fused features that integrate the visual domain and the wavelet domain through wavelet transform and the calculation of the attention mechanism.

[0081] In Step 2, the DWT-ATT module performs the calculation of the attention mechanism on the visual features and wavelet features of the current layer. The obtained fused features are then subjected to self-attention mechanism calculation to obtain high-weight attention to important features, while suppressing irrelevant or redundant information.

[0082] In Step 2, each DWT-ATT module takes the visual features and wavelet features of the corresponding scale layer as inputs. After passing through a linear layer, further wavelet transform is performed by the Haar wavelet basis to obtain four wavelet subbands: X LL , X LH , X HL and X HH ; X LL represents the encoding of the low-frequency component and contains coarse-grained structural information; X LH , X HL , X HH represent high-frequency components and contain fine-grained textures; The four subbands are concatenated along the channel dimension to obtain denoted as,

[0083]

[0084] Feature concatenation output is converted into a key K DWT matrix, a value V DWT matrix, and a query Q DWT matrix through a convolutional layer, and performs cross-attention mechanism calculation with the Q vis , K vis , V vis formed by the visual features, denoted as,

[0085]

[0086] where, d k represents the scale factor, and respectively represent the features after cross-attention in the wavelet domain and the visual domain; Through Concatenate the results obtained from cross-attention calculation to get the fused features, and retain the features after each layer of wavelet transform as the input for the next layer of wavelet transform to achieve further wavelet extraction; perform self-attention calculation on the features output by the DWT-ATT module Obtain the reweighting of the features, and perform cross-attention calculation with the fused features updated layer by layer to get the finally updated fused features of this layer.

[0087] Step 3: By initializing to generate learnable fused features and introducing a cross-attention mechanism for updating in subsequent layers, global features can be extracted at the initial stage of the network and high-level features can be gradually fused;

[0088] Step 3 processes the obtained fused features using a feature fusion module, ensures the consistency of the fused features of the current layer and the previous layer in scale by using a 1×1 convolution block, performs cross-attention calculation with the fused features updated from the previous layer, and generates the updated fused features for subsequent processing.

[0089] In Step 3, a learnable fused feature is initialized in the first layer of the multi-scale convolutional neural network, and is fused with the features extracted layer by layer through the feature fusion module The fusion process is expressed as

[0090]

[0091] where n represents the number of layers of the multi-scale convolutional neural network, and CrossAttention(·) is the cross-attention mechanism calculation. This process enables the learnable fused features to fuse the global information of the previous layer and the local features of this layer, and the sufficient interaction between the upper and lower layer features can capture richer multi-scale information and improve the expression ability of the features.

[0092] Step 4: Concatenate the visual features, wavelet features, and fused features to obtain a feature vector with rich ship wake features, and input it into the fully connected layer to get the final detection result.

[0093] Step 4 adopts a loss function that is the weighted sum of the object classification loss, directional bounding box regression loss, wake feature regression loss, and fused feature optimization loss, and jointly optimizes the classification ability, localization accuracy, and fused expression of multi-modal features of the network, which is specifically expressed as

[0094]

[0095] where represents the object classification loss, represents the directional bounding box regression loss, represents the wake feature regression loss, Denote the fusion feature optimization loss, and λ1, λ2, λ3, λ4 denote the weight hyperparameters.

[0096] Target classification loss Adopt the cross - entropy loss, denoted as

[0097]

[0098] where y i denotes the target class label of the i - th sample (1 represents the positive sample, 0 represents the negative sample), p i denotes the predicted probability of the i - th sample, and N denotes the total number of positive and negative samples. Directional bounding box regression loss is used to optimize the position, size and directionality of the target bounding box. It covers five parameters: length h, width w, center coordinates x, y of the bounding box and rotation angle θ, denoted as

[0099]

[0100] where p=(x s , y s , w s , h s , θ s ) represents the offset between the prediction result and the preset anchor box, and p′=(x′ s , y′ s , w′ s , h′ s , θ′ s ) represents the offset between the predicted box and the ground truth. and represent the smooth loss with a difference of π / 2. and are denoted as

[0101]

[0102] where is defined as

[0103]

[0104] For the wake feature regression loss, it is necessary to perform regression by combining the Kelvin arms and vertices of the wake. Similar to the directional bounding box regression loss, it is denoted as

[0105]

[0106] where represents the result offset, represents the offset between the predicted box and the actual value, represents the balance loss. Fusion feature optimization loss To ensure the semantic expression consistency of three features, it is constrained in the way of mean square error, expressed as

[0107]

[0108] where represents the square of the Euclidean distance between features. The above four losses are weighted and adjusted by hyperparameters, so that the model can achieve an effective balance among classification accuracy, target localization accuracy, and multi-modal feature fusion expression. Finally, the concatenated features pass through the fully connected layer to obtain the mapping of detection task information and give the ship wake detection result.

[0109] The present invention also provides an electronic device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the method for detecting ship wakes in SAR images based on a multi-modal fusion attention mechanism SiDAT model are implemented.

[0110] The present invention also provides a computer-readable storage medium for storing computer instructions, and when the computer instructions are executed by the processor, the steps of the method for detecting ship wakes in SAR images based on a multi-modal fusion attention mechanism SiDAT model are implemented.

[0111] The memory in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DRRAM). It should be noted that the memory of the method described in the present invention is intended to include but not limited to these and any other suitable types of memories.

[0112] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a high-density digital video disc (DVD)), or a semiconductor medium (such as a solid state disc (SSD)), etc.

[0113] In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor or the instructions in the form of software. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware processor, or executed and completed by a combination of the hardware and software modules in the processor. The software module can be located in a mature storage medium in the art such as a random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.

[0114] It should be noted that the processor in the embodiments of the present application may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method embodiments can be completed by the integrated logic circuit in the hardware of the processor or instructions in software form. The above-mentioned processor may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory or electrically erasable programmable memory, register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.

[0115] The above has introduced in detail a method for detecting ship wakes in SAR images of a SiDAT model based on a multi-modal fusion attention mechanism proposed by the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A method for detecting ship wakes in SAR images using a SiDAT model based on a multi-modal fusion attention mechanism, characterized in that, The method includes the following steps: Step 1: Input the SAR image into the siamese network for feature extraction. Select the multi-scale convolutional neural network as the feature extraction model. After feature extraction by the visual branch and the wavelet branch, the input image is transformed into a feature vector and concatenated with the subsequent fusion features; Step 2: Perform feature fusion on the visual features and the wavelet features. Input the features of each layer with different scales into the DWT-ATT module, and generate the fusion features that integrate the visual domain and the wavelet domain through wavelet transform and the calculation of the attention mechanism; Step 3: Initialize to generate learnable fusion features and introduce the cross-attention mechanism for update in subsequent layers, which can extract global features at the initial stage of the network and gradually fuse high-level features; Step 4: Concatenate the visual features, the wavelet features, and the fusion features to obtain a feature vector with rich ship wake features, and input it into the fully connected layer to obtain the final detection result.

2. The method according to claim 1, wherein Based on the siamese network structure, a feature extraction network model is constructed, which consists of two multi-scale sub-networks sharing parameters; each sub-network extracts rich feature representations from different scales; assume the network inputs are X1 and X2, and the outputs after network convolution are the feature vectors f(X1) and f(X2), and the similarity between them is calculated through the Euclidean distance, and the similarity metric is expressed as d(X1,X2) = ||f(X1) - f(X2)||2 Since the two branches share the same parameters, when calculating the similarity loss, the gradients of the two branches will act on the shared parameters simultaneously.

3. The method according to claim 2, wherein The loss function is expressed as: where y is the annotation, and m is the distance threshold; the gradient update of the shared parameters is expressed as: The two branches share the parameter W, and they act on the same set of weights together during gradient update.

4. The method according to claim 1, wherein In Step 2, the DWT-ATT module performs the attention mechanism calculation on the visual features and the wavelet features of the current layer. The obtained fusion features are then subjected to the self-attention mechanism calculation to obtain high-weight attention to important features, while suppressing irrelevant or redundant information.

5. The method according to claim 4, wherein In step 2, each DWT-ATT module takes the visual features and wavelet features of the corresponding scale layer as inputs. After passing through a linear layer, a further wavelet transform is performed by the Haar wavelet basis to obtain four wavelet subbands: X LL , X LH , X HL and X HH ; X LL represents the encoding of the low-frequency component, which contains coarse-grained structural information; X LH , X HL , X HH represent high-frequency components, which contain fine-grained textures; the four subbands are concatenated along the channel dimension to obtain denoted as, Feature splicing output Is converted to key K through the convolutional layer DWT Matrix, value V DWT Matrix and query Q DWT Matrix, and Q formed by visual features vis ,K vis ,V vis Perform cross-attention mechanism calculation, expressed as Among them, d k represents the scale factor, and respectively represent the features after cross-attention in the wavelet domain and the visual domain; the results obtained by cross-attention calculation are concatenated through to obtain the fused features, and the features after each layer of wavelet transform are retained as the input of the next layer of wavelet transform to achieve further wavelet extraction; the features output by the DWT-ATT module are subjected to self-attention calculation to obtain the re-weighting of the features, and cross-attention calculation is performed with the fused features updated layer by layer to obtain the finally updated fused features of this layer.

6. The method according to claim 1, wherein Step 3 uses the feature fusion module to process the obtained fusion features, ensures the consistency of the fusion features of the current layer and the previous layer in scale by using a 1×1 convolutional block, performs cross-attention calculation with the fusion features updated from the previous layer, and generates the updated fusion features for subsequent processing.

7. The method according to claim 6, wherein In step 3, a learnable fusion feature is initialized in the first layer of the multi-scale convolutional neural network and fused with the features extracted layer by layer through the feature fusion module. The fusion process is expressed as fused, and the fusion process is expressed as where n represents the number of layers of the multi-scale convolutional neural network, and CrossAttention(·) is the cross-attention mechanism calculation.

8. The method according to claim 1, wherein Step 4 adopts a loss function that is the weighted sum of the object classification loss, the directional bounding box regression loss, the wake feature regression loss, and the fusion feature optimization loss, which is expressed as Among them, represents the target classification loss, represents the directional bounding box regression loss, represents the wake feature regression loss, represents the fused feature optimization loss, and λ1, λ2, λ3, λ4 represent weight hyperparameters.

9. An electronic device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1-8.

10. A computer-readable storage medium for storing computer instructions, characterized in that, When the computer instructions are executed by the processor, it implements the steps of the method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Improved visual Transform seabed sediment sonar image classification method based on transfer learning

    CN115170943A

  • Multi-modal fusion wavelet knowledge distillation video behavior identification method and system based on cross attention

    CN115294498A

  • Dynamic illumination face image quality enhancement method based on multi-scale attention mechanism

    CN115880225A

  • SAR image ship wake detection method based on frequency domain attention

    CN116778176A

  • SAR (Synthetic Aperture Radar) image ship identification method based on latent diffusion model technology

    CN117173562A

Cited By

  • Remote sensing ship orientation identification method based on Mama and convolution dynamic fusion

    CN120953834A

  • Remote sensing ship heading identification method based on mamba and convolution dynamic fusion

    CN120953834B

  • Multi-scale feature fusion ship attitude prediction method

    CN121524557A

  • A multi-scale feature fusion ship attitude prediction method

    CN121524557B