Underwater target detection method based on multi-modal features and domain adaptation
Through the underwater object detection method of multimodal feature fusion and domain adaptation, the fusion ratio of sonar and optical feature is dynamically adjusted, combined with environmental data and spatial attention calculation, the problem of insufficient utilization of multimodal data in traditional methods is solved, and the accuracy and adaptability of underwater object detection is improved.
Patent Information
- Application Number
- CN202510342530.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-03-21
AI Technical Summary
Traditional underwater object detection methods cannot effectively utilize the complementarity of multimodal data, fixed weight fusion methods lack flexibility, single modal detection methods have deteriorated performance in complex underwater environments, and the synthetic data is very different from real scenarios, resulting in insufficient detection accuracy and generalization capabilities.
Through the multimodal feature fusion and domain adaptation method, the fusion ratio of sonar and optical features is dynamically adjusted, combined with environmental data, spatial attention calculation and target edge alignment are performed, and the difference between synthesis and real domain is narrowed through the asymptotic domain alignment technology to build an object detection model.
It significantly improves the accuracy and reliability of underwater target detection, enhances the adaptability and generalization capabilities of the model in complex and changing environments, and provides stronger technical support for marine exploration, military defense and underwater resource development.
Smart Images

Figure CN120259864A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of target detection, and particularly to an underwater target detection method based on multi-modal features and domain adaptation. Background Art
[0002] Underwater target detection is extremely important in fields such as ocean exploration, military defense, environmental monitoring, and underwater resource development. However, the complex underwater environment poses great challenges to target detection. Light attenuates, scatters, and is absorbed in water, resulting in noisy underwater images, blurred textures, low contrast, and color distortion. Suspended particles and plankton also scatter light, reducing visibility and blurring target edges and details. Although sonar can penetrate turbid waters, it is vulnerable to ocean environmental noise interference, such as water flow noise and biological noise, resulting in a decline in detection accuracy. Therefore, underwater target detection often requires the combination of multiple modal data such as optical images and sonar images. However, the features and attributes of different modal data are different, and it is difficult to effectively utilize their complementarity. For example, optical images are rich in visual information but have poor quality in turbid waters, while sonar images have strong penetration ability but low resolution. Traditional fixed-weight fusion methods cannot fully exploit the potential value of heterogeneous data, and the fusion effect is not good, affecting the detection performance. In addition, the difference between synthetic data and real scenes is large, and the models trained based on synthetic data have poor effects in practical applications. The lighting conditions, water quality characteristics, etc. of synthetic data are different from the real underwater environment, resulting in the models being unable to accurately identify and locate targets. The dynamic changes in the underwater environment, such as water flow, lighting intensity, and direction changes, also pose higher requirements for the generalization ability of the models.
[0003] Traditional target detection methods rely on fixed-weight fusion or single-modal detection. Fixed-weight fusion methods lack flexibility and cannot be adaptively adjusted; single-modal detection methods are limited by the information of a single data source and are difficult to utilize the advantages of multi-modal data. In turbid waters or under dynamic lighting conditions, the performance of these traditional methods drops sharply and cannot meet the actual needs. Summary of the Invention
[0004] The purpose of the present invention is to provide an underwater target detection method based on multi-modal features and domain adaptation, which significantly improves the accuracy of underwater target monitoring by combining multi-modal data acquisition, dynamic feature fusion and decoupling, and embedded real-time detection.
[0005] To achieve the above-mentioned invention purpose, the present invention provides an underwater target detection method based on multi-modal features and domain adaptation, and the method includes:
[0006] S11. Collect underwater target data, where the target data includes sonar images, optical images, and environmental data;
[0007] S12. Extract the features of the sonar image and the optical image respectively to obtain sonar features and optical features. Encode the environmental data into environmental channel weights, and dynamically adjust the fusion ratio of the sonar features and the optical features through the environmental channel weights to obtain fusion features;
[0008] S13. Calculate the spatial attention of the sonar features to obtain a spatial weight map. After enhancing the optical features using the spatial weight map, force the alignment of the enhanced optical features and the sonar features at the target edge through contrastive learning constraints and the Sobel operator;
[0009] S14. Decouple the fusion features into synthetic domain features, and decouple the fusion features and the environmental data into real domain features. Gradually align the synthetic domain features and the real domain features through asymptotic domain alignment to complete the construction of the target detection model.
[0010] Further, extracting the features of the sonar image and the optical image respectively specifically includes:
[0011] Extract the features of the sonar image through Wavelet-CNN. Use the Haar wavelet to decompose the sonar image I sonar into high-frequency components, and the expression is:
[0012] F sonar = HL(I sonar ) + LH(I sonar ) + HH(I sonar )
[0013] where I sonar is the sonar image, I sonar ∈R l×H×W , R is the set of real numbers, l is a single channel, H is the height, W is the width, H×W is the spatial size, F sonar is the sonar feature, and HL, LH, and HH are the high-low, low-high, and high-high frequency components of the Haar wavelet decomposition respectively;
[0014] Retain the high-frequency components and further extract features through CNN, and output the completed sonar feature F sonar ∈R C1 ×H×W , where C1 is the number of channels of the sonar feature;
[0015] Extract the features of the optical image through ResNet-50, and output the optical feature F optical ∈R C2×H×W , where C2 is the number of channels of the optical feature.
[0016] Further, encoding the environmental data into environmental channel weights specifically includes:
[0017] Convert the environmental data into an environmental parameter vector E = [τ, l, d], and encode E through MLP1. The expression is:
[0018] W env = MLP1(E)
[0019] where E ∈ R 3 , W env is the channel weight generated for the environmental parameters, and W env ∈ R C . τ is turbidity, l is light intensity, and d is water depth.
[0020] Furthermore, dynamically adjust the fusion ratio of sonar features and optical features through the environmental channel weight, specifically including:
[0021] S21. Merge the optical features and sonar features in the channel dimension to obtain the concatenated features. The expression is:
[0022]
[0023] where F concat is the concatenated feature of the optical feature and the sonar feature, and F concat ∈
[0024] R (CI+C2)×H×W ;
[0025] S22. After performing global average pooling on F concat , compress it through MLP2 to obtain the weight control vector. The expression is:
[0026]
[0027] W feat = MLP2(F pooled )
[0028] where F pooled is the compressed feature, and F pooled ∈ R C1+C2 , W feat is the weight control vector, and W feat ∈ R C ;
[0029] S23. After multiplying W env element-wise with W feat , compress it through the Sigmoid function to obtain the channel weight matrix. The expression is:
[0030] W combined = W env ⊙ W feat
[0031] W = σ(W combined )
[0032] where W combined is the result obtained by element-wise multiplication of W env and W feat . W is the channel weight matrix, σ() is the Sigmoid function, and ⊙ is the element-wise multiplication operation;
[0033] S24. After expanding the channel weight matrix to the spatial dimension through the broadcasting mechanism, fuse the sonar feature and the optical feature to obtain the fused feature. The expression is:
[0034] W spatial = W ∈ R C×H×W
[0035] F fusion = W spatial · F optical + (1 - W spatial )· F sonar
[0036] where W spatial is the result obtained after expanding W to the spatial dimension through the broadcasting mechanism, W spatial ∈ R C×H×W , and F fusion is the fused feature.
[0037] Furthermore, after calculating the spatial attention of the sonar feature, perform a normalization operation to generate the spatial attention map. The expression is:
[0038] M att = Softmax(Conv 1×1 (F sonar )) ∈ [0, 1] H×W
[0039]
[0040] where M att is the spatial weight map, Softmax() is the normalization operation, Conv 1×1 () is the one-dimensional convolution operation, exp() is the exponential function, Matt[h, w] is the element value at the h-th row and w-th column in the spatial attention map M att , [h, w] is a specific position in the spatial attention map M att or the sonar feature map, where h is the row index and w is the column index, and [h′, w′] is the position in the spatial attention map or the sonar feature map, which is a variable used for the summation operation.
[0041] Furthermore, enhance the optical feature using the spatial weight map, specifically including:
[0042] Enhance the optical features using a spatial weight map, and the expression is:
[0043]
[0044] Where, is the enhanced optical feature, ⊙ represents the element-wise multiplication operation,
[0045] M att ⊙F optical means multiplying the spatial weight map M att element-wise with the corresponding positions of each channel of the optical feature F optical , and M att ⊙F optical +F optical means adding the weighted optical feature to the original optical feature.
[0046] Furthermore, force the alignment of the enhanced optical features and sonar features at the target edge through contrastive learning constraints and the Sobel operator, specifically including:
[0047] S31. Respectively use the horizontal Sobel operator and the vertical Sobel operator to perform convolution operations on to obtain the gradient map in the horizontal direction and the gradient map in the vertical direction
[0048] S32. Respectively use the horizontal Sobel operator and the vertical Sobel operator to perform convolution operations on F sonar to obtain the gradient map in the horizontal direction and the gradient map in the vertical direction
[0049] S33. For each position (i, j) on the feature map, calculate the L2 norm between the gradient vector of the enhanced optical feature and the gradient vector of the sonar feature, and the L2 norm is the Euclidean distance;
[0050] S34. Sum the gradient differences at all positions (i, j) on the feature map to obtain the contrastive learning constraint loss L cont , and the expression is:
[0051]
[0052] S35. Update the parameters by passing L cont through the loss function and backpropagation, so that Lcont Gradually decrease to achieve the alignment of optical features and sonar features at the target edge.
[0053] Furthermore, decouple the fused feature into synthetic domain features, specifically including:
[0054] S41. Propagate the fused feature forward and backward, and the expression is:
[0055] F inv = GRL(F fusion )
[0056]
[0057] where F inv is the feature generated by forward propagation, GRL() is the gradient reversal operation, is the gradient operator, L adv is the adversarial loss function, and -λ is a hyperparameter, i.e., the gradient reversal coefficient;
[0058] S42. Input F inv into the discriminator network to calculate the adversarial loss L adv , forcing the feature extractor to generate synthetic domain features F syn that cannot be distinguished by the discriminator;
[0059] Furthermore, decouple the fused feature and environmental data into real domain features, specifically including:
[0060] Embed the fused feature and environmental data into the feature space to obtain real domain features, and the expression is:
[0061]
[0062] where F real is the real domain feature, γ is the scaling factor, and β is the offset factor;
[0063] The calculation expression of u is:
[0064]
[0065] The calculation expression of σ is:
[0066]
[0067] Input the environmental parameter vector E = [τ, l, d] into MLP3, and split the neurons output by MLP3 into γ and β with C dimensions each to obtain γ(E) and β(E), and the expression is:
[0068] γ,β = MLP3(E)
[0069] where γ,β ∈ Rc ×R c 。
[0070] Furthermore, the synthetic domain features are gradually aligned with the real domain features through asymptotic domain alignment, which specifically includes:
[0071] S51. Calculate the distribution difference between the synthetic domain feature F syn and the real domain feature F real using the MMD metric to obtain MMD(F syn , F real );
[0072] S52. Detect the real domain data and calculate the detection loss L det according to the detection results and the real labels;
[0073] S53. Calculate the total loss of asymptotic domain alignment, and the expression is:
[0074] L DA = e -0.05t ·MMD(F syn , F real ) + (1 - e -0.05t )·L det
[0075] where L DA is the total loss of asymptotic domain alignment, t is the number of training iterations, e -0.05t is the weight coefficient of the domain alignment loss, and 1 - e -0.05t is the weight coefficient of the detection loss;
[0076] S54. Use the backpropagation algorithm to calculate the gradient of the total loss L DA with respect to the parameters of the object detection model, and update the parameters of the object detection model according to the gradient to make the total loss L DA gradually decrease;
[0077] S55. Iteratively train steps S51 to S54 until the preset number of training times is reached to complete the construction of the object detection model.
[0078] Compared with the prior art, the beneficial effects of the present invention are:
[0079] The underwater target detection method based on multi-modal features and domain adaptation provided by the present invention overcomes the limitations of single-modal detection methods by fusing the features of sonar images and optical images and dynamically adjusting the fusion ratio of sonar features and optical features in combination with environmental data. This enables the model to adaptively optimize the feature fusion effect according to different underwater environments and target characteristics, thereby more accurately identifying and locating targets, and significantly improving the accuracy and reliability of underwater target detection. At the same time, through spatial attention calculation and target edge forced alignment technology, the differences in target edges and details between multi-modal data are solved, further improving the model's ability to capture target edges and details. In addition, the asymptotic domain alignment strategy can gradually reduce the differences between synthetic data and real scenes, reduce the model's dependence on synthetic data, and enhance the adaptability and generalization ability of the model in different underwater environments. The present invention can still maintain excellent detection performance in complex and variable underwater environments, providing more powerful technical support for fields such as ocean exploration, military defense, environmental monitoring, and underwater resource development. BRIEF DESCRIPTION OF THE DRAWINGS
[0080] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application, and for those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0081] Figure 1 It is a schematic flow chart of the underwater target detection method based on multi-modal features and domain adaptation provided by the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0082] The present invention will be further described in detail below with reference to the drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present invention, rather than limiting the present invention. In addition, it should be noted that for the sake of description, only parts related to the present invention are shown in the drawings, not all structures.
[0083] Refer to Figure 1 , this embodiment provides an underwater target detection method based on multi-modal features and domain adaptation, and the method includes:
[0084] S11. Collect underwater target data, where the target data includes sonar images, optical images, and environmental data. In this embodiment, the collection of underwater targets can synchronously collect data through an optical camera (RGB + multispectral), a side-scan sonar, and an environmental sensor (turbidity, light, depth), or any collection device capable of collecting sonar images, optical images, and environmental data can be applied to the present invention, and the present invention does not limit this.
[0085] S12. Extract the features of the sonar image and the optical image respectively to obtain sonar features and optical features. After encoding the environmental data into environmental channel weights, dynamically adjust the fusion ratio of the sonar features and the optical features through the environmental channel weights to obtain fusion features. In this embodiment, the fusion ratio of the optics and sonar is dynamically adjusted according to the environmental parameter vector, solving the problem of insufficient adaptability of fixed-weight fusion, that is, driving the pixel-level fusion of the optics and sonar through the environmental parameter vector.
[0086] Extract the features of the sonar image and the optical image respectively, specifically including:
[0087] Extract the features of the sonar image through Wavelet-CNN. Use the Haar wavelet to decompose the sonar image I sonar into high-frequency components, and the expression is:
[0088] F sonar =HL(I sonar )+LH(I sonar )+HH(I sonar )
[0089] where I sonar is the sonar image, I sonar ∈R l×H×W , R is the set of real numbers, l is a single channel, H is the height, W is the width, H×W is the spatial size, F sonar is the sonar feature, and HL, LH, and HH are the high-low, low-high, and high-high frequency components of the Haar wavelet decomposition respectively.
[0090] Retain the high-frequency components and further extract features through CNN, and output the completed sonar feature F sonar ∈R C1 ×H×W , where C1 is the number of channels of the sonar feature.
[0091] In this embodiment, wavelet decomposition can effectively remove the high-frequency noise of the sonar image, retain the high-frequency components to highlight the target contour, and suppress the noise.
[0092] Extract the features of the optical image through ResNet-50, and output the optical feature F optical ∈R C2×H×W , where C2 is the number of channels of the optical feature.
[0093] In this embodiment, the input optical image is processed through a 50-layer residual network (ResNet-50), and finally a feature map F optical ∈R C2×H×W, this feature map contains the high-level semantic features of the input image and can be used for subsequent multi-modal fusion and object detection tasks. Among them, ResNet-50 contains 4 residual blocks, each block contains multiple convolutional layers (such as 3×3 convolution), batch normalization (BN) operations and ReLU activation functions.
[0094] R C×H×W The parameter meanings are as follows. R represents the set of real numbers, and R is used to describe the value range of the feature tensor. C represents the number of channels. In the output of ResNet-50, the number of channels C depends on the specific structure and configuration of the network. Usually, the number of channels output by the residual block in the last stage is the number of channels of the final output features. For example, if the number of output channels of the residual block in the last stage is 256, then C = 256. The number of channels reflects the richness of the features, and each channel can be regarded as a representation of different features of the image. H represents the height of the feature map. After a series of convolution and pooling operations, the height of the feature map will gradually decrease. The final height H depends on parameters such as the size of the input image, the size of the convolution kernel, the stride, and the padding. W represents the width of the feature map, which is the result obtained after convolution and pooling operations, and it reflects the size of the feature map in the horizontal direction.
[0095] Encode the environmental data into environmental channel weights, specifically including:
[0096] Convert the environmental data into an environmental parameter vector E = [τ, l, d], and encode E through MLP1. The expression is:
[0097] W env = MLP1(E)
[0098] Among them, E ∈ R 3 W env is the channel weight generated by the environmental parameters, W env ∈ R C , τ is the turbidity, the range is [0, 1], l is the light intensity, the unit is lux, and d is the water depth, the unit is meter.
[0099] In this embodiment, map the environmental parameter vector to a weight vector in the channel dimension. W env is the channel weight generated by the environmental parameters to control the contribution ratio of sonar / optics. For example, when the turbidity τ > 0.7, W env [C = 0] (low-frequency channel) may be encoded as 0.9. As the input, the turbidity τ may generate a higher value in the low-frequency channel (such as C = 0) after the non-linear transformation of MLP1. Example: When τ = 0.8, MLP1 may output W env[C = 0] = 0.9. The structure of MLP1 is as follows: input layer: 3 - dimensional (τ, l, d); hidden layer: 64 neurons, ReLU activation; output layer: C - dimensional (consistent with the number of feature channels).
[0100] Among them, the channel weight W generated by the environmental parameters env is a set of weight values obtained by processing the environmental parameter vector E, and is used to adjust the importance of different channels in the subsequent fusion process of multi - modal data (such as optical and sonar data). Generation process: Convert the environmental data into an environmental parameter vector, and then input it into MLP1 for encoding. MLP1 is a multi - layer perceptron composed of an input layer, a hidden layer, and an output layer. The input layer receives a three - dimensional environmental parameter vector, the hidden layer performs a non - linear transformation, and the output layer outputs a vector with a dimension of C, that is, the channel weight W generated by the environmental parameters env .
[0101] Dynamically adjust the fusion ratio of sonar features and optical features through the environmental channel weight, specifically including:
[0102] S21. Merge the optical features and sonar features in the channel dimension to obtain the concatenated features. The expression is:
[0103]
[0104] Among them, F concat is the feature after concatenating the optical feature and the sonar feature, F concat ∈
[0105] R (CI+C2)×H×W .
[0106] In this embodiment, when merging the optical features and sonar features in the channel dimension, the prerequisite is that the spatial dimensions (height H and width W) of F optical and F sonar must be the same, that is, F optical ∈R C2×H×W , F sonar ∈R C1×H×W . Concatenate them in the channel dimension, and the dimension of the concatenated feature F concat is R (CI+C2)×H×W . In a deep learning framework (such as PyTorch), the torch.cat() function can be used to achieve this.
[0107] S22. After performing global average pooling on F concat , compress it through MLP2 to obtain the weight control vector. The expression is:
[0108]
[0109] Wfeat = MLP2)F pooled )
[0110] Where F pooled is the compressed feature, F pooled ∈ R C1+C2 , and W feat is the weight control vector, W feat ∈ R C .
[0111] In this embodiment, after the concatenated feature F concat is subjected to global average pooling, the compressed feature F pooled is obtained. In multi-modal data processing, the optical feature and the sonar feature are concatenated after extraction, and the resulting F concat contains rich information but has a high dimension. To process these features more efficiently, dimensionality reduction is performed on it through global average pooling operation. The compressed feature F pooled will be compressed by a multi-layer perceptron MLP2 to obtain the weight control vector W feat , which is used to control the weights of each channel during multi-modal data fusion. MLP2 usually consists of an input layer, several hidden layers, and an output layer. The number of neurons in the input layer is equal to the number of channels of the concatenated feature C1 + C2, and the number of neurons in the output layer is C, that is, the dimension of W feat ; the number of neurons and the number of layers in the hidden layer can be adjusted according to specific tasks and datasets.
[0112] Forward propagation process, input layer: The feature F concat after global average pooling is flattened into a one-dimensional vector as the input of MLP2. Assume F concat ∈ R (CI+C2)×H×W , and the dimension of the flattened vector is R (CI+C2)×H×W . Hidden layer: The input vector passes through each hidden layer in turn. Each hidden layer contains several neurons, and the neurons are connected in a fully connected manner. In each hidden layer, the input vector is multiplied by the weight matrix of that layer, then added with the bias vector, and finally undergoes a non-linear transformation through an activation function (such as ReLU). Output layer: After being processed by several hidden layers, the output of the last hidden layer is input to the output layer. The output layer also performs weighted summation and bias operations, but usually does not use an activation function (unless there are special requirements) to obtain the final output vector W feat ∈ R C .
[0113] S23. After multiplying W env element-wise with W feat , it is compressed through the Sigmoid function to obtain the channel weight matrix, and the expression is:
[0114] W combined = W env ⊙W feat
[0115] W = σ(W combined )
[0116] In this embodiment, W combined is the result obtained by element-wise multiplying the channel weight W generated by environmental parameters env and the weight control vector W feat . It synthesizes the influence of environmental factors and the characteristics of multimodal data itself on the channel weight, initially fusing the weight information from two different sources, preparing for the subsequent generation of the final channel weight matrix. W is the finally obtained channel weight matrix, which is obtained by compressing W combined through the Sigmoid function. The element values in W are restricted between 0 and 1, so that it can be used as a weight to reasonably fuse each channel of multimodal data. σ() is the Sigmoid function, and its expression is Its role is to compress the input value into the interval of 0 to 1.
[0117] Element-wise multiply W env and W feat to combine environmental parameters and feature statistical information. For example: if W env [C = 0] = 0.9 (high turbidity) and W feat [C = 0] = 0.5 (feature shows low-frequency importance), then W combined [C = 0] = 0.45. When τ > 0.7, the value of W env in the low-frequency channel is relatively high, resulting in a lower W, thereby increasing the sonar weight 1 - W. W feat reflects the importance of the current input feature. For example: if the sonar feature has a strong response in the low-frequency channel, then W feat [C = 0] is relatively high, further increasing the sonar weight. The sonar weight = 1 - W. For example: if W[C = 0] = 0.2, then the sonar weight in this channel is 1 - 0.2 = 0.8.
[0118] Low-frequency channel characteristics. The low-frequency components of sonar (such as contour information) are less affected by turbidity. When the turbidity is high, the low-frequency channels of optical features may contain more noise. By increasing the sonar weight, these noises can be suppressed. Weight fusion dynamically generates fine-grained fusion weights through the synergistic effect of environmental parameters and feature statistics. When the turbidity exceeds the threshold, the sonar weight in the low-frequency channel increases significantly, effectively suppressing optical noise and realizing the adaptive fusion of multimodal data.
[0119] S24. After expanding the channel weight matrix to the spatial dimension through the broadcasting mechanism, fuse the sonar feature and the optical feature to obtain the fused feature. The expression is as follows:
[0120] W spatial = W ∈ R C×H×W
[0121] F fusion = W spatial · F optical +(1 - W spatial )· F sonar
[0122] In this embodiment, W spatial is the result obtained by expanding the channel weight matrix W to the spatial dimension through the broadcasting mechanism. Its main function is to perform weighted fusion on the sonar feature F sonar and the optical feature F optical in the spatial dimension to generate the fused feature F fusion , so as to make full use of the complementary information of the two different modality data and improve the accuracy and robustness of target detection. The broadcasting mechanism is a mechanism that enables tensors of different shapes to perform element-wise operations without actually replicating data. Specifically, it replicates the weight value of each channel in W in the height and width directions so that it covers the entire spatial range of the feature map, thereby obtaining W spatial .
[0123] Expanding the channel weight to the spatial dimension realizes per-pixel dynamic adjustment. Gated feature fusion is a spatial refinement step of dynamic weight fusion. By expanding the channel weight to the spatial dimension, per-pixel adjustment of modal contributions is realized. This fine-grained control mechanism can effectively suppress environmental noise (such as turbidity interference) and enhance modal complementarity, ultimately improving the robustness of underwater target detection. If C = 256, H = 256, W = 256, then the weight of each channel will be replicated to all spatial positions. For each channel C and each spatial position [h, w], element-wise operations are performed. The expression is as follows:
[0124] F fusion [c, h, w] = W[c, h, w]· F optical [c, h, w]+(1 - W[c, h, w])· F sonar [c, h, w]
[0125] Balance the two modalities through linear interpolation to avoid feature conflicts. The element-wise multiplication operation supports backpropagation, facilitating end-to-end training. When sonar detects the target contour (such as a metal pipe) but the optical image is blurred, gated fusion increases the contribution of the sonar channel in this area to strengthen the contour features. For example: in areas with high turbidity (such as τ>0.7), the sonar weight 1-W is increased to 0.8 in the low-frequency channel to suppress the optical noise in this area.
[0126] S13. Perform spatial attention calculation on the sonar features to obtain a spatial weight map. After enhancing the optical features using the spatial weight map, force-align the enhanced optical features and sonar features at the target edge through contrastive learning constraints and the Sobel operator. Use the contour information of sonar to enhance the target area of the optical image and solve the problem of edge blurring. Sonar data is sensitive to the target contour (stable low-frequency components), but has low resolution; optical images are rich in details but prone to blurring. Transfer the contour knowledge of sonar to the optical features through the attention mechanism to strengthen the target edge.
[0127] After performing spatial attention calculation on the sonar features, perform a normalization operation to generate a spatial attention map, and the expression is:
[0128] M att =Softmax(Conv 1×1 (F sonar ))∈[0,1] H×W
[0129]
[0130] Among them, M att is the spatial weight map, Softmax() is the normalization operation, Conv 1×1 () is a one-dimensional convolution operation, exp() is the exponential function, Matt[h,w] is the element value at the h-th row and w-th column in the spatial attention map M att Among them, [h,w] is a specific position in the spatial attention map M att or the sonar feature map, where h is the row index, w is the column index, and [h′,w′] represents a position in the spatial attention map or the sonar feature map, which is a variable used for the summation operation.
[0131] In this embodiment, the spatial attention map M att is a two-dimensional matrix with a dimension of (H×W). The value of Matt[h,w] is in the range of [0,1], which reflects the importance of the sonar feature map at the position (h,w). The closer the value is to 1, the more important the features at this position are for target detection; the closer the value is to 0, the lower the importance of the features at this position. [h′,w′] is in the denominator ∑ of the Softmax normalization operationh′,w′ exp(Conv 1×1 (F sonar )[h′,w′]), [h',w'] traverses all positions in the spatial attention map, that is, h' ranges from 0 to H - 1 and w' ranges from 0 to W - 1. Then, the values of exp(Conv 1×1 (F sonar )[h′,w′]) at all positions are summed up. This summation result is used to normalize the values of exp(Conv 1×1 (F sonar )[h,w]) at each position, so as to ensure that the sum of all elements in M att is 1.
[0132] Conv 1×1 () is a 1×1 convolution operation. The 1×1 convolution kernel is a special convolution operation with a kernel size of 1×1, which compresses the sonar feature F sonar from C channels to 1 channel. For each position [h,w] in the input feature map F sonar , the convolution will perform a weighted sum of the values of the C channels at this position to obtain a new value. This process can be regarded as a linear combination of features from different channels, thus compressing the multi-channel features into single-channel features. The Softmax function converts each element in the input vector or tensor into a probability value such that the sum of all elements is 1. The Softmax function processes the single-channel feature map compressed by the 1×1 convolution to generate a spatial weight map M att to highlight the target contour area detected by the sonar.
[0133] The spatial weight map is a two-dimensional matrix with a dimension of H×W. Each element in the spatial weight map represents the importance of the sonar feature at the corresponding spatial position. The larger the value, the more important the sonar feature at that position. In the subsequent optical feature enhancement process, the optical features at that position will receive more attention and enhancement. Due to the Softmax normalization, each element value in M att is within the range of [0, 1], and the sum of all elements is 1. For example, in underwater target detection, if the value of a certain area in M att is large, it indicates that the sonar has detected an obvious target contour in this area, and this information can be used later to enhance the features of the corresponding area in the optical image.
[0134] The optical features are enhanced using the spatial weight map, and the expression is:
[0135]
[0136] Among them, For the enhanced optical feature, ⊙ represents the element-wise multiplication operation, M att ⊙F optical It is to multiply the spatial weight map M att element-wise with the corresponding positions of each channel of the optical feature F optical , M att ⊙F optical +F optical It is to add the weighted optical feature to the original optical feature.
[0137] In this embodiment, M att ⊙F optical is to strengthen the target area indicated by sonar, superimpose the original feature to retain background information, and avoid over-enhancement. The optical feature is weighted according to the importance of sonar features. In the area where the sonar detects the target contour (i.e., the area with a larger median value in M att ), the response of the optical feature at these positions will be enhanced; while in the area where the sonar deems unimportant M att (the area with a smaller median value), the response of the optical feature will be suppressed. M att ⊙
[0138] F optical +F optical In order to retain the background information and other details in the original optical feature, avoid information loss caused by over-enhancement, and at the same time highlight the target area features indicated by sonar. For example: when the sonar detects the contour of a metal pipe, M att is 0.9 in this area, and the response of the corresponding area of the optical feature is increased by 90%.
[0139] Sonar data is sensitive to the target contour and can relatively stably detect the approximate contour of the target even in a complex environment; while the optical image, although rich in details, is easily affected by environmental factors (such as illumination, turbidity, etc.), resulting in blurred target edges. By using the spatial weight map M att generated from sonar features to weight the optical feature, the contour information of sonar can be transferred to the optical feature, thereby strengthening the features of the target area in the optical image and solving the problem of blurred edges.
[0140] By using contrastive learning constraints and Sobel operators to force the alignment of the enhanced optical feature and sonar feature at the target edge, specifically including:
[0141] S31. Respectively use the horizontal Sobel operator and the vertical Sobel operator to perform a convolution operation to obtain the gradient map in the horizontal direction and the gradient map in the vertical direction
[0142] S32. Use the horizontal Sobel operator and the vertical Sobel operator to perform a convolution operation on F sonar to obtain the gradient map in the horizontal direction and the gradient map in the vertical direction
[0143] S33. For each position (i, j) on the feature map, calculate the enhanced optical feature gradient vector and the sonar feature gradient vector The L2 norm between them, and the L2 norm is the Euclidean distance;
[0144] S34. Sum the gradient differences at all positions (i, j) on the feature map to obtain the contrastive learning constraint loss L cont , and the expression is:
[0145]
[0146] S35. Update the parameters by passing L cont through the loss function and backpropagation, so that L cont gradually decreases to achieve the alignment of the optical feature and the sonar feature at the target edge.
[0147] In this embodiment, by using the contrastive learning constraint and the Sobel operator, it is possible to force the enhanced optical feature and the sonar feature to be consistent at the target edge, improving the effect of multi-modal feature fusion. L cont is the contrastive learning constraint loss, which is used to measure the enhanced optical feature and the difference in edge information between the original sonar feature F sonar . Its goal is to minimize this loss so that the edges of the enhanced optical feature are as aligned as possible with the edges of the sonar feature, thereby solving the modal conflict and making the features of the two modalities consistent at the target edge. If L cont increases, the sonar weight is automatically increased to further suppress optical noise.
[0148] is the gradient vector of the enhanced optical feature at the position (i, j), representing the edge strength and direction information of the enhanced optical feature at this position. is the gradient vector of the sonar feature at the position (i, j), which also represents the edge strength and direction information of the sonar feature at this position. Σ i,j is to sum over all positions (i, j) on the feature map, that is, to accumulate the gradient differences at each position on the feature map to obtain an overall loss value. The Sobel operator in the horizontal direction is a 3×3 convolution kernel used to detect edges in the horizontal direction of an image or feature map. The Sobel operator in the vertical direction is the transpose of the Sobel operator in the horizontal direction and is used to detect edges in the vertical direction of an image or feature map.
[0149] Through the sonar-guided attention mechanism, edge enhancement is performed on the optical features before fusion, solving the possible blur problem of the fused features. The enhanced optical features output not only optimize the weight distribution of the fused features but also directly improve the localization and classification performance of the detection head, forming a closed-loop optimization of multi-modal features. The edges of the enhanced optical features are sharper, improving the localization accuracy of the detection box, and the sonar-optical edge alignment reduces feature ambiguity and improves the classification confidence.
[0150] S14: Decouple the fused features into synthetic domain features, and decouple the fused features and environmental data into real domain features. Gradually align the synthetic domain features with the real domain features through asymptotic domain alignment to complete the construction of the object detection model. In this embodiment, the difference in the distribution of synthetic data and real water areas leads to a decline in the overall prediction performance. By adversarial training, the essential features of the target are separated so that they are not affected by environmental changes, separating the synthetic domain features (essential attributes of the target or environment-invariant features) from the real domain features (illumination, turbidity, i.e., environment-related features), and improving the cross-domain generalization ability.
[0151] Decoupling the fused features into synthetic domain features specifically includes:
[0152] S41: Perform forward propagation and backward propagation on the fused features, and the expression is:
[0153] F inv = GRL(F fusion )
[0154]
[0155] where F inv is the feature generated by forward propagation, GRL() is the gradient reversal operation, is the gradient operator, L adv is the adversarial loss function, and -λ is a hyperparameter, i.e., the gradient reversal coefficient.
[0156] In this embodiment, in the neural network, forward propagation means that data starts from the input layer,
[0157] The process of passing through each hidden layer in sequence and finally reaching the output layer. The purpose of forward propagation is to extract and process input features through a series of calculations and transformations in the network to obtain a more meaningful feature representation that can be used for subsequent tasks (such as classification, detection, etc.). In normal neural network training, the gradient is the direction information used to update network parameters to minimize the loss function. The Gradient Reversal Layer (GRL) reverses the direction of the gradient during backpropagation. When data passes through the GRL layer during forward propagation, it simply outputs the input as it is, i.e., F inv = GRL(F fusion ) where the F inv and F fusion are numerically the same. But during backpropagation, the gradient passing through the GRL layer is multiplied by a negative coefficient (usually -1), causing the direction of the gradient to be reversed.
[0158] is used to calculate the gradient of a certain function with respect to the parameter θ. The gradient represents the gradient of the adversarial loss L adv with respect to the network parameter θ, which reflects the rate and direction of change of the loss function L adv as the parameter θ changes. -λ is a positive value that is used to control the intensity of gradient reversal. During backpropagation, the gradient is multiplied by -λ, i.e., The value of λ affects the balance between domain alignment and object detection in the model. If λ is large, the effect of gradient reversal is stronger, and the model will pay more attention to domain alignment, that is, making the feature distributions of data in different domains closer; if λ is small, the model will focus more on the accuracy of object detection. Therefore, the choice of λ needs to be adjusted according to the specific task and dataset to achieve the best performance. L adv is used to measure the ability of the discriminator to distinguish data from different domains.
[0159] S42. Input F inv into the discriminator network to calculate the adversarial loss L adv , forcing the feature extractor to generate synthetic domain features F syn that cannot be distinguished by the discriminator.
[0160] In this embodiment, adversarial training forces Feature extractor to generate synthetic domain features F syn that cannot be distinguished by the discriminator through backpropagation gradient reversal. These features capture the essential attributes of the target (such as shape, material), and together with the fused features and enhanced optical features, significantly improve the generalization ability of the model in the synthetic → real domain.
[0161] The calculation expression of the adversarial loss L adv is:
[0162] L adv = -E real [logD(F inv )] - E syn [log(1 - D(F inv ))]
[0163] where E real is the expectation of the real domain features, from the real data processing flow, E syn is the expectation of the synthetic domain features, from the real data processing flow, D is the discriminator network, with input F inv , and output the domain probability (0 = synthetic domain, 1 = real domain).
[0164] The discriminator D is a binary classification network for judging the domain label (synthetic / real) of the features. The synthetic domain is the fused features (from synthetic data, such as images generated by simulating the underwater environment), and the real domain is the fused features and environmental data (from real - collected data, such as optical and sonar images of the actual water area). The feature extractor is the generator, which is used to generate synthetic domain features.
[0165] The synthetic domain refers to the domain where synthetic data generated by simulating the underwater environment, etc. is located. "The synthetic domain is the fused features" means the features extracted from synthetic data and after the fusion operation. These fused features are formed by fusing sonar features and optical features, etc. through a series of operations mentioned above, and are used to represent the synthetic data at the feature level, so that the discriminator D can judge whether the feature comes from the synthetic domain.
[0166] The real domain refers to the domain where the actually collected data is located, such as the optical and sonar images obtained in the actual water area. "The real domain is the fused features and environmental data" means that in the real domain, the data used for the discriminator D to judge not only includes the fused features extracted and fused from the actually collected optical and sonar images, etc., but also includes the relevant environmental data. The environmental data may include some background information, lighting conditions, water quality, etc. of the actual water area. These information, together with the fused features, serve as the complete representation of the real - domain data, enabling the discriminator D to more accurately judge whether the feature comes from the real domain and distinguish the data - feature differences between the real domain and the synthetic domain.
[0167] Decoupling the fused features and environmental data into real - domain features specifically includes:
[0168] Embedding the fused features and environmental data into the feature space to obtain the real - domain features, and the expression is:
[0169]
[0170] Among them, F real is the real domain feature, γ is the scaling factor, and β is the offset factor;
[0171] The calculation expression of u is:
[0172]
[0173] The calculation expression of σ is:
[0174]
[0175] Input the environmental parameter vector E = [τ, l, d] into MLP3, split the neurons output by MLP3 into γ and β with each C dimensions, and obtain γ(E) and β(E). The expression is:
[0176] γ, β = MLP3(E)
[0177] In this embodiment, the environmental parameters such as light and turbidity in different waters vary greatly, resulting in changes in feature distribution. When the turbidity is high, the scaling factor γ of the low-frequency channel decreases; when the light is low, the offset β of the luminance channel increases, and γ, β ∈ R c ×R c . For example, when the turbidity τ = 0.9, the noise in the low-frequency channel of the optical feature increases, and the gain needs to be reduced. The feature values are scaled to a similar range through normalization for subsequent processing.
[0178] The structure of MLP3 includes: input layer: 3 neurons (corresponding to τ, l, d); hidden layer: 64 neurons, ReLU activation; output layer: 2C neurons (split into γ and β with each C dimensions). For each channel C and each spatial position [h, w], element-wise operations are performed. The expression is:
[0179]
[0180] When the environment is in high turbidity (τ = 0.9), MLP3 outputs γ[C = 0] = 0.5 (the gain of the low-frequency channel is reduced by 50%) to suppress the noise in the sonar feature. When the environment is in low light (l = 10 lux), MLP3 outputs β[C = 2] = 0.3 (the offset of the luminance channel increases by 0.3) to enhance the visibility of the optical feature.
[0181] Align the synthetic domain features with the real domain features step by step through asymptotic domain alignment, specifically including:
[0182] S51. Use the MMD metric to calculate the distribution difference between the synthetic domain feature F syn and the real domain feature F real to obtain MMD(F syn , Freal )。
[0183] S52. Detect the real - domain data, and calculate the detection loss L according to the detection result and the real label det 。
[0184] S53. Calculate the total loss of progressive domain alignment, and the expression is:
[0185] L DA =e -0.05t ·MMD(F syn ,F real )+(1 - e -0.05t )·L det
[0186] Among them, L DA is the total loss of progressive domain alignment, t is the number of training iterations, e -0.05t is the weight coefficient of the domain - alignment loss, and 1 - e -0.05t is the weight coefficient of the detection loss.
[0187] In this embodiment, synthetic - domain data and real - domain data are collected. The synthetic - domain data can generate optical and sonar images by simulating the underwater environment, and the real - domain data are optical and sonar images collected from the actual water area. After processing the above data, fused features and environmental data are obtained. The synthetic - domain features are obtained by decoupling the fused features, and the real - domain features are obtained by decoupling the fused features and environmental data.
[0188] L DA is the weighted sum of the domain - alignment loss and the detection loss, which is used to consider both domain alignment and detection accuracy during the training process; the value of t increases continuously as the training progresses; e -0.05t As the number of training iterations t increases, the value of e -0.05t gradually decreases, which means that at the beginning of training, this coefficient is larger and the model focuses more on domain alignment. MMD(F syn ,F real ) is the Maximum Mean Discrepancy (MMD) loss, which is used to measure the distribution difference between the synthetic - domain feature F syn and the real - domain feature F real . The smaller the MMD value, the closer the feature distributions of the two domains are. The value of 1 - e -0.05t increases gradually as the number of training iterations t increases, which means that in the later stage of training, this coefficient is larger and the model focuses more on detection accuracy; L --0.05t det is used to measure the detection accuracy of the model.
[0189] S54. Use the back - propagation algorithm to calculate the total loss L DA Regarding the gradients of the parameters of the object detection model, update the parameters of the object detection model according to the gradients, so that the total loss L DA gradually decreases.
[0190] S55. Iteratively train steps S51 to S54 until a preset number of training times is reached, and complete the construction of the object detection model.
[0191] In this embodiment, in the cross-domain object detection task, the environment-invariant features (synthetic domain data) are the mathematical representations of the essential attributes of the object. For example, the invariant features of a metal object may include high-frequency reflectivity patterns; the invariant features of a cylindrical object may include circular edge responses. The environment-related features (real domain data) are the features affected by the environment. For example, the illumination intensity of an optical image, such as the pixel values in overexposed areas; the noise level of a sonar image, which is positively correlated with turbidity; the color shift of a multispectral image, which is affected by water depth. There are distribution differences between the synthetic domain data and the real domain data. When directly applying the model trained in the synthetic domain to the real domain, the performance of the model often drops significantly. The purpose of domain alignment is to narrow the feature distribution differences between the synthetic domain and the real domain, so that the model can perform well on different domains.
[0192] However, simply performing domain alignment may ignore the goal of the detection task itself, resulting in a situation where although the model is aligned in terms of the domain, the detection accuracy is not high. Therefore, a progressive domain alignment strategy is adopted. In the initial stage of training, more emphasis is placed on domain alignment, allowing the model to first learn the common features between the synthetic domain and the real domain; in the later stage of training, more emphasis is placed on detection accuracy, ensuring that the model can accurately detect the object. This can balance the relationship between domain alignment and detection accuracy, and improve the generalization ability and detection performance of the model in the real domain.
[0193] Through adversarial training and conditional normalization, the fused features are decoupled into the essential attributes of the object (invariant features) and environmental interferences (related features). This separation mechanism not only retains the core features shared across domains but also provides an explicit compensation channel for environmental changes, thereby achieving robust underwater object detection. The metal object reflection pattern learned by the invariant features in synthetic data can be directly transferred to real waters; the related features encode factors such as illumination and turbidity as adjustable scaling factors. For example, at high turbidity, the gain of optical features is reduced to suppress noise. Through the above progressive domain alignment method, the domain alignment and detection accuracy are dynamically balanced during the training process, improving the performance of the model in cross-domain scenarios.
[0194] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. An underwater target detection method based on multi-modal features and domain adaptation, characterized in that, The method includes: S11. Collect underwater target data, where the target data includes sonar images, optical images, and environmental data; S12. Extract the features of the sonar image and the optical image respectively to obtain sonar features and optical features, encode the environmental data as environmental channel weights, and dynamically adjust the fusion ratio of the sonar features and the optical features through the environmental channel weights to obtain fusion features; S13. Calculate the spatial attention of the sonar features to obtain a spatial weight map, use the spatial weight map to enhance the optical features, and then force-align the enhanced optical features and the sonar features at the target edge through contrastive learning constraints and the Sobel operator; S14. Decouple the fusion features into synthetic domain features, decouple the fusion features and the environmental data into real domain features, and gradually align the synthetic domain features and the real domain features through asymptotic domain alignment to complete the construction of the target detection model.
2. The underwater target detection method based on multi-modal features and domain adaptation according to claim 1, characterized in that, The extraction of the features of the sonar image and the optical image respectively specifically includes: Feature extraction of sonar images is performed by Wavelet-CNN, and the sonar image I is decomposed into high-frequency components using Haar wavelets. The expression is as follows: sonar The decomposition is as follows: F sonar = HL(I sonar ) + LH(I sonar ) + HH(I sonar ) Among them, I sonar is a sonar image, I sonar ∈R l×H×W , where R is the set of real numbers, l is a single channel, H is the height, W is the width, H×W is the spatial dimension, and F sonar is a sonar feature, and HL, LH, and HH are the high-low, low-high, and high-high frequency components of Haar wavelet decomposition, respectively; Retain the high-frequency components and further extract features through the CNN, and output the completed sonar feature F sonar ∈R C1×H×W , where C1 is the number of channels of the sonar feature; Extract the features of the optical image through ResNet-50 and output the optical feature F optical ∈R C2×H×W , where C2 is the number of channels of the optical feature.
3. The underwater target detection method based on multi-modal features and domain adaptation according to claim 2, wherein, The encoding of the environmental data as environmental channel weights specifically includes: Convert the environmental data into an environmental parameter vector E = [τ, l, d], and encode E through MLP1. The expression is: W env = MLP1(E) where \(E\in R\) 3 , \(W\) env is the channel weight generated by environmental parameters, \(W\) env \(\in R\) C , \(\tau\) is turbidity, \(l\) is light intensity, and \(d\) is water depth.
4. The underwater target detection method based on multi-modal features and domain adaptation according to claim 3, wherein The dynamic adjustment of the fusion ratio of the sonar features and the optical features through the environmental channel weights specifically includes: S21. Merge the optical features and the sonar features in the channel dimension to obtain the concatenated features. The expression is: Among them, F concat is the feature after splicing the optical feature and the sonar feature, F concat ∈R (CI+C2)×H×W ; S22. Global average pooling is performed on F concat After that, it is compressed by MLP2 to obtain a weight control vector, and the expression is as follows: W feat = MLP2(F pooled ) Among them, F pooled is a compression feature, F pooled ∈R C1+C2 , W feat is a weight control vector, W feat ∈R C ; S23. Multiply W env element-wise with W feat and then compress it through the Sigmoid function to obtain the channel weight matrix. The expression is as follows: W combined = W env ⊙W feat W = σ(W combined ) Among them, W combined is the result obtained by element-wise multiplication of W env and W feat The result obtained by element-wise multiplication, W is the channel weight matrix, σ() is the Sigmoid function, and ⊙ is the element-wise multiplication operation; S24. After expanding the channel weight matrix to the spatial dimension through the broadcast mechanism, fuse the sonar features and the optical features to obtain fusion features. The expression is: W spatial = W ∈ R C×H×W F fusion = W spatial · F optical +(1 - W spatial )· F sonar Among them, W spatial is the result obtained by expanding W to the spatial dimension through the broadcasting mechanism. W spatial ∈R C×H×W , and F fusion is the fused feature.
5. The underwater target detection method based on multi-modal features and domain adaptation according to claim 2, characterized in that, After calculating the spatial attention of the sonar features, perform a normalization operation to generate a spatial attention map. The expression is: M att = Softmax(Conv 1×1 (F sonar )) ∈ [0, 1] H×W Among them, M att is the spatial weight map, Softmax() is the normalization operation, Conv 1×1 () is the one-dimensional convolution operation, exp() is the exponential function, and Matt[h,w] is the element value at the h-th row and w-th column in the spatial attention map M att where [h,w] is a specific position in the spatial attention map M att or a specific position in the sonar feature map, where h is the row index and w is the column index, and [h ′ ,w ′ represents the position in the spatial attention map or the sonar feature map, which is a variable used for the summation operation.
6. The underwater target detection method based on multi-modal features and domain adaptation according to claim 5, wherein The enhancement of the optical features using the spatial weight map specifically includes: The enhancement of the optical features using the spatial weight map. The expression is: Among them, is the enhanced optical feature, ⊙ is the element-wise multiplication operation, M att ⊙F optical means to multiply the spatial weight map M att element-wise with the corresponding positions of each channel of the optical feature F optical , and M att ⊙F optical +F optical means to add the weighted optical feature to the original optical feature.
7. The underwater target detection method based on multi-modal features and domain adaptation according to claim 6, wherein The forced alignment of the enhanced optical features and the sonar features at the target edge through contrastive learning constraints and the Sobel operator specifically includes: S31. Respectively use the horizontal Sobel operator and the vertical Sobel operator to perform a convolution operation to obtain the gradient map in the horizontal direction and the gradient map in the vertical direction S32. Respectively use the horizontal Sobel operator and the vertical Sobel operator to perform a convolution operation on F sonar to obtain the gradient map in the horizontal direction and the gradient map in the vertical direction S33. For each position (i, j) on the feature map, calculate the enhanced optical feature gradient vector and the sonar feature gradient vector The L2 norm between them, and the L2 norm is the Euclidean distance; S34. Sum the gradient differences at all positions (i, j) on the feature map to obtain the contrastive learning constraint loss L cont , and the expression is as follows: S35. Set L cont Update the parameters through the loss function and backpropagation so that L cont gradually decreases to achieve the alignment of the optical features and sonar features at the target edge.
8. The underwater target detection method based on multi-modal features and domain adaptation according to claim 4, wherein The decoupling of the fusion features into synthetic domain features specifically includes: S41. Perform forward propagation and backward propagation on the fusion features. The expression is: F inv = GRL(F fusion ) Among them, F inv is the feature generated by forward propagation, GRL() is the gradient reversal operation, is the gradient operator, L adv is the adversarial loss function, -λ is a hyperparameter, namely the gradient reversal coefficient; S42. Input F inv into the discriminator network to calculate the adversarial loss L adv , forcing the feature extractor to generate synthetic domain features F that cannot be distinguished by the discriminator syn .
9. The underwater target detection method based on multi-modal features and domain adaptation according to claim 8, characterized in that, The decoupling of the fusion features and the environmental data into real domain features specifically includes: Embed the fusion features and the environmental data into the feature space to obtain real domain features. The expression is: where F real is the real domain feature, γ is the scaling factor, and β is the offset factor; The calculation expression of u is: The calculation expression of σ is: Input the environmental parameter vector E = [τ, l, d] into MLP3, and split the neurons output by MLP3 into γ and β, each with C dimensions, to obtain γ(E) and β(E). The expression is: γ, β = MLP3(E) where γ, β ∈ R c ×R c .
10. The underwater target detection method based on multi-modal features and domain adaptation according to claim 9, wherein, The gradual alignment of the synthetic domain features and the real domain features through asymptotic domain alignment specifically includes: S51. Calculate the distribution difference between the synthetic domain feature F syn and the real domain feature F real to obtain MMD(F syn , F real ); S52. Detect the real-domain data, and calculate the detection loss L according to the detection result and the real label det ; S53. Calculate the total loss of the asymptotic domain alignment. The expression is: L DA = e -0.05t ·MMD(F syn , F real ) + (1 - e -0.05t )·L det Among them, L DA is the total loss of progressive domain alignment, t is the number of training iterations, and e -0.05t is the weight coefficient of the domain alignment loss, and 1 - e -0.05t is the weight coefficient of the detection loss; S54. Calculate the total loss L using the backpropagation algorithm DA For the gradients of the object detection model parameters, update the parameters of the object detection model according to the gradients so that the total loss L DA gradually decreases; S55. Iteratively train steps S51 to S54 until the preset number of training times is reached to complete the construction of the target detection model.
Citation Information
Patent Citations
Underwater target detection and identification method based on acousto-optic fusion
CN116452965A
Model training method and device based on multi-modal data, equipment and storage medium
CN117807495A
Underwater target detection method
CN117876856A
Underwater scene-oriented decoupling representation domain adaptive sonar image classification method
CN118212458A
Underwater target identification method based on multi-modal fusion
CN118485908A
Cited By
Multi-modal bridge pier damage detection method and system based on image sonar adjustment
CN120765643A
Underwater target tracking system based on DBN
CN120831669A
Tree species identification method and system based on multi-source remote sensing data fusion
CN120873827A
A tree species identification method and system based on multi-source remote sensing data fusion
CN120873827B
Underwater structure defect dynamic grading method based on multi-modal perception fusion
CN121234286A