A sonar and camera image fusion denoising method based on cross-modal frequency domain modulation and deformable cross attention
Patent Information
- Application Number
- CN202611001328.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-07
- Publication Date
- 2026-09-29
AI Technical Summary
[0006]本发明目的是为了解决水下极端环境中可见光图像受散射与浑浊影响导致细节丢失、声呐图像存在散斑噪声且空间分辨率较低以及跨模态视角差与几何失配难以对齐等技术问题,提供了一种基于跨模态频域调制与可变形交叉注意力的声呐和摄像头图像融合去噪方法
[0053]本发明提出了一种能够在复杂工况下有效融合声呐与可见光互补信息、兼顾去噪与细节保真的图像重建方法,具体为面向水下极端环境,基于可学习小波分解、SwinTransformer 全局建模与可变形跨模态注意力融合的声呐与可见光图像融合与去噪增强方法。
Smart Images

Figure CN122841191A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of digital image processing technology, and in particular to underwater image processing and multimodal sensing. Background Technology
[0002] Unmanned autonomous underwater vehicles (UUVs) have been widely used in various underwater operations, such as hydrological surveying, environmental assessment, offensive and defensive operations, and underwater salvage missions. Underwater imagery, as a key means for UUVs to acquire environmental information, plays a crucial role in these missions. However, underwater images are often affected by various factors, such as turbidity in lakes or shallow seas, water scattering at long distances, and the influence of lighting, resulting in loss of detail, insufficient usable information, and a decline in image quality indicators such as contrast. Therefore, restoring underwater images to high quality, especially image restoration and enhancement under low-light, turbid, and long-distance conditions, has become a critical problem that urgently needs to be solved.
[0003] Currently, underwater image restoration methods can be divided into two categories: those based on traditional mathematical modeling and those based on deep learning. Traditional UIR methods rely on physical models or digital image processing techniques. The core of these methods is to use prior knowledge or reasonable assumptions to reverse the image degradation process, thereby restoring a clear underwater image, such as histogram equalization, white balance, and wavelet transform. Deep learning-based underwater image restoration methods utilize their powerful nonlinear mapping capabilities to automatically learn the nonlinear mapping relationship between degraded and clear images. They can adaptively learn the image degradation patterns of different underwater environments and improve image quality.
[0004] In deep learning methods, single-modal methods focus on enhancing contrast and suppressing noise in visible light or sonar images, but they lack sufficient utilization of complementary information. In contrast, multimodal fusion methods can utilize information from other modalities, such as sonar image information, and perform better under extreme conditions. Fusion methods include handcrafted methods based on traditional registration and weighting, feature-level / pixel-level fusion based on convolutional neural networks, and long-range dependency modeling methods utilizing self-attention mechanisms. These methods still face challenges in cross-modal geometric mismatch, strong noise interference, and generalization to complex scenes. For example, registration is sensitive to prior assumptions, simple stitching is difficult to achieve sufficient alignment and noise suppression, and some deep learning methods suffer from insufficient detail texture recovery or real-time limitations.
[0005] Therefore, there is an urgent need for a fusion denoising technology that can effectively align and fuse complementary sonar and visible light information in complex underwater conditions, while suppressing noise and improving the fidelity of structural and texture details. Summary of the Invention
[0006] The purpose of this invention is to solve the technical problems of visible light images being affected by scattering and turbidity in extreme underwater environments, resulting in loss of detail; sonar images having speckle noise and low spatial resolution; and difficulty in aligning cross-modal viewpoint differences and geometric mismatches. This invention provides a sonar and camera image fusion and denoising method based on cross-modal frequency domain modulation and deformable cross-attention.
[0007] The technical solution adopted by this invention to solve the above problems is: a sonar and camera image fusion denoising method based on cross-modal frequency domain modulation and deformable cross-attention, the method comprising:
[0008] Step 1: Perform pairwise organization and consistency processing on sonar images and visible light images;
[0009] Step 2: Extract scale-consistent deep sonar features from the sonar images processed in Step 1;
[0010] Step 3: Extract multi-scale semantic and texture features from the visible light image processed in Step 1 to obtain semantically enhanced deep visible light features;
[0011] Step 4: The sonar deep features and visible light deep features obtained in Steps 2 and 3 are unified in terms of channel and spatial scale, and geometrically consistent cross-modal feature alignment is achieved.
[0012] Step 5: The visible light image features and sonar image features obtained in Step 4 are fused using a deformable cross-attention mechanism to achieve multi-scale adaptive fusion and information complementarity, including:
[0013] Using visible light features as queries and sonar features as keys, deformable cross-attention is used to adaptively sample and aggregate sonar features at learnable reference points, thereby achieving cross-modal feature alignment and information complementarity.
[0014] The deformable cross attention generates normalized reference points on the visible light feature grid, predicts the offset and weight for each query position, samples a preset number of points from the sonar features for weighted aggregation, and then outputs the fusion result through linear projection, residual connection and layer normalization.
[0015] Step 6: The fusion result obtained in Step 5 is gradually upsampled and reconstructed by the decoder to output a clear image with enhanced noise reduction; during the training phase, reconstruction, adversarial and frequency domain modulation multi-target loss are jointly optimized.
[0016] Further, step 2 specifically includes: extracting scale-consistent deep sonar features from the sonar image using learnable wavelet decomposition. The learnable wavelet decomposition adopts a two-dimensional separable lifting structure, and is decomposed sequentially along both column and row directions to output a low-frequency LL subband and three directional subbands of high frequencies LH, HL, and HH. The filtering weights of the prediction and update operators in the learnable wavelet decomposition are parameterized by a neural network and adaptively optimized through backpropagation during end-to-end training.
[0017] The low-frequency LL subband is encoded and decoded and skip connections are used to enhance structural representation, outputting low-frequency features dominated by structure; the high-frequency LH, HL, and HH directional subbands are respectively passed through two convolutional layers, one BN layer and one ReLU activation function, and then a spatial attention module is used to suppress noise and extract feature texture, outputting three high-frequency features.
[0018] Low-frequency features and three high-frequency features are spliced together in the channel dimension and compressed to the preset fusion dimension by 1×1 convolution to obtain sonar deep features with consistent scale.
[0019] Further, step 3 specifically includes: using a visual Transformer with window self-attention and hierarchical representation capabilities to perform hierarchical encoding on the visible light image, wherein the visual Transformer includes patch embedding and progressive downsampling modules; modeling within a local window while preserving global dependencies across windows, obtaining deeper features with stronger semantics through progressive downsampling and channel expansion; and selecting an appropriate hierarchical output as the fusion input according to the target task to take into account both texture details and global context.
[0020] Further, step 4 specifically includes: projecting the two-way features of sonar deep features and visible light deep features obtained in steps 2 and 3 onto a unified fusion channel dimension, and unifying the spatial resolution to the same scale through interpolation or lightweight alignment modules; introducing coordinate or position encoding to enhance geometric perception, while maintaining consistency between the numerical domain and dynamic range, and reducing error propagation caused by distribution differences and scale mismatches of cross-modal features before fusion.
[0021] Further, step 5 specifically includes:
[0022] Step 5.1: The visible light image features and sonar image features obtained in Step 4 are serialized and arranged through the position encoding coordinate channel module; the spatial dimension H×W is stretched to the sequence length N=H×W, and the channel vector of each position is used as the embedding representation of a single token;
[0023] Step 5.2: To establish cross-modal associations in a unified feature space, perform independent linear projection operations on the input sequences to generate query vector Q, key vector K, and value vector V;
[0024] Step 5.3: To realize the feature fusion mechanism of geometric perception, a normalized reference point coordinate grid is generated according to the spatial resolution (H,W) of the feature map; the two-dimensional coordinates are mapped to the interval [0,1] with the pixel center as the reference, as a spatial reference for attention sampling; finally, the reference point tensor is obtained, and the spatial shape parameters and hierarchical index are recorded simultaneously.
[0025] Step 5.4: Use a multi-head deformable cross-attention module to fuse the visible light image features and sonar image features obtained in Step 4. For each query point, extract local value vectors from the sonar features at several sampling locations near the reference point, and aggregate them through attention weights to achieve cross-modal dynamic fusion.
[0026] Step 5.5: The intermediate result F obtained by fusion is processed by the output linear layer and random deactivation, then added to the input query Q through a residual connection, and layer normalization is performed. The calculation formula is as follows:
[0027]
[0028] The fused sequence F′ is then rearranged back into a two-dimensional feature form according to its original spatial dimensions to restore the spatial structure and adapt it to the subsequent decoding module, outputting the fused feature map.
[0029] Furthermore, step 6 specifically includes:
[0030] Step 6.1: The fusion result obtained in Step 5 is first upsampled through a deconvolution layer to expand the feature space size, and then batch normalization and ReLU activation function are applied after convolution to obtain upsampled features;
[0031] Subsequently, the upsampled features are further processed through a convolutional reconstruction layer to generate an intermediate reconstructed image;
[0032] Step 6.2: The intermediate reconstructed image generated in step 6.1 is adjusted to the resolution of visible light features through bilinear interpolation. During the interpolation process, the input pixels are regarded as the center points of continuous cells, and the output pixels correspond to the relative proportional positions of the input cells. The pixel values of the final reconstructed image output are in the numerical domain. The numerical domain is mapped back to the image domain to obtain a visualization image, which is used for quantitative index calculation and subjective quality assessment.
[0033] Step 6.3: During the training phase, reconstruction loss, adversarial loss, and frequency domain modulation loss are used to optimize the generator, thereby improving performance in terms of structure fidelity, visual naturalness, and spectral consistency.
[0034] Further, step 6.3 includes:
[0035] The overall optimization objective of the generator is defined as:
[0036]
[0037] Among them, reconstruction loss L recon By calculating the differences between the generated image and the real image at the pixel level or perceptual feature level, the adversarial loss L... adv Based on the LSGAN framework, the output of the constraint generator can deceive the discriminator, λ recon , λ adv and λ cmfm These are the weight parameters for reconstruction loss, adversarial loss, and frequency domain modulation loss, respectively.
[0038] Frequency domain modulation loss :
[0039] Where α and β are the weighting parameters for the low-frequency and high-frequency terms, respectively; and These are low-frequency and high-frequency terms, respectively.
[0040] and The calculation method is as follows: the input tensor X is subjected to a two-dimensional FFT and center translation in the Tanh domain, using a concentric circular mask M with radius r. low Perform low-pass and high-pass separation, and take the real part as the low-frequency component and high-frequency component:
[0041] ,
[0042] Among them, low frequency items The low-frequency components of the graph generated using L1 distance constraints are consistent with the low-frequency components of the camera input, while the high-frequency components... InfoNCE contrast loss is used, with the generated high frequency as the anchor; frequency domain modulation loss L cmfm Structural and texture constraints are introduced at the low and high frequency levels respectively to ensure that the generated image is consistent with the real sample in terms of spectral characteristics such as illumination and detail.
[0043] The total loss of the discriminator is defined as:
[0044]
[0045] in, It is the discrimination loss of the discriminator when judging real, clean images. It is the discriminator's loss in distinguishing fake images generated by the generator. This is the gradient penalty, used to stabilize the convergence process of adversarial training and prevent gradient explosion or vanishing.
[0046] Subsequently, parameter updates employ an alternating optimization approach. In each training iteration, the system first freezes the discriminator parameters to update the generator, and then freezes the generator parameters in reverse to optimize the discriminator. The update rule is as follows:
[0047]
[0048] Where, η G With η D The learning rates for the generator and discriminator are respectively. and This represents the trainable parameters of the network, which are the generator and discriminator. and This represents the gradient of the loss function with respect to its respective parameters.
[0049] Secondly, the present invention provides a sonar and camera image fusion denoising system based on cross-modal frequency domain modulation and deformable cross-attention. The system has a program module corresponding to the steps of the sonar and camera image fusion denoising method based on cross-modal frequency domain modulation and deformable cross-attention as described above. When running, the system executes the steps in the sonar and camera image fusion denoising method based on cross-modal frequency domain modulation and deformable cross-attention.
[0050] Thirdly, the present invention provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the processor runs the computer program stored in the memory, it performs the steps of a sonar and camera image fusion and denoising method based on cross-modal frequency domain modulation and deformable cross-attention as described above.
[0051] Fourthly, the present invention provides a computer-readable storage medium for storing a computer program that executes a sonar and camera image fusion denoising method based on cross-modal frequency domain modulation and deformable cross-attention as described above.
[0052] The beneficial effects of this invention are:
[0053] This invention proposes an image reconstruction method that can effectively fuse complementary information from sonar and visible light under complex working conditions, while taking into account both denoising and detail preservation. Specifically, it is a sonar and visible light image fusion and denoising enhancement method based on learnable wavelet decomposition, SwinTransformer global modeling, and deformable cross-modal attention fusion for extreme underwater environments.
[0054] This invention effectively suppresses noise and fogging in visible light images under typical underwater degradation conditions such as strong turbidity, low illumination, scattering, and speckle. Guided by the prior knowledge of the stable sonar structure, it restores contour layers and contrast, reduces color cast, and improves the naturalness of the visual perception, making it more in line with the subjective visual evaluation of the human eye. The paired data organization and unified preprocessing adopted are simple and efficient, and the model completes fusion and reconstruction end-to-end in a single forward pass, balancing speed and accuracy, and has engineering deployment value. At the same time, it can maintain structural continuity and clear texture even in the presence of registration errors, scale deviations, or slight attitude changes, providing clearer and more reliable input for underwater mapping, detection, identification, and other tasks that depend on image quality.
[0055] This invention employs learnable wavelet decomposition of sonar images to explicitly obtain and encode low / high frequency and directional sub-bands. Then, a visual Transformer extracts hierarchical global semantic features from the visible light image. These two feature paths are aligned by channel and spatial scale before entering a deformable cross-attention module. Using visible light as the query and sonar as the key, adaptive sampling and aggregation are performed at learnable reference points. Cross-modal frequency domain modulation constraints are used to enhance low-frequency structure and high-frequency texture, and suppress noise propagation. This design achieves more comprehensive, robust, and realistic fusion results without significantly increasing computational overhead. It significantly alleviates ghosting and artifact problems caused by misalignment, comprehensively improving objective indicators such as PSNR and SSIM, as well as subjective perception, and provides stable performance gains for downstream tasks such as target detection and segmentation.
[0056] This invention is applicable to achieving high-quality acoustic-optical fusion and image reconstruction in complex underwater environments. Attached Figure Description
[0057] To more clearly illustrate the technical solution of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0058] Figure 1 This is a schematic diagram of the system structure of the present invention.
[0059] Figure 2 This is a schematic diagram of the training process of the present invention.
[0060] Figure 3 This is a schematic diagram of the sonar feature extraction module.
[0061] Figure 4 This is a schematic diagram of feature alignment and deformable cross attention module.
[0062] Figure 5 This is a comparison chart of ablation experiment results.
[0063] Figure 6 This is a comparison chart of high and low frequency decomposition. Detailed Implementation
[0064] The specific implementation method of this embodiment, a sonar and camera image fusion denoising method based on cross-modal frequency domain modulation and deformable cross-attention, includes:
[0065] Step 1: Data Preprocessing: Register, normalize, and resize sonar and visible light images; perform paired organization and consistency processing; including:
[0066] Sonar and visible light images are processed in pairs for organization and standardization: they are unified to a preset resolution, the sonar is converted to a single channel and intensity is normalized, and the visible light is standardized according to common statistics to adapt to the pre-trained distribution of the visual backbone; if necessary, the two images are subjected to synchronous geometric processing to reduce registration error, and stable reuse is ensured during training and validation stages through cache indexing; at the same time, the input size is ensured to meet the requirements of even height and width for subsequent wavelet decomposition in order to obtain a regular subband grid.
[0067] Step 2: Sonar Feature Extraction: Learnable wavelet decomposition is performed on the sonar image. After channel fusion and compression, high-information-density, scale-consistent deep sonar features are generated to provide structural priors for subsequent cross-modal alignment; including:
[0068] Learnable wavelet decomposition is performed on sonar images, completing the "prediction-update" process along both column and row directions, decoupling the original image into low-frequency and three-way directional high-frequency sub-bands. Differentiated coding strategies are used to characterize and denoise the different sub-bands: low-frequency focuses on global structure modeling, while high-frequency focuses on texture edge and speckle suppression. Subsequently, fusion and compression are performed in the channel dimension to form scale-consistent, high-information-density sonar deep features, providing robust structural priors for cross-modal alignment and fusion.
[0069] Specifically, the learnable wavelet decomposition is implemented using a two-dimensional separable lifting structure, decomposing sequentially along both column and row directions to output a low-frequency LL subband and three directional subbands (LH, HL, and HH) of high frequencies. The filter weights of the prediction and update operators in the wavelet decomposition are parameterized by a neural network and adaptively optimized through backpropagation during end-to-end training. This allows the filter kernel weights of the subband decomposition to automatically adjust according to the input data distribution, achieving the learnable characteristics of wavelet transform. The low-frequency subband undergoes a deeper encoder-decoder branch and uses skip connections to enhance structural representation. The three high-frequency subbands undergo shallow convolution and integrate channel and spatial attention modules to suppress noise and preserve texture. The features of the four subbands are concatenated in the channel dimension and compressed to a preset fusion dimension via 1×1 convolution.
[0070] Step 3: Visible Light Image Feature Extraction: Multi-scale semantic and texture features of the visible light image are extracted layer by layer to obtain semantically enhanced deep features; including:
[0071] A visual Transformer with window self-attention and hierarchical representation capabilities is used to encode visible light images hierarchically. It efficiently models within local windows while preserving global dependencies across windows. It obtains deeper features with stronger semantics through progressive downsampling and channel expansion. The appropriate hierarchical output is selected as the fusion input according to the target task to take into account both texture details and global context. If necessary, it can be replaced with a lightweight convolutional backbone to meet real-time requirements based on computing power.
[0072] Specifically, the visual Transformer is a hierarchical backbone based on window self-attention, including patch embedding and progressive downsampling modules. Preferably, the patch stride is 4, the window size is 7, and the output selects the features from the final stage. Meanwhile, the visible light feature extraction can be replaced with a lightweight network composed of depthwise separable convolutions stacked together, while maintaining the same number of output channels as the fusion dimension.
[0073] It should be noted that a lightweight network with depthwise separable convolution stacks can also be used to extract visible light features while maintaining the number of output channels consistent with the fusion dimension.
[0074] Step 4: Feature Alignment: Unify the two feature paths in terms of channel and spatial scale, and introduce coordinate encoding to achieve geometrically consistent cross-modal feature alignment; including:
[0075] Sonar and visible light features are projected onto a unified fusion channel dimension, and the spatial resolution is unified to the same scale through interpolation or lightweight alignment modules. Coordinate or position encoding is introduced as needed to enhance geometric perception, while maintaining consistency between the numerical domain and dynamic range, reducing error propagation caused by distribution differences and scale mismatches of cross-modal features before fusion.
[0076] Step 5: Cross-modal fusion: Utilizing a deformable cross-attention mechanism, visible light image features and sonar image features are fused to achieve multi-scale adaptive fusion and information complementarity; including:
[0077] Using visible light features as queries and sonar features as keys, a deformable cross-attention approach is used to adaptively sample and aggregate sonar features at learnable reference points, thereby achieving robust alignment and information fusion even in the presence of viewpoint differences, scale biases, or slight misalignments. The fusion layer combines residuals and normalization to stabilize training, and multiple layers can be chained together as needed or combined with simple concatenation + convolution strategies to achieve a balance between performance and overhead.
[0078] Specifically, deformable cross attention generates normalized reference points on the visible light feature grid and predicts the offset and weight for each query position. A preset number of points are sampled from the sonar features and weighted and aggregated. The fused output is then subjected to linear projection, residual connection and layer normalization to stabilize the training.
[0079] The deformable cross-attention is a multi-head structure, and the number of sampling points and scale level are configurable parameters. Preferably, the number of multi-heads is 4-16, the number of sampling points per query is 2-8, and the scale level is 1 or multiple scales.
[0080] Step Six: Reconstruction and Training: The fused features are reconstructed into a denoised and enhanced image using a decoder, and structural fidelity and visual quality are improved through joint optimization using multi-target loss; including:
[0081] The fused features are fed into the decoder, and the spatial resolution is gradually restored through upsampling and convolution to generate a denoised and enhanced target image. During the training phase, multi-objective loss is jointly optimized, including pixel-level reconstruction and perceptual consistency, structural similarity and its multi-scale form, and frequency domain modulation to constrain low and high frequency consistency respectively. Adversarial loss and gradient regularization can be optionally introduced to improve subjective quality and stability. This scheme is trained end-to-end, and fusion and reconstruction can be completed in a single forward pass during the inference phase, which is convenient for engineering deployment and expansion.
[0082] Specifically, the overall model of this implementation is a conditional GAN, which uses a multi-branch fusion generator and a PatchGAN discriminator for adversarial training, and additionally superimposes frequency domain and perceptual class constraints.
[0083] Multi-objective loss function L total Reconstruction loss L recon , countering losses L adv With frequency domain modulation loss L cmfm Linear weighted composition, satisfying:
[0084]
[0085] Where λ recon , λ adv and λ cmfm These are the weight parameters for reconstruction loss, adversarial loss, and frequency domain modulation loss, preferably set to 70.0, 7.0, and 5.0, respectively.
[0086] Cross-modal frequency domain modulation loss separates the low-frequency and high-frequency components of the generated image and the input / reference image using a frequency domain mask.
[0087]
[0088] Where α and β are the weighting parameters for the low-frequency and high-frequency terms, respectively, and the low-frequency term L... lowThe low-frequency graph generated using L1 distance constraints matches the low-frequency input from the camera, while the high-frequency term L... high InfoNCE contrast loss is employed, using generated high frequencies as anchors, with clean camera high frequencies as positive and sonar high frequencies as negative to improve local texture representation. Frequency domain modulation loss L... cmfm Structural and texture constraints are introduced at the low and high frequency levels respectively to ensure that the generated image is consistent with the real sample in terms of spectral characteristics such as illumination and detail. Pixel difference constraints are applied to the low frequency, and contrast loss is applied to the high frequency to make the generated high frequency image close to a clean image and far away from the sonar high frequency.
[0089] The generator model employs a sonar and visible light fusion denoising network: WaveletCNN is used to extract sonar structure and texture features on the sonar side, while Tiny-SwinTransformer is used to extract global semantic features on the visible light side. After the two features are aligned by 1×1 convolution and interpolation, they are input into a deformable cross-attention module for cross-modal fusion. Finally, upsampling and convolutional decoding are used to generate a denoised and enhanced RGB image.
[0090] The discriminator model uses a conditional PatchGANDiscriminator: the image to be judged (a real, clean image or a generated image) and the conditional image (a noisy camera image) are concatenated along the channel dimension, then fed into a multi-layer 4×4 convolutional CNN with BN layers and LeakyReLU, outputting an N×N real / fake rating image, where each position corresponds to the original image. Figure 1 Determine the truth value of a local patch.
[0091] The overall training paradigm is the conditional LSGAN framework: the generator minimizes the reconstruction / perception loss and cross-modal frequency domain modulation loss (low frequencies align with camera input, high frequencies move towards the clean image and away from sonar high frequencies) while "fooling" the PatchGAN discriminator through adversarial loss; the discriminator distinguishes between real clean images and generated images and adds R1 / R2 gradient penalties to keep training stable.
[0092] Example:
[0093] like Figure 1 and Figure 2 As shown, a sonar and camera image fusion denoising method based on cross-modal frequency domain modulation and deformable cross-attention specifically includes:
[0094] Step 1: Data preprocessing steps include: performing slight affine transformation and color jitter on the camera-captured images; matching the corresponding visible light images and sonar images; scaling the visible light images and sonar images to a uniform size, and then mapping them to their respective domains.
[0095] Step 2: Input the sonar image and extract its deep features, such as... Figure 3 As shown, it includes the following steps:
[0096] Step 2-1: Represent the input sonar image in the form of the following formula:
[0097]
[0098] Where H is the number of pixels in the vertical direction of the image, and W is the number of pixels in the horizontal direction of the image.
[0099] Step 2-2: Perform column-direction lifting decomposition on the sonar image formula, as shown in the following formula:
[0100]
[0101] in It is a prediction operator, a learnable local operator of Conv1d→ReLU→Conv1d, which uses even sampling points to predict odd sampling points; It is an update operator, also a learnable operator of Conv1d→ReLU→Conv1d, which can be defined as correcting even sampling points with odd sampling points; and These are the trainable parameters in the network; their low-frequency components are obtained. and high frequency components k and l are the local window radii.
[0102] The filter weights of the prediction and update operators in wavelet decomposition are parameterized by a neural network and adaptively optimized through backpropagation during end-to-end training. This allows the filter kernel weights of the subband decomposition to be automatically adjusted according to the distribution of the input data, thus realizing the learnable characteristics of wavelet transform.
[0103] Steps 2-3: Then, perform row-direction lifting decomposition on the low-frequency component A(x, y) and high-frequency component D(x, y) of each row, as shown in the following formula:
[0104]
[0105]
[0106] in It is a prediction operator (Predict step), which predicts the odd-numbered columns of pixels along the horizontal direction; It is the update operator (Update step), which corrects even-numbered columns of pixels using the prediction residual; k and l are the local window radii, which control the size of the neighborhood. In the algorithm, k=l=2, which corresponds to a 5×1 convolution.
[0107] Obtain the low-frequency branch and high-frequency branches , , Among them, the low-frequency branch This represents the smooth components, contours, and large-scale intensity distribution in the image; High-frequency details representing the direction of travel are "horizontal textures / edges"; It features high-frequency details along the column direction and relatively smoothness along the row direction, mainly corresponding to "vertical texture / edge"; It is a two-way high-frequency detail, namely the oblique / corner / texture noise components.
[0108] Steps 2-4: Extract features for each branch separately. Low-frequency branch. After the first two layers of ResNet18 encoding, followed by two stages of upsampling and skip-layer connections, the output is low-frequency features dominated by structure; three high-frequency branches , , Each feature map is processed through two convolutional layers, one batch normalization (BN) layer, and one ReLU activation function. Then, a spatial attention module is used to suppress noise and extract feature textures, outputting each feature map.
[0109] Steps 2-5: Input low-frequency features and three high-frequency features, and after channel-dimensional splicing, perform a 1×1 convolution fusion for channel compression and information interaction to obtain a sonar feature map with unified semantic expression.
[0110] Steps 2-6: The sonar feature map is processed by a channel-aligned 1×1 convolution (unified to the fusion dimension) and used as a spatial reference to align with the camera feature size, resulting in the sonar aligned feature (B×fusion channel×target height×target width, default fusion channel=256, target height and width aligned with the camera) for subsequent cross-modal fusion.
[0111] Step 3: Visible light image feature extraction, which includes the following steps:
[0112] Step 3-1: Process the preprocessed visible light image Input patch embedding module The image is divided into non-overlapping patches using a convolution operation with a kernel size of 4 and a stride of 4. This operation simultaneously performs spatial downsampling and channel mapping.
[0113]
[0114] Among them, output features The dimension consists of the batch size dimension B, the sequence length N generated after downsampling the input image resolution H×W, and the embedding channel dimension C0. Where:
[0115]
[0116] This corresponds to the initial channel dimension of the Swin Transformer-Tiny (Swin-T) structure. This process is equivalent to dividing the input image into 4×4 image patches and projecting each patch onto a 96-dimensional feature space to form the initial token sequence of the Transformer.
[0117] Step 3-2: Input the features of the l-th layer Divide the window into sections of size K × K, and for each window K, calculate the query (Q), key (K), and value (V) matrices:
[0118]
[0119] in, Input feature matrix The feature matrix of the w-th local window obtained after partitioning; , , These represent the learnable linear projection weight matrices used to generate Q, K, and V, respectively. Multi-head self-attention calculation is then performed within the window:
[0120]
[0121] Where d represents the single-head attention dimension; 𝐵 rel For relative position offset; 𝑀 mask Used to mask features outside the current window region. Within each layer, a standard Transformer layer structure is formed by residual connections and a feedforward network (MLP):
[0122] , ,
[0123] Where LN(⋅) represents Layer Normalization. To enhance cross-window information interaction, a Shifted Window Multi-head Self-Attention (SW-MSA) mechanism is introduced between adjacent layers. Before performing attention calculation, a cyclic shift operation is performed on the input features, with a shift step size 𝑠=𝑀 / 2, where M is the local window divided during multi-head self-attention calculation (default value is 7 in this method):
[0124] ,
[0125] The W-MSA layer and SW-MSA layer are stacked alternately to achieve dual feature modeling, both local and cross-window.
[0126] Step 3-3: At the end of each stage, perform a patch merging operation to halve the resolution and double the number of channels. Let the stage number be s∈{1,2,3}, then the patch merging process is as follows:
[0127]
[0128] Immediately adjacent 2×2 tokens (i.e. After splicing, linear layer mapping reduces the spatial size by a factor of 2, while increasing the channel dimension by a factor of 2. In the Swin-T structure, the correspondence between channels and scale is as follows:
[0129]
[0130] Steps 3-4: The final output feature, denoted as Z4, is restored to a two-dimensional mesh in spatial order.
[0131]
[0132] To achieve fusion with the sonar branch, a 1×1 convolution is first used for channel projection, compressing the dimension from 768 to the fusion dimension D:
[0133]
[0134] Then, bilinear interpolation is used to adjust the feature map to the same spatial resolution (H) as the sonar feature map. s W s ):
[0135]
[0136] at this time, This is a visible light branch feature representation aligned with sonar features, used as input for subsequent acousto-optic fusion networks.
[0137] Step 4: Alignment of visible light image features with sonar image features, including:
[0138] Step 4-1: Determine the target scale and define the sonar feature tensor. for:
[0139]
[0140] Where B is the batch size, Cs is the number of sonar feature channels, and Hs and Ws are the spatial resolution of the sonar features.
[0141] First, based on the spatial resolution (H) of sonar features... s W s As a unified target alignment scale, its output is:
[0142]
[0143] Step 4-2: Align the sonar and visible light sides in the channel dimension. First, to ensure the sonar feature channel dimensions are consistent with the fusion dimension D, perform a 1×1 convolutional channel projection on the sonar features:
[0144]
[0145] After this step, the output sonar-side feature shape is B×D×H. s ×W s , where D is the fusion dimension.
[0146] The characteristic shape of the visible light backbone network is B×H. c ×W c ×C c To maintain consistency with sonar features, the feature format is first converted to NCHW arrangement, then 1×1 convolution is used for channel projection, and the ReLU activation function is selected.
[0147]
[0148] To enhance the nonlinear representation of features, the output feature dimension is B×D×Hc×Wc, resulting in visible light features with consistent channel dimensions.
[0149] Step 4-3: Align the sonar features and visible light features in terms of spatial scale. Since there is a difference in resolution between the camera and sonar features, the camera features are adjusted to the target scale (Hs, Ws) using bilinear interpolation. This method is computationally efficient, provides smooth interpolation, effectively reduces edge distortion, and ensures that the interpolated camera and sonar features are completely consistent in spatial dimensions.
[0150] Step 5: Cross-modal fusion, such as... Figure 4 As shown, the steps are as follows:
[0151] Step 5-1: The input is processed through the position encoding coordinate channel module for serialization and feature arrangement. The input consists of visible light features and sonar features obtained from the aforementioned alignment module:
[0152]
[0153] Where B is the batch size, D is the number of channels (embedding dimension), and H and W are the spatial dimensions. Before fusion, the two-dimensional feature maps are flattened and rearranged into a sequence to meet the input requirements for attention computation. The spatial dimension H×W is stretched to a sequence length N=H×W, and the channel vector at each position serves as the embedding representation of a single token.
[0154]
[0155] Step 5-2: To establish cross-modal associations in a unified feature space, perform independent linear projection operations on the input sequences to generate query vector Q, key vector K, and value vector V. This process can be represented as:
[0156]
[0157] Among them W Q W K W V ∈R D×D This is a learnable parameter matrix. Establishing cross-modal relationships allows the model to adaptively learn mapping weights between features from different modalities, achieving finer-grained feature matching and fusion. In the subsequent deformable cross-attention calculation, the module mainly utilizes Q and V for sampling and weighted aggregation.
[0158] Step 5-3: To implement the geometrically perceptive feature fusion mechanism, the system generates a normalized reference point coordinate grid based on the spatial resolution (H, W) of the feature map. Using the pixel center as a reference, the two-dimensional coordinates are mapped to the interval [0,1], serving as the spatial reference for attention sampling. The calculation of the reference points can be expressed as:
[0159]
[0160] Finally, the reference point tensor is obtained. Simultaneously record the spatial shape parameters spatial_shapes=[(H,W)] and the level index level_start_index=[0].
[0161] Step 5-4: Use a multi-head deformable cross-attention module to fuse the aligned visible light image features with the sonar image features. For each query point q... i The system at the reference point From sonar features at several nearby sampling locations (number P) Extracting the local value vector v i,j And through attention weight α i,j Weighted aggregation is performed to achieve dynamic fusion across modalities. This calculation process can be simplified as follows:
[0162]
[0163] Where P is the number of sampling points for each query. This module can achieve flexible matching of acoustic and optical features at the spatial structure and semantic levels while maintaining high computational efficiency.
[0164] Step 5-5: The intermediate result F∈R obtained by fusionB×N×D Local fusion vector of all query points The sequence is cascaded along the sequence axis, and after passing through an output linear layer and random deactivation processing, it is added to the input query Q via a residual connection, followed by layer normalization. This structure maintains feature flow while improving training stability. Its computational form is as follows:
[0165]
[0166] The fused sequence F′∈R is then... B×N×D The features are rearranged back into a two-dimensional form according to the original spatial dimensions (H,W) to restore the spatial structure and adapt it to subsequent decoding modules. This process can be represented as:
[0167]
[0168] Output fusion feature map FusionFeature∈R B×D×H×W It integrates multimodal information from visible light and sonar, and can be directly used as input for decoding and reconstruction or task recognition.
[0169] Step Six: The steps for fusing the feature decoding and reconstruction module and its training methods include the following:
[0170] Step 6-1: Input is the fused feature map FusionFeature∈R B×D×H×W First, upsampling is performed through a deconvolution (transposed convolution) layer to expand the feature space size. The deconvolution kernel size is 4×4, stride is 2, padding is 1, and the number of output channels is 128. Batch normalization (BN) and ReLU activation function are added after convolution. After this process, the upsampled feature UpFeature∈𝑅𝐵×128×2𝐻×2𝑊 is obtained, which can be represented as:
[0171]
[0172] This section performs spatial expansion and preliminary texture restoration of the fused features. Subsequently, the upsampled feature map is further processed through a convolutional reconstruction layer (Conv2d) to generate an intermediate reconstructed image. The parameters of this convolutional layer are: kernel 7×7, stride 1, padding 3, and number of output channels 3. At the output, a hyperbolic tangent function (Tanh) is used for non-linear compression, limiting the output value to the range [−1, 1]. This process can be represented as:
[0173]
[0174] The output image size is R B×3×2H×2W It is twice the resolution of the original input image.
[0175] Step 6-2: Intermediate Reconstructed Image Imgmid Adjusted to the camera input resolution (H) using bilinear interpolation. c W c This allows for subsequent quantitative evaluation and visualization. During interpolation, input pixels are treated as the center points of consecutive cells, rather than strict corner points, and output pixels correspond to the relative proportional positions of input cells to reduce edge distortion and scale errors.
[0176]
[0177] The final reconstructed image output is Imgfinal∈R B×3×Hc×Wc The pixel values remain in the [−1,1] range. To facilitate display and evaluation, the [−1,1] range is mapped back to the commonly used image range [0,1] through a linear transformation:
[0178]
[0179] Obtain the image Img for visualization vis ∈R B×3×Hc×Wc It is used for quantitative indicator calculation and subjective quality assessment.
[0180] Step 6-3: In the final stage of training, the entire algorithm comprehensively optimizes the generator by using reconstruction loss, adversarial loss, and frequency domain modulation loss, thereby achieving performance improvements in three aspects: structural fidelity, visual naturalness, and spectral consistency.
[0181]
[0182]
[0183] in, Let G represent the denoised image generated by the generator based on the input sonar image A and visible light image B, and let C represent the clear underwater visible light image as supervision. This represents the probability output score by which the discriminator evaluates the authenticity of the forged image G(A, B) generated by the generator, using the sonar image A as a condition. Reconstruction loss L recon By calculating the differences between the generated image and the real image at the pixel level or perceptual feature level, the spatial consistency of the reconstruction result with the reference target is ensured. Adversarial loss L adv Built on the LSGAN framework, the output of the constraint generator can deceive the discriminator, making the generated results more realistic.
[0184]
[0185] Here, α and β are the weighting parameters for the low-frequency and high-frequency terms, respectively.
[0186] As shown in Figure I, the input tensor X undergoes a two-dimensional FFT with center translation over the Tanh domain [-1,1], using a concentric circular mask M with radius r. low Low-pass and high-pass filters are performed, and the real part is taken as the low / high frequency component. The calculation method is as follows:
[0187] ,
[0188] Among them, low frequency items The low-frequency components of the graph generated using L1 distance constraints are consistent with the low-frequency components of the camera input, while the high-frequency components... InfoNCE contrast loss is employed, using generated high frequencies as anchors, with clean camera high frequencies as positive and sonar high frequencies as negative to improve local texture representation. Frequency domain modulation loss L... cmfm Structural and texture constraints are introduced at the low and high frequency levels respectively to ensure that the generated image is consistent with the real sample in terms of spectral characteristics such as illumination and detail.
[0189] The overall optimization objective of the generator and the total loss of the discriminator can be defined as follows:
[0190]
[0191] Where, λ recon , λ adv and λ cmfm These are the weight parameters for reconstruction loss, adversarial loss, and frequency domain modulation loss, preferably set to 70.0, 7.0, and 5.0, respectively.
[0192]
[0193] in, It is the discrimination loss of the discriminator when judging real, clean images. It is the discriminator's loss in distinguishing fake images generated by the generator. This is the gradient penalty, used to stabilize the convergence process of adversarial training and prevent gradient explosion or vanishing. The discriminator differs from the generator in that it stabilizes the convergence process by minimizing the mean squared error between real and fake samples and introducing a gradient penalty term. Subsequently, parameter updates employ an alternating optimization approach. In each training iteration, the system first freezes the discriminator parameters to update the generator, and then freezes the generator parameters in reverse to optimize the discriminator. The update rule is as follows:
[0194]
[0195] Where η G With η D The learning rates for the generator and discriminator are respectively. and This represents the trainable parameters of the network, which are the generator and discriminator. and This represents the gradient of the loss function with respect to its respective parameters. To ensure numerical stability and gradient continuity, mixed-precision backpropagation and gradient clipping techniques are used during training. After each training round, the system calculates quantitative metrics including PSNR, SSIM, MS-SSIM, UCIQE, and UIQM on the validation set and saves the current model state to a checkpoint file for subsequent recovery and continuous optimization. Through this joint optimization process, the model gradually learns the reconstruction mapping relationship under acoustic and optical dual-modal feature constraints, achieving high-fidelity restoration of underwater images and enhanced multimodal consistency.
[0196] like Figure 5 and Figure 6 As shown, the method proposed in this invention has achieved significant technical effects in both underwater image restoration quality and physical feature decoupling. Figure 5 Quantitative comparative results of ablation experiments on the three core modules are presented on a standard dataset. PSNR, SSIM, and MS-SSIM were used as objective evaluation metrics. Removing CMFM resulted in a decline in all metrics (especially PSNR), indicating that joint frequency domain constraints are crucial for correcting color cast and suppressing turbidity degradation. Removing the deformable cross-attention fusion module led to a significant decrease in structural similarity metrics due to the inability to dynamically compensate for geometric viewpoint biases between multiple sensors. Replacing the Wavelet with a regular convolutional branch weakened the model's ability to suppress sonar speckle noise, suppressing both the signal-to-noise ratio and texture detail fidelity of the reconstructed image.
[0197] Figure 6 This section presents the visualization results of the high-low frequency step-by-step decoupling of underwater images in the cross-modal frequency domain modulation module. The input degraded underwater image is processed by 2D FFT, and then spectral slicing is performed using a concentric circle Low Mask with a set radius (r=10) and its complement High Mask. The low-frequency image obtained after the inverse transform perfectly removes high-frequency speckle and impurities, preserving smooth macroscopic illumination and large-scale background contours; while the high-frequency image accurately captures the fine edges of underwater targets, rigging lines, and other key topological textures. This demonstrates that the algorithm can achieve independent representation and targeted supervised optimization of structure and details at the frequency domain level.
[0198] It should be noted that in step 2, the learnable wavelet convolution that the sonar input tensor passes through can be replaced with a lightweight three-layer convolutional neural network with a kernel size k of 3, a stride s of 2, and including BN layers and ReLU activation functions. This can improve the speed of the sonar image feature extraction part, and will have better performance when facing hardware devices with lower computing power. Other steps and parameters remain unchanged.
[0199] It should be noted that in step 3, the Tiny-SwinTransformer network through which the camera tensor input from the visible light side passes can be replaced with a more lightweight depthwise separable convolutional network. This significantly simplifies the computation graph, resulting in lower latency and smaller memory usage, making it easier to deploy and run stably on edge devices; at the same time, it can converge without relying on large-scale pre-trained weights, facilitating rapid iteration in small data scenarios.
[0200] This application discloses a sonar and camera image fusion denoising method based on cross-modal frequency domain modulation and deformable cross-attention. The method first performs pairwise preprocessing and normalization on the sonar and visible light images to ensure consistent data resolution and distribution. Then, learnable wavelet decomposition is performed on the sonar image to separate low-frequency structural and high-frequency detail features, which are encoded separately to achieve noise suppression and structural enhancement. For the visible light image, a visual Transformer with multi-head self-attention using a window is employed for hierarchical feature extraction, obtaining a multi-layered representation with global semantics and local texture. After unified projection, the two feature paths are used in the fusion stage with visible light features as queries and sonar features as keys. Deformable cross-attention is used for adaptive sampling and aggregation at learnable reference points, thereby achieving accurate alignment and information complementarity of cross-modal features. The fusion result is progressively upsampled and reconstructed through a decoder, outputting a clear, denoised, and enhanced image. The training phase employs joint optimization using reconstruction, adversarial, and frequency domain modulation multi-objective losses to balance low-frequency consistency and high-frequency fidelity. This method is compact, robust, and can achieve high-quality acoustic-optical fusion and image reconstruction in complex underwater environments.
Claims
1. A sonar and camera image fusion denoising method based on cross-modal frequency domain modulation and deformable cross-attention, characterized in that, The method includes: Step 1: Perform pairwise organization and consistency processing on sonar images and visible light images; Step 2: Extract scale-consistent deep sonar features from the sonar images processed in Step 1; Step 3: Extract multi-scale semantic and texture features from the visible light image processed in Step 1 to obtain semantically enhanced deep visible light features; Step 4: The sonar deep features and visible light deep features obtained in Steps 2 and 3 are unified in terms of channel and spatial scale, and geometrically consistent cross-modal feature alignment is achieved. Step 5: The visible light image features and sonar image features obtained in Step 4 are fused using a deformable cross-attention mechanism to achieve multi-scale adaptive fusion and information complementarity, including: Using visible light features as queries and sonar features as keys, deformable cross-attention is used to adaptively sample and aggregate sonar features at learnable reference points, thereby achieving cross-modal feature alignment and information complementarity. The deformable cross attention generates normalized reference points on the visible light feature grid, predicts the offset and weight for each query position, samples a preset number of points from the sonar features for weighted aggregation, and then outputs the fusion result through linear projection, residual connection and layer normalization. Step 6: The fusion result obtained in Step 5 is gradually upsampled and reconstructed by the decoder to output a clear image with enhanced noise reduction; during the training phase, reconstruction, adversarial and frequency domain modulation multi-target loss are jointly optimized.
2. The sonar and camera image fusion denoising method based on cross-modal frequency domain modulation and deformable cross-attention as described in claim 1, characterized in that: Step 2 specifically includes: extracting scale-consistent deep sonar features from sonar images using learnable wavelet decomposition. The learnable wavelet decomposition adopts a two-dimensional separable lifting structure, decomposing sequentially along both column and row directions to output a low-frequency LL subband and three directional subbands of high frequencies: LH, HL, and HH. The filtering weights of the prediction and update operators in the learnable wavelet decomposition are parameterized by a neural network and adaptively optimized through backpropagation during end-to-end training. The low-frequency LL subband is encoded and decoded and skip connections are used to enhance structural representation, outputting low-frequency features dominated by structure; the high-frequency LH, HL, and HH directional subbands are respectively passed through two convolutional layers, one BN layer and one ReLU activation function, and then a spatial attention module is used to suppress noise and extract feature texture, outputting three high-frequency features. Low-frequency features and three high-frequency features are spliced together in the channel dimension and compressed to the preset fusion dimension by 1×1 convolution to obtain sonar deep features with consistent scale.
3. The sonar and camera image fusion denoising method based on cross-modal frequency domain modulation and deformable cross-attention as described in claim 1, characterized in that: Step 3 specifically includes: using a visual Transformer with window self-attention and hierarchical representation capabilities to perform hierarchical encoding on the visible light image, wherein the visual Transformer includes patch embedding and progressive downsampling modules; modeling within a local window while preserving global dependencies across windows, obtaining deeper features with stronger semantics through progressive downsampling and channel expansion; and selecting an appropriate hierarchical output as the fusion input according to the target task to take into account both texture details and global context.
4. The sonar and camera image fusion denoising method based on cross-modal frequency domain modulation and deformable cross-attention as described in claim 1, characterized in that: Step 4 specifically includes: projecting the two-way features of sonar deep features and visible light deep features obtained in steps 2 and 3 onto a unified fusion channel dimension, and unifying the spatial resolution to the same scale through interpolation or lightweight alignment modules; introducing coordinate or position encoding to enhance geometric perception, while maintaining consistency between the numerical domain and dynamic range, and reducing error propagation caused by distribution differences and scale mismatches of cross-modal features before fusion.
5. The sonar and camera image fusion denoising method based on cross-modal frequency domain modulation and deformable cross-attention as described in claim 1, characterized in that: Step 5 specifically includes: Step 5.1: The visible light image features and sonar image features obtained in Step 4 are serialized and arranged through the position encoding coordinate channel module; the spatial dimension H×W is stretched to the sequence length N=H×W, and the channel vector of each position is used as the embedding representation of a single token; Step 5.2: To establish cross-modal associations in a unified feature space, perform independent linear projection operations on the input sequences to generate query vector Q, key vector K, and value vector V; Step 5.3: To realize the feature fusion mechanism of geometric perception, a normalized reference point coordinate grid is generated according to the spatial resolution (H,W) of the feature map; the two-dimensional coordinates are mapped to the interval [0,1] with the pixel center as the reference, as a spatial reference for attention sampling; finally, the reference point tensor is obtained, and the spatial shape parameters and hierarchical index are recorded simultaneously. Step 5.4: Use a multi-head deformable cross-attention module to fuse the visible light image features and sonar image features obtained in Step 4. For each query point, extract local value vectors from the sonar features at several sampling locations near the reference point, and aggregate them through attention weights to achieve cross-modal dynamic fusion. Step 5.5: The intermediate result F obtained by fusion is processed by the output linear layer and random deactivation, then added to the input query Q through a residual connection, and layer normalization is performed. The calculation formula is as follows: The fused sequence F′ is then rearranged back into a two-dimensional feature form according to its original spatial dimensions to restore the spatial structure and adapt it to the subsequent decoding module, outputting the fused feature map.
6. The sonar and camera image fusion denoising method based on cross-modal frequency domain modulation and deformable cross-attention as described in claim 1, characterized in that: Step 6 specifically includes: Step 6.1: The fusion result obtained in Step 5 is first upsampled through a deconvolution layer to expand the feature space size, and then batch normalization and ReLU activation function are applied after convolution to obtain upsampled features; Subsequently, the upsampled features are further processed through a convolutional reconstruction layer to generate an intermediate reconstructed image; Step 6.2: The intermediate reconstructed image generated in step 6.1 is adjusted to the resolution of visible light features through bilinear interpolation. During the interpolation process, the input pixels are regarded as the center points of continuous cells, and the output pixels correspond to the relative proportional positions of the input cells. The pixel values of the final reconstructed image output are in the numerical domain. The numerical domain is mapped back to the image domain to obtain a visualization image, which is used for quantitative index calculation and subjective quality assessment. Step 6.3: During the training phase, reconstruction loss, adversarial loss, and frequency domain modulation loss are used to optimize the generator, thereby improving performance in terms of structure fidelity, visual naturalness, and spectral consistency.
7. The sonar and camera image fusion denoising method based on cross-modal frequency domain modulation and deformable cross-attention as described in claim 6, characterized in that: Step 6.3 includes: The overall optimization objective of the generator is defined as: Among them, reconstruction loss L recon By calculating the differences between the generated image and the real image at the pixel level or perceptual feature level, the adversarial loss L... adv Based on the LSGAN framework, the output of the constraint generator can deceive the discriminator, λ recon , λ adv and λ cmfm These are the weight parameters for reconstruction loss, adversarial loss, and frequency domain modulation loss, respectively. Frequency domain modulation loss : Where α and β are the weighting parameters for the low-frequency and high-frequency terms, respectively; and These are low-frequency and high-frequency terms, respectively. and The calculation method is as follows: the input tensor X is subjected to a two-dimensional FFT and center translation in the Tanh domain, using a concentric circular mask M with radius r. low Perform low-pass and high-pass separation, and take the real part as the low-frequency component and high-frequency component: , Among them, low frequency items The low-frequency components of the graph generated using L1 distance constraints are consistent with the low-frequency components of the camera input, while the high-frequency components... InfoNCE contrast loss is used, with the generated high frequency as the anchor; frequency domain modulation loss L cmfm Structural and texture constraints are introduced at the low and high frequency levels respectively to ensure that the generated image is consistent with the real sample in terms of spectral characteristics such as illumination and detail. The total loss of the discriminator is defined as: in, It is the discrimination loss of the discriminator when judging real, clean images. It is the discriminator's loss in distinguishing fake images generated by the generator. This is the gradient penalty, used to stabilize the convergence process of adversarial training and prevent gradient explosion or vanishing. Subsequently, parameter updates employ an alternating optimization approach. In each training iteration, the system first freezes the discriminator parameters to update the generator, and then freezes the generator parameters in reverse to optimize the discriminator. The update rule is as follows: Where, η G With η D These are the learning rates for the generator and the discriminator, respectively. and This represents the trainable parameters of the network, which are the generator and discriminator. and This represents the gradient of the loss function with respect to its respective parameters.
8. A sonar and camera image fusion denoising system based on cross-modal frequency domain modulation and deformable cross-attention, characterized in that: The system has a program module corresponding to the steps of any one of claims 1-7, and executes the steps in the sonar and camera image fusion and denoising method based on cross-modal frequency domain modulation and deformable cross attention when it is run.
9. A computer device, characterized in that: It includes a memory and a processor, wherein the memory stores a computer program, and when the processor runs the computer program stored in the memory, it performs the steps of the sonar and camera image fusion denoising method based on cross-modal frequency domain modulation and deformable cross-attention as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that: The storage medium is used to store a computer program that executes a sonar and camera image fusion and denoising method based on cross-modal frequency domain modulation and deformable cross-attention, as described in any one of claims 1-7.