Feature-preserving gaussian blind denoising method and system

CN122820481APending Publication Date: 2026-09-25XI AN JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610952173.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-29
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0012]本发明所要解决的技术问题在于针对上述现有技术中的不足,提供一种特征保真的高斯盲去噪方法及系统,用于解决现有图像去噪任务中长距离依赖建模与局部细节提取难以兼顾、多尺度特征融合不充分,以及过度平滑导致局部不变特征丢失的技术问题

Benefits of technology

一种特征保真的高斯盲去噪方法,构建多尺度U型高斯盲去噪器、通过跨层注意力融合模块融合多尺度特征、在跳跃连接处设置空频特征保真适配器、执行双阶段训练以及利用训练完成的去噪器进行单次前向推理。能够将基础去噪能力、多尺度特征融合能力和局部特征保真能力统一到同一去噪框架中。多尺度U型结构有利于同时利用浅层纹理信息和深层结构信息;并行轴向Transformer-卷积模块用于兼顾局部特征处理和长距离依赖建模;空频特征保真适配器在跳跃连接处对传递特征进行处理,有助于减少残余噪声并保留局部结构;双阶段训练则使主干去噪能力与特征保真能力分阶段优化,从而提高去噪图像在视觉质量和机器视觉任务中的适用性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122820481A_ABST
    Figure CN122820481A_ABST
Patent Text Reader

Abstract

The application discloses a feature-fidelity-preserving Gaussian blind denoising method and system, and belongs to the technical field of image processing and computer vision. The method constructs a Gaussian blind denoiser with a multi-scale U-shaped structure, which comprises a feature encoder, a cross-layer attention fusion module and a feature decoder, and a parallel axial Transformer-convolution module is adopted for the feature extraction unit; multi-scale features are fused through the cross-layer attention fusion module; a space-frequency feature fidelity adapter is arranged at the skip connection of the feature encoder to the feature decoder to obtain space-frequency fidelity features; two-stage training is performed on the Gaussian blind denoiser; an input image containing unknown level Gaussian noise is input into the trained Gaussian blind denoiser, and a denoised image is output through single forward inference. The application can remove Gaussian noise while preserving local key features such as edges, textures and corners, and improve the applicability of the denoised image in feature matching, image registration and image stitching tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of image processing and computer vision technology, and specifically relates to a Gaussian blind denoising method and system that preserves feature fidelity. Background Technology

[0002] Image denoising is a fundamental task in image processing and computer vision, aiming to recover a clear image from a noisy image that is as close as possible to the original noise-free image. In actual imaging processes, due to factors such as sensor thermal noise, insufficient illumination, transmission interference, and limitations of imaging equipment performance, images are often mixed with Gaussian noise of varying intensities. This type of noise destroys edges, textures, corners, and local structural information in the image, thus affecting subsequent machine vision tasks such as feature extraction, feature matching, image registration, image stitching, and object recognition. Therefore, how to remove Gaussian noise of unknown levels while preserving key local features of the image is a crucial problem that needs to be solved in the field of image denoising.

[0003] In recent years, with the development of deep neural networks, image denoising methods based on convolutional neural networks (CNNs) have gradually replaced traditional filtering methods, becoming the mainstream technology in image denoising tasks. CNNs can learn the mapping relationship between noisy and clean images through multiple convolutional operations, and to a certain extent improve the visual quality of denoised images. However, traditional CNNs mainly rely on fixed-size convolutional kernels for local neighborhood feature extraction. Their receptive field is limited by the number of network layers and the kernel size, resulting in insufficient ability to model long-distance spatial dependencies and non-local structural associations in images. When images contain large-scale texture continuation, long edge structures, or repetitive structures, relying solely on local convolutions is insufficient to fully express these global or long-distance correlation features, easily leading to structural breaks, texture discontinuities, or insufficient detail recovery in the denoised image.

[0004] To enhance long-range dependency modeling capabilities, some existing techniques have attempted to incorporate the Transformer architecture into image denoising tasks. The self-attention mechanism in the Transformer can establish relationships between different spatial locations within an image, helping to capture global structural information. However, standard two-dimensional self-attention typically requires calculating the attention matrix over a large spatial area, resulting in high computational and memory overhead, especially in high-resolution image denoising scenarios, which can lead to slow inference speeds and high deployment costs. Furthermore, relying solely on global attention mechanisms may weaken high-frequency details such as local textures, edges, and corners, making the denoising result visually smoother but unsuitable for subsequent machine vision tasks.

[0005] In terms of network structure, existing deep denoising models typically employ a U-shaped encoder-decoder architecture. This architecture extracts multi-scale features through an encoder, then restores image resolution level by level through a decoder, using skip connections to pass shallow detail features. This type of structure can, to some extent, balance low-level texture features and high-level semantic features. However, traditional U-shaped networks often use simple concatenation, element-wise addition, or direct skip connections when fusing multi-scale features, failing to fully consider the information correlation and contribution differences between features at different levels. Since shallow features are more focused on local details such as texture and edges, while high-level features are more focused on abstract structures and global semantics, direct, simple fusion may result in redundant noise entering the decoding stage with skip connections, or the complementary information between features at different levels may not be fully utilized, thus limiting denoising performance and structural recovery capabilities.

[0006] Furthermore, most existing denoising networks prioritize pixel-level reconstruction quality as their primary optimization objective, employing methods such as L1 loss, L2 loss, or mean squared error loss to constrain the pixel differences between the denoised and clean images. While this optimization approach effectively reduces noise residue and improves the overall visual smoothness of the denoised image, in scenarios with strong noise or blind denoising, the model often tends to reduce pixel errors by smoothing high-frequency regions. As a result, although the denoised image appears cleaner subjectively, edge textures, corner points, local gradients, and fine-grained structures may be weakened or even lost, leading to a decrease in the number and stability of extracted local invariant feature points such as SIFT and ORB.

[0007] For images intended solely for human viewing, the aforementioned smoothing phenomenon may still be acceptable; however, for image denoising tasks aimed at machine vision, preserving key local features is crucial. For example, in image registration and image stitching tasks, algorithms typically rely on corner points, edges, texture blocks, and their corresponding feature descriptors for matching. If the denoising process disrupts these local structures, it can lead to a reduction in the number of effective matching points, an increase in the mismatch rate, and a decrease in the number of interior points, thereby affecting registration accuracy and stitching stability. Therefore, existing denoising methods that primarily target visual quality or pixel-level errors are insufficient to fully meet the demands of subsequent machine vision tasks for faithful preservation of local features.

[0008] Furthermore, during the feature transmission process of skip connections, while shallow features contain rich edges, textures, and local details, they may also carry a significant amount of residual noise. Existing U-shaped denoising networks typically pass shallow features directly to the decoder, lacking a fine-grained filtering and reshaping mechanism for skip connection features. On the one hand, if too much shallow information is retained, residual noise may be carried into the decoding process; on the other hand, if shallow information is excessively suppressed, useful local high-frequency information for feature matching and structure recovery may be lost. Therefore, how to distinguish noise components from useful structural features at skip connections and how to co-modulate spatial structural information and frequency components remains an unresolved problem in existing technologies.

[0009] In terms of training strategies, existing denoising models typically train pixel-level denoising loss directly in conjunction with perceptual loss and feature loss. Since pixel-level loss primarily drives the model to reduce overall pixel differences, while feature-level loss focuses more on edges, textures, and high-frequency structures, these two types of losses may conflict in their optimization directions. Directly optimizing both simultaneously throughout the entire network can lead to the backbone network failing to learn a stable denoising map while simultaneously failing to effectively preserve key local features. Therefore, existing training methods struggle to establish a clear division between basic denoising capabilities and local feature preservation capabilities, easily resulting in an unstable trade-off between denoising effectiveness and feature preservation.

[0010] In summary, the existing technology has at least the following problems: Traditional convolutional denoising networks are limited by local receptive fields, making it difficult to fully model long-range dependencies and non-local structural associations in images; While Transformer-based denoising structures have global modeling capabilities, standard self-attention computation has high overhead and may weaken local high-frequency details. Existing U-shaped denoising networks have relatively simple multi-scale feature fusion methods, making it difficult to adaptively aggregate effective information from features at different levels; The shallow features transmitted by skip connections contain both local details and residual noise. Existing methods lack a mechanism for joint spatial and frequency domain modulation of skip connection features. Existing denoising models mostly focus on pixel-level error or visual quality as the main optimization target, which can easily lead to excessive smoothing of local key features such as edges, textures, and corners. When pixel-level denoising targets and feature-preserving targets are directly trained together, gradient conflicts may occur, making it difficult to stably balance basic denoising capabilities and local key feature preservation capabilities. While denoised images may have better subjective visual effects, their usability in machine vision tasks such as SIFT, ORB feature extraction, image registration, and image stitching may still decrease.

[0011] Therefore, there is an urgent need for a feature-preserving Gaussian blind denoising method and denoiser that can handle unknown levels of Gaussian noise while taking into account long-distance dependency modeling, local detail extraction, multi-scale feature adaptive fusion, and spatio-frequency joint reshaping of skip connection features. Furthermore, a reasonable training strategy should be used to reduce the conflict between pixel-level denoising and feature-preserving objectives, thereby improving the applicability of denoised images in subsequent machine vision tasks. Summary of the Invention

[0012] The technical problem to be solved by the present invention is to provide a Gaussian blind denoising method and system that preserves feature fidelity, in order to address the shortcomings of the prior art. This method and system solve the technical problems in existing image denoising tasks, such as the difficulty in balancing long-distance dependency modeling and local detail extraction, insufficient multi-scale feature fusion, and loss of local invariant features due to excessive smoothing.

[0013] The present invention adopts the following technical solution: A feature-preserving Gaussian blind denoising method, characterized by comprising the following steps: S1. Construct a Gaussian blind denoiser. The Gaussian blind denoiser adopts a multi-scale U-shaped structure. The Gaussian blind denoiser includes a feature encoder, a cross-layer attention fusion module, and a feature decoder. The feature extraction units of the feature encoder and the feature decoder adopt parallel axial Transformer-convolution modules. S2. The multi-scale features output by the feature encoder are fused using the cross-layer attention fusion module to obtain fused features; S3. A space-frequency feature fidelity adapter is set at the skip connection where the feature encoder transmits features to the feature decoder. The space-frequency feature fidelity adapter processes the skip connection features to obtain space-frequency fidelity features, and then transmits the space-frequency fidelity features to the feature decoder. The space-frequency feature fidelity adapter includes a spatial modulation branch, a frequency modulation branch, and a space-frequency feature fusion branch. The spatial modulation branch generates a spatial attention mask based on the skip connection features and obtains spatial modulation features based on the spatial attention mask. The frequency modulation branch converts the skip connection features to the frequency domain and obtains frequency modulation features based on the frequency domain features. The space-frequency feature fusion branch fuses the spatial modulation features and the frequency modulation features to obtain the space-frequency fidelity features. S4. Perform two-stage training on the Gaussian blind denoiser. In the first stage, the spatial frequency feature fidelity adapter is not introduced, and the backbone network of the Gaussian blind denoiser is pre-trained using L1 loss. In the second stage, the backbone network weights obtained in the first stage are loaded and frozen. The spatial frequency feature fidelity adapter is embedded at the skip connections, and the spatial frequency feature fidelity adapter is trained using a loss function that combines L1 reconstruction loss and perceptual loss based on shallow feature maps of the pre-trained VGG network. S5. Input the input image containing unknown level Gaussian noise into the trained Gaussian blind denoiser, and output the denoised image through a single forward inference.

[0014] Preferably, the parallel axial Transformer-convolution module reduces the dimensionality of the module input features and splits the dimensionality-reduced module input features into convolutional sub-features and Transformer sub-features in the channel dimension. The convolutional sub-features are processed by the convolutional branch, the Transformer sub-features are processed by the Transformer branch, and the output features of the convolutional branch and the output features of the Transformer branch are fused.

[0015] Preferably, the convolutional branch uses deformable convolution to process the convolutional sub-features, and the deformable convolution adjusts the convolutional sampling position by learning coordinate offset and modulation scalar.

[0016] Preferably, the Transformer branch uses an axial attention mechanism to process the Transformer sub-features, the axial attention mechanism including row direction attention calculation and column direction attention calculation; the Transformer branch also performs convolution processing on the Transformer sub-features through depthwise convolution, and performs gating processing on the information flow of the Transformer sub-features through a dual-gate feedforward network.

[0017] Preferably, in step S2, input features at multiple scales are concatenated along the channel dimension, and a query vector, a key vector, and a value vector are generated through convolution operations. A matrix product is calculated based on the query vector and the key vector. The matrix product is normalized by a scaling factor and then multiplied by the value vector. The multiplication result is added to the input features of the cross-layer attention fusion module through a residual connection to obtain the fused features.

[0018] Preferably, the query vector, the key vector, and the value vector are generated by 1×1 convolution and 3×3 depthwise separable convolution.

[0019] Preferably, in step S3, the spatial modulation branch linearly projects the skip connection feature and combines it with depthwise separable convolution to generate the spatial attention mask. After the spatial attention mask is activated by Sigmoid, it is multiplied element-wise with the feature to be modulated in the spatial modulation branch to obtain the spatial modulation feature.

[0020] Preferably, in step S3, the frequency modulation branch transforms the skip connection feature to the frequency domain through a fast real Fourier transform to obtain a frequency domain complex tensor as the frequency domain feature. The amplitude of the frequency domain complex tensor is extracted, and a spectral modulation mask is generated based on the amplitude through a convolutional network. After modulating the frequency domain complex tensor with the spectral modulation mask, the frequency modulation feature is obtained through an inverse Fourier transform.

[0021] Preferably, in step S4, in the first stage, a progressive learning strategy is used to pre-train the backbone network, the progressive learning strategy including gradually transitioning from smaller image patches and larger batch sizes to larger image patches and smaller batch sizes; in the second stage, the pre-trained VGG network is a VGG-19 network, and the perceptual loss is calculated based on the shallow output feature map of the VGG-19 network.

[0022] Secondly, embodiments of the present invention provide a feature-fidelity Gaussian blind denoising system, comprising: A Gaussian blind denoiser construction module is used to construct a Gaussian blind denoiser. The Gaussian blind denoiser adopts a multi-scale U-shaped structure and includes a feature encoder, a cross-layer attention fusion module, and a feature decoder. The feature extraction units of the feature encoder and the feature decoder adopt parallel axial Transformer-convolution modules. A multi-scale feature fusion module is used to fuse the multi-scale features output by the feature encoder through the cross-layer attention fusion module to obtain fused features; The space frequency feature fidelity processing module includes a space frequency feature fidelity adapter, which is disposed at the jump connection where the feature encoder transmits features to the feature decoder. The adapter is used to process the jump connection features to obtain space frequency fidelity features and transmit the space frequency fidelity features to the feature decoder. The space-frequency feature fidelity adapter includes a spatial modulation branch, a frequency modulation branch, and a space-frequency feature fusion branch. The spatial modulation branch is used to generate a spatial attention mask based on the skip connection feature and obtain a spatial modulation feature based on the spatial attention mask. The frequency modulation branch is used to convert the skip connection feature to the frequency domain and obtain a frequency modulation feature based on the frequency domain feature. The space-frequency feature fusion branch is used to fuse the spatial modulation feature and the frequency modulation feature to obtain the space-frequency fidelity feature. A two-stage training module is used to perform two-stage training on the Gaussian blind denoiser. In the first stage, the spatial frequency feature fidelity adapter is not introduced, and the backbone network of the Gaussian blind denoiser is pre-trained using L1 loss. In the second stage, the backbone network weights obtained in the first stage are loaded and frozen. The spatial frequency feature fidelity adapter is embedded at the skip connections, and the spatial frequency feature fidelity adapter is trained using a loss function that combines L1 reconstruction loss and perceptual loss based on shallow feature maps of the pre-trained VGG network. The denoising inference module is used to input an input image containing unknown levels of Gaussian noise into the trained Gaussian blind denoiser, and output a denoised image through a single forward inference.

[0023] Thirdly, a computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the aforementioned feature-fidelity Gaussian blind denoising method.

[0024] Fourthly, embodiments of the present invention provide a computer-readable storage medium including a computer program, which, when executed by a processor, implements the steps of the aforementioned feature-fidelity Gaussian blind denoising method.

[0025] Fifthly, a chip includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the aforementioned feature-fidelity Gaussian blind denoising method.

[0026] In a sixth aspect, embodiments of the present invention provide an electronic device, including a computer program, wherein when the computer program is executed by the electronic device, it implements the steps of the above-described feature-fidelity Gaussian blind denoising method.

[0027] Compared with the prior art, the present invention has at least the following beneficial effects: A feature-preserving Gaussian blind denoising method is proposed, which constructs a multi-scale U-shaped Gaussian blind denoiser, fuses multi-scale features through a cross-layer attention fusion module, sets a space-frequency feature-preserving adapter at skip connections, performs two-stage training, and uses the trained denoiser for a single forward inference. This method unifies basic denoising capabilities, multi-scale feature fusion capabilities, and local feature preservation capabilities into a single denoising framework. The multi-scale U-shaped structure facilitates the simultaneous utilization of shallow texture information and deep structural information; the parallel axial Transformer-convolution module is used to balance local feature processing and long-range dependency modeling; the space-frequency feature-preserving adapter processes transitive features at skip connections, helping to reduce residual noise and preserve local structure; and the two-stage training optimizes the backbone denoising capability and feature preservation capability in stages, thereby improving the applicability of denoised images in visual quality and machine vision tasks.

[0028] Furthermore, the same feature extraction unit can simultaneously include local processing paths and global dependency modeling paths. Feature dimensionality reduction reduces the amount of data and computational overhead in subsequent branch processing; channel splitting allows different sub-features to enter appropriate processing branches, avoiding a single path simultaneously undertaking all feature representation tasks. Convolutional branches are suitable for processing local information such as neighborhood structure, edges, and textures, while Transformer branches are suitable for processing spatial relationships over a larger range. The fusion of these two branches yields feature representations that combine local structure and long-distance correlations, thus providing a more comprehensive feature foundation for subsequent denoising and reconstruction.

[0029] Furthermore, deformable convolution can adaptively learn sampling offsets based on input features, allowing the sampling position to adjust with changes in edge direction, texture morphology, and local structure; the modulation scalar can also adjust the response intensity at different sampling positions. This enables more accurate acquisition of local information such as edges, textures, and corners, reducing the loss of detail caused by fixed convolution sampling during denoising.

[0030] Furthermore, while controlling computational complexity, it achieves better spatial dependency modeling capabilities. Standard 2D self-attention requires calculating positional relationships over a large spatial range, resulting in high computational costs; axial attention decomposes 2D spatial relationships into row and column directions, reducing the computational overhead of attention while still modeling structural relationships over long distances. Deep convolution supplements local context processing capabilities, preventing Transformer branches from overly favoring global relationships and neglecting neighborhood details. Dual-gate feedforward networks regulate information flow through gating paths, facilitating the filtering and modulation of responses to different features. Together, these improvements enhance the Transformer branches' ability to express image structure, texture continuity, and local details.

[0031] Furthermore, by calculating the correlation between Q and K, the hierarchical correlation between features at different scales can be obtained; then, by weighting V using the attention results, the fusion process can be adaptively adjusted according to the degree of correlation between features. Residual connections help preserve the original input features and avoid the loss of effective information during the fusion process. Multi-scale feature fusion is a dynamic fusion based on hierarchical correlation.

[0032] Furthermore, the generation of feature representations required for attention computation takes into account both channel transformation and local spatial processing. 1×1 convolutions are primarily used for channel-dimensional information transformation and compression, allowing adjustment of feature channel representations with low computational cost; 3×3 depthwise separable convolutions introduce local spatial neighborhood information while preserving low parameter and computational costs. The resulting Q, K, and V values ​​not only contain inter-channel recombination information but also local spatial context, providing a more reasonable feature basis for the cross-layer attention fusion module when calculating hierarchical correlations. This reduces parameter and computational costs and enhances the ability to perceive local structures during multi-scale feature fusion.

[0033] Furthermore, feature representation can be adjusted through linear projection, and local context is introduced through depthwise separable convolution. The generated spatial attention mask is used to represent the importance of different spatial locations. After Sigmoid normalization, it is multiplied element-wise with the features to be modulated, and the feature responses at different locations are weighted, achieving selective filtering and reshaping of the spatial domain in the skip connection path.

[0034] Furthermore, image noise and effective structure often differ in frequency distribution, with key local structures such as edges, textures, and corners often corresponding to specific high-frequency or frequency band responses. Through Fourier transform, the model can analyze feature components in the frequency domain; by generating a spectral modulation mask using amplitude, it can adaptively adjust different frequency components; and then return to the spatial domain via inverse Fourier transform, enabling S3FA to process features not only in the spatial domain but also to reshape feature representations in the frequency domain.

[0035] Furthermore, the progressive learning strategy allows the backbone network to first learn basic denoising mappings on smaller image patches, and then gradually adapt to long-range spatial dependencies in larger image patches, which is beneficial for a smooth transition in the training process. The second stage uses VGG-19 shallow feature maps to calculate the perceptual loss, enabling S3FA to focus on shallow visual features such as edges, textures, and local structures. Since the backbone network is frozen, the training focus is concentrated on the S3FA parameters, which helps reduce the conflict between pixel-level denoising goals and feature fidelity goals.

[0036] It is understood that the beneficial effects of the second to sixth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here.

[0037] In summary, this invention achieves the removal of unknown-level Gaussian noise and preservation of local key features through a U-shaped Gaussian blind denoiser, a parallel axial Transformer-convolution module, a cross-layer attention fusion module, and a spatial frequency feature fidelity adapter. It can reduce the computational complexity of the model, enhance multi-scale feature fusion, and alleviate the conflict between denoising and feature fidelity goals through two-stage training, making the denoised images more suitable for feature matching, registration, and stitching tasks.

[0038] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0039] Figure 1 Diagram of the U-PATC denoising network framework; Figure 2 This is a structural diagram of the cross-layer attention fusion module; Figure 3 This is a structural diagram of a space-frequency characteristic fidelity adapter; Figure 4 Flowchart of the two-stage training process for the denoiser; Figure 5 A schematic diagram of a computer device provided in an embodiment of the present invention; Figure 6 This is a block diagram of a chip provided according to an embodiment of the present invention.

[0040] Among them, 60. Computer equipment; 61. Processor; 62. Memory; 63. Computer program; 600. Electronic device; 610. Processing unit; 620. Storage unit; 6201. Random access memory unit; 6202. Cache memory unit; 6203. Read-only memory unit; 6204. Program / utility; 6205. Program module; 630. Bus; 640. Display unit; 650. Input / output interface; 660. Network adapter; 700. External device. Detailed Implementation

[0041] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0042] In the description of this invention, it should be understood that the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.

[0043] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.

[0044] It should also be further understood that the term "and / or" as used in this specification and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes such combinations. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone. Additionally, the character " / " in this invention generally indicates that the preceding and following objects have an "or" relationship.

[0045] It should be understood that although terms such as first, second, third, etc., may be used in the embodiments of the present invention to describe the preset range, these preset ranges should not be limited to these terms. These terms are only used to distinguish the preset ranges from one another. For example, without departing from the scope of the embodiments of the present invention, the first preset range may also be referred to as the second preset range, and similarly, the second preset range may also be referred to as the first preset range.

[0046] Depending on the context, the word "if" as used here can be interpreted as "when," "when," "in response to determination," or "in response to detection." Similarly, depending on the context, the phrase "if determination" or "if detection (of the stated condition or event)" can be interpreted as "when determination," "in response to determination," "when detection (of the stated condition or event)," or "in response to detection (of the stated condition or event)."

[0047] The accompanying drawings illustrate various structural schematic diagrams according to embodiments disclosed in this invention. These drawings are not to scale, and some details have been enlarged for clarity, and some details may have been omitted. The shapes of the various regions and layers shown in the drawings, as well as their relative sizes and positional relationships, are merely exemplary and may deviate from reality due to manufacturing tolerances or technical limitations. Furthermore, those skilled in the art can design regions / layers with different shapes, sizes, and relative positions as needed.

[0048] This invention provides a Gaussian blind denoising method that preserves feature fidelity. First, a Gaussian blind denoiser with a multi-scale U-shaped structure is constructed, and parallel axial Transformer-convolutional modules are introduced into the feature encoder and decoder. This module combines the local feature processing capabilities of convolutional structures with the long-distance dependency modeling capabilities of axial attention mechanisms. Compared to simple convolutional networks, it improves the problem of insufficient global structural representation; compared to the standard Transformer structure, it reduces attention computation overhead and improves the lightweight nature and deployment adaptability of the denoiser. Second, this invention sets up a cross-layer attention fusion module to adaptively fuse multi-scale features. Unlike the simple splicing or direct skip connections in traditional U-shaped networks, this method can more fully utilize the complementary information between features at different levels, improving the structural recovery capability during denoising. Third, this invention sets up a spatial-frequency feature fidelity adapter at the skip connections, reshaping the skip connection features through spatial and frequency modulation. This suppresses residual noise propagation while preserving key local features such as edges, textures, corners, and high-frequency gradients. Finally, this invention employs a two-stage training strategy: first, the backbone network is trained to acquire basic denoising capabilities; then, the backbone network is frozen and a space-frequency feature fidelity adapter is trained, thereby mitigating the conflict between pixel-level denoising and feature fidelity objectives. Through these improvements, this invention can remove unknown levels of Gaussian noise while simultaneously enhancing the visual quality, local feature stability, and usability of denoised images in machine vision tasks such as feature matching, image registration, and image stitching.

[0049] This invention provides a Gaussian blind denoising method that preserves feature fidelity, comprising the following steps: S1. Construct a Gaussian blind denoiser (U-PATC) based on a multi-scale U-shaped architecture. The denoiser adopts a multi-scale U-shaped structure, which includes a feature encoder, a cross-layer attention fusion module (CAFB), and a feature decoder.

[0050] The core feature extraction unit of the encoder and decoder adopts a parallel axial Transformer-convolution module (PATC).

[0051] S2. Adaptive fusion of multi-scale features using the Cross-Layer Attention Fusion Module (CAFB). In the hierarchical feature extraction process, the input features of multiple dimensions are concatenated and flattened along the channel dimension. .

[0052] Combination Convolution and Depthwise separable convolutions generate corresponding query vectors (Q), key vectors (K), and value vectors (V).

[0053] The hierarchical relevance attention matrix is ​​obtained by calculating the matrix product of the query vector and the key vector, then multiplied by the value vector and normalized by introducing a scaling factor. Finally, it is added to the original input features through residual connections to output the fusion result.

[0054]

[0055] in, For query vector based Key vector Sum value vector The calculated cross-layer attention output results The fusion features output by the cross-layer attention fusion module. For normalized exponential functions, Scaling factor The convolution weights are for a 1×1 convolution. These are the input features for the cross-layer attention fusion module.

[0056] Please see Figure 1 The feature-preserving Gaussian blind denoiser receives an image containing unknown levels of Gaussian noise and performs multi-scale feature extraction using an encoder composed of PATC modules, progressively reducing the feature resolution and increasing the number of channels. Since features extracted from different network layers often respond to structures or specific local details at different scales, the denoiser introduces [a certain feature] during the multi-branch interaction of feature downsampling. Figure 2 The Cross-Layer Attention Fusion (CAFB) module is shown. This module concatenates and reconstructs the input features from each dimension along the channel dimension, utilizing... and Depthwise separable convolutions generate query, key, and value vectors. Then, a relevance attention matrix is ​​calculated in the channel and layer spaces, enabling adaptive dynamic weighting and efficient fusion of multi-level features. These features are then fed into the corresponding decoder to reconstruct the denoised image at the original resolution layer by layer.

[0057] S3. Introduce a spatial frequency feature fidelity adapter (S3FA) at the jump connection of the noise denoiser for feature reshaping. Inside the denoiser, an S3FA module is introduced at the jump connection where the encoder passes features to the decoder. This module employs a dual-branch parallel architecture. Spatial Modulation Module (SMM): Utilizes input features Linear projection layers combined with a large receptive field Depthwise separable convolutions generate a spatial attention mask matrix. This mask, after sigmoid activation, is multiplied element-wise with intermediate features to suppress noise in smooth regions and enhance structural regions. The formula is expressed as:

[0058] in, The first dimension reduction transform weight in the spatial modulation branch. It is the Sigmoid activation function. The convolution weights are for a 5×5 depth separable convolution. This is the second dimension reduction transform weight in the spatial modulation branch. This refers to the upsampling transform weights in the spatial modulation branch.

[0059] Frequency Modulation Module (FMM): Converts spatial features into frequency domain complex tensors using a Fast Real Fourier Transform (FFT). It extracts the amplitude of the frequency domain tensor and learns a spectral modulation mask through a lightweight convolutional network. This mask is used to amplify frequency components strongly correlated with feature matching and attenuate noise bands. Finally, an inverse Fourier transform restores the spatial features.

[0060]

[0061] Spatial-frequency feature fusion: This involves concatenating the outputs of spatial modulation and frequency modulation along the channel dimension, utilizing... The convolutional layer learns a weighted combination and performs a residual connection with the original input features to output the following:

[0062] in, The spatial modulation characteristics of the output of the spatial modulation branch. The frequency modulation characteristics of the frequency modulation branch output. The input characteristics of the space frequency feature fidelity adapter. The spectral modulation mask generated for the frequency modulation branch. For frequency modulation branches based on frequency domain characteristic amplitude The calculated spectral modulation result.

[0063] Please see Figure 3 At the jump connection between shallow and deep features in the U-shaped structure of the denoiser, an S3FA module is deployed as an intelligent filtering module to prevent high-frequency noise during downsampling from being directly mixed into the decoding process. This module performs dual modulation on the input features: In the spatial domain, using the Spatial Modulation Module (SMM), features are reduced through linear dimensionality reduction and summation. The deep separable convolution with a large receptive field extracts a broader local context, and the spatial mask matrix generated by sigmoid activation suppresses noise in smooth background regions and enhances physical edge features.

[0064] In the frequency domain, using a frequency modulation module (FMM), features are transformed to the frequency domain via FFT. By extracting amplitude and learning a spectral mask, random noise across the entire frequency band is precisely attenuated, and specific high-frequency bands that facilitate downstream feature matching are amplified. Finally, the features are restored to the spatial domain via IFFT. The two precisely modulated features are concatenated along the channel dimension. After convolutional fusion, the residual structure is superimposed into the jump connection main path and sent to the decoding side of the denoiser.

[0065] S4. Optimize the denoiser by implementing a two-stage training strategy that decouples denoising and fidelity preservation. Phase 1 (Denoising Prior Pre-training): The denoiser does not introduce the S3FA module; instead, it uses only L1 loss, which aims to minimize pixel value differences, to perform end-to-end pre-training on the backbone network. A progressive learning strategy is employed in this phase, gradually transitioning from smaller image patches and larger batch sizes to larger image patches and smaller batch sizes, enabling the backbone network to converge quickly and fully learn large-scale spatial dependencies.

[0066] The second stage (feature-fidelity fine-tuning) involves loading the backbone network weights converged in the first stage and freezing their parameters. Randomly initialized S3FA modules are embedded at each skip connection, and only the parameters of these adapters are set to trainable states. A loss function combining pixel-level fidelity and feature-level perception is employed. Optimize and update S3FA:

[0067] in, The weighting coefficients for perceived loss, This represents the shallow feature extraction operation of the pre-trained VGG network, used to calculate the feature-aware loss.

[0068] To address the severe gradient conflict caused by directly combining traditional pixel-level reconstruction loss with high-frequency feature-focused perceptual loss, the decoupled two-stage training implementation of the denoiser is described below. Figure 4 The decoupled two-stage training is specifically as follows: Phase 1: The S3FA module is temporarily disabled in the denoiser; only the L1 loss function is used to optimize the backbone network. A blind denoising setting is adopted (the standard deviation of the superimposed noise ranges from [0, 50]), and a progressive learning strategy is implemented (for example, during the iteration process, the image patch size and batch size are progressively adjusted from 128×128 and 64 to 384×384 and 8), which prompts the backbone network to fully learn large-scale spatial dependencies and local pixel mappings, thus solidifying the basic Gaussian blind denoising capabilities.

[0069] Phase 2: Load and freeze the backbone network weights from the converged phase 1, embed S3FA modules at the skip connections of the network and make them trainable. Use L1 reconstruction loss and feature matching loss based on pre-trained VGG-19 shallow output feature maps (the weight ratio of the two is configured as follows). Fine-tuning is performed on the joint loss function composed of ( ).

[0070] During this stage, due to the freezing of the backbone network parameters, all gradient information is precisely used to update the S3FA module, which accounts for only about 0.1% of the total parameters. This allows the module to learn to actively compensate for and reshape high-frequency local invariant features (such as SIFT and ORB feature points), which are crucial for machine vision tasks, without compromising the backbone's basic denoising capabilities. The denoiser trained using this method can directly output a clean image with excellent visual quality and high fidelity when faced with Gaussian noise of unknown intensity through a single forward computation.

[0071] S5. Perform Gaussian blind denoising using the optimized denoiser. Feed the input image containing unknown noise into the trained denoiser, and through a single forward inference, directly output a high-quality, feature-rich, denoised, and clean image.

[0072] In another embodiment of the present invention, a Gaussian blind denoising system with feature fidelity is provided. This system can be used to implement the above-mentioned Gaussian blind denoising method with feature fidelity. Specifically, the Gaussian blind denoising system with feature fidelity includes a Gaussian blind denoiser construction module, a multi-scale feature fusion module, a space-frequency feature fidelity processing module, a two-stage training module, and a denoising inference module.

[0073] Among them, the Gaussian blind denoiser construction module is used to construct the Gaussian blind denoiser. The Gaussian blind denoiser adopts a multi-scale U-shaped structure. The Gaussian blind denoiser includes a feature encoder, a cross-layer attention fusion module and a feature decoder. The feature extraction units of the feature encoder and the feature decoder adopt parallel axial Transformer-convolution modules. A multi-scale feature fusion module is used to fuse the multi-scale features output by the feature encoder through the cross-layer attention fusion module to obtain fused features; The space frequency feature fidelity processing module includes a space frequency feature fidelity adapter, which is disposed at the jump connection where the feature encoder transmits features to the feature decoder. The adapter is used to process the jump connection features to obtain space frequency fidelity features and transmit the space frequency fidelity features to the feature decoder. The space-frequency feature fidelity adapter includes a spatial modulation branch, a frequency modulation branch, and a space-frequency feature fusion branch. The spatial modulation branch is used to generate a spatial attention mask based on the skip connection feature and obtain a spatial modulation feature based on the spatial attention mask. The frequency modulation branch is used to convert the skip connection feature to the frequency domain and obtain a frequency modulation feature based on the frequency domain feature. The space-frequency feature fusion branch is used to fuse the spatial modulation feature and the frequency modulation feature to obtain the space-frequency fidelity feature. A two-stage training module is used to perform two-stage training on the Gaussian blind denoiser. In the first stage, the spatial frequency feature fidelity adapter is not introduced, and the backbone network of the Gaussian blind denoiser is pre-trained using L1 loss. In the second stage, the backbone network weights obtained in the first stage are loaded and frozen. The spatial frequency feature fidelity adapter is embedded at the skip connections, and the spatial frequency feature fidelity adapter is trained using a loss function that combines L1 reconstruction loss and perceptual loss based on shallow feature maps of the pre-trained VGG network. The denoising inference module is used to input an input image containing unknown levels of Gaussian noise into the trained Gaussian blind denoiser, and output a denoised image through a single forward inference.

[0074] This invention provides a terminal device comprising a processor and a memory. The memory stores a computer program, which includes program instructions. The processor executes the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or other general-purpose processors, graphics processing units (GPUs), tensor processing units (TPUs), digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing and control core of the terminal, suitable for implementing one or more instructions, specifically suitable for loading and executing one or more instructions to implement a corresponding method flow or corresponding function. The processor described in this embodiment can be used for the operation of a feature-fidelity Gaussian blind denoising method, including: A Gaussian blind denoiser is constructed, employing a multi-scale U-shaped structure. The Gaussian blind denoiser includes a feature encoder, a cross-layer attention fusion module, and a feature decoder. The feature extraction units of the feature encoder and the feature decoder utilize parallel axial Transformer-convolution modules. The cross-layer attention fusion module fuses the multi-scale features output by the feature encoder to obtain fused features. A spatial frequency feature fidelity adapter is placed at the skip connection points where the feature encoder transmits features to the feature decoder. The spatial frequency feature fidelity adapter processes the skip connection features to obtain spatial frequency fidelity features, which are then transmitted to the feature decoder. The spatial frequency feature fidelity adapter includes a spatial modulation branch, a frequency modulation branch, and a spatial frequency feature fusion branch. The spatial modulation branch generates a spatial attention mask based on the skip connection features and, based on the spatial attention mask... The code obtains spatial modulation features. The frequency modulation branch transforms the skip connection features to the frequency domain and obtains frequency modulation features based on the frequency domain features. The spatial-frequency feature fusion branch fuses the spatial modulation features and the frequency modulation features to obtain the spatial-frequency fidelity features. The Gaussian blind denoiser is trained in two stages. In the first stage, the spatial-frequency feature fidelity adapter is not introduced, and the backbone network of the Gaussian blind denoiser is pre-trained using L1 loss. In the second stage, the backbone network weights obtained in the first stage are loaded and frozen. The spatial-frequency feature fidelity adapter is embedded at the skip connection, and the spatial-frequency feature fidelity adapter is trained using a loss function that combines L1 reconstruction loss and perceptual loss based on the shallow feature map of the pre-trained VGG network. An input image containing unknown levels of Gaussian noise is input into the trained Gaussian blind denoiser, and a denoised image is output through a single forward inference.

[0075] Please see Figure 5 The terminal device is a computer device. In this embodiment, the computer device 60 includes a processor 61, a memory 62, and a computer program 63 stored in the memory 62 and executable on the processor 61. When executed by the processor 61, the computer program 63 implements the feature-fidelity Gaussian blind denoising method of the embodiment. To avoid repetition, these details are not elaborated here. Alternatively, when executed by the processor 61, the computer program 63 implements the functions of each model / unit in the feature-fidelity Gaussian blind denoising system of the embodiment. To avoid repetition, these details are not elaborated here.

[0076] Computer device 60 can be a desktop computer, laptop, handheld computer, cloud server, or other computing device. Computer device 60 may include, but is not limited to, a processor 61 and a memory 62. Those skilled in the art will understand that... Figure 5This is merely an example of computer device 60 and does not constitute a limitation on computer device 60. It may include more or fewer components than shown, or combine certain components, or different components. For example, computer device may also include input / output devices, network access devices, buses, etc.

[0077] The processor 61 may be a Central Processing Unit (CPU), or other general-purpose processors, graphics processing units (GPUs), tensor processing units (TPUs), digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.

[0078] The memory 62 can be an internal storage unit of the computer device 60, such as a hard disk or memory of the computer device 60. The memory 62 can also be an external storage device of the computer device 60, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc. equipped on the computer device 60.

[0079] Furthermore, the memory 62 may include both internal storage units of the computer device 60 and external storage devices. The memory 62 is used to store computer programs and other programs and data required by the computer device. The memory 62 can also be used to temporarily store data that has been output or will be output.

[0080] Please see Figure 6 The terminal device is an electronic device 600, which is manifested in the form of a general-purpose computing device. The components of the electronic device may include, but are not limited to: at least one processing unit 610, at least one storage unit 620, a bus 630 connecting different platform components (including storage unit 620 and processing unit 610), a display unit 640, etc.

[0081] The storage unit stores program code, which can be executed by the processing unit 610 to perform the steps described in the method section of this specification according to various exemplary embodiments of the present invention. For example, the processing unit 610 can perform actions such as... Figure 1 The steps are shown in the figure.

[0082] Storage unit 620 may include readable media in the form of volatile storage units, such as random access memory (RAM) 6201 and / or cache memory 6202, and may further include read-only memory (ROM) 6203.

[0083] Storage unit 620 may also include a program / utility 6204 having a set (at least one) program module 6205, such program module 6205 including but not limited to: operating system, one or more application programs, other program modules and program data, each or some combination of these examples may include an implementation of a network environment.

[0084] Bus 630 can represent one or more of several types of bus structures, including a memory cell bus or memory cell controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of the multiple bus structures.

[0085] Electronic device 600 can also communicate with one or more external devices 700 (e.g., keyboard, pointing device, Bluetooth device, etc.), and with one or more devices that enable a user to interact with electronic device 600, and / or with any device that enables electronic device 600 to communicate with one or more other computing devices (e.g., router, modem). This communication can be performed via input / output interface 650. Furthermore, electronic device 600 can also communicate with one or more networks (e.g., local area network, wide area network, and / or public network, such as the Internet) via network adapter 660. Network adapter 660 can communicate with other modules of electronic device 600 via bus 630. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with electronic device 600, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage platforms.

[0086] This invention also provides a storage medium, specifically a computer-readable storage medium, which is a memory device in a terminal device for storing programs and data. It is understood that the computer-readable storage medium here can include both built-in storage media in the terminal device and extended storage media supported by the terminal device; it can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. The computer-readable storage medium provides storage space that stores the terminal's operating system. Furthermore, the storage space also stores one or more instructions suitable for loading and execution by a processor, which can be one or more computer programs (including program code). More specific examples of the computer-readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, random access memory, read-only memory, erasable programmable read-only memory, optical fiber, portable compact disk read-only memory, optical storage device, magnetic storage device, or any suitable combination thereof.

[0087] Computer-readable storage media also include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable storage medium can also be any readable medium other than a readable storage medium that can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium can be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, radio frequency, etc., or any suitable combination thereof.

[0088] Program code for performing the operations of this invention can be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java and C++, and conventional procedural programming languages ​​such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0089] One or more instructions stored in a computer-readable storage medium can be loaded and executed by a processor to implement the corresponding steps of the Gaussian blind denoising method related to feature fidelity in the above embodiments; one or more instructions in the computer-readable storage medium are loaded and executed by the processor to perform the following steps: A Gaussian blind denoiser is constructed, employing a multi-scale U-shaped structure. The Gaussian blind denoiser includes a feature encoder, a cross-layer attention fusion module, and a feature decoder. The feature extraction units of the feature encoder and the feature decoder utilize parallel axial Transformer-convolution modules. The cross-layer attention fusion module fuses the multi-scale features output by the feature encoder to obtain fused features. A spatial frequency feature fidelity adapter is placed at the skip connection points where the feature encoder transmits features to the feature decoder. The spatial frequency feature fidelity adapter processes the skip connection features to obtain spatial frequency fidelity features, which are then transmitted to the feature decoder. The spatial frequency feature fidelity adapter includes a spatial modulation branch, a frequency modulation branch, and a spatial frequency feature fusion branch. The spatial modulation branch generates a spatial attention mask based on the skip connection features and, based on the spatial attention mask... The code obtains spatial modulation features. The frequency modulation branch transforms the skip connection features to the frequency domain and obtains frequency modulation features based on the frequency domain features. The spatial-frequency feature fusion branch fuses the spatial modulation features and the frequency modulation features to obtain the spatial-frequency fidelity features. The Gaussian blind denoiser is trained in two stages. In the first stage, the spatial-frequency feature fidelity adapter is not introduced, and the backbone network of the Gaussian blind denoiser is pre-trained using L1 loss. In the second stage, the backbone network weights obtained in the first stage are loaded and frozen. The spatial-frequency feature fidelity adapter is embedded at the skip connection, and the spatial-frequency feature fidelity adapter is trained using a loss function that combines L1 reconstruction loss and perceptual loss based on the shallow feature map of the pre-trained VGG network. An input image containing unknown levels of Gaussian noise is input into the trained Gaussian blind denoiser, and a denoised image is output through a single forward inference.

[0090] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0091] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0092] To verify the technical effectiveness of this invention, the U-PATC denoiser of this invention can be compared with existing baseline denoising networks. The test results show that the FLOPs of the denoiser of this invention are approximately 9.785G, which is about one-fourteenth of that of the RFPCNN baseline denoising network. This indicates that the invention significantly reduces computational complexity while maintaining denoising processing capabilities, making it suitable for deployment in application scenarios with high requirements for inference efficiency and computational resources.

[0093] Furthermore, this invention incorporates a spatial-frequency feature fidelity adapter (S3FA) at the skip connection points, with the number of parameters in this adapter accounting for approximately 0.1% of the total network parameters. Since the S3FA introduces only a few additional parameters, it can modulate the skip connection features from both the spatial and frequency domains, thereby compensating for and reshaping key local features such as edges, textures, corners, and high-frequency gradients without significantly increasing the model size.

[0094] Furthermore, the usability of denoised images in machine vision tasks can be evaluated using metrics such as the number of SIFT keypoints, the number of ORB keypoints, the number of effective matching points, image registration error, and image stitching success rate. Compared to denoising networks that do not incorporate S3FA or only use ordinary skip connections, this invention, through joint space-frequency modulation and two-stage training, enables the denoiser to remove unknown levels of Gaussian noise while better preserving local invariant features, thereby improving the applicability of denoised images in downstream tasks such as feature matching, image registration, and image stitching.

[0095] In summary, this invention presents a feature-preserving Gaussian blind denoising method and system, addressing the denoising needs of images with unknown levels of Gaussian noise. It constructs a feature-preserving Gaussian blind denoiser, improving the visual quality of denoised images while preserving key local features. By employing parallel axial Transformer-convolution modules in the feature encoder and decoder, this invention combines the local feature processing capabilities of convolutional structures with the long-distance dependency modeling capabilities of axial attention mechanisms, improving image structure restoration while reducing the computational overhead of standard self-attention. Through the inclusion of a cross-layer attention fusion module, this invention can adaptively fuse multi-scale features, avoiding the insufficient utilization of hierarchical information caused by simple splicing of traditional U-shaped networks. Furthermore, this invention sets a spatial-frequency feature-preserving adapter at skip connections, modulating the skip connection features in both the spatial and frequency domains to reduce residual noise propagation while preserving local structural information such as edges, textures, corners, and high-frequency gradients. By employing a two-stage training strategy, the basic denoising capabilities of the backbone network are first trained, and then the backbone is frozen and the spatial frequency feature fidelity adapter is trained. This reduces the conflict between pixel-level denoising targets and feature fidelity targets, making denoised images more suitable for machine vision tasks such as feature matching, image registration, and image stitching.

[0096] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0097] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0098] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed in this invention can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0099] In the embodiments provided by this invention, it should be understood that the disclosed devices / terminals and methods can be implemented in other ways. For example, the device / terminal embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0100] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0101] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0102] If the integrated module / unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of the present invention can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random-access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.

[0103] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus, and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0104] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0105] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0106] The above content is only for illustrating the technical concept of the present invention and should not be construed as limiting the scope of protection of the present invention. Any modifications made to the technical solution based on the technical concept proposed in this invention shall fall within the scope of protection of the claims of this invention.

Claims

1. A Gaussian blind denoising method that preserves feature fidelity, characterized in that, Includes the following steps: S1. Construct a Gaussian blind denoiser. The Gaussian blind denoiser adopts a multi-scale U-shaped structure. The Gaussian blind denoiser includes a feature encoder, a cross-layer attention fusion module, and a feature decoder. The feature extraction units of the feature encoder and the feature decoder adopt parallel axial Transformer-convolution modules. S2. The multi-scale features output by the feature encoder are fused using the cross-layer attention fusion module to obtain fused features; S3. A space-frequency feature fidelity adapter is set at the skip connection where the feature encoder transmits features to the feature decoder. The space-frequency feature fidelity adapter processes the skip connection features to obtain space-frequency fidelity features, and then transmits the space-frequency fidelity features to the feature decoder. The space-frequency feature fidelity adapter includes a spatial modulation branch, a frequency modulation branch, and a space-frequency feature fusion branch. The spatial modulation branch generates a spatial attention mask based on the skip connection features and obtains spatial modulation features based on the spatial attention mask. The frequency modulation branch converts the skip connection features to the frequency domain and obtains frequency modulation features based on the frequency domain features. The space-frequency feature fusion branch fuses the spatial modulation features and the frequency modulation features to obtain the space-frequency fidelity features. S4. Perform two-stage training on the Gaussian blind denoiser. In the first stage, the spatial frequency feature fidelity adapter is not introduced, and the backbone network of the Gaussian blind denoiser is pre-trained using L1 loss. In the second stage, the backbone network weights obtained in the first stage are loaded and frozen. The spatial frequency feature fidelity adapter is embedded at the skip connections, and the spatial frequency feature fidelity adapter is trained using a loss function that combines L1 reconstruction loss and perceptual loss based on shallow feature maps of the pre-trained VGG network. S5. Input the input image containing unknown level Gaussian noise into the trained Gaussian blind denoiser, and output the denoised image through a single forward inference.

2. The Gaussian blind denoising method with feature preservation according to claim 1, characterized in that, The parallel axial Transformer-convolution module reduces the dimensionality of the module input features and splits the dimensionality-reduced module input features into convolutional sub-features and Transformer sub-features in the channel dimension. The convolutional sub-features are processed through convolutional branches, the Transformer sub-features are processed through Transformer branches, and the output features of the convolutional branches and the Transformer branches are fused together.

3. The Gaussian blind denoising method with feature preservation according to claim 2, characterized in that, The convolutional branch uses deformable convolution to process the convolutional sub-features, and the deformable convolution adjusts the convolutional sampling position by learning coordinate offset and modulation scalar.

4. The Gaussian blind denoising method with feature preservation according to claim 2, characterized in that, The Transformer branch uses an axial attention mechanism to process the Transformer sub-features, which includes row-direction attention calculation and column-direction attention calculation. The Transformer branch also performs convolution processing on the Transformer sub-features through depthwise convolution and performs gating processing on the information flow of the Transformer sub-features through a dual-gate feedforward network.

5. The Gaussian blind denoising method with feature preservation according to claim 1, characterized in that, In step S2, input features at multiple scales are concatenated along the channel dimension, and a query vector, key vector, and value vector are generated through convolution operations. A matrix product is calculated based on the query vector and the key vector. After normalizing the matrix product by a scaling factor, it is multiplied by the value vector. The result of the multiplication is added to the input features of the cross-layer attention fusion module through a residual connection to obtain the fused features.

6. The Gaussian blind denoising method with feature preservation according to claim 5, characterized in that, The query vector, the key vector, and the value vector are generated by 1×1 convolution and 3×3 depthwise separable convolution.

7. The Gaussian blind denoising method with feature preservation according to claim 1, characterized in that, In step S3, the spatial modulation branch linearly projects the skip connection feature and combines it with depthwise separable convolution to generate the spatial attention mask. After the spatial attention mask is activated by Sigmoid, it is multiplied element-wise with the feature to be modulated in the spatial modulation branch to obtain the spatial modulation feature.

8. The Gaussian blind denoising method with feature preservation according to claim 1, characterized in that, In step S3, the frequency modulation branch transforms the skip connection feature to the frequency domain through a fast real Fourier transform to obtain a frequency domain complex tensor as the frequency domain feature. The amplitude of the frequency domain complex tensor is extracted, and a spectral modulation mask is generated based on the amplitude through a convolutional network. After the frequency domain complex tensor is modulated using the spectral modulation mask, the frequency modulation feature is obtained through an inverse Fourier transform.

9. The Gaussian blind denoising method with feature preservation according to claim 1, characterized in that, In step S4, in the first stage, a progressive learning strategy is used to pre-train the backbone network. The progressive learning strategy includes gradually transitioning from smaller image patches and larger batch sizes to larger image patches and smaller batch sizes. In the second stage, the pre-trained VGG network is a VGG-19 network, and the perceptual loss is calculated based on the shallow output feature map of the VGG-19 network.

10. A feature-fidelity Gaussian blind denoising system, characterized in that, include: A Gaussian blind denoiser construction module is used to construct a Gaussian blind denoiser. The Gaussian blind denoiser adopts a multi-scale U-shaped structure and includes a feature encoder, a cross-layer attention fusion module, and a feature decoder. The feature extraction units of the feature encoder and the feature decoder adopt parallel axial Transformer-convolution modules. A multi-scale feature fusion module is used to fuse the multi-scale features output by the feature encoder through the cross-layer attention fusion module to obtain fused features; The space frequency feature fidelity processing module includes a space frequency feature fidelity adapter, which is disposed at the jump connection where the feature encoder transmits features to the feature decoder. The adapter is used to process the jump connection features to obtain space frequency fidelity features and transmit the space frequency fidelity features to the feature decoder. The space-frequency feature fidelity adapter includes a spatial modulation branch, a frequency modulation branch, and a space-frequency feature fusion branch. The spatial modulation branch is used to generate a spatial attention mask based on the skip connection feature and obtain a spatial modulation feature based on the spatial attention mask. The frequency modulation branch is used to convert the skip connection feature to the frequency domain and obtain a frequency modulation feature based on the frequency domain feature. The space-frequency feature fusion branch is used to fuse the spatial modulation feature and the frequency modulation feature to obtain the space-frequency fidelity feature. A two-stage training module is used to perform two-stage training on the Gaussian blind denoiser. In the first stage, the spatial frequency feature fidelity adapter is not introduced, and the backbone network of the Gaussian blind denoiser is pre-trained using L1 loss. In the second stage, the backbone network weights obtained in the first stage are loaded and frozen. The spatial frequency feature fidelity adapter is embedded at the skip connections, and the spatial frequency feature fidelity adapter is trained using a loss function that combines L1 reconstruction loss and perceptual loss based on shallow feature maps of the pre-trained VGG network. The denoising inference module is used to input an input image containing unknown levels of Gaussian noise into the trained Gaussian blind denoiser, and output a denoised image through a single forward inference.