A self-supervised denoising network and method based on a U-shaped encoding-decoding topology
Patent Information
- Application Number
- CN202610345135.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-20
- Publication Date
- 2026-09-25
AI Technical Summary
[0007]本发明的目的在于解决现有基于U型拓扑结构的地震数据去噪网络因采用固定参数卷积块作为基础计算单元、缺乏与输入特征统计特性及噪声水平动态关联的门控调节机制,且提示学习信号与去噪算子参数生成路径耦合不足,在强噪声与非平稳噪声(如空间/时间变异噪声)条件下引起的过度平滑、弱反射事件与薄层结构细节丢失、边缘模糊以及去噪强度无法自适应调控的问题
[0060](1)动态门控去噪:BN-GTVD模块通过门控系数对候选滤波响应进行逐元素调制,使网络能够依据输入特征内容与局部噪声差异自适应调整去噪强度,提高对非平稳噪声及不同区域噪声方差差异的适应性;
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] This invention relates to the field of seismic exploration and provides a self-supervised denoising network and method based on a U-shaped encoder-decoder topology. Background Technology
[0002] In applications such as seismic exploration, mechanical vibration monitoring, and acoustic imaging, the acquired two-dimensional signals (e.g., seismic record slices, time-frequency maps, or spectrograms) are often affected by multi-source interference, including random noise, coherent noise, and instrument noise. This noise can mask weak reflection events or weak fault components and may cause error propagation and accumulation in subsequent interpretation, inversion, and diagnostic results. To improve data quality and interpretation accuracy, denoising of noisy data is usually required in the preprocessing stage. Existing denoising methods, such as predictive filtering, sparse representation, and variational optimization, often rely on physical models or statistical priors, and their effectiveness is often sensitive to parameter settings and prior assumptions. When the statistical characteristics of noise change with time or spatial location, or when there are differences in noise variance across different regions, prior mismatch can easily occur, leading to a decline in denoising performance. In recent years, deep learning-based denoising methods have been used in related tasks, among which the encoder-decoder U-shaped structure is widely adopted due to its multi-scale feature extraction and skip connection reconstruction mechanism.
[0003] However, existing denoising networks based on U-shaped structures still have certain limitations in engineering applications:
[0004] First, many networks use fixed convolutional blocks as the basic computational unit, and the convolutional kernel parameters remain unchanged during the inference phase, making it difficult to adaptively adjust the filtering / denoising intensity based on the changes in noise intensity and structural differences in different regions of the input data.
[0005] Secondly, in noisy scenarios, in order to reduce the overall reconstruction error during the decoding and reconstruction process, the network output often exhibits a smoothing trend, which leads to excessive suppression of high-frequency details and local structural information, resulting in problems such as blurred edges, discontinuous thin-layer structures, or loss of weak signal components.
[0006] Third, existing cue learning or conditional control mechanisms are mostly used in high-level semantic tasks such as classification and segmentation. When transferred to denoising and reconstruction tasks, they often take the form of attention weighting or feature selection. There is a lack of clear coupling path between the cue signal and the parameter adjustment of the denoising operator, making it difficult to provide controllable denoising strategy modulation. Based on this, it is necessary to provide a denoising method that can achieve adaptive gating adjustment during U-shaped reconstruction and introduce controllable modulation signals in the decoding upsampling stage to improve the ability to preserve details. Summary of the Invention
[0007] The purpose of this invention is to solve the problems of existing U-shaped topology-based seismic data denoising networks, which use fixed-parameter convolutional blocks as basic computational units, lack a gating adjustment mechanism that dynamically correlates with the statistical characteristics of input features and noise levels, and have insufficient coupling between the learning signal and the parameter generation path of the denoising operator. These problems lead to over-smoothing, loss of details of weak reflection events and thin-layer structures, edge blurring, and inability to adaptively adjust the denoising intensity under strong noise and non-stationary noise (such as spatial / temporal variation noise).
[0008] To achieve the above objectives, the present invention employs the following technical means:
[0009] This invention provides a self-supervised denoising network based on a U-shaped encoder-decoder topology, comprising:
[0010] The encoding path, bottleneck layer, decoding path, and skip connection for passing features between the encoding path and the decoding path are connected in sequence.
[0011] The encoding path includes multiple levels of encoding units, and each level of encoding unit includes:
[0012] The downsampling operator and at least one batch normalized gated time-varying denoising module (BN-GTVD) are used to reduce the spatial resolution of the input features.
[0013] The bottleneck layer includes at least one BN-GTVD module for global context enhancement of the low-resolution features output by the encoding path.
[0014] The decoding path includes multiple levels of decoding units, each level of decoding unit including:
[0015] Upsampling operator: Configured to upsample the output of the previous decoding unit or the bottleneck layer;
[0016] Skip-connect feature fusion module: configured to fuse the upsampled features output by the upsampled module with the encoded features passed from the corresponding level of the encoding path through skip connections to generate decoded fusion features. ;
[0017] Visual Cue Learning Module (VPLM): Located after the skip connection feature fusion module and before the BN-GTVD module, it provides cue parameters. and the prompt parameters Mapping to prompt embedding Using the aforementioned prompt embedding For the decoding fusion features Conditional modulation is performed to obtain modulation characteristics. ;
[0018] And at least one BN-GTVD: the BN-GTVD module receives the modulation features after conditional modulation. As input, its gated generation branch incorporates the hint embedding in the decoding path. By dynamically adjusting the gating coefficient, the embedded input prompts are ignored in the encoding path and bottleneck layer. ;
[0019] The network also includes an output layer for outputting denoising results or noise residual estimates.
[0020] In the above scheme, the Visual Cue Learning (VPLM) module decodes and fuses features. The conditional modulation satisfies one of the following equations:
[0021] (a) Additive modulation is as follows:
[0022]
[0023] in This indicates that the prompt is embedded. express;
[0024] (b) The splicing and fusion modulation is shown below: where, Indicates channel-dimensional splicing. To fuse convolutions:
[0025]
[0026] (c) Affine modulation is shown below:
[0027]
[0028] in This indicates element-wise multiplication. and For the prompt parameters Affine modulation parameters obtained through mapping.
[0029] In the above scheme, the BN-GTVD includes a gated generation branch and a candidate filtering branch, and the gated generation branch outputs gating coefficients. The candidate filter branch outputs the candidate filter response. The gated modulation response is obtained through element-wise modulation. And by modulating element by element, the following formula is obtained:
[0030]
[0031] The gating coefficient of the BN-GTVD The generation satisfies the following equation:
[0032]
[0033] in, These are intermediate features obtained from input features through convolution and nonlinear activation. For gated branch convolution operators, Used to embed prompts Mapped to an additional item consistent with the gated branch channel. Use the Sigmoid activation function;
[0034] The candidate filter response of the BN-GTVD satisfies the following equation:
[0035]
[0036] in For candidate filter branch convolution operators;
[0037] The output of the BN-GTVD satisfies the formula:
[0038]
[0039]
[0040] in, For the input features of BN-GTVD, This is the response after gating modulation. For batch normalization operators.
[0041] In the above scheme, the prompting parameters It can be a learnable cue vector or a learnable cue graph; or it can be generated by a noise intensity estimation branch and correlated with the input noise level.
[0042] In the above scheme, the prompting parameters Mapped to cue embedding via 1×1 convolution or multilayer perceptron and make The number of channels is consistent with the number of channels in the decoding fusion feature.
[0043] In the above scheme, the skip connection feature fusion module uses channel splicing followed by 1×1 convolution fusion, or uses element-wise summation fusion.
[0044] In the above scheme, the output layer output noise residual estimation The denoising result is obtained through residual learning:
[0045]
[0046] in To input noisy data, This is the result of noise reduction.
[0047] This invention provides a denoising method, characterized in that it uses the self-supervised denoising network described in any one of claims 1-7 to process two-dimensional noisy data and outputs denoising results or noise residual estimates.
[0048] The present invention also provides a self-supervised training method, which trains a self-supervised denoising network using noisy data, the method comprising:
[0049] S1: Acquire noisy two-dimensional data ;
[0050] S2: Generate a binary mask And construct the network input:
[0051]
[0052] in Fill with zero or noise;
[0053] S3: Will Input network to get output Or noise residual estimation ;
[0054] S4: Calculate the mask reconstruction loss:
[0055]
[0056] Training period Because there is a mask; there is no mask during inference. Returning naturally ;
[0057] S5: Update the network parameters through backpropagation based on the loss;
[0058] The training may optionally further include consistency loss and / or noise constraint loss.
[0059] Because the present invention employs the above-mentioned technical means, it has the following beneficial effects:
[0060] (1) Dynamic gated denoising: The BN-GTVD module modulates the candidate filter response element by element through the gate coefficient, so that the network can adaptively adjust the denoising intensity according to the difference between the input feature content and the local noise, thereby improving the adaptability to non-stationary noise and the difference in noise variance in different regions.
[0061] (2) Achieving consistency with training: By using a combination of batch normalization and smooth activation function in the BN-GTVD module, the stability of feature distribution can be improved and the risk of gradient fluctuation can be reduced under batch training conditions;
[0062] (3) Conditional reconstruction in the decoding stage: The VPLM module introduces a cue signal in the upsampling reconstruction stage to conditionally modulate the reconstruction features of different scales or different regions, thereby reducing excessive smoothing and loss of details in strong noise scenes;
[0063] (4) Hint-gated cooperative modulation: Hint embedding can be used to fuse the feature distribution of modulation and decoding, and can also be coupled to the gating branch to affect the gating coefficient generation process, so as to realize the joint control of policy hints and operator gating;
[0064] (5) Engineering feasibility: The network maintains a U-shaped topology and the computing units are uniform, which facilitates reuse and deployment in different two-dimensional signal denoising tasks. Attached Figure Description
[0065] Figure 1 This is a flowchart of the method of the present invention;
[0066] Figure 2 This is a schematic diagram of the overall structure of the U-shaped self-supervised denoising network of the present invention;
[0067] Figure 3 This is a schematic diagram of the BN-GTVD module.
[0068] Figure 4 This is a schematic diagram of the VPLM module.
[0069] Figure 5 This is a schematic diagram illustrating the relationship between "upsampling—fusion—cue modulation—BN-GTVD refinement" in the decoding stage;
[0070] Figure 6 This diagram illustrates the effect of denoising test samples using the model trained with the four-level U-shaped self-supervised denoising network structure described in Example 1 and the self-supervised training strategy described in Example 2. Detailed Implementation
[0071] The embodiments of the present invention will be described in detail below. Although the present invention will be described and illustrated in conjunction with some specific embodiments, it should be noted that the present invention is not limited to these embodiments. On the contrary, any modifications or equivalent substitutions made to the present invention should be covered within the scope of the claims of the present invention.
[0072] Furthermore, to better illustrate the present invention, numerous specific details are set forth in the following detailed embodiments. Those skilled in the art will understand that the present invention can be practiced without these specific details.
[0073] This invention aims to provide a self-supervised denoising network and method based on a U-shaped encoder-decoder topology. While maintaining the U-shaped multi-scale feature extraction and skip connection reconstruction mechanism, this network replaces the traditional U-Net convolutional computation units with a batch normalized gated denoising module (BN-GTVD module). This allows the denoising weights to adaptively change according to the statistical characteristics and content differences of the input features, thereby achieving dynamic gated filtering in the feature domain. Furthermore, after upsampling and skip connection fusion at each stage of the decoding phase, this invention introduces a visual cue learning module (VPLM module). This module uses cue signals to conditionally modulate the decoded and reconstructed features and embeds the cue signals into the gated branches of the BN-GTVD module to implement a cue-guided dynamic denoising strategy, thereby reducing over-smoothing and improving structural fidelity under strong noise conditions.
[0074] To achieve the above objectives, the present invention adopts the following technical solution:
[0075] (1) Construct a U-shaped topology denoising network, the network including an encoding path, a bottleneck layer, a decoding path and a jump connection for cross-layer information transmission;
[0076] (2) In the encoding path, bottleneck layer and decoding path, the BN-GTVD module is uniformly set as the basic computing unit to replace the traditional U-Net convolutional block, so that the denoising filtering intensity of each layer feature can be adaptively adjusted according to the input feature content.
[0077] (3) In the decoding path, after each level of upsampling and fusion with the corresponding coding features, the VPLM module is set to generate a cue embedding, and the cue embedding is injected into the decoding fusion feature to achieve conditional modulation;
[0078] (4) Input the features after the prompt embedding modulation into the subsequent BN-GTVD module for refinement and reconstruction; In the implementation, the prompt embedding is mapped to an additional bias term and scale term consistent with the gating generation branch channel, and coupled to the gating coefficient generation process to form a prompt-gating cooperative modulation mechanism, so that the prompt signal directly affects the generation of the gating coefficient.
[0079] (5) The network is trained using a self-supervised training method to obtain a denoising model, and the trained model is used to perform denoising inference on actual noisy data.
[0080] Example 1: Construction of a four-level U-shaped self-supervised denoising network based on BN-GTVD and VPLM
[0081] Step 1: Dataset Construction
[0082] Noisy two-dimensional data collected in real-world scenarios are selected as training samples. This two-dimensional data includes seismic record slices, mechanical vibration time-frequency maps, or acoustic spectrograms. Optionally, Gaussian white noise, impulse noise, or coherent noise is superimposed on the samples according to a preset noise model to generate synthetically noisy data. The two-dimensional data is sliced into data blocks of uniform size (e.g., 128×128 or 256×256) according to a preset window, and divided into training, validation, and test sets. The data blocks are normalized and subjected to one or more data augmentation operations, including rotation, flipping, and random pruning.
[0083] Step 2: Network Construction (Encoding Path and Bottleneck Layer)
[0084] like Figure 2 As shown, a four-level U-shaped topology network is constructed, including encoding layers E1 to E4, bottleneck layer B, and decoding layers D4 to D1. Each level in the encoding path includes a downsampling operator. With at least one BN-GTVD module, in this embodiment two BN-GTVD modules are set for each stage; the downsampling operator 3×3 convolutional downsampling with a stride of 2 or 2×2 average pooling downsampling can be used. Bottleneck layer B is set between the encoding and decoding paths. In this embodiment, four BN-GTVD modules are cascaded to enhance the ability to represent low-resolution global structures.
[0085] Step 3: Decoding path construction (including VPLM location)
[0086] Each stage in the decoding path includes an upsampling operator. The system includes skip connection feature fusion, a VPLM module, and at least one BN-GTVD module. In this embodiment, two BN-GTVD modules are set for each stage. The upsampling operator... Bilinear interpolation upsampling or transposed convolutional upsampling can be used. The skip connection feature fusion is used to fuse the features of the corresponding coding level with the upsampled decoded features; in one implementation, channel concatenation followed by 1×1 convolutional fusion is used, while in another implementation, element-wise summation fusion is used. The VPLM module is set "after upsampling and skip connection fusion, and before entering BN-GTVD refinement" to perform cue-driven conditional modulation of the decoded fusion features.
[0087] Step 4: Output Layer
[0088] At the decoding end, a 1×1 convolution is set to map the features to the target number of output channels, resulting in the network output. The network output can be configured to directly output the denoised image or output a noise residual estimate. When configured to output a noise residual estimate, the denoising result can be obtained based on the input noisy data and the noise residual estimate, thereby reducing false suppression of structural components.
[0089] Symbol and terminology conventions:
[0090] In this specification, the input noisy two-dimensional data or its feature representation is denoted as... .
[0091] Network output configured for output noise residual estimation The denoising result is obtained through residual learning. During the decoding stage, the features fused after skip connections are denoted as... The features after embedding and modulation are denoted as follows: The prompt parameter is denoted as... The cue parameter P can be either an independent cue for each layer or a shared cue across scales. The mapped hint embedding is denoted as In the BN-GTVD module, intermediate activation features are denoted as... The gating coefficient is denoted as The candidate filter response is denoted as The gated modulation response is denoted as The module output is denoted as .
[0092] Structure and Calculation Steps of the BN-GTVD Module
[0093] Design motivation: The noise in 2D seismic records is non-stationary, and the "average filtering intensity" of the fixed convolution block causes the strong noise area to be over-smoothed, and weak reflections or thin layer edges to be erased.
[0094] Structural approach: The candidate filtering branch is given to the "denoised response", and the gating branch is given to the [0,1] gating coefficients for element-wise modulation; batch normalization is used to stabilize the distribution and reduce gradient fluctuations, making the gating more reliable in distinguishing between noise-dominated and structure-dominated systems.
[0095] The achieved results are: gating changes with the input to achieve "strong noise enhancement and suppression, and structural edge reduction and suppression"; and it emphasizes that the gating at the decoding end is further coupled with the embedding of prompts to form a key combination of "BN stability + dynamic gating filtering + (decoding end) prompt coupling".
[0096] The BN-GTVD module is used to replace traditional convolutional blocks and implement gated modulation denoising computation. Let the input feature map be X, and the convolution operator be... Its convolution kernel parameters are The bias is set to The BN-GTVD module includes a feature mapping submodule, a nonlinear activation submodule, a gated generation submodule, a candidate filtering submodule, and a residual fusion and normalization submodule. In one implementation, the BN-GTVD module is calculated according to the following relationship:
[0097] (1) Feature mapping:
[0098]
[0099] (2) Nonlinear activation:
[0100] The feature mapping result Applying the SiLU activation function yields intermediate activation features. :
[0101]
[0102] in , This refers to the Sigmoid function.
[0103] (3) Gating generation:
[0104]
[0105] (4) Candidate filtering:
[0106]
[0107] (5) Gated modulation:
[0108]
[0109] (6) Residual fusion and normalization:
[0110]
[0111] In the case where no hint embedding is introduced in the encoding path and bottleneck layer, let .
[0112] Structure, principle and placement of VPLM module
[0113] The "visual cue" described in this invention represents a conditional modulation cue for a two-dimensional signal tensor in the spatial dimension. It is not limited to natural images and is applicable to two-dimensional data such as seismic slices, time-frequency maps, spectrograms, and medical imaging. The VPLM module is positioned at each stage of the decoding path: after upsampling, after fusion with the corresponding coded features, and before entering the BN-GTVD module. It provides external cue signals to the decoding stage to achieve conditional reconstruction. Let the decoding fusion feature be... The prompt parameter is The VPLM module can be implemented using the following steps:
[0114] (1) Prompt parameter learning: In this embodiment, the prompt parameters It can be set as a learnable cue vector or a learnable cue graph, which can generate cues for each layer by using an independent cue vector P or by sharing a cue vector and scale embedding; optionally, the cue parameters It can also be generated by the noise intensity estimation branch and correlated with the input noise level;
[0115] (2) Prompt embedding mapping: The prompt parameters Mapped to features via 1×1 convolution or multilayer perceptron. Consistent channel count prompts embedded ;
[0116] (3) Feature modulation: embedding the cue Injecting feature T to obtain modulated features In this embodiment, Additive modulation can be used. Alternatively, splicing and fusion can also be adopted. Conditional modulation can be achieved through methods such as affine modulation.
[0117] Hint—Gated Co-modulation
[0118] To achieve cue-gated cooperative modulation, in one implementation, the cue is embedded... This is further mapped to additional bias or scaling terms in the gating branch and coupled to the gating generation process; the gating generation is then rewritten as:
[0119]
[0120] in Embedded by prompt The scale term obtained from the mapping, This is the bias term obtained from the hint embedding map.
[0121] Step 5: Verify the noise reduction effect (corresponding to) Figure 6 )
[0122] In this embodiment, the four-level U-shaped network structure described in Embodiment 1 is used in conjunction with the self-supervised training strategy described in Embodiment 2 for training and validation. For ease of reproduction, the method for generating... Figure 6 The resulting engineering configuration is as follows: the number of channels at the four levels of the encoder is 32, 64, 128, and 256 respectively; the number of channels at the bottleneck layer is 512; and the number of channels at the decoder is symmetrically set to 256, 128, 64, and 32. Each level of the encoder has two BN-GTVD modules, and the bottleneck layer has four BN-GTVD modules connected in series. The BatchNormalization momentum is set to 0.1, and the epsilon is set to... The suggested channel number is set to 16 or 32 (optional), and the multi-scale shared temporal embedding dimension is set to 8 or 16 (optional). The batch size during training is set to 16, and the number of training epochs is set to 100 to 300; the optimizer is AdamW, and the initial learning rate is 1*10^6. -3 The optimal model parameters were selected using cosine annealing or step decay and early stopping strategies. The trained model was then used to denoise test samples, and the resulting visualized denoising effect is shown below. Figure 6 As shown.
[0123] Example 2: Self-supervised training strategy and loss function design
[0124] The self-supervised training phase uses only noisy data as training samples to reduce reliance on clean labeled data. In one implementation, at least a mask reconstruction loss is used for training; and optionally, a consistency loss and a noise constraint loss are superimposed to improve stability.
[0125] (a) Mask reconstruction loss: for input data Generate a binary mask Constructing network input ,in Fill with zero or noise. Inputting the network yields noise residual estimates. , and by The denoising result is obtained, and the loss is
[0126] =
[0127] (ii) Consistency Loss: Obtained by applying two random enhancements to the same sample. and Inputting them into the network yields the following results: and and in response to , After applying the corresponding enhanced inverse transform, the constraints are consistent between the two.
[0128]
[0129] in This represents the inverse transform corresponding to data augmentation.
[0130] (iii) Noise constraint loss: When the network is configured to output noise residual At that time, it can be used for Apply statistical or frequency domain regularization constraints to reduce the risk of misclassifying structural components as noise.
[0131] The training process uses backpropagation to update network parameters, and the optimizer is selected. We also set learning rate decay and early stopping strategies to obtain a model with stable convergence performance.
[0132] Example 3: Multi-cue strategy and multi-scale cue sharing mechanism
[0133] To improve the adaptability of the cue signal to different decoding scales, cue vectors can be set separately at different decoding levels, or a shared cue vector can be set. And combined with scale embedding Generate scale-related prompts. In the implementation, ,in Embed the scale number; then... Mapped to this level Modulation required and and for the first Level decoding fusion feature execution This approach allows shallow decoding to focus more on local detail cues, while deep decoding focuses more on structural continuity cues, thus preserving structural edges while suppressing random noise.
[0134] Additionally, a noise level map R can be optionally constructed from the output of the noise intensity estimation subnetwork, and generated using prompts driven by R, for example... The MLP is then used to obtain the cue vector, which is then correlated with the input noise level to form an adaptive denoising strategy.
[0135] Reasoning steps
[0136] During the inference phase, the trained model parameters are loaded, and the noisy 2D data to be denoised is input into the network. The denoised result is output after passing through the encoding path, bottleneck layer, and decoding path. When the input is large-format 2D data, a sliding window block inference method can be used, and weighted fusion of overlapping areas can be performed to reduce block boundary artifacts.
[0137] Optional deformation and equivalent replacement
[0138] Without departing from the core concept of this invention, the following equivalent substitutions can be made:
[0139] (1) The convolution transformation in the BN-GTVD module can be a combination of 1×1 and 3×3, or depthwise separable convolution can be used to reduce the amount of computation.
[0140] (2) The upsampling operator can be replaced by interpolation or transpose convolution with subpixel upsampling;
[0141] (3) The modulation method can be replaced by splicing and fusion, indicating that additive modulation can be replaced by splicing and fusion. Affine modulation or gated modulation, but the prompt signal should be injected as external strategy information into the decoding feature or gated branch;
[0142] (4) The training loss function can be selected according to the task, such as L1, L2, Charbonnier loss or frequency domain consistency loss.
[0143] Additional notes: Engineering parameter configuration and effect description
[0144] To ensure full transparency and ease of engineering implementation, a set of optional network parameters and training configurations is provided:
[0145] (1) Number of channels: The number of channels at the four levels of the encoding end can be set to 32, 64, 128, and 256 respectively, and the bottleneck layer can be set to 512; the number of channels at the decoding end is symmetrically set to 256, 128, 64, and 32. The input and output channels of each level of the BN-GTVD module are consistent to facilitate residual connection.
[0146] (2) Convolution kernel and padding: When using 3×3 convolution, padding is 1 to maintain spatial size; when using 1×1 convolution, spatial size remains unchanged.
[0147] (3) Batch Norm parameters: Momentum can be set to 0.1, epsilon to... When the batch size is small, synchronous batch normalization or frozen statistics can be used.
[0148] (4) Prompt Dimension: Number of prompt channels It can be set to 16 or 32; when using a multi-scale sharing mechanism, the scale embedding dimension can be set to 8 or 16.
[0149] (5) Hint injection location: In addition to modulating the decoding fusion features in the VPLM module, the hint can also be embedded into the additional bias term mapped to the gating branch and coupled to the gating generation process to improve the sensitivity of the gating coefficient to changes in noise level.
[0150] (6) Training configuration: Under the condition of 128×128 slices, the batch size can be set to 8 to 32; the number of training rounds can be set to 100 to 300, and the optimal model parameters are selected based on the validation set loss; the learning rate is set to... And it is used in conjunction with cosine annealing or step decay.
[0151] (7) Evaluation indicators: PSNR, SSIM and other indicators can be used; in the seismic data scenario, signal-to-noise ratio improvement or event continuity indicators can be further adopted, and in the mechanical vibration scenario, task-related indicators such as fault feature retention can be adopted.
[0152] (8) Effect description: The gated modulation mechanism of BN-GTVD can suppress the noise-dominant component and retain the structural-dominant component in the feature domain; VPLM provides external cues, enabling the decoding upsampling stage to adopt differentiated reconstruction intensity in different noise regions, thus making it less prone to large-area smoothing compared to traditional convolutional U-shaped networks, and more conducive to preserving structural edges and weak reflection components.
[0153] (9) Deployment: The model can be deployed on GPUs or edge devices; when resources are limited, the number of BN-GTVDs per stage can be reduced or depthwise separable convolutions can be used, and inference latency can be reduced through quantization or pruning.
[0154] (10) Robustness: Through data augmentation, consistency constraints and cue modulation mechanisms, the model’s adaptability to different acquisition conditions and noise distribution changes can be improved.
Claims
1. A self-supervised denoising network based on a U-shaped encoder-decoder topology, characterized in that, include: The encoding path, bottleneck layer, decoding path, and skip connection for passing features between the encoding path and the decoding path are connected in sequence. The encoding path includes multiple levels of encoding units, and each level of encoding unit includes: The downsampling operator and at least one batch normalized gated time-varying denoising module (BN-GTVD) are used to reduce the spatial resolution of the input features. The bottleneck layer includes at least one BN-GTVD module for global context enhancement of the low-resolution features output by the encoding path. The decoding path includes multiple levels of decoding units, each level of decoding unit including: Upsampling operator: Configured to upsample the output of the previous decoding unit or the bottleneck layer; Skip-connect feature fusion module: configured to fuse the upsampled features output by the upsampled module with the encoded features passed from the corresponding level of the encoding path through skip connections to generate decoded fusion features. ; Visual Cue Learning Module (VPLM): Located after the skip connection feature fusion module and before the BN-GTVD module, it provides cue parameters. and the prompt parameters Mapping to prompt embedding Using the aforementioned prompt embedding For the decoding fusion features Conditional modulation is performed to obtain modulation characteristics. ; And at least one BN-GTVD: the BN-GTVD module receives the modulation features after conditional modulation. As input, its gated generation branch incorporates the hint embedding in the decoding path. By dynamically adjusting the gating coefficient, the embedded input prompts are ignored in the encoding path and bottleneck layer. ; The network also includes an output layer for outputting denoising results or noise residual estimates.
2. The self-supervised denoising network according to claim 1, characterized in that, The Visual Cue Learning Module (VPLM) decodes and fuses features. The conditional modulation satisfies one of the following equations: (a) Additive modulation is as follows: in This indicates that the prompt is embedded. express; (b) The splicing and fusion modulation is shown below: where, Indicates channel-dimensional splicing. To fuse convolutions: (c) Affine modulation is shown below: in This indicates element-wise multiplication. and For the prompt parameters Affine modulation parameters obtained through mapping.
3. The self-supervised denoising network according to claim 2, characterized in that, The BN-GTVD includes a gated generation branch and a candidate filtering branch, with the gated generation branch outputting gating coefficients. The candidate filter branch outputs the candidate filter response. The gated modulation response is obtained through element-wise modulation. And by modulating element by element, the following formula is obtained: The gating coefficient of the BN-GTVD The generation satisfies the following equation: in, These are intermediate features obtained from input features through convolution and nonlinear activation. For gated branch convolution operators, Used to embed prompts Mapped to an additional item consistent with the gated branch channel. Use the Sigmoid activation function; The candidate filter response of the BN-GTVD satisfies the following equation: in For candidate filter branch convolution operators; The output of the BN-GTVD satisfies the formula: in, For the input features of BN-GTVD, This is the response after gating modulation. For batch normalization operators.
4. The self-supervised denoising network according to claim 1, characterized in that, The prompt parameters It can be a learnable cue vector or a learnable cue graph; or it can be generated by a noise intensity estimation branch and correlated with the input noise level.
5. The self-supervised denoising network according to claim 1, characterized in that, The prompt parameters Mapped to cue embedding via 1×1 convolution or multilayer perceptron and make The number of channels is consistent with the number of channels in the decoding fusion feature.
6. The self-supervised denoising network according to claim 1, characterized in that, The skip connection feature fusion module uses channel splicing followed by 1×1 convolution fusion, or element-wise summation fusion.
7. The self-supervised denoising network according to claim 1, characterized in that, Output layer output noise residual estimation The denoising result is obtained through residual learning: in To input noisy data, This is the result of noise reduction.
8. A noise reduction method, characterized in that, The self-supervised denoising network according to any one of claims 1-7 is used to process two-dimensional noisy data and output denoising results or noise residual estimates.
9. A self-supervised training method, characterized in that, The method of training the self-supervised denoising network according to any one of claims 1 to 10 using noisy data includes: S1: Acquire noisy two-dimensional data ; S2: Generate a binary mask And construct the network input: in Fill with zero or noise; S3: Will Input network to get output Or noise residual estimation ; S4: Calculate the mask reconstruction loss: Training period Because there is a mask; there is no mask during inference. Returning naturally ; S5: Update the network parameters through backpropagation based on the loss; The training may optionally further include consistency loss and / or noise constraint loss.