A D-NeRF image denoising method and system based on MSAF-DT
By introducing the MSAF-DT method in D-NeRF, the multi-scale fusion attention mechanism and the Transformer module are used for image noise reduction, the noise problem of D-NeRF in sparse view images is solved, and efficient and accurate image noise reduction effect is achieved.
Patent Information
- Application Number
- CN202411838558.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-13
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2044-12-13
AI Technical Summary
D-NeRF has serious noise problems when processing sparse view images, which affects the realism of the image and the performance of downstream tasks. Traditional noise reduction methods cannot effectively deal with complex noise patterns.
Using an image noise reduction method based on MSAF-DT, RGB images are generated through the D-NeRF model, and feature extraction, fusion and noise reduction are used to use the multi-scale fusion attention mechanism (MSAF) and Transformer modules. Combined with BatchNorm and multi-head self-attention mechanism, noise characteristics are accurately modeled and complex noise patterns are suppressed.
It significantly improves the noise reduction quality and processing efficiency of D-NeRF images, generates high-quality images with rich details and real visual effects, reduces training time and optimizes noise reduction accuracy.
Smart Images

Figure CN119295342B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image noise reduction, and in particular to a D-NeRF image noise reduction method and system based on MSAF-DT. Background Art
[0002] As an innovative 3D scene reconstruction technology, NeRF can learn and reconstruct a complete 3D scene from a large number of 2D images, and render high-quality RGB images from different perspectives through a virtual camera. NeRF uses a volume rendering method based on deep learning to reconstruct realistic 3D visual effects based on scene properties such as lighting and volume density, and has shown great potential in multiple fields, including virtual reality, augmented reality, movie special effects production, and autonomous driving.
[0003] Although D-NeRF has shown excellent performance in processing dynamic scenes, it still faces a long-standing problem - image noise. Due to the sparsity of the input data set, especially when there is a lack of multi-view data, the images generated by D-NeRF often have severe noise, which affects the accuracy and quality of the final rendering effect. This noise not only reduces the realism of the image, but also may affect the performance of downstream tasks, such as obstacle detection in autonomous driving and immersion in VR. Therefore, the design of excitation methods, systems and devices faces the following challenges: (1) When processing non-rigid dynamic scenes, due to the sparsity of the input data, the images generated by D-NeRF are often affected by noise. Traditional image denoising methods such as transform domain processing and spatial filtering cannot effectively handle complex noise patterns, especially in the case of multi-view and high-dimensional data. (2) Traditional convolutional neural networks still have limitations in noise modeling and processing capabilities in high-dimensional feature spaces. How to improve denoising effects while maintaining high efficiency is still a bottleneck in the development of technology. (3) Conventional denoising methods usually require a large number of training iterations on high-resolution images, which is computationally intensive and leads to long training time. Summary of the invention
[0004] In order to solve the above-mentioned problems, the present invention provides a D-NeRF image denoising method and system based on MSAF-DT.
[0005] In the first aspect, the present invention provides a D-NeRF image denoising method based on MSAF-DT, which adopts the following technical solution:
[0006] A D-NeRF image denoising method based on MSAF-DT, comprising:
[0007] Get image dataset;
[0008] The acquired image dataset is processed based on the D-NeRF model to generate RGB images;
[0009] Feature extraction of RGB images based on MSAF-DT method;
[0010] Perform feature fusion on the extracted features based on BatchNorm;
[0011] Perform image denoising on the adapted features based on Transformer;
[0012] Perform dimension conversion on the denoised image data;
[0013] Output image.
[0014] Furthermore, the acquired image data set is processed based on the D-NeRF model to generate an RGB image, including using spatial mapping , the displacement changes of pixels at different time points in the data image are converted into the standardized scene configuration based on the adjusted coordinates ,The volume rendering technology is used to calculate the illumination intensity and volume density information and generate RGB images.
[0015] Furthermore, the method processes the acquired image data set based on the D-NeRF model to generate an RGB image, and also includes using a deformation network of the D-NeRF model Estimate the deformation field between the scene at a specific time point and the canonical spatial scene; use volume rendering equations to calculate non-rigid deformations in the 6D neural radiation field,
[0016] Among them, the expected color C of pixel p at time t is:
[0017]
[0018]
[0019]
[0020]
[0021] in, : Cumulative transmittance, indicating the transparency of the light from the starting point to h; p(h, t): 3D point coordinates in a non-rigid deformation scene, p(s, t): 3D coordinates of a point on the ray, d is the unit vector of the ray direction, s is the integral variable, indicating the depth parameter on the ray; σ(p(h, t)): indicates the density function, which is the volume density of the ray at point p(h, t); σ(p(s, t)): indicates the volume density of point p(s, t) at a depth of s on the ray, x(h) is a point emitted from the projection center O to the pixel P along the camera ray, It is the closest place to the camera. is the farthest point from the camera; the 3D point p(h, t) represents the point on the camera ray x(h), which passes through the deformation network Transformed into the canonical space, is from arrive The cumulative probability that the emitted ray does not hit any other particle, density and color c is given by the canonical network Predicted.
[0022] Furthermore, the feature extraction of the RGB image based on the MSAF-DT model includes converting the pixel information of the original image into an abstract feature representation through two convolutional layers Conv1 and Conv2 based on the MSAF-DT model, extracting global features through two 1x1 convolutional layers, and extracting feature information of different scales by adding BatchNorm and ReLU activation functions.
[0023] Furthermore, the extracted features are fused based on BatchNorm, including performing a convolution operation on the image features through a 1×1 convolution kernel and normalizing using BatchNorm to accelerate training and improve model stability, thereby obtaining a feature map. The BatchNorm formula is as follows:
[0024]
[0025] Here, x is the output of a 1×1 convolution, μ is the batch mean, σ is the batch standard deviation, and γ and β are learnable scaling and offset parameters.
[0026] Furthermore, the image denoising based on the adapted features of the Transformer includes projecting the feature map to the dimension required by the Transformer module, and obtaining a denoised feature sequence after processing using the Transformer Block.
[0027] Furthermore, the dimensionality conversion of the denoised image data includes restoring the dimensional order from (H×W, B, C) to (B, C, H×W) through a permutation operation; then splitting the H×W dimension back into H and W to restore the spatial structure of the feature map, and finally obtaining a feature map with a shape of (B, C, H, W).
[0028] In the second aspect, a D-NeRF image denoising system based on MSAF-DT includes:
[0029] The data acquisition module is configured to acquire an image data set;
[0030] The image processing module is configured to process the acquired image data set based on the D-NeRF model to generate an RGB image;
[0031] The feature extraction module is configured to extract features from the RGB image based on the MSAF-DT method;
[0032] The feature fusion module is configured to perform feature fusion on the extracted features based on BatchNorm;
[0033] A denoising module is configured to perform image denoising on the adapted features based on Transformer;
[0034] A conversion module is configured to perform dimension conversion on the denoised image data;
[0035] The output module is configured to output an image.
[0036] In a third aspect, the present invention provides a computer-readable storage medium storing a plurality of instructions, wherein the instructions are suitable for being loaded and executed by a processor of a terminal device for performing the MSAF-DT-based D-NeRF image denoising method.
[0037] In a fourth aspect, the present invention provides a terminal device including a processor and a computer-readable storage medium, wherein the processor is used to implement various instructions; the computer-readable storage medium is used to store multiple instructions, wherein the instructions are suitable for being loaded by the processor and executing the D-NeRF image denoising method based on MSAF-DT.
[0038] In summary, the present invention has the following beneficial technical effects:
[0039] The present invention significantly improves the denoising quality and processing efficiency of D-NeRF sparse view images by combining the multi-scale fusion attention mechanism (MSAF) with the Transformer-based denoising module (DenoiseTransformer). When processing sparsely sampled dynamic non-rigid scenes, this method can accurately model noise features, capture local and global features through multi-scale contextual attention, suppress a variety of complex noise patterns, and generate high-quality images with richer details and more realistic visual effects. Compared with traditional denoising technology, the MSAF-DT method has significant advantages in reducing training time and optimizing denoising accuracy. It can not only effectively reduce noise residue, but also greatly improve the dynamic performance of three-dimensional scenes, providing technical support for applications such as virtual reality (VR) and augmented reality (AR).
[0040] The present invention is highly adaptable and does not rely on precise ground truth geometry or multi-view data. It can achieve efficient operation only through sparse monocular views. The resulting optimization mechanism can automatically adapt to a variety of noise environments, enhancing the robustness of the algorithm. Verification on a variety of test data sets shows that this method not only surpasses traditional methods in image quality assessment indicators, but also significantly improves the stability and operating efficiency of the denoising model, providing a new technical path and application prospect for 3D scene reconstruction and dynamic rendering.
[0041] In addition, in the present invention, the Multi-Head Self-Attention mechanism is introduced into the DenoiseTransformer module to enhance the global feature modeling capability in the sparse view image denoising process. The Multi-Head Self-Attention mechanism divides the input features into several subspaces by calculating multiple self-attention functions in parallel, and captures the correlation between features in each subspace, thereby obtaining richer and more comprehensive feature information. Finally, the features in these subspaces are integrated to form a more powerful feature expression capability. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 is a schematic diagram of noise reduction of the overall method of embodiment 1 of the present invention;
[0043] Figure 2 is an architecture diagram of the core noise reduction layer Transformer Block of Embodiment 1 of the present invention;
[0044] Figure 3 is a schematic diagram of the evaluation index (PSNR) of Example 1 of the present invention;
[0045] Figure 4 It is a schematic diagram of the evaluation index (Loss) of Example 1 of the present invention. DETAILED DESCRIPTION
[0046] The present invention is further described in detail below in conjunction with the accompanying drawings.
[0047] Embodiment 1
[0048] Glossary:
[0049] MSAF-DT: Multiple-Scale-Adaptive-Feature Denoise Transformer, MSAF-DT is a novel denoising network specifically designed to address the noise problem inherent in DNeRF (Dynamic Neural Radiance Field) image rendering, especially when dealing with sparse input views.
[0050] D-NeRF: Dynamic-Neural Radiance Fields, D-NeRF (Dynamic-Neural Radiance Fields) is an extension of NeRF (Neural Radiance Fields) that is capable of modeling and rendering dynamic scenes. NeRF itself is good at representing static scenes, but it is powerless for scenes containing motion. D-NeRF overcomes this limitation, enabling it to reconstruct and render realistic dynamic scenes from sparse monocular videos.
[0051] Reference Figure 1 , a D-NeRF image denoising method based on MSAF-DT in this embodiment includes:
[0052] Get image dataset;
[0053] The acquired image dataset is processed based on the D-NeRF model to generate RGB images;
[0054] Feature extraction of RGB images based on MSAF-DT method;
[0055] Perform feature fusion on the extracted features based on BatchNorm;
[0056] Perform image denoising on the adapted features based on Transformer;
[0057] Perform dimension conversion on the denoised image data;
[0058] Output image.
[0059] Specifically, the following methods are included:
[0060] S1. Obtain image dataset;
[0061] The dataset is the public image dataset Nerf_synthetic.
[0062] S2. Process the acquired image dataset based on the D-NeRF model to generate an RGB image;
[0063] Among them, D-NeRF technology is used to synthesize sparse views of dynamic non-rigid scenes to generate images at different time points in the three-dimensional scene. By learning the displacement relationship of each point in the scene at different time points, the spatiotemporal representation of the dynamic scene is achieved. , transforming the displacement changes at different time points into the canonical scene configuration based on the adjusted coordinates ,Use volume rendering technology to calculate information such as light intensity and volume density to generate RGB images;
[0064] The specific steps of the D-NeRF image generation stage are:
[0065] The D-NeRF model consists of two networks: the canonical network and deformation network Composition. Deformation Network It is optimized to estimate the deformation field between the scene at a specific time point and the canonical space scene. To simplify the processing, in all experiments, the canonical scene is set to the scene at t = 0, and the calculation formula is as follows:
[0066]
[0067] in is The displacement of the trained output.
[0068] In D-NeRF, the definition of the canonical space is fixed to the static state of the scene at time t=0. The deformation network Ψt is responsible for calculating the position deformation of each point in the scene over time. Specifically, the deformation network outputs a deformation field Ψt according to the current time t, which represents the displacement Δx of each point relative to the canonical space. If time t=0, the output of the deformation network is zero.
[0069] The core process of image generation involves calculating the expected color C(p, t) of each pixel through volume rendering. This process requires simulating the passage of light through various layers in the scene, taking into account the effects of light absorption, scattering, etc.
[0070] Volume rendering is used to explain the non-rigid deformation in the proposed 6D neural radiation field. Non-rigid deformation means that the object does not maintain its original rigid structure during the deformation process, that is, the various parts of the object move or stretch relative to each other, resulting in a change in the shape of the object, but does not involve the rupture or breakage of the material.
[0071] Then the expected color C of pixel p at time t is:
[0072]
[0073]
[0074]
[0075]
[0076] in, : Cumulative transmittance, indicating the transparency of the light from the starting point to h; p(h, t): 3D point coordinates in a non-rigid deformation scene, p(s, t): 3D coordinates of a point on the ray, d is the unit vector of the ray direction, s is the integral variable, indicating the depth parameter on the ray; σ(p(h, t)): indicates the density function, which is the volume density of the ray at point p(h, t); σ(p(s, t)): indicates the volume density of point p(s, t) at a depth of s on the ray, x(h) is a point emitted from the projection center O to the pixel P along the camera ray, It is the closest place to the camera. is the farthest point from the camera; the 3D point p(h, t) represents the point on the camera ray x(h), which passes through the deformation network Transformed into the canonical space, is from arrive The cumulative probability that the emitted ray does not hit any other particle, density and color c is given by the canonical network Predicted.
[0077] S3. Feature extraction of RGB images based on MSAF-DT method;
[0078] In this stage, a denoising network (MSAF-DT) based on multi-scale fusion attention and Transformer architecture is designed to reduce the noise caused by sparse sampling and enhance the realism and detail expression of the image.
[0079] Through the MSAF module, local context and multi-level features are captured from the image to enhance the recognition ability of complex noise patterns. The Transformer-based denoising module uses a multi-head self-attention mechanism to extract image noise features from a global scale, eliminate global inconsistency noise, and retain details in the image.
[0080] The specific steps of the image denoising stage are:
[0081] The proposed MSAF-DT method is composed of four modules: feature extraction layer, feature enhancement layer, feature adaptation layer and core denoising layer.
[0082] Images are usually represented in the form of tensors with dimensions (B, C, H, W), where B is the batch size, C is the number of channels, H is the image height, and W is the image width. The two convolutional layers Conv1 and Conv2 extract features from the image, such as edges, textures, object shapes, contours, etc. The two convolutional layers convert the pixel information of the original image into a more abstract feature representation in preparation for subsequent denoising operations. Both convolutional layers use the ReLU activation function to introduce nonlinearity, as shown in the following formula:
[0083] ReLU(x)=max(0,x).
[0084] S4. Perform feature fusion on the extracted features based on BatchNorm;
[0085] Among them, the features are globally averaged pooled, global features are extracted through two 1x1 convolutional layers, and BatchNorm and ReLU activation functions are added. Then, the features are multi-scale fused using the MSAF_2D module. The MSAF_2D module contains multiple branches, each of which downsamples the features to different degrees, and then processes them using two 1x1 convolutional layers, and adds BatchNorm and ReLU activation functions. These branches are responsible for extracting information at different scales, such as local details, medium-scale structures, and global information. Finally, the outputs of all branches, including global features, are added, and then passed through the Sigmoid function to obtain the attention weight. This weight is used to fuse the original features and multi-scale contextual information. The formula for the Sigmoid activation function is given below:
[0086]
[0087] This makes the output in the range of (0, 1), which can be used as a weight for weighted calculation of feature maps.
[0088] The convolution kernel is used to operate without changing the spatial dimension of the feature map. Each convolution kernel has a set of weights and biases. It is usually normalized using BatchNorm after the 1×1 convolution kernel to speed up training and improve model stability. The BatchNorm formula is as follows:
[0089]
[0090] Here, x is the output of a 1×1 convolution, μ is the batch mean, σ is the batch standard deviation, and γ and β are learnable scaling and offset parameters.
[0091] S5. Perform image denoising on the adapted features based on Transformer;
[0092] Combine the two dimensions H and W into one dimension H×W, which represents the location of all pixels in the image. At this point, the shape of the feature map becomes (B, C, H×W), which is equivalent to flattening the image into a pixel sequence. Next, use the permutation operation to convert the dimension order from (B, C, H×W) to (H×W, B, C).
[0093] The encoder projects the feature map to the dimensions required by the Transformer module. In each TransformerBlock, there are multi-head self-attention mechanisms, feedforward neural networks, layer normalization, and residual connections. These blocks use the self-attention mechanism to capture the dependencies between pixels, learn, and perform noise reduction. After processing through multiple TransformerBlocks, the denoised feature sequence is obtained, and the shape is still (H×W, B, C).
[0094] First, the feature map of the image is flattened and converted into a pixel sequence. Given the original feature map shape of (B, C, H, W) (where B is the batch size, C is the number of channels, H and W are the height and width of the image respectively), through the flattening operation, the spatial dimensions H and W of the image are merged into one dimension, and the new feature map shape is (B, C, H×W). At this time, the position of each pixel has been converted into a sequence element. In this way, the feature map becomes a sequence of length H×W, ready to be input into the Transformer module.
[0095] Next, the dimension order of the feature map is converted from (B, C, H×W) to (H×W, B, C) through a permutation operation, that is, the batch dimension and the channel dimension are swapped. The purpose of this step is to treat each pixel as a time step, which is convenient for the Transformer model to process. At this point, each pixel position of the feature map is input into the Transformer module as an independent sequence element, maintaining the spatial dependency between pixels, and the input feature dimension has been adjusted to the format required by the Transformer.
[0096] In the Transformer encoder, multiple Transformer Blocks process these features in sequence. Each Transformer Block includes a multi-head self-attention mechanism, a feedforward neural network, layer normalization, and residual connections. The role of the multi-head self-attention mechanism is to capture the interdependence between different pixels. It calculates the relationship between each pixel and other pixels through three matrices: query, key, and value, and then learns the potential patterns and features in the image. Through the attention calculation of multiple heads, the model can focus on different pixel information at multiple spatial positions, gradually remove noise and extract useful features. Each Transformer Block also includes a feedforward neural network to further enhance the nonlinear expression ability of the model for better noise reduction.
[0097] After being processed by multiple Transformer Blocks, a denoised feature sequence with a shape of (H×W, B, C) is obtained. The next step is to rearrange the denoised features to restore them to the shape of the original image. Through the inverse permutation operation, the feature dimensions are restored to (B, C, H, W), thus obtaining the final denoised image. This process can effectively remove noise while retaining the details and structure of the image. Through such operations, the Transformer model can not only capture the dependencies between pixels, but also learn how to denoise and improve the quality of the image.
[0098] S6. performing dimension conversion on the denoised image data;
[0099] After obtaining the output result, the dimension order is restored from (H×W, B, C) to (B, C, H×W) through permutation operation. Then, the H×W dimension is split back into H and W to restore the spatial structure of the feature map. The final feature map shape is (B, C, H, W).
[0100] Effect evaluation stage: After the noise reduction is completed, a specially designed quality evaluation mechanism is used to quantitatively analyze the image denoising effect;
[0101] The image clarity, detail restoration and noise suppression capabilities are evaluated through subjective and objective indicators. The MSAF-DT method is compared with existing mainstream denoising methods (such as DnCNN and FFDNet) to verify the superiority of the model in processing D-NeRF scenarios.
[0102] Among them, after being processed by multiple encoder blocks in the Transformer module (including multi-head self-attention, feedforward neural network, layer normalization and residual connection), the shape of the output feature sequence is (H×W, B, C), and the features of each pixel are denoised. In order to restore the spatial structure of the image, the dimension of the feature map needs to be converted back from (H×W, B, C) to (B, C, H×W). This process is achieved through a permutation operation, the purpose of which is to restore the batch dimension and channel dimension to the correct position so that the features of each pixel are recombined into a complete image. Then, through a further Reshape operation, the flattened H×W dimension is split back into the original image height and width dimensions H and W, the spatial structure of the feature map is restored, and finally a denoised image with a shape of (B, C, H, W) is obtained.
[0103] As a further embodiment,
[0104] like Figure 1 As shown, the present invention designs a D-NeRF image denoising method, system and device based on MSAF-DT. The specific steps are as follows:
[0105] Step 1: The input image data enters the feature enhancement module after extracting preliminary features through two convolutional layers (Conv1 and Conv2). The feature enhancement module consists of multiple submodules, including the local attention module (Local att), the multi-scale contextual attention module (Context1, Context2, Context3) and the global attention module (Global att). These modules work together to capture the local and multi-scale contextual features in the input image, thereby strengthening the feature representation and improving the ability to model complex noise patterns.
[0106] Step 2: The multi-scale feature vector extracted by the feature enhancement module is passed to the feature adapter, which is used to reintegrate and adjust multiple features to make them more suitable for the processing requirements of the subsequent Transformer module. The feature adaptation stage not only ensures that the features extracted by different modules can be effectively integrated, but also improves the uniformity of feature input and computational efficiency.
[0107] Step 3: The optimized features output by the feature adapter are input into the Denoise Transformer module. Denoise Transformer is composed of multiple stacked Transformer Blocks. Each block models the global context of the image through a multi-head self-attention mechanism and a feedforward neural network module, captures the noise characteristics caused by sparse sampling, and gradually restores the detailed information of the image. With the stacking of Transformer Blocks, the denoising effect is gradually enhanced, achieving accurate processing of noise of different scales.
[0108] Step 4: The features processed by Denoise Transformer are passed to the Image Reconstructor module to restore a high-quality image after noise reduction optimization. This image not only has a significant noise suppression effect, but also retains the detailed features of the input data, providing a clearer and more realistic visual effect for the subsequent dynamic rendering of the 3D scene.
[0109] Figure 2 This is the architecture diagram of the Transformer Block of this method. The specific process is as follows:
[0110] In this method, the feature map of the input image is first deformed and permuted. The shape of the input feature map is (B, C, H, W) (where B is the batch size, C is the number of channels, and H and W are the height and width of the feature map). By merging the two dimensions H and W into a sequence length of H×W and adjusting the dimension order to (H×W, B, C), the conversion from two-dimensional spatial features to one-dimensional sequence features is achieved. This transformation formally embeds the spatial information of the image into the sequence, laying the foundation for the subsequent processing of the Transformer module, enabling it to efficiently capture global contextual relationships.
[0111] Next, the transformed sequence is input into one or more Transformer blocks for processing. Each Transformer block consists of modules such as Multi-Head Self-Attention, Feed-Forward Network, and Layer Normalization. These components can capture the dependencies between different positions in the sequence through the attention mechanism, and perform complex nonlinear transformations on the features, significantly enhancing the feature representation capabilities. The shape of the processed output sequence is (H×W, B, C), which contains the feature information enhanced by the Transformer. According to actual needs, the output sequence can be restored to a two-dimensional feature map through another "Reshape & Permute" module to adapt to subsequent image processing tasks.
[0112] Figure 3 The figure shows how the PSNR value changes with the increase of training during the entire 800K iterations of training using this model technology.
[0113] Figure 4 The figure shows how the loss value changes as the training increases during the entire 800K iterations of training using this model technology.
[0114] Embodiment 2
[0115] This embodiment provides a D-NeRF image denoising system based on MSAF-DT, including:
[0116] The data acquisition module is configured to acquire an image data set;
[0117] The image processing module is configured to process the acquired image data set based on the D-NeRF model to generate an RGB image;
[0118] The feature extraction module is configured to extract features from the RGB image based on the MSAF-DT method;
[0119] The feature fusion module is configured to perform feature fusion on the extracted features based on BatchNorm;
[0120] A denoising module is configured to perform image denoising on the adapted features based on Transformer;
[0121] A conversion module is configured to perform dimension conversion on the denoised image data;
[0122] The output module is configured to output an image.
[0123] A computer-readable storage medium stores a plurality of instructions, wherein the instructions are suitable for being loaded and executed by a processor of a terminal device, a D-NeRF image denoising method based on MSAF-DT.
[0124] A terminal device includes a processor and a computer-readable storage medium, wherein the processor is used to implement various instructions; the computer-readable storage medium is used to store multiple instructions, wherein the instructions are suitable for being loaded by the processor and executing the D-NeRF image denoising method based on MSAF-DT.
[0125] The above are all preferred embodiments of the present invention, and are not intended to limit the protection scope of the present invention. Therefore, any equivalent changes made based on the structure, shape, and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A D-NeRF image denoising method based on MSAF-DT, characterized in that: include: Get image dataset; The acquired image dataset is processed based on the D-NeRF model to generate RGB images; Based on the MSAF-DT method, the RGB image is subjected to image denoising. Specifically, the RGB image data is input, and preliminary features are extracted through two convolutional layers Conv1 and Conv2, and then enter the feature enhancement module; The feature enhancement module consists of multiple sub-modules, including the local attention module, the multi-scale contextual attention module, and the global attention module; The multi-scale feature vectors extracted by the feature enhancement module are passed to the feature adapter to reintegrate and adjust multiple features to adapt them to the processing requirements of the subsequent Transformer module; the optimized features output by the feature adapter are input to the DenoiseTransformer module. The Denoise Transformer is composed of multiple Transformer Blocks stacked together. Each block models the global contextual relationship of the image through a multi-head self-attention mechanism and a feedforward neural network module, captures the noise characteristics caused by sparse sampling, and gradually restores the detailed information of the image through denoising; the features processed by the Denoise Transformer are passed to the image reconstruction module to restore the high-quality image after denoising optimization; the image reconstruction module includes: dimensional conversion of the image data processed by the Denoise Transformer, and restores the dimension order from (H×W, B, C) to (B, C, H×W) through a permutation operation; then the H×W dimension is split back into H and W to restore the spatial structure of the feature map, and the final feature map shape is (B, C, H, W); the denoised image is output, where B is the batch size, C is the number of channels, H is the image height, and W is the image width.
2. The D-NeRF image denoising method based on MSAF-DT according to claim 1, characterized in that: The D-NeRF model is used to process the acquired image data set to generate an RGB image, including using a spatial mapping (x, y, z, t)→( ), converting the displacement changes of pixels at different time points in the data image into the standard scene configuration based on the adjusted coordinates ( ), using volume rendering technology to calculate the light intensity and volume density information and generate RGB images.
3. The D-NeRF image denoising method based on MSAF-DT according to claim 2, characterized in that: The method processes the acquired image data set based on the D-NeRF model to generate an RGB image, and also includes using a deformation network of the D-NeRF model Estimate the deformation field between the scene at a specific time point and the canonical spatial scene; use volume rendering equations to calculate non-rigid deformations in the 6D neural radiation field, Among them, the expected color C of pixel p at time t is: in, : Cumulative transmittance, indicating the transparency of the light from the starting point to h; p(h, t): 3D point coordinates in a non-rigid deformation scene, p(s, t): 3D coordinates of a point on the ray, d is the direction unit vector of the ray, s is the integral variable, indicating the depth parameter on the ray; σ(p(h, t)): indicates the density function, which is the volume density of the ray at point p(h, t); σ(p(s, t)): indicates the volume density of point p(s, t) at a depth of s on the ray, x(h) is a point emitted from the projection center O to the pixel P along the camera ray, It is the closest place to the camera. is the farthest point from the camera; the 3D point p(h, t) represents the point on the camera ray, which passes through the deformation network Transformed into the canonical space, is from arrive The cumulative probability that the emitted ray does not hit any other particle, density and color c is given by the canonical network Predicted.
4. A D-NeRF image denoising system based on MSAF-DT, executing a D-NeRF image denoising method based on MSAF-DT as claimed in claim 1, characterized in that: include: The data acquisition module is configured to acquire an image data set; The image processing module is configured to process the acquired image data set based on the D-NeRF model to generate an RGB image; A denoising module is configured to perform image denoising on the RGB image based on the MSAF-DT method; The output module is configured to output the denoised image.
5. A computer-readable storage medium storing a plurality of instructions, characterized in that: The instructions are suitable for being loaded by a processor of a terminal device and executing the method according to claim 1 .
6. A terminal device, comprising a processor and a computer-readable storage medium, wherein the processor is used to implement each instruction; and the computer-readable storage medium is used to store multiple instructions, characterized in that: The instructions are suitable for being loaded by a processor and executing the method as claimed in claim 1 .
Citation Information
Patent Citations
Transform-based multi-scale feature representation image defogging method
CN115953311A
Junk image denoising method based on multi-dimensional image information fusion
CN116543168A