A super-resolution dynamic neural radiation field method enhanced by six-plane feature representation

By constructing a super-resolution dynamic neural radiation field method enhanced with six-plane feature representation and combining spatiotemporal information, the problem of image blurring in dynamic neural radiation fields at high resolution is solved, achieving high-quality image generation and memory optimization.

CN119624775BActive Publication Date: 2025-10-28HANGZHOU DIANZI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411725650.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-28
Publication Date
2025-10-28
Estimated Expiration
2044-11-28

AI Technical Summary

Technical Problem

Existing dynamic neural radiation field methods suffer from blurring issues when synthesizing views with high-resolution input, especially in dynamic scenes.

Method used

A super-resolution dynamic neural radiation field method with six-plane feature representation enhancement is adopted. By constructing a low-resolution dynamic neural radiation field model and a feature enhancement network model, and combining the spatiotemporal information of previous and subsequent frames, a high-resolution viewpoint image is generated. The feature enhancement network is used to learn and enhance the low-resolution feature map, thereby improving image detail and realism.

Benefits of technology

It significantly improves the image generation quality in dynamic scenes, maintains high consistency and robust detail representation, reduces memory usage, and generates high-quality images in complex and dynamically changing scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119624775B_ABST
    Figure CN119624775B_ABST
Patent Text Reader

Abstract

This invention discloses a super-resolution dynamic neural radiation field method with enhanced six-plane feature representation. The method includes the following steps: 1. Constructing a low-resolution dynamic neural radiation field model to generate an initial six-plane feature map; 2. Constructing a feature enhancement network model to generate an enhanced six-plane feature map; 3. Performing high-resolution viewpoint synthesis based on the enhanced six-plane feature map; 4. Training the super-resolution dynamic neural radiation field model; 5. Testing the trained super-resolution dynamic neural radiation field model on a test dataset. This invention integrates temporal and spatial information into a super-resolution dynamic neural radiation field framework. The six-plane feature representation-enhanced super-resolution dynamic neural radiation field model constructed in this invention significantly improves the detail and realism of the generated image, ensuring the overall structure and color distribution of the generated image. Especially when handling complex dynamic scenes, it can generate high-quality images while reducing memory usage.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of deep learning and computer vision, and specifically relates to a super-resolution dynamic neural radiation field method with enhanced six-plane feature representation. Background Technology

[0002] With the advent of neural radiation fields (NeRF), arbitrary view synthesis has made significant progress. Compared with static NeRF, dynamic NeRF (D-NeRF) is more difficult, but has greater potential for practical applications.

[0003] Dynamic NeRF is a neural radiation field modeling method for dynamic scenes, primarily used for 3D reconstruction in dynamic environments. Traditional NeRF methods are typically limited to static scenes, while Dynamic NeRF can handle time-varying dynamic scenes, such as moving objects or dynamic environments. It introduces a temporal dimension to model these dynamic changes, resulting in more accurate 3D reconstruction and rendering. Dynamic NeRF is generally categorized into time-encoding-based, object deformation-based, fluid or motion field-based, scene segmentation and deformable component-based, generative adversarial networks (GANs)-based, and controllable methods. Time-encoding-based methods render rigid and non-rigid objects by adding time as an input parameter. This method adds temporal parameter encoding to spatial encoding, allowing the scene to change over time. Object deformation-based methods learn time-dependent deformation fields, predicting not only scene color but also changes in scene shape. Fluid or motion field-based methods learn motion fields to capture the motion trajectories or fluidity of objects in the scene, generating more accurate dynamic rendering. Scene segmentation and deformable component-based methods learn multiple independent deformable components that can dynamically change to better capture local motion within the scene. Generative Adversarial Network (GAN)-based methods predict the rendering of dynamic scenes through a generator, while a discriminator evaluates the realism of the rendering. This adversarial mechanism can enhance the detail and visual realism of dynamic scenes. Controllable methods introduce controllable parameters, such as object motion paths, speeds, and shape changes, allowing users to dynamically manipulate the generated scene by adjusting these parameters. However, most of these methods perform well only with low-resolution inputs. As the input resolution increases, limitations such as blurred synthesized views emerge, especially in dynamic scenes. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention provides a super-resolution dynamic neural radiation field method with enhanced six-plane feature representation.

[0005] This method first encodes a set of low-resolution (LR) images of a dynamic scene into six feature planes for training, learning local details and temporal variations of the scene to obtain low-resolution six-plane feature maps. Then, these low-resolution six-plane feature maps are passed through a feature enhancement network to obtain feature-enhanced six-plane feature maps, which are used to render high-resolution (HR) viewpoint images, ultimately improving the quality of viewpoint synthesis. The core of this method is to learn and enhance the low-resolution six-plane feature maps through a neural network model, and then synthesize high-resolution viewpoint images based on the feature-enhanced six-plane feature maps, achieving high-quality rendering of dynamic scenes. Furthermore, this method introduces a training strategy that adapts to the dynamic characteristics of the video, enabling the network to more accurately learn dynamic super-resolution detail information and calculate the motion intensity of different blocks. By employing different training strategies for high-dynamic and low-dynamic regions, the efficiency of the entire super-resolution rendering model can be significantly improved. Experiments on the challenging Plenoptic Video dataset show that, compared with traditional methods combining dynamic NeRF with super-resolution networks, the proposed method has superior performance. The proposed method is applicable to reconstruction at any viewpoint and any time point, and maintains high consistency and robust detail representation in dynamic scenes.

[0006] The technical solution adopted by this invention to solve its technical problem includes the following steps:

[0007] Step S1: Construct a low-resolution dynamic neural radiation field model to generate an initial six-plane feature map.

[0008] Step S2: Construct a feature enhancement network model to generate a six-plane feature map with enhanced features.

[0009] Step S3: Based on the feature-enhanced six-plane feature map, perform high-resolution viewpoint synthesis to obtain a high-resolution image.

[0010] Step S4: Train the super-resolution dynamic neural radiation field model.

[0011] Step S5: Use the trained super-resolution dynamic neural radiation field model to perform tests on the test dataset.

[0012] The beneficial effects of this invention are as follows:

[0013] This invention integrates temporal and spatial information into a super-resolution dynamic neural radiation field framework, which consists of a dynamic neural radiation field model for generating low-resolution new view synthesis and a feature enhancement network model. The model combines spatiotemporal information from previous and subsequent frames to improve spatiotemporal perception.

[0014] The six-plane feature representation-enhanced super-resolution dynamic neural radiation field model constructed in this invention significantly improves the detail and realism of the generated images, ensuring the overall structure and color distribution of the generated images. Especially when dealing with complex dynamic scenes, it can reduce memory usage while generating high-quality images. Attached Figure Description

[0015] Figure 1 This is a flowchart of the main method of the present invention;

[0016] Figure 2 This is a schematic diagram of the main steps of the present invention. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings.

[0018] This invention discloses a super-resolution dynamic neural radiation field method with six-plane feature representation enhancement. The invention includes the following steps: 1. Constructing a low-resolution dynamic neural radiation field model to generate an initial six-plane feature map; 2. Constructing a feature enhancement network model to generate a feature-enhanced six-plane feature map; 3. Generating a high-resolution viewpoint synthesis based on the feature-enhanced six-plane feature map; 4. Training the super-resolution dynamic neural radiation field model; 5. Testing the trained super-resolution dynamic neural radiation field model on a test dataset. This invention integrates temporal and spatial information into a super-resolution dynamic neural radiation field framework, which consists of a dynamic neural radiation field model for generating low-resolution new view synthesis and a feature enhancement network model. The model combines spatiotemporal information from previous and subsequent frames to improve spatiotemporal perception. The six-plane feature representation enhanced super-resolution dynamic neural radiation field model constructed in this invention significantly improves the detail and realism of the generated image, ensuring the overall structure and color distribution of the generated image. Especially when dealing with complex dynamic scenes, it can reduce memory usage while generating high-quality images.

[0019] like Figure 1 As shown, a super-resolution dynamic neural radiation field method with enhanced six-plane feature representation includes the following steps:

[0020] Step S1: Construct a low-resolution dynamic neural radiation field model to generate an initial six-plane feature map:

[0021] According to canonical polynomial decomposition, a given 3D tensor can be decomposed into the sum of the outer products of three vectors. Therefore, for a 3D scene vector... It can be decomposed into the sum of the outer products of vectors XY, Z, and F'. Since the dynamic neural radiation field is related to time T, by fusing time T with spatial XYZ, the 3D volume vector V at time t is obtained. t It can be broken down into:

[0022]

[0023] Where X, Y, and Z represent the X-axis, Y-axis, and Z-axis of the three-dimensional coordinate system, respectively, and T represents the time axis. Let R1, R2, and R3 represent the r-th characteristic matrices of the 3D volume vector after decomposition in the XY, XZ, ZY, ZT, YT, and XT planes, respectively. R1, R2, and R3 are the number of matrices decomposed in groups of XY-ZT, XZ-YT, and YZ-XT planes, respectively. n represents a dimension vector. This is the weight vector at time t, used to capture the model's dependence on time.

[0024] The feature matrices of the same plane after decomposition are concatenated to form M. XY M XZ M ZY M ZT M YT , The feature maps of the six planes are represented respectively, resulting in the initial six-plane feature map M = {M}. XY M XZ M ZY M ZT M YT M XT}

[0025] Step S2: Construct a feature enhancement network model to generate enhanced six-plane feature maps:

[0026] The initial six-plane feature map obtained in step S1 is fed into the feature enhancement network model to generate an enhanced six-plane feature map.

[0027] The feature enhancement network model consists of a preprocessing module and five nested residual dense modules with depth adjustment, which are cascaded in sequence. Each nested residual dense module contains three residual dense modules and a scalable and shiftable convolutional layer (SFT layer). The residual dense module in turn contains five two-dimensional convolutional layers and two scalable and shiftable convolutional layers (SFT layers).

[0028] The preprocessing module performs initial processing and feature extraction on the input features, preparing them for feature enhancement in subsequent modules. This module contains a two-dimensional convolutional layer with a stride of 1, padding of 1, and a kernel size of (3,3), designed to extract low-level features from the input. The input data is first initialized for feature extraction by the preprocessing module, and then further processed by the preprocessing module for deeper feature extraction, ensuring the network can learn multi-level feature representations.

[0029] Each residual dense module and each nested residual dense module contains several SFT layers. Each SFT layer consists of two scaled convolutional layers with a stride of 1 and a kernel size of (1,1) and two shifted convolutional layers with a stride of 1 and a kernel size of (1,1). Specifically, the SFT layers effectively adjust the spatial distribution of the input features by scaling and shifting them, thereby enhancing the model's ability to represent the input image. In this way, the network can automatically learn to flexibly transform between feature maps of different depths, further improving the expressive power of the features.

[0030] The residual dense module is the core component of this network. Each module contains five standard convolutional layers with a stride of 1, padding of 1, and a kernel size of (3,3), as well as two SFT layers. The convolutional layers are stacked to progressively extract spatial information from the input feature map. The output of each convolutional layer is directly added to the input to form a residual connection, effectively mitigating the gradient vanishing problem during deep network training. Each convolutional layer is followed by an SFT layer for further spatial adjustment of the features. Finally, a nonlinear transformation is performed using the LeakyReLU activation function to enhance the nonlinear expressive power of the feature map.

[0031] The feature enhancement network model extracts and refines the initial six-plane feature map M layer by layer through a preprocessing module and multiple nested residual dense modules. This allows the feature enhancement network model to mine more information at different scales and spatial dimensions. It can not only capture the detailed information in the initial six-plane feature map, but also express the global semantics of the scene at a deeper level, thereby effectively improving the performance of subsequent tasks.

[0032] Step S3: Generate high-resolution viewpoint synthesis based on the feature-enhanced six-plane feature map;

[0033] Based on the enhanced six-plane feature map obtained in step S2, for any sampling point ρ(x,y,z,t) in space, obtain the mapping features F of sampling point ρ on the six planes. r The six-plane dimensional features are obtained. Next, perform a Hadamard product of vectors for every two corresponding planar features. Then concatenate the three vectors resulting from the Hadamard product, and then combine them with the corresponding vectors. Dot multiplication Obtain the output vector F out The output vector F out The data is fed into the density branch and the color branch respectively; in the density branch, the output vector F is taken. out The first value is used as the density value σ of the sampling point ρ; in the color branch, the output vector F out The appearance features are analyzed using MLP decoding to obtain the RGB color values ​​of sampling points ρ. Then, volume rendering is performed on the sampled rays, which involves acquiring the density and color values ​​of all sampling points on any sampling ray in space and then averaging them to calculate the final pixel color value. Specifically, this process is equivalent to integrating the radiation field of voxels in the scene and accumulating the contributions of each sampling point to generate the final image. In this way, NeRF can generate detailed, deep, and realistic 3D reconstructed images from a given viewpoint.

[0034] Step S4: Train the super-resolution dynamic neural radiation field model;

[0035] When using the Neural Radiation Field (NeRF) framework for training low-resolution (LR) to high-resolution (HR) image reconstruction, the entire process involves several key steps.

[0036] Specifically, the process begins by using a low-resolution image (LR) as input. The image's spatial information is then transformed into six planes (corresponding to different viewpoints or directions) through a six-plane projection. These planes represent the feature information when viewing the image from different angles. At this stage, feature information for each sampling point is obtained by sampling rays and calculating their representation on these planes. Next, through volume rendering, the feature information from these sampling points is aggregated to obtain a predicted low-resolution image.

[0037] Subsequently, the predicted low-resolution image is compared with the real low-resolution image, and the loss between the two is calculated. Since the rendering of dynamic NeRF is relatively slow, only L2 loss is used when training the low-resolution dynamic neural radiation field model, and its training loss function is:

[0038]

[0039] Among them, L reg , λ reg Let C(τ) represent the regularization loss and its weights, respectively. R is the set of rays. These are the predicted color and the actual color of the light, respectively.

[0040] This loss function is used for backpropagation training of the neural network to optimize the model's parameters, thereby making predictions of low-resolution images more accurate.

[0041] In training high-resolution six-plane features, the parameters of the low-resolution six-plane features are first frozen. This means that the weights of these features are not updated at this stage. Instead, the pre-trained low-resolution six-plane features are used as input to a feature enhancement network model. In this feature enhancement network, the neural network learns the details of the low-resolution six-plane features, enhancing and supplementing their missing information to generate richer and more refined high-resolution six-plane features. These enhanced high-resolution six-plane features can better represent high-frequency details in the image, thus providing more accurate feature information for the generation of high-resolution images.

[0042] In the training stage of high-resolution six-plane features, a depth smoothing loss L is used. DS This enhances the detail and realism of the generated images.

[0043]

[0044] Where R' represents a set of rays sampled from the camera pose. R represents the desired depth. i,j S represents the ray that passes through the pixel (i, j) of the patch centered at r. patch It is the size of the rendering patch.

[0045] Total Variation (TV) loss is used to improve the spatial smoothness of an image. It effectively reduces high-frequency noise and preserves edge information by penalizing large differences between adjacent pixels.

[0046]

[0047] Where, x i,j This represents the pixel value at position (i, j) in the image. The first term in the formula calculates the difference between each pixel and its rightmost pixel, and the second term calculates the difference between each pixel and its bottommost pixel. By summing these differences, the TV loss penalizes large variations between adjacent pixels, thus promoting image smoothness.

[0048] Finally, by employing L2 loss, depth smoothing loss, and TV loss, the detail and realism of the generated images are significantly improved, ensuring the overall structure and color distribution of the images. This is particularly effective when handling complex, dynamically changing scenes, enabling the generation of high-quality images. Therefore, the overall training objective is defined as:

[0049] L=λ1L2+λ2L DS +λ3L TV (5)

[0050] Wherein, λ1, λ2, and λ3 are hyperparameters that control the loss weights.

[0051] Once the high-resolution six-plane features are obtained, they can be input into a ray sampling-based network model. This model continues to use volumetric rendering techniques, performing volumetric rendering on the high-resolution six-plane features to obtain the predicted high-resolution image. Similar to the low-resolution image analysis, the predicted high-resolution image is also compared with the true high-resolution image, and a loss function is calculated. By optimizing the loss function, the model can gradually learn how to accurately map the low-resolution image to the high-resolution image, thus achieving super-resolution reconstruction.

[0052] Step S5: Use the trained dynamic neural radiation field neural network model to perform tests on the test dataset:

[0053] For each scene in the test dataset, the relevant information of the ray corresponding to each pixel is extracted from the high-resolution image based on the camera parameters. After selecting sampling points, the information is input into the dynamic NeRF six-plane network model to obtain the information of the corresponding sampling points. Finally, the final high-resolution image is obtained through the volume rendering formula. Finally, the predicted image and the ground truth are compared to calculate Peak Signal-to-Noise Ratio (PSNR), Structural Similarity (SSIM), and Perceptual Similarity (LPIPS). PSNR is used to quantify the error between the reconstructed image and the original image, SSIM is used to evaluate the similarity between the images, and LPIPS is used to measure image details, texture, and high-level structural features.

[0054] Table 1 compares the results with the Hexplane method.

[0055]

[0056] Table 1 presents the experimental results on the Plenoptic Video dataset. As can be seen from the data in the table, the model in this study performs excellently in PSNR, SSIM, and LPIPS metrics. These figures represent improvements of 1.92%, 0.02%, and 0.05% respectively compared to the best performance of other models, demonstrating the enhancement effect of the feature enhancement network on the six-plane features, further improving the detail of the generated images.

[0057] The above description, in conjunction with specific / preferred embodiments, provides a further detailed explanation of the present invention. It should not be construed that the specific implementation of the present invention is limited to these descriptions. Those skilled in the art can make various substitutions or modifications to these described embodiments without departing from the inventive concept, and all such substitutions or modifications should be considered within the scope of protection of the present invention.

Claims

1. A super-resolution dynamic neural radiation field method with enhanced six-plane feature representation, characterized in that... Includes the following steps: Step S1: Construct a low-resolution dynamic neural radiation field model and generate an initial six-plane feature map; Step S2: Construct a feature enhancement network model to generate a six-plane feature map with enhanced features; Step S3: Based on the feature-enhanced six-plane feature map, perform high-resolution viewpoint synthesis to obtain a high-resolution image; Step S4: Train the super-resolution dynamic neural radiation field model; Step S5: Use the trained super-resolution dynamic neural radiation field model to perform tests on the test dataset; Step S2 involves constructing a feature enhancement network model by feeding the initial six-plane feature map obtained in step S1 into the feature enhancement network model to generate an enhanced six-plane feature map; the details are as follows: The feature enhancement network model consists of a preprocessing module and five nested residual dense modules with depth adjustment, which are cascaded in sequence. Each nested residual dense module contains three residual dense modules and a scalable and shiftable convolutional layer SFT. The residual dense module contains five two-dimensional convolutional layers and two scalable and shiftable convolutional layer SFT layers. The feature enhancement network model extracts and refines the features of the initial six-plane feature map layer by layer through a preprocessing module and multiple residual nested dense modules; Each residual dense module and each residual nested dense module contains several SFT layers; each SFT layer consists of two scaling convolutional layers with a stride of 1 and a kernel size of (1,1) and two shifting convolutional layers with a stride of 1 and a kernel size of (1,1); the SFT layer effectively adjusts the spatial distribution of the input features by scaling and shifting them, thereby enhancing the model's ability to represent the input image. Each residual dense module contains five standard convolutional layers with a stride of 1, padding of 1, and a kernel size of (3,3), and two SFT layers. The convolutional layers are stacked to progressively extract spatial information from the input feature map. The output of each convolutional layer is directly added to the input to form a residual connection, thereby effectively alleviating the gradient vanishing problem during deep network training. Each convolutional layer is followed by an SFT layer to further adjust the spatial properties of the features. Finally, the LeakyReLU activation function is used for nonlinear transformation to enhance the nonlinear expressive power of the feature map.

2. The super-resolution dynamic neural radiation field method with enhanced six-plane feature representation according to claim 1, characterized in that... Step S1 involves constructing a low-resolution dynamic neural radiation field model to generate an initial six-plane feature map. The specific method is as follows: According to canonical polynomial decomposition, a given 3D tensor can be decomposed into the sum of the outer products of three vectors. Therefore, for a 3D scene vector... It can be decomposed into the sum of the outer products of vectors XY, Z, and F'; since the dynamic neural radiation field is related to time T, by fusing time T with spatial XYZ respectively, the 3D volume vector at time T is decomposed as follows: Where X, Y, and Z represent the X-axis, Y-axis, and Z-axis of the three-dimensional coordinate system, respectively, and T represents the time axis. R1, R2, and R3 represent the r-th characteristic matrices of the 3D volume vector after decomposition in the XY, XZ, YZ, ZT, YT, and XT planes, respectively; R1, R2, and R3 are the number of matrices decomposed in groups of XY-ZT, XZ-YT, and YZ-XT planes, respectively. i = 1, 2 or 3, where n represents the dimension vector; Let be the weight vector at time t; The feature matrices of the same plane after decomposition are concatenated together to form... Representing the feature maps of the six planes respectively, we obtain the initial six-plane feature map M = {M}. XY M XZ M ZY M ZT M YT M XT } 3. The super-resolution dynamic neural radiation field method with enhanced six-plane feature representation according to claim 2, characterized in that... The preprocessing module performs preliminary processing and feature extraction on the input features. This module contains a two-dimensional convolutional layer with a stride of 1, padding of 1, and a kernel size of (3,3), which is designed to extract low-level features of the input. The input data is first initialized and extracted through the preprocessing module, and then further processed through the preprocessing module to extract deeper features, ensuring that the network can learn multi-level feature representations.

4. The super-resolution dynamic neural radiation field method with enhanced six-plane feature representation according to claim 1 or 3, characterized in that... Step S3 involves synthesizing high-resolution viewpoints based on the feature-enhanced six-plane feature map. The specific method is as follows: Based on the enhanced six-plane feature map obtained in step S2, for any sampling point ρ(x,y,z,t) in space, obtain the mapping features F of sampling point ρ on the six planes. r The six-plane dimensional features are obtained. Next, perform a Hadamard product of vectors for every two corresponding planar features. Then concatenate the three vectors resulting from the Hadamard product, and then combine them with the corresponding vectors. Dot multiplication Obtain the output vector F out ; output vector F out The data is fed into the density branch and the color branch respectively; in the density branch, the output vector F is taken. out The first value is used as the density value σ of the sampling point ρ; in the color branch, the output vector F out The appearance features are obtained by using MLP decoding to obtain the color value (rgb) of the sampling point ρ; then, volume rendering is performed on the sampled rays, that is, the density and color values ​​of all sampling points on any sampling ray in space are obtained and then weighted averaged to calculate the final pixel color value.

5. The super-resolution dynamic neural radiation field method with enhanced six-plane feature representation according to claim 4, characterized in that... Step S4 involves training the super-resolution dynamic neural radiation field model, as detailed below: First, a low-resolution image is used as input. By performing a six-plane projection on the image, the spatial information of the image is transformed into six planes. These planes are used to represent the feature information of the image when viewed from different angles. In this stage, feature information of each sampling point is obtained by sampling light rays and calculating the performance of sampling points on these planes; Secondly, the feature information of these sampling points is aggregated through volume rendering to obtain a predicted low-resolution image; Then, the predicted low-resolution image is compared with the real low-resolution image, and the loss between the two is calculated. Since the rendering of dynamic NeRF is relatively slow, only L2 loss is used when training the low-resolution dynamic neural radiation field model. The training loss function is: Among them, L reg , λ reg Let C(τ) represent the regularization loss and its weights, respectively. R is the set of rays. These are the predicted color and the actual color of the light, respectively.

6. The super-resolution dynamic neural radiation field method with enhanced six-plane feature representation according to claim 5, characterized in that... The training process for high-resolution six-plane features is as follows: First, the parameters of the low-resolution six-plane features are frozen. The trained low-resolution six-plane features are then used as input to a feature enhancement network model. In this feature enhancement network, the neural network learns the details of the low-resolution six-plane features, enhances and supplements their missing information, thereby generating richer and more refined high-resolution six-plane features. In the training stage of high-resolution six-plane features, a depth smoothing loss L is used. DS To improve the detail and realism of the generated images; Where R' represents a set of rays sampled from the camera pose. R represents the desired depth. i,j S represents the ray that passes through the pixel (i, j) of the patch centered at r. patch It is the size of the rendering patch; Total variation loss L TV Used to improve the spatial smoothness of images: Where, x i,j This represents the pixel value at position (i, j) in the image; the first term in the formula calculates the difference between each pixel and the pixel to its right, and the second term calculates the difference between each pixel and the pixel below it; Finally, combining L2 loss and deep smoothing loss L... DS Total variation loss L TV The overall training loss L is defined as: L=λ1L2+λ2L DS +λ3L TV (5) Where λ1, λ2, and λ3 are hyperparameters that control the loss weights; The obtained high-resolution six-plane features are input into a ray sampling-based network model. By performing volumetric rendering on the high-resolution six-plane features, the predicted high-resolution image is obtained. The predicted high-resolution image is compared with the real high-resolution image, and the loss function is calculated. By optimizing the loss function, the model can gradually learn how to accurately map the low-resolution image to the high-resolution image, thereby achieving super-resolution reconstruction.

7. The super-resolution dynamic neural radiation field method with enhanced six-plane feature representation according to claim 6, characterized in that... The specific method for step S5 is as follows: The high-resolution images of each scene in the test dataset are used to extract the relevant information of the ray corresponding to each pixel according to the camera parameters. After selecting the sampling points, they are input into the dynamic NeRF six-plane network model to obtain the information of the corresponding sampling points. Finally, the final high-resolution image is obtained through the volume rendering formula.

Citation Information

Patent Citations

  • Large-scale scene three-dimensional reconstruction algorithm based on neural radiation field and dynamic decoding mechanism

    CN118279520A

  • Cross-modal three-dimensional light field rendering method and three-dimensional rendering device

    CN118736084A