3D Scene Rendering Method Based on Spatiotemporal Super-Resolution and Neural Radiance Fields
By adopting spatiotemporal super-resolution technology in NeRF rendering, using time domain interpolation and spatial supersampling, the problem that NeRF rendering in the prior art is difficult to achieve real-time performance, and real-time rendering effect with high quality and low computing cost is achieved.
Patent Information
- Application Number
- CN202411452920.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-17
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-10-17
AI Technical Summary
The prior art is difficult to achieve real-time performance in neural radiation field (NeRF) rendering, especially while improving image quality, the high computing demand leads to a long rendering time and the real-time rendering cannot be performed.
Using a spatial and temporal super-resolution method, the time required for NeRF rendering and super-sampling is reduced and the rendering frame rate is improved through time domain interpolation and spatial supersampling. The specific methods include inputting multi-view images into neural radiation fields, upsampling and video interpolation network processing, and generating target three-dimensional scenes in combination with a lightweight neural renderer.
It realizes high-quality real-time rendering effect under low computing cost conditions, significantly improving the rendering frame rate of NeRF, while maintaining high-quality rendering effect.
Smart Images

Figure CN119399337B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of three-dimensional scene reconstruction, and particularly relates to a three-dimensional scene rendering method based on spatio-temporal super-resolution and neural radiance fields. Background Art
[0002] The novel view synthesis task refers to rendering a target image corresponding to a target pose given a source image, a source pose, and a target pose. Novel view synthesis has a wide range of applications in fields such as 3D reconstruction, AR / VR, etc. As users increasingly seek realistic and immersive experiences, end-users strongly demand efficient and high-quality methods for synthesizing novel views. Among various representative methods, Neural Radiance Fields (NeRF) have received extensive attention due to their ability to generate high-fidelity images. However, with the increasing demand for realistic and high-quality visual experiences, the challenge of achieving real-time performance in NeRF rendering still exists. Balancing the trade-off between quality and performance is becoming increasingly challenging because improving one often comes at the expense of the other. In addition, the computational requirements of NeRF result in long rendering times and the inability to perform real-time rendering. Therefore, the present invention provides a three-dimensional scene rendering method based on spatio-temporal super-resolution and neural radiance fields. Summary of the Invention
[0003] To solve the above technical problems, the present invention proposes a three-dimensional scene rendering method based on spatio-temporal super-resolution and neural radiance fields to solve the problems existing in the above prior art.
[0004] To achieve the above object, the present invention provides a three-dimensional scene rendering method based on spatio-temporal super-resolution and neural radiance fields, including:
[0005] Inputting multi-view images into a neural radiance field to obtain a low-resolution target frame image and a low-resolution support frame image;
[0006] Performing upsampling processing on the low-resolution target frame image to obtain a high-resolution target frame image of the first branch;
[0007] Based on the low-resolution target frame image and the low-resolution support frame image, processing using a video interpolation network and performing image projection to obtain a high-resolution target frame image of the second branch;
[0008] Inputting the high-resolution target frame image of the first branch and the high-resolution target frame image of the second branch into a lightweight neural renderer to obtain a target three-dimensional scene.
[0009] Optionally, the process of the target frame high-resolution image of the second branch includes:
[0010] Extract the feature maps of the target frame low-resolution image and the support frame low-resolution image;
[0011] Create an intermediate frame feature map using a learned feature interpolation function based on the feature maps;
[0012] Reproject the intermediate frame feature map onto the point cloud, and project the intermediate frame feature map on the point cloud back to the high-resolution image from the perspective of the target frame to obtain the target frame high-resolution image of the second branch.
[0013] Optionally, the expression for generating the intermediate frame feature map is:
[0014] L t ,L t-1 ,L t-2 =V(L t-1 ,L t-2 )
[0015] In the formula, L t represents the target frame, L t-2 represents the support frame, L t-1 represents the intermediate frame, and V represents the learned feature interpolation function.
[0016] Optionally, the expression for the process of reprojecting the intermediate frame feature map onto the point cloud and projecting the intermediate frame feature map on the point cloud back to the high-resolution image from the perspective of the target frame is:
[0017] H t→t ,H t-1→t ,H t-2→t =W(L t ,L t-1 ,L t-2 )
[0018] =P W (P t ,P t-1 ,P t-2 )
[0019] =P W (P U (L t ,L t-1 ,L t-2 ))
[0020] In the formula, P t ,P t-1 ,P t-2 respectively represent the point clouds obtained by the reprojection of L t ,L t-1 ,L t-2 ; Ht→t , H t-1→t , H t-2→t respectively represent the high-resolution images projected from the point cloud P t , P t-1 , P t-2 onto the L t view; P U represents a function that reprojects the low-resolution image onto the point cloud; P W represents a function that projects the point cloud from the view of the t-th frame onto the high-resolution image; W represents a function that warps the low-resolution image to the high-resolution image corresponding to the view of the t-th frame.
[0021] Optionally, the process of constructing the target three-dimensional scene based on the high-resolution image of the target frame of the first branch and the high-resolution image of the target frame of the second branch further includes a process of fusing a plurality of high-resolution images, wherein the expression for fusing a plurality of high-resolution images is:
[0022] N t = O(H t→t , H t-1→t , H t-2→t , U t )
[0023] In the formula, N t represents a new view synthesis image generated from multiple high-resolution images using U-Net.
[0024] Optionally, in the process of inputting the multi-view image into the neural radiance field to generate the low-resolution image, the image quality of the neural radiance field is further constrained by using photometric loss and depth loss;
[0025] Among them, the expressions for photometric loss and depth loss are:
[0026]
[0027] In the formula, L c represents photometric loss, and L d represents depth loss.
[0028] The present invention also provides a computer terminal device, including:
[0029] One or more processors;
[0030] A memory, coupled to the processor, for storing one or more programs;
[0031] When the one or more programs are executed by the one or more processors, the one or more processors implement a three-dimensional scene rendering method based on spatio-temporal super-resolution and neural radiance field.
[0032] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, a three-dimensional scene rendering method based on spatio-temporal super-resolution and neural radiance fields is implemented.
[0033] Compared with the prior art, the present invention has the following advantages and technical effects:
[0034] The present invention proposes a NeRF acceleration method for spatio-temporal super-resolution, enabling NeRF to achieve high-quality real-time rendering effects under low computational costs.
[0035] The present invention proposes a new method using spatio-temporal super-resolution to achieve high-quality rendering of NeRF at high FPS through temporal interpolation and spatial supersampling. Specifically, through temporal interpolation, the present invention can intelligently interpolate information between adjacent frames in the time dimension, thereby achieving a more continuous effect. In combination with a lightweight super-resolution network, the time required for rendering and super-resolution of neural radiance fields can be significantly reduced while maintaining high-quality rendering effects. At the same time, the present invention makes full use of depth information to better realize the process of fusing multi-frame feature maps into the target frame. The present invention uses spatial supersampling technology to increase the density of sampling points in the spatial dimension, greatly improving the details and clarity of the rendered image. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] The drawings constituting a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation of this application. In the drawings:
[0037] Figure 1 is a flowchart of the three-dimensional scene construction method according to an embodiment of the present invention;
[0038] Figure 2 is an overall flowchart of an embodiment of the present invention;
[0039] Figure 3 is a frame feature time interpolation method based on deformable sampling according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0040] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The following will refer to the drawings and combine the embodiments to detail this application.
[0041] It should be noted that the steps shown in the flowchart of the drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0042] Example 1
[0043] Existing solutions do not make full use of spatio-temporal information. To accelerate rendering, a large amount of memory consumption has to be sacrificed, and the rendering speed for real scenes is still poor. Therefore, in this embodiment, a three-dimensional scene rendering method based on spatio-temporal super-resolution and neural radiance field is provided. This method achieves high-quality rendering of NeRF at high FPS through temporal interpolation and spatial supersampling. The present invention includes the following steps:
[0044] As shown in Figure 1 , the multi-view images are input into the neural radiance field to obtain the target-frame low-resolution image and the support-frame low-resolution image; the target-frame low-resolution image is upsampled to obtain the target-frame high-resolution image of the first branch; the target-frame low-resolution image and the support-frame low-resolution image are processed by a video interpolation network and image projection to obtain the target-frame high-resolution image of the second branch; the target-frame high-resolution image of the first branch and the target-frame high-resolution image of the second branch are input into a lightweight neural renderer to obtain the target three-dimensional scene.
[0045] As a specific implementation of this embodiment, as shown in Figure 2 , a super-resolution method based on space and viewpoints is proposed. Specifically, first, multi-view images are used as input to iteratively train the neural radiance field. Through volume rendering, low-resolution RGB and depth images L t-2 and L t are obtained respectively, where the subscript t represents the image at frame t, and L represents the low-resolution image. Next, spatio-temporal super-resolution is performed using two branches to obtain the final output O t at the target resolution, where O represents the output image. The first branch focuses on the direct upsampling operation applied to the target frame to obtain the high-resolution image H t of the target frame, where H represents the high-resolution image. The second branch implements the spatio-temporal super-resolution method: the video frame interpolation network generates RGB and depth images (L t-1 ) for the intermediate frames, then the RGB image is reprojected onto the point cloud using the depth information and projected back to the high-resolution image from the perspective of the target frame. Finally, the high-resolution target-frame images obtained from the two branches are input into the lightweight neural renderer of this embodiment to generate the final output.
[0046] Spatio-temporal super-resolution NeRF
[0047] The space and viewpoint super-resolution module of this embodiment includes two main components: viewpoint interpolation and spatial upsampling, as shown in Figure 2as shown in the spatial and view - point super - resolution module. By integrating these components, this embodiment improves the resolution and coherence of the target frame, thus enhancing the visual quality. The initial stage of the module in this embodiment focuses on view - point interpolation, including extracting feature maps from the input target frame (L t ) and the support frame (L t-2 ), and then learning the feature interpolation function V to directly create intermediate feature maps. The main goal is to accurately generate the feature map of the intermediate frame (L t-1 ), which can be expressed as:
[0048] L t ,L t-1 ,L t-2 = V(L t-1 ,L t-2 )
[0049] To approximate the forward and backward motion, this embodiment utilizes the motion information between the target frame and the support frame. Inspired by ZoomingSlow - Mo, as Figure 3 shown, this embodiment integrates deformable convolutions into the sampling function, enabling this embodiment to effectively explore complex local temporal contexts. This function enables the view - point interpolation of this embodiment to proficiently adapt to a large amount of movement between two frames. In addition, combined with deformable convolutions, this embodiment learns the offsets as sampling parameters, implicitly capturing the forward and backward motion information. This offset helps with precise alignment, ensuring temporal coherence and smooth transitions. Through these techniques, this embodiment can generate intermediate feature maps that seamlessly connect the target and support frames.
[0050] In addition to view - point interpolation, this embodiment also employs spatial up - sampling to increase the resolution of the target frame. After performing view - point interpolation, this embodiment obtains the feature maps of the support frame, the intermediate frame, and the target frame. Utilizing the depth information, this embodiment projects the RGB image onto a point cloud and then projects it onto a high - resolution image from the perspective of the target frame:
[0051] H t→t ,H t-1→t ,H t-2→t = W(L t ,L t-1 ,L t-2 )
[0052] = P W (P t ,P t-1 ,P t-2 )
[0053] = P W (P U (L t ,L t-1 ,L t-2 ))
[0054] Among them, P t , P t-1 , P t-2 respectively represent the point clouds obtained by the reprojection through L t , L t-1 , L t-2 . H t→t , H t-1→t , H t-2→t respectively represent the high-resolution images projected from the point clouds P t , P t-1 , P t-2 onto the L t view. The functions P U , P W and W are respectively used to reproject the low-resolution image onto the point cloud, project the point cloud from the view of the t-th frame onto the high-resolution image, and warp the low-resolution image to the high-resolution image corresponding to the view of the t-th frame. By integrating these high-resolution frames obtained by warping from different views and the upsampling result of the target frame, this embodiment can obtain the upsampled target frame:
[0055] U t = U p (L t )
[0056] N t = O(H t→t , H t-1→t , H t-2→t , U t )
[0057] Among them, U t represents the upsampled frame from the low-resolution image L t , and N t represents the new view synthesis image generated from multiple high-resolution images using U-Net. The functions U P and O respectively represent upsampling and fusing multiple high-resolution images through U-Net.
[0058] In order to minimize time loss, this embodiment can improve efficiency and optimize memory utilization by using the LR rendering buffer. Although these buffers incur time and memory costs, they provide a more efficient alternative for rendering HR frames. In addition, the LR images from the rendering buffer can be used to warp into different views, which helps the spatial upsampling of this embodiment. The method of this embodiment optimizes real-time rendering by integrating the rendering buffer, improving efficiency, memory usage, and spatial upsampling to obtain better visual quality.
[0059] The rendering buffer stores the low-resolution feature maps {F t-L, ,, ,, F t-1} and depth map {D t-L , ,, ,, D t-1}. At the current viewpoint t, the low-resolution feature map F t and depth map D t .
[0060] Loss function
[0061] In this embodiment, photometric loss and depth loss are used to constrain the quality of the images generated by NeRF. The depth maps generated during the NeRF training process are often inaccurate, which poses challenges when performing spatial oversampling. Accurate depth information is crucial for ensuring that other frames can be correctly warped to the target frame. To address this issue and improve the accuracy of depth information, this embodiment utilizes the trained network to generate pseudo GT of the depth map and uses it to supervise the depth information.
[0062] The calculation methods of photometric loss and depth loss are as follows:
[0063]
[0064]
[0065] The total loss function of the entire network is:
[0066] L total = L c + λ d L d
[0067] where λ d is the trade-off parameter for depth loss.
[0068] Compared with the existing methods, the present invention significantly improves the rendering frame rate of NeRF while maintaining high-quality rendering.
[0069] Embodiment 2
[0070] This embodiment also provides a computer terminal device, including:
[0071] One or more processors;
[0072] A memory, coupled to the processor, for storing one or more programs;
[0073] When the one or more programs are executed by the one or more processors, the one or more processors implement a three-dimensional scene rendering method based on spatio-temporal super-resolution and neural radiance fields.
[0074] This embodiment also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, a three-dimensional scene rendering method based on spatio-temporal super-resolution and neural radiance fields is implemented.
[0075] The above are only the preferred specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A three-dimensional scene rendering method based on spatiotemporal super-resolution and neural radiation field, characterized in that: The following steps are involved: Inputting the multi-view images into the neural radiance field to obtain a target frame low-resolution image and a support frame low-resolution image; Performing upsampling processing on the target frame low-resolution image to obtain a target frame high-resolution image of the first branch; Based on the target frame low-resolution image and the support frame low-resolution image, a video interpolation network is used for processing and image projection is performed to obtain a target frame high-resolution image of a second branch; The process of obtaining the target frame high-resolution image of the second branch includes: extracting feature maps of the target frame low-resolution image and the support frame low-resolution image; creating an intermediate frame feature map based on the feature map by using a learning feature interpolation function; reprojecting the intermediate frame feature map onto a point cloud, and projecting the intermediate frame feature map on the point cloud back to the high-resolution image from the perspective of the target frame to obtain the target frame high-resolution image of the second branch; The expression for generating the intermediate frame feature map is: L t ,L t-1 ,L t-2 =V(L t-1 ,L t-2 ) Where Lt represents the target frame, L t-2 Indicates support for frames, L t-1 represents the intermediate frame, V represents the learned feature interpolation function; The expression of the process of reprojecting the intermediate frame feature map onto the point cloud and projecting the intermediate frame feature map on the point cloud back to the high-resolution image from the perspective of the target frame is: H t→t ,H t-1→t ,H t-2→t =W(L t ,L t-1 ,L t-2 ) =P W (P t ,P t-1 ,P t-2 ) =P W (P U (L t ,L t-1 ,L t-2 )) Where P t ,P t- 1,P t-2 Respectively, through L t ,L t-1 ,L t-2 The point cloud obtained by reprojection of t→t ,H t-1→t ,H t-2→t Respectively represent the point cloud P t ,P t-1 ,P t-2 Projection to L t High-resolution images on the view; P U represents the function of reprojecting the low-resolution image onto the point cloud; P W represents the function of projecting the point cloud from the view of the t-th frame to the high-resolution image; W represents the function of warping the low-resolution image to the high-resolution image corresponding to the t-th frame view; Inputting the target frame high-resolution image of the first branch and the target frame high-resolution image of the second branch into a lightweight neural renderer to obtain a target three-dimensional scene; The process of inputting multi-view images into the neural radiance field to generate a low-resolution image also includes using photometric loss and depth loss to constrain the image quality of the neural radiance field; Among them, the expressions of photometric loss and depth loss are: Where, L c Represents the luminosity loss, L d represents the depth loss.
2. The three-dimensional scene rendering method based on spatiotemporal super-resolution and neural radiation field according to claim 1, characterized in that: The process of constructing a target three-dimensional scene based on the target frame high-resolution image of the first branch and the target frame high-resolution image of the second branch also includes a process of fusing a plurality of high-resolution images, wherein the expression for fusing a plurality of high-resolution images is: N t =O(H t→t ,H t-1→t ,H t-2→t ,U t ) Where Nt represents a new view synthetic image generated from multiple high-resolution images using U-Net.
3. A computer terminal device, characterized in that: include: one or more processors; A memory, coupled to the processor, for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the three-dimensional scene rendering method based on spatiotemporal super-resolution and neural radiation field as described in any one of claims 1-2.
4. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the three-dimensional scene rendering method based on spatiotemporal super-resolution and neural radiation field as described in any one of claims 1-2 is implemented.