A NeRF 3D reconstruction method based on deformable pose blur kernel

By introducing deformable pose fuzzy kernel and camera pose sequence interpolation technology in NeRF 3D reconstruction method, the three-dimensional reconstruction problems in motion blur images and unknown camera poses are solved, and a higher quality 3D reconstruction effect is achieved.

CN119169180BActive Publication Date: 2025-06-03HANGZHOU DIANZI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411115443.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-14
Publication Date
2025-06-03
Estimated Expiration
2044-08-14

AI Technical Summary

Technical Problem

In the prior art, when processing images with motion blur, especially when the camera position is unknown, it is difficult to effectively reconstruct a three-dimensional scene, resulting in poor reconstruction effect.

Method used

A NeRF three-dimensional reconstruction method based on deformable pose fuzzy core is proposed. By interpolation of the estimated camera pose, a camera pose sequence is generated, and a deformable fuzzy core is combined to simulate the generation of motion blur, minimizing the generated motion blur and real blur loss to optimize the deformable fuzzy core and camera pose.

Benefits of technology

The quality of three-dimensional reconstruction is significantly improved when the blurred image input and camera pose are unknown, and compared with the existing methods, the PSNR, SSIM and LPIPS indicators are significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119169180B_ABST
    Figure CN119169180B_ABST
Patent Text Reader

Abstract

The present invention discloses a NeRF three-dimensional reconstruction method based on a deformable pose blur kernel. First, information extraction is performed on the blurred images captured from different perspectives in the same scene to obtain a pixel matrix and estimated pose information. A blur kernel and a view encoding representation are created for each blurred image, and a multi-layer perceptron is used to deform the blur kernel. The camera poses at the start and end of the estimated exposure are parameterized, and linear interpolation is performed on the camera poses according to the number of kernel points. Then, the blur kernel and the estimated sequence of camera poses are input into the ray encoding layer, and multiple other rays in adjacent frames are modeled by estimating the poses during the camera movement. Next, the weights corresponding to the rays are adjusted according to the target blur kernel, enabling the target pixels to effectively extract the pixel information of adjacent frames, so that the reconstructed blurred pixels are closer to the real image. In the scene reconstruction stage, the rays of the new view are directly rendered by removing the blur kernel, and the output of a clear scene is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image data processing, relates to a three-dimensional reconstruction method, and particularly relates to a NeRF three-dimensional reconstruction method based on a deformable pose blur kernel. Background Art

[0002] Three-dimensional reconstruction is the process of creating or restoring three-dimensional objects and scenes by obtaining information from two-dimensional images or other data sources (such as lidar data, depth sensors, etc.). With the continuous improvement of computing power, three-dimensional reconstruction has become a key technology in many real-world applications, such as virtual reality and augmented reality, medical image processing, etc.

[0003] Traditional methods of three-dimensional reconstruction usually explicitly represent a three-dimensional object or scene using three-dimensional point clouds, triangular meshes, or volume networks. Different from traditional three-dimensional reconstruction methods, Neural Radiance Field (NeRF) is a novel and effective implicit scene representation. NeRF represents the color and density of each point in the scene in the form of a continuous volume function through a multi-layer perceptron (MLP). The three-dimensional reconstruction based on NeRF is a technology that trains a multi-layer perceptron (MLP) with three-dimensional scene information through two-dimensional images and corresponding camera poses, and then renders the trained implicit scene into an explicit static three-dimensional scene through volume rendering.

[0004] During the process of image acquisition in real-world applications (such as medical imaging, drone photography), information loss of varying degrees will inevitably occur. A relatively common situation is motion blur, that is, due to the movement of the camera or the object during the exposure process, the image produces blurred artifacts. Using images with motion blur to train NeRF will greatly affect the final three-dimensional scene reconstruction effect. Existing technologies generally use a blur kernel to simulate motion blur during the training process and remove the blur kernel during the inference stage, which to a certain extent solves the impact of images with motion blur on NeRF reconstruction. This method requires the true camera pose of the image to be preset, but in real-world scenarios, the camera pose is generally estimated through the COLMAP software. However, images with motion blur will cause a large error in the camera pose obtained by this software compared to clear images, resulting in poor performance of existing technologies in processing motion blur images when the camera pose is unknown. Summary of the Invention

[0005] In view of the deficiencies of the prior art, the present invention proposes a NeRF three-dimensional reconstruction method based on a deformable pose blur kernel. For the case where the blurred image and the camera pose are unknown, a camera pose sequence is generated by interpolating the estimated camera pose, and a deformable blur kernel is combined to simulate the generation of motion blur. The deformable blur kernel and the camera pose are optimized by minimizing the loss between the generated motion blur and the real blur. In the scene reconstruction stage, the blur kernel is removed to solve the problem of unknown blurred image input and camera pose.

[0006] A NeRF three-dimensional reconstruction method based on a deformable pose blur kernel specifically includes the following steps:

[0007] Step 1, Blurred image data extraction and processing

[0008] Blurred image data extraction and processing is to extract image information and estimated pose information from multiple input blurred images to facilitate subsequent processing of simulated motion blur, specifically including the following steps:

[0009] Step 1.1, Input dataset extraction.

[0010] For the blurred images taken from different perspectives of the same scene, pixel information is extracted and represented by a vector , where represents the pixel matrix of the a-th input image. The estimated pose information corresponding to the image is read , where represents the camera pose matrix corresponding to the a-th image, including a 3×3 rotation matrix , and a 3D column vector . Among them, the rotation matrix describes the rotation information of the camera in the world coordinate system, and the vector represents the translation information of the camera in the world coordinate system.

[0011] Step 1.2, Pose information processing.

[0012] Random perturbations are added to the camera pose matrix of each image, and the camera pose at the start of the estimated exposure and the pose at the end of the exposure are parameterized, where is the Lie algebra and is a 6D vector. The obtained camera pose estimation information is concatenated to obtain the sets and , where, and respectively represent the sets of estimated information on the starting and ending positions of the camera pose during a single camera exposure.

[0013] Step 1.3. Initialization of blur kernel information.

[0014] Create m initial blur kernels for m blurred images and their corresponding view encoding representations . Among them , is the number of kernel points of a blur kernel, represents the normalized coordinates of the kernel points in the blur kernel, and they are randomly initialized using a normal distribution. Initialize the vector value of the view encoding representation to 0.

[0015] Step 2. Blur kernel deformation

[0016] After obtaining the blurred image information and camera pose information, for the generation of blurred pixels at the th image coordinates, according to the view encoding and pixel coordinates deform the initial blur kernel , and the specific steps are as follows:

[0017] Step 2.1. Select the corresponding initial blur kernel according to the blurred image . First, the kernel points in are obtained through denormalization as , and then the new kernel points are randomly perturbed using a normal distribution to obtain , which form a new blur kernel . Sample and encode the pixel coordinates of the kernel points in the blur kernel to generate a two-dimensional vector .

[0018] Step 2.2. Select the corresponding view encoding according to the blurred image . Duplicate and expand the view encoding according to the number of kernel points n of the blur kernel to obtain a two-dimensional vector , and then splice and to obtain a two-dimensional vector representation . Calculate the one-dimensional vector representation of the coordinates of the blurred pixels to be generated, and then perform a duplicate and expand operation to obtain a two-dimensional vector , and splice it with to obtain a two-dimensional vector .

[0019] Step 2.3. The vector Input the first MLP layer into it, and then take the obtained output and vector to splice and get vector . Then input it into the second MLP layer :

[0020]

[0021]

[0022] where is the offset of the kernel point, is the weight when the kernel point synthesizes blurred pixels.

[0023] Step 2.4. Use the kernel point offset obtained in Step 2.3 to deform the blur kernel .

[0024] Step 3. Blur pixel rendering

[0025] According to the blur kernel obtained in Step 2, simulate the generation of motion-blurred pixels and complete model training at the same time. The specific steps are as follows:

[0026] Before ray encoding acquisition in Step 3.1, it is necessary to estimate the exposure camera pose. According to the number of kernel points n in the deformed blur kernel obtained in Step 2, linearly interpolate the camera poses , obtained in Step 1:

[0027]

[0028] represents the exposure time. represents the estimated camera pose obtained by interpolation at time t during the exposure process. Splice to get the estimated camera pose sequence corresponding to the blur kernel .

[0029] Step 3.2. According to the camera intrinsic matrix and the estimated camera pose sequence obtained in Step 3.1, convert the corresponding kernel point information in the blur kernel output in Step 2 into ray representation through ray encoding:

[0030]

[0031] where They are vectors representing the starting point and direction of the ray. After separately processing n kernel points and the estimated camera pose, a set of rays is obtained. .

[0032] Step 3.3: Input the set of rays Rays obtained in Step 3.2 into NeRF and render to obtain a set of pixel RGB representations :

[0033]

[0034] where , represents the RGB value rendered by each ray, and are the parameters of the NeRF model.

[0035] Step 3.4: According to the weights obtained in Step 2.3, perform weighted summation on the results output by the NeRF model to obtain a simulated blurred pixel RGB representation :

[0036]

[0037] Step 3.5 adds a structural loss to the loss function and constructs a loss function : :

[0038]

[0039]

[0040] where represents the RGB value of the real blurred pixel, B is the set of blurred pixels, is the weight parameter of the structural loss , and is a hyperparameter. represents randomly selecting a set of pixels from the set of predicted pixels and the corresponding set of pixels in the training set to calculate the SSIM quality criterion. The purpose is to consider the structural information when calculating each pixel one by one to improve the reconstruction effect.

[0041] Minimize the loss function and optimize the model parameters , , , , . Where represents the parameters of , 。

[0042] Step 4, Scene Reconstruction

[0043] After completing the training in Steps 2 and 3, select consecutive generation perspectives, and input the corresponding camera pose matrices into NeRF in sequence, and then obtain a set of pixel groups . For a pixel group , where H and W are the height and width of the frame image respectively, and c represents the pixel color. Use an image processing tool to operate on the pixel groups in S to obtain a continuous generated view representation , which represents a reconstructed two-dimensional view of a frame.

[0044] Use the consecutive new views IM as the frame data of the clear scene output, and output the reconstructed clear scene.

[0045] The present invention has the following beneficial effects:

[0046] 1. Aiming at the problem that there is a large deviation between the camera pose estimation of motion-blurred images and the true pose, the camera pose sequences at different moments during the exposure process are estimated by an interpolation method, the blur kernel is combined with the estimated camera pose sequences, and the blur kernel and the estimated camera pose sequences are continuously optimized during the training phase.

[0047] 2. The initial blur kernel is deformed to obtain the blur kernel corresponding to the target pixel, then each kernel point of the blur kernel is bound to an estimated camera pose, and then multiple rays of adjacent frames are calculated according to the bound estimated camera poses. Finally, the weights corresponding to the rays are adjusted by the target blur kernel to effectively extract the pixel information of adjacent frames for the generation of blurred pixels, so that the reconstructed blurred pixels are closer to the true blurred pixels.

[0048] 3. Experiments are carried out on 5 static scenes. Compared with the existing methods for NeRF with blurred input, LPIPS is improved by 2% to 6.5% in all 5 scenes; PSNR and SSIM are improved by an average of 3.42 and 3.09% in the 5-scene experiments, and the effect is significant. Description of the Drawings

[0049] Figure 1 It is a flowchart of the method for 3D reconstruction of NeRF based on a deformable pose blur kernel;

[0050] Figure 2 It is an example diagram of the training set used in the embodiment;

[0051] Figure 3 It is an example diagram of the training set used in the embodiment;

[0052] Figure 4 Schematic diagram of the multi-layer perceptron (MLP) structure used in the embodiment;

[0053] Figure 5 Comparison diagram of the effects of static scene reconstruction in the embodiment. Detailed implementation manners

[0054] The present invention will be further explained below with reference to the accompanying drawings:

[0055] As Figure 1 shown, a NeRF three-dimensional reconstruction method based on a deformed pose blur kernel includes a training stage and a classification stage. The specific steps can be divided into four parts: blurred image data extraction and processing, blur kernel deformation, blurred pixel rendering, and scene reconstruction. The blurred image data used in this embodiment comes from the synthetic dataset with motion blur in the DeblurNeRF public dataset, and the camera poses are estimated and generated by the COLMAP software. Reconstruction is performed through a NeRF three-dimensional reconstruction method based on a deformed pose blur kernel as shown in Figure 1 below, including blurred image data extraction and processing, blur kernel deformation, blurred pixel rendering, and scene reconstruction. The specific steps are as follows:

[0056] Step 1. Target data extraction and processing

[0057] Target data extraction includes extracting input image data and camera pose data. For the input image data, the image with motion blur is used as the training sample, and the clear image is used as the corresponding label. The specific steps are as follows:

[0058] Step 1.1. Use the scene COZYROOM generated by render as the static scene to be reconstructed. Render 29 images with motion blur from different perspectives, as shown in Figure 2 below; 5 clear images, as shown in Figure 3 below. The size of all images is 400×600. Input the obtained 34 images into COLMAP to estimate the camera pose information corresponding to each image.

[0059] Step 1.2. Normalize the pixel RGB data of the image data obtained in Step 1.1 to obtain the image vector. The image vectors of the training set and the test set are respectively , .

[0060] Step 1.3. Initialize the starting camera pose and the ending camera pose during the camera exposure according to the camera pose information obtained in Step 1.1. The pseudocode is as follows:

[0061]

[0062] Among them, represents the camera pose matrix corresponding to the a-th image. and respectively represent the set of estimated information on the starting position and the ending position of the camera pose during a single camera exposure. represents a matrix initialized randomly with the same shape as and represents the jitter of the calculated matrix.

[0063] Step 1.4: Set the number of kernel points , create an initial blur kernel corresponding to the blurred image according to the normal distribution . Then initialize the corresponding view encoding set representation according to the number of images 29 , and each initial view encoding is a vector with a length of 32.

[0064] Step 2: Blur kernel deformation

[0065] Perform a deformation operation on the blur kernel according to the coordinates of the target pixel. Taking the pixel at the 2nd picture as an example, the specific steps are as follows:

[0066] Step 2.1: Select the corresponding blur kernel according to the target pixel k 2 =[ −0.0409,−0.0222 ,…,[0.0550,0.0690]] . Perform an inverse normalization process on , add a random perturbation of the normal distribution to the result of the inverse normalization to obtain a new blur kernel = [ −0.3790,−0.2102 ,…,[0.6493,0.5643]] . Calculate the encoding vector of .

[0067] Step 2.2: Select the corresponding view encoding according to the target pixel e 2 =[0,…,0] Perform a replication and extension operation according to the number of kernel points of the blur kernel to obtain . Concatenate obtained in Step 2.1 and to obtain a new vector with a shape of . For the coordinates of the blurred pixel to be generated, calculate its vector representation , after dimension expansion of sp, concatenate it with to obtain a vector .

[0068] Step 2.3, as Figure 4 shown, input the vector into the first MLP layer:

[0069]

[0070] The output vector y ∈ , and through the residual operation, splice and to obtain the new input vector . Then input into the second MLP layer:

[0071]

[0072] The output vector , and through the cutting operation, obtain the offset of the kernel point and the one-dimensional weight of length 5 . Finally, use the offset to deform the blur kernel to obtain k 2 = [ −0.6463,−1.7088 1 ,…, [2.1307,0.6183] 5 ] .

[0073] Step 3: Blurred pixel rendering

[0074] According to the blur kernel obtained in Step 2 and the camera pose information in Step 1, render the motion-blurred pixels. Taking the pixel at the second picture as an example, the specific steps are as follows:

[0075] Step 3.1: According to the number of kernel points of the blur kernel , perform linear interpolation on the estimated poses before and after exposure in Step 1. Set the duration of one exposure to 1 second and the interpolation interval to 0.2 seconds to obtain the estimated camera poses during one exposure process.

[0076] Step 3.2: Combine the blur kernel k 2 = [ −0.6463,−1.7088 1 ,…, [2.1307,0.6183] 5 ] with the target pixel coordinates Add them up and then estimate the camera pose Convert it into a ray representation in 3D space:

[0077]

[0078] where is the camera intrinsic matrix, is the coordinate information of the kernel point in the blur kernel, ∈ respectively represent the starting position and direction vector of the ray.

[0079] Step 3.3: Input the ray group into NeRF and render to obtain the corresponding RGB pixel set

[0080] Step 3.4: Perform weighted summation on the RGB pixel set using the weight vector to obtain the RGB information of the pixel at the target pixel of the second image. Finally, obtain the true RGB information at the target pixel of the second image in the training set, minimize the corresponding loss function and continuously optimize the model parameters.

[0081] Step 4: Scene reconstruction

[0082] For the 29 pose information in the training set, select 120 consecutive new view poses and input the new view poses into NeRF in sequence, and then obtain the newly rendered view images ∈ . Then use an image and video processing tool to process the vector and output the reconstructed clear scene.

[0083] Finally, for the images with motion blur in 5 different scenes, use different methods to generate the reconstructed images Figure 5 such as Figure 5As shown, where (a) is the original NeRF reconstructed image, (b) is the DeblurNeRF reconstructed image, (c) is the reconstructed image of this method, and (d) is the original scene. And three metrics, PSNR, SSIM, and LPIPS of the reconstructed images are calculated. Among them, PSNR is the peak signal-to-noise ratio (dB), which is used to evaluate the quality of the image; SSIM is the structural similarity index, which evaluates the image quality by comparing the brightness, contrast, and structural information of the images; LPIPS is a deep learning-based image quality evaluation metric that extracts image features through a pre-trained convolutional neural network and calculates the feature difference between two images. The results are shown in Table 1:

[0084] Table 1

[0085]

[0086] As can be seen from Table 1: For the unprocessed native NeRF, the quality of scene reconstruction is very poor. For the methods of pre-deblurring the pictures, MPR+NeRF and PVD+NeRF, since the information of the blurred images is limited, the real images cannot be well restored either. The evaluation metrics obtained by this method are better than those of DeblurNeRF and DP-NeRF in most scenarios. PSNR and SSIM have varying degrees of improvement in other scenarios except in the "pool" scenario. On average, both have significant improvements. Compared with DeblurNeRF and DP-NeRF, PSNR has increased by 2.43 and 2.28 respectively, and SSIM has increased by 3.09% and 6.34% respectively; LPIPS has been improved in all five scenarios compared with the original methods, and has increased by 4.39% and 2.52% on average compared with DeblurNeRF and DP-NeRF respectively.

Claims

1. A NeRF 3D reconstruction method based on deformable pose blur kernel, characterized by: The specific steps include: Step 1: Fuzzy image data extraction and processing Shots from different angles of the same scene Blurred image, extract its pixel matrix And estimated pose information , ; Create a corresponding blur kernel for each blurred image And the corresponding view encoding representation , and initialize; for the camera pose matrix Add random perturbations to parameterize the camera pose at the beginning and end of exposure , ; Step 2: Blur Kernel Deformation According to the perspective encoding and the coordinates of the blurred pixels , use the multi-layer perceptron to get the kernel point offset of the blur kernel deformation And the weights when synthesizing blurred pixels , for the initial blur kernel Transform to get , is the number of kernel points of a blur kernel; Step 3: Blurred pixel rendering According to the number of kernel points n in the blur kernel, the camera pose , Perform linear interpolation: ; represents the estimated camera pose obtained by interpolation at time t during the exposure process. Splice and get the blur kernel The corresponding estimated camera pose sequence ; Blur Kernel The core point information in is converted into ray representation through ray coding, and the ray group is obtained. : ; in are the vectors representing the starting point and direction of the ray respectively; represents the camera intrinsic parameter matrix, Blur kernel The coordinates of the t0th kernel point in the image; input the ray group Rays into NeRF and render to get the pixel RGB set representation ; Then according to the weight right Perform weighted summation to obtain the simulated blurred pixel RGB representation , construct the loss function , with the goal of minimizing the loss function, optimize the multilayer perceptron parameters NeRF parameters , blur kernel, view encoding and estimated pose information; Step 4: Scene reconstruction choose The corresponding camera pose matrix is ​​input into NeRF to obtain the reconstructed clear scene.

2. The NeRF 3D reconstruction method based on a deformable pose blur kernel as claimed in claim 1, characterized in that: The camera pose matrix , including a rotation matrix of size 3×3 , and a 3D column vector , where the rotation matrix Describes the rotation information of the camera in the world coordinate system, vector Represents the translation information of the camera in the world coordinate system.

3. The NeRF 3D reconstruction method based on deformable pose blur kernel as claimed in claim 1, characterized in that: Create m initial blur kernels for m blurred images And the corresponding view encoding representation ;in , Represents the normalized coordinates of the kernel point in the blur kernel, which is randomly initialized using a normal distribution; the perspective encoding is represented The vector value is initialized to 0 directly.

4. The NeRF 3D reconstruction method based on deformable pose blur kernel as claimed in claim 1, characterized in that: Initial blur kernel To transform, the specific steps are: Step 2.1: First, The core point By denormalizing , and then for the new core point Perform a random perturbation of the normal distribution and obtain , forming a new blur kernel ; For blur kernel In The pixel coordinates of the kernel points are sampled and encoded to generate a two-dimensional vector ; Step 2.2: Select the corresponding viewing angle encoding based on the blurred image , encode the perspective According to the number of blur kernel points n, the copy and expansion operation is performed to obtain a two-dimensional vector , then and Splicing to get a two-dimensional vector representation ; Calculate the coordinates of the blurred pixels to be generated One-dimensional vector representation of , and then perform a copy and expansion operation to obtain a two-dimensional vector ,and The concatenation process results in a two-dimensional vector ; Step 2.3: The vector obtained in step 2.2 Enter the first MLP layer Then the output With vector Concatenate to get vector , and then Input to the second MLP layer middle: ; ; in, ; Step 2.4: Use the kernel point offset obtained in step 2.3 Blur Kernel Deformation .

5. The NeRF 3D reconstruction method based on deformable pose blur kernel as claimed in claim 1, characterized in that: The loss function for: ; ; in, represents the RGB value of the real fuzzy pixel, B is the fuzzy pixel set, It is a structural loss The weight parameter, represents the parameters of the multilayer perceptron, ;in is a hyperparameter, Indicates randomly selecting a pixel set from the predicted pixel set And the corresponding pixel set in the training set Calculate the SSIM quality criterion.

6. The NeRF 3D reconstruction method based on deformable pose blur kernel as claimed in claim 1, characterized in that: choose A continuous generated perspective, the corresponding camera pose matrix Input the trained NeRF in sequence to get the pixel group set , , H and W are the height and width of the frame image, respectively, and c represents the pixel color; use image processing tools to operate the pixel groups in S to obtain a continuous generated view representation , Represents a reconstructed two-dimensional view; the continuous new view IM is used as the frame data of the clear scene output to obtain the reconstructed clear scene.

7. A computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to execute the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Method for reconstructing neural radiation field by using few blurred images

    CN118247418A

  • Neural radiation field new view angle synthesis method for fuzzy scene

    CN118279168A