A Generalizable Neural Radiance Field Reconstruction Method Based on Multimodal Information Fusion

By constructing multimodal feature bodies and fused light context features in neural radiation field reconstruction, combined with photometric and geometric supervision, the problem of low reconstruction accuracy of neural radiation field at a small number of perspective angles is solved, high-quality three-dimensional reconstruction and two-dimensional rendering are achieved, and generalization ability is improved.

CN119359934BActive Publication Date: 2025-05-27ZHEJIANG UNIV CITY COLLEGE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411926361.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-25
Publication Date
2025-05-27
Estimated Expiration
2044-12-25

AI Technical Summary

Technical Problem

The prior art has low surface reconstruction accuracy in neural radiation field reconstruction based on a small number of viewing angles, especially in boundaries and reflective areas, and does not have the ability to generalize across scenes.

Method used

By constructing photometric feature bodies and geometric feature bodies, multimodal information is gradually fused, rays are sampled and fusion light context features are integrated, bulk density and radiation brightness are decoded, free viewing RGB-D images are rendered, and dense reconstruction of low-texture scenes is combined with photometric supervision and sparse geometric supervision.

Benefits of technology

High-quality three-dimensional reconstruction and two-dimensional rendering are realized, the shape radiation ambiguity problem is solved, the surface reconstruction accuracy of generalized neural radiation fields is improved, and the cross-scene reconstruction is shown to be higher robustness and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119359934B_ABST
    Figure CN119359934B_ABST
Patent Text Reader

Abstract

The present invention discloses a generalizable neural radiance field reconstruction method based on multi-modal information fusion, which specifically includes the following steps: Step 1, construct a photometric feature volume and a geometric feature volume based on unstructured multi-views, and construct a multi-modal neural encoding volume through progressive complementary fusion; Step 2, convert the multi-modal neural encoding volume and the original RGB pixel volume of the unstructured multi-views into volume density and radiance; Step 3, sample rays, fuse the context features of the sampled rays to obtain ray context features; Step 4, use the ray context features to decode the volume density and radiance, render and generate free-view RGB-D images, and then combine photometric supervision and sparse geometric supervision to guide the dense reconstruction of low-texture scenes. The present invention can solve the shape-radiance ambiguity problem, achieve high-quality 3D reconstruction and 2D rendering, and improve the surface reconstruction accuracy of the generalizable neural radiance field.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The field of the present invention is the field of visual image generation technology, and specifically relates to a method for reconstructing a generalizable neural radiance field based on multi-modal information fusion. Background Art

[0002] Scene reconstruction and novel view generation based on multi-view visual information have been long-standing challenges in computer vision and graphics. In recent years, neural rendering methods have significantly promoted the integration and development of these two fields. Neural Radiance Fields (NeRF) based on implicit 3D reconstruction and its subsequent works can already produce realistic novel view synthesis results, but they heavily rely on using a large number of multi-view images, combined with a long single-scene perspective optimization process. A significant drawback of these NeRF technologies is the shape-radiance ambiguity problem, that is, the inability to explicitly reconstruct the 3D structure of the scene with high quality. Another is the lack of generalization, and it is unable to quickly render across scenes.

[0003] While the works based on a small number of views have great advantages in terms of efficiency and have demonstrated the generalization potential of neural radiance fields in inferring cross-scene geometry and appearance information. These methods utilize the recent success in deep multi-view stereo (MVS), by warping 2D image features (inferred by 2D CNN) from nearby input views onto the swept plane in the frustum of the reference view, implicitly reasoning about the scene geometry by constructing a cost volume at the input reference view, quickly regressing the true image of the new viewpoint, with strong generalizability and avoiding the cumbersome per-scene optimization. However, when only a small number of views are available, the surface reconstruction accuracy of existing generalizable neural radiance fields is low, especially for complex regions such as boundaries and specular reflections, and incomplete geometric perception often leads to low-quality 2D view rendering results. Summary of the Invention

[0004] The purpose of the present invention is to provide a method for reconstructing a generalizable neural radiance field based on multi-modal information fusion. The present invention can solve the shape-radiance ambiguity problem while achieving high-quality 3D reconstruction and 2D rendering, and improving the surface reconstruction accuracy of the generalizable neural radiance field.

[0005] The technical solution of the present invention: A method for reconstructing a generalizable neural radiance field based on multi-modal information fusion specifically includes the following steps:

[0006] Step 1: Based on unstructured multi-views, construct a photometric feature volume and a geometric feature volume, and construct a multi-modal neural encoding volume by gradually complementary fusion of the photometric feature volume and the geometric feature volume;

[0007] Step 2: Convert the multi-modal neural encoding volume and the original RGB pixel volume of the unstructured multi-views into volume density and radiance;

[0008] Step 3: Sample rays based on the constructed multimodal neural encoding volume, and fuse the context features of the sampled rays based on the Transformer network to obtain ray context features;

[0009] Step 4: Use the ray context features to decode the volume density and radiance, render and generate a free-view RGB-D image based on the decoded volume density and radiance, and then combine photometric supervision and sparse geometric supervision to guide the dense reconstruction of low-texture scenes.

[0010] In the aforementioned generalizable neural radiance field reconstruction method based on multimodal information fusion, in Step 1, the process of constructing the photometric feature volume is as follows: First, according to the unstructured multi-view, use a bidirectional fusion backbone network to extract image features, use ConvNeXt to extract multi-scale semantic information providing regional and target overall surface features from the downsampled 4x, 8x, 16x, and 32x levels, extract shallow local appearance features downsampled 4 times, and then perform bidirectional feature fusion to encode the unstructured multi-view into a semantically enhanced photometric feature volume :

[0011] ;

[0012] where: represents the unstructured multi-view.

[0013] In the aforementioned generalizable neural radiance field reconstruction method based on multimodal information fusion, the process of constructing the geometric feature volume is as follows:

[0014] First, obtain the camera parameters , where is the camera intrinsic matrix, is the rotation matrix of the camera relative to the world coordinate system, is the translation vector of the camera relative to the world coordinate system. Then, use the homography transformation matrix to transform the 2D features of the th auxiliary view to the reference view, and obtain the depth at the photometric feature volume :

[0015] ;

[0016] where , and are the camera parameters of the th auxiliary view, , and are the camera parameters of the reference view, is the normal direction of the reference view;

[0017] ;

[0018] wherein, is the standardized device coordinates under the reference view;

[0019] is to explicitly encode the difference degree between the spatial point photometric features of the unstructured multi-view, and calculate the photometric variance feature volume , and the formula is as follows:

[0020] ;

[0021] wherein, Var is the variance function for calculating the photometric feature of each spatial point under the reference view;

[0022] Then use 3D CNN to encode and decode the photometric variance feature volume , and after sigmoid operation, convert it into an explicit geometric feature volume .

[0023] In the aforementioned method for reconstructing a generalizable neural radiance field based on multi-modal information fusion, in step one, the photometric feature volume and the geometric feature volume are used to construct a multi-modal neural encoding volume through progressive complementary fusion as follows:

[0024] Calculate the photometric peak feature volume :

[0025] ;

[0026] wherein, MVSMaxPooling calculates the maximum value of the photometric feature of each spatial point under the reference view;

[0027] Use the 3D convolutional layer to fuse the original RGB pixel volume and the photometric variance feature volume , use the 3D convolutional layer to fuse the photometric peak feature volume and the geometric feature volume , and finally generate the final multi-modal neural encoding volume under the control of the trainable scaling factor , and the formula is as follows:

[0028] ;

[0029] ;

[0030] ;

[0031] Among them, is the fused feature volume of the original RGB pixel volume and the photometric variance feature volume . is the fused feature volume of the photometric peak feature volume and the geometric feature volume .

[0032] In the aforementioned method for reconstructing a generalizable neural radiance field based on multimodal information fusion, in the photometric peak feature volume contains voxels, which are fused separately with the geometric feature volume to form a gating mechanism to filter out the noise introduced by non-surface elements in the feature space.

[0033] In the aforementioned method for reconstructing a generalizable neural radiance field based on multimodal information fusion, in step two, the process of converting the multimodal neural encoding volume and the original RGB pixel volume of the unstructured multi-view into volume density and radiance is to construct a neural radiance field . Given any 3D position , viewing direction unit vector , learn the three-dimensional environmental geometry and appearance information encoded in the multimodal neural encoding volume . The neural radiance field is converted from the multimodal neural encoding volume and the original RGB pixel volume of the unstructured multi-view into the corresponding volume density and radiance through continuous interpolation. It is expressed by the formula as follows:

[0034] ;

[0035] Among them, is the normalized device coordinate under the reference view, is the unit vector of the reference view coordinate system. The multimodal neural encoding volume performs trilinear interpolation according to the coordinate .

[0036] In the aforementioned method for reconstructing a generalizable neural radiance field based on multimodal information fusion, in step three, the process of fusing the ray context features is to first represent the camera ray as:

[0037] ;

[0038] Among them, is the ray origin, is the distance along the ray direction, is the unit vector along the ray direction;

[0039] Then, perform hierarchical sampling on the camera rays, with the range being the outermost boundary of the ray origin and the innermost boundary , and divide to into intervals. Randomly select a sample point from each interval. The th sampling point is expressed by the formula:

[0040] ;

[0041] Then, based on the transformer residual network fuse the multimodal neural encoding body in the ray context information and the original RGB pixel body to obtain the ray context feature body :

[0042] .

[0043] In the aforementioned method for reconstructing a generalizable neural radiance field based on multimodal information fusion, in step four, the process of decoding the volume density and radiance using the ray context features is based on a multi-layer perceptron , and to decode the volume density and radiance point by point; among them, the decoding process of the th sampling point is expressed as:

[0044] ;

[0045] ;

[0046] ;

[0047] The process of rendering and generating a free-view RGB-D image is to render the RGB image C and the depth image D through the differentiable ray marching algorithm for the decoded volume density and radiance ; among them, the RGB value and depth value of the kth pixel are calculated as follows:

[0048] ;

[0049] ;

[0050] ;

[0051] Among them, represents the volume transmittance, represents the total number of sampling points on a single ray, represents the th distance from the sampling point to the origin, represents the th distance between the sampling point and the next point.

[0052] In the aforementioned method for reconstructing a generalizable neural radiance field based on multi-modal information fusion, the sparse geometric supervision converts the high-confidence sparse point cloud output by the screening algorithm into a voxel representation, anchors the surrounding area to restrict the generation of a smooth radiance field, and constructs geometric constraints in the depth image space using the sparse depth map output by COLMAP; the sparse geometric supervision expands the area with voxels as anchor points to form a geometric supervision signal in the form of a heat map ; applying the geometric supervision signal to the geometric feature volume to construct the sparse point cloud loss .

[0053] In the aforementioned method for reconstructing a generalizable neural radiance field based on multi-modal information fusion, the process of photometric supervision is to train the neural radiance field and calculate the RGB loss ;

[0054] The process of combining photometric supervision and sparse geometric supervision is to calculate the sparse depth loss , and perform weighted summation with the sparse point cloud loss to obtain the neural radiance field loss :

[0055] ;

[0056] Among them, and are weight coefficients.

[0057] Compared with the prior art, the present invention has the following beneficial effects:

[0058] The present invention constructs a photometric feature volume and a geometric feature volume, gradually and orderly fuses multi-modal information, samples light rays and fuses the light ray context features, decodes the volume density and radiance to render a free-view RGB-D image, realizes high-quality 3D reconstruction and 2D rendering, solves the shape-radiance ambiguity problem under as few as three unstructured multi-views, and combines the advantages of explicit and implicit scene reconstruction. The present invention introduces a sparse geometric supervision signal and combines it with photometric supervision, thereby enhancing the robustness and accuracy of joint appearance and geometric reconstruction. In addition, the present invention introduces a multi-modal information fusion module for geometric information, photometric information and semantic information, which gradually fuses the context features of light rays, thereby enhancing the generalization ability of the multi-modal neural encoding volume under limited viewpoints. Description of the Drawings

[0059] Figure 1 is the main idea of the method of the present invention;

[0060] Figure 2 is the overall process of the method of the present invention;

[0061] Figure 3 is the specific structure of the bidirectional fusion backbone network of the present invention;

[0062] Figure 4 is the construction process of the photometric variance feature volume and the photometric peak feature volume of the present invention;

[0063] Figure 5 is the construction process of the geometric feature volume of the present invention;

[0064] Figure 6 is the construction process of the multi-modal neural encoding volume fusion in the low-texture environment of the present invention;

[0065] Figure 7 is the structure of the light ray context fusion module of the present invention;

[0066] Figure 8 is an example of the RGB appearance reconstruction result on the LLFF dataset of the present invention;

[0067] Figure 9 is an example of the RGB appearance reconstruction result on the DTU dataset of the present invention;

[0068] Figure 10 is an example of the depth reconstruction result on the DTU dataset of the present invention. Detailed Embodiment

[0069] The present invention will be further described below in conjunction with the drawings and embodiments, but it shall not be used as a basis for limiting the present invention.

[0070] Embodiment

[0071] A method for reconstructing a generalizable neural radiance field based on multi-modal information fusion, the main idea is as follows Figure 1 As shown, given unstructured multi-views and the corresponding camera parameters , the 3D scene reconstruction task is to generate the corresponding RGB image C and depth image D for any view . For the geometric feature volume and photometric feature volume explicitly constructed for unstructured multi-views, a high-information multi-modal neural encoding volume is constructed through progressive complementary fusion. The geometric information is fused in the geometric feature volume , and the photometric information and semantic information of semantic priors are fused in the photometric feature volume . Based on the constructed multi-modal neural encoding volume , rays are sampled, the ray context features are fused based on the transformer network, and free-view RGB-D images are generated through rendering. Combining photometric supervision and sparse geometric supervision to guide the dense reconstruction of low-texture scenes, the overall process of this method is as Figure 2 shown.

[0072] Specifically, it is the following steps:

[0073] Step 1: Construct a photometric feature volume and a geometric feature volume based on unstructured multi-views, and construct a multi-modal neural encoding volume by progressively complementary fusion of the photometric feature volume and the geometric feature volume;

[0074] In this step, the current scene reconstruction technology relies heavily on highly distinguishable feature points and is not sensitive enough to low-texture areas. The reason is that the underlying neural network cannot use high-level semantic information to understand regions or objects from top to bottom, but can only mine feature point clues from low-level texture information from bottom to top. The instability of a small number of view feature point matches increases the reconstruction difficulty. Therefore, in order to encode low-texture scene features from images with as few as 3 views, this embodiment introduces a pre-trained deep neural network to supplement the overall semantic consistency prior, and transforms the features of as few as 2 auxiliary images into the frustum of the reference view through the plane sweep mode, and fuses the variance information and peak information of multi-view features to construct the photometric feature volume .

[0075] Use the Figure 3 bidirectional fusion backbone network shown in to extract image features. The bidirectional fusion backbone network in this embodiment For ConvNext, ConvNeXt is a modern convolutional neural network (CNN) architecture that, while maintaining the advantages of traditional CNNs, incorporates some features of the Transformer model to achieve higher accuracy and efficiency. The ConvNeXt architecture enhances the performance of the Transformer model by using larger convolutional kernels, improved network block designs, simplified network structures, etc. A separate downsampling layer is used in the network to achieve more flexible and effective spatial dimensionality reduction and feature fusion. ConvNeXt pre-trained on ImageNet22K (ImageNet22K is a large image recognition dataset containing 22,000 categories) extracts multi-scale semantic information from downsampling factors of 4, 8, 16, and 32 times, providing the overall surface features of regions and objects. Additional 2D CNNs are used to extract shallow local appearance features at a downsampling factor of 4 times. Through bidirectional feature fusion, the unstructured multi-view is encoded into a semantically enhanced photometric feature volume :

[0076] ;

[0077] In the formula: represents the unstructured multi-view.

[0078] Obtain the camera parameters , where is the camera intrinsic matrix, is the rotation matrix of the camera relative to the world coordinate system, is the translation vector of the camera relative to the world coordinate system, and then the homography transformation matrix is used to transform the 2D features of the th auxiliary view to the reference view, obtaining the photometric feature volume at depth :

[0079] ;

[0080] Among them, , and are the camera parameters of the th auxiliary view, , and are the camera parameters of the reference view, is the normal direction of the reference view;

[0081] ;

[0082] Among them, is the normalized device coordinate in the reference view.

[0083] To explicitly encode the degree of difference between the photometric features of spatial points in an unstructured multi-view, a photometric variance feature volume is calculated , and the formula is as follows:

[0084] ;

[0085] where Var is the variance function for calculating the photometric feature of each spatial point in the reference view.

[0086] The photometric variance feature volume is essentially constructed through pixel matching of multi-views and already implicitly contains some geometric information. On this basis, 3D CNN is used for encoding and decoding, and through sigmoid operation, it can be converted into an explicit geometric feature volume , as Figure 5 shown.

[0087] The sigmoid operation is an S-shaped function, which is common in biology and is also known as the S-shaped growth curve. In information science, due to its properties such as being monotonically increasing and having a monotonically increasing inverse function, the sigmoid function is often used as the activation function of neural networks to map variables between 0 and 1.

[0088] Adopt the process as Figure 6 shown, and gradually fuse the semantically enhanced photometric feature volume with the geometric feature volume to form a multi-modal neural encoding volume for the low-texture environment . The structure of the 3D CNN used, that is, the neural network module with a downsampling convolutional layer, an upsampling convolutional layer, and skip connections, can effectively infer and propagate the global information of the environment.

[0089] In the appearance and geometry joint reconstruction task, the variance cost feature volume only contains the relative differences between views and cannot provide absolute scale information. Perform multi-view maximum pooling operation, and by extracting the effective peak information in each view, construct a photometric peak feature volume for encoding the fused information of complete semantic information and photometric information:

[0090] ;

[0091] where MVSMaxPooling calculates the maximum value of the photometric feature of each spatial point in the reference view.

[0092] The photometric variance feature volume and the photometric peak feature volume are constructed as shown in Figure 4 .

[0093] 3D CNN (Three-Dimensional Convolutional Neural Network) is a deep learning model that applies 3D convolutional kernels to extract features from image or video data.

[0094] 3D convolutional layer is used to fuse the original RGB pixel volume of unstructured multi-views and the enhanced photometric variance feature volume , 3D convolutional layer is used to fuse the enhanced photometric peak feature volume and the geometric feature volume , and finally generates the final multi-modal neural coding volume under the control of the trainable scale coefficient , specifically calculated by the formula:

[0095] ;

[0096] ;

[0097] ;

[0098] Enhanced photometric peak feature volume contains a large number of meaningless voxels. When the voxel point is on the object surface, the peak feature can be used to supplement the photometric cues of the absolute scale, while when the element point is at other non-surface locations in space, the peaks from multi-view features are not meaningful. Fusing it with the geometric feature volume carrying object surface information alone can form a gating mechanism to filter the noise introduced by non-surface elements in the feature space. However, at the initial stage of model training, the learning of sparse point clouds is not yet complete, and at this time, the geometric feature volume has too much self-noise. Therefore, the initial value of the scale coefficient is set to 0, and in subsequent training, the value ratio of the two photometric peak feature volumes and the geometric feature volume can be adaptively evaluated through backpropagation.

[0099] Step 2: Convert the multi-modal neural coding volume and the original RGB pixel volume of unstructured multi-views into volume density and radiance;

[0100] Traditional scene reconstruction work explicitly represents the three-dimensional scene as point clouds or voxels, etc., which are not fine enough and not texture-friendly. Neural Radiance Field implicitly represents the geometry and appearance of the environment with a neural network, and can render a dense RGB-D map of the three-dimensional environment according to physical principles, but the accuracy of the depth map is extremely low. And the original Neural Radiance Field A large number of multi-view images are required for training, on the order of dozens to hundreds, and 3D reconstruction can only be performed in a single environment. The network obtained through long-term optimization cannot be generalized to new environments. Some research works have extended the neural radiance field using the general image features encoded by convolutional networks. However, the nature of its implicit learning results in poor quality in cross-scene geometric reconstruction.

[0101] In this step, after pre-training using image data from multiple scenes, there is no need to perform a time-consuming optimization process in the new environment, and only as few as 3 images are required for fast 3D reconstruction. The difference from previous works is that the proposed model combines the advantages of the high information content of explicit geometric information, photometric information, and semantic information and the smooth characteristics of implicit learning, and can generate more accurate representations for low-texture environments with a small number of viewpoints.

[0102] Specifically, given any 3D position and the unit vector of the viewing direction , the neural radiance field regresses the corresponding volume density and radiance from the multi-modal neural encoding volume and the raw RGB pixel volume of unstructured multi-views , which is expressed by the formula:

[0103] ;

[0104] where is the normalized device coordinate under the reference view, is the unit vector of the reference view coordinate system, and performs trilinear interpolation according to the coordinate .

[0105] To enhance the high-frequency details in the reconstruction result, positional encoding is applied to to convert it into a high-frequency representation. In addition, to provide the relative positions of the sampling points for the subsequent transformer network, positional encoding is also applied to . The multi-modal neural encoding volume has a relatively low resolution after 4-fold downsampling, and the raw RGB pixel volume of unstructured multi-views can provide more high-frequency appearance information.

[0106] Step 3: Sample rays based on the constructed multi-modal neural encoding volume, and fuse the context features of the sampled rays based on the transformer network to obtain the ray context features;

[0107] Since the 3D CNN used has a large regional receptive field, the multi-modal neural encoding volume During the fusion process, the reconstructed result is too smooth. The rendering processes of the RGB image and the depth image are achieved by tracing the virtual camera rays, and the rays are discretized and aligned with the pixel points. Therefore, independent fusion of the ray context information at this stage provides a finer perception and improves the volume density. The accuracy of regression.

[0108] Such as Figure 7 The ray context fusion module shown constructs a neural radiance field , learns the 3D environmental geometry and appearance information encoded in the multi-modal neural encoding volume , and converts it into volume density through continuous interpolation and view-related radiance . Represent the camera ray as:

[0109] ;

[0110] where, is the ray origin, is the distance along the ray direction, is the unit vector along the ray direction; the boundaries of the farthest and nearest points of the ray origin are and respectively.

[0111] To fuse the ray context information, first perform hierarchical sampling on the camera ray, divide to into intervals, randomly select a sample point from each interval, and the th sampling point can be expressed by the formula:

[0112] ;

[0113] Then, based on the Transformer residual network fuse the environmental feature volume in the ray context and the original RGB pixel volume to obtain the ray context feature :

[0114] ;

[0115] Step 4: Use the ray context feature to decode the volume density and radiance, render a free-view RGB-D image based on the decoded volume density and radiance, and then combine photometric supervision and sparse geometric supervision to guide the dense reconstruction of low-texture scenes.

[0116] In this step, based on the multi-layer perceptron 、 and decode the volume density and radiance point by point. The decoding process of the th sampling point can be expressed as:

[0117] ;

[0118] ;

[0119] .

[0120] Due to the existence of shape-radiance ambiguity, a radiance field that simply satisfies the RGB constraint cannot guarantee the correct geometric structure. Especially in a low-texture environment, the reconstructed geometric structure often deviates severely. A geometric constraint established in the 3D voxel space is introduced to limit the generation of a smooth radiance field by anchoring the surrounding area with a sparse key point cloud. Since a series of downsamplings have been performed, the resolution of the 3D voxel space is lower than that of the original image, so the geometric constraint force is limited. To obtain an accurate reconstruction result with high resolution, the sparse depth map output by COLMAP is used to further construct the geometric constraint in the depth image space.

[0121] Neural Radiance Field adopts a physically based voxel rendering process to render the RGB image and the depth image through the differentiable ray marching algorithm. The RGB value and the depth value of the

[0122] th pixel can be calculated by the formula:

[0123] ;

[0124] ;

[0125] Among them, represents the volume transmittance, represents the total number of sampling points on a single ray, represents the distance from the th sampling point to the origin, represents the distance from the th sampling point to the next point.

[0126] The overall network model is trained in an end-to-end manner. The L2 loss function is used to calculate the RGB loss , and the smooth L1 loss function is used to calculate the sparse depth loss , and then weighted sum with the sparse point cloud loss to obtain the final neural radiance field loss :

[0127] ;

[0128] where and are weight coefficients.

[0129] The scene appearance reconstruction experiments were conducted on the DTU and LLFF datasets, and the depth reconstruction experiments were conducted on the DTU dataset. The DTU dataset includes some low-texture scenes, each scene has 49 viewpoints, and each viewpoint contains 7 kinds of illuminations. The data settings are consistent with the literature. 88 scene data on the DTU dataset were used to train the proposed generalizable neural radiance field, 16 scene data were used for testing, and RGB-D images with a resolution of 512×640 were used. The LLFF dataset includes some low-texture and reflective environments, and each environment has 20 viewpoints of RGB images. This dataset has a different distribution from the DTU training set, and a total of 8 environments were used for testing. For each test environment, 7 nearby views were selected, among which 3 views were used as inputs, and the remaining 4 views were used to evaluate the model performance.

[0130] Given the ground truth image and the generated image , in the experiment, PSNR, SSIM, and LPIPS metrics were used to evaluate the RGB image generation performance of the model under new viewpoints.

[0131] PSNR is the peak signal-to-noise ratio, which is defined based on the mean squared error MSE. The larger the PSNR, the less the image distortion. The specific calculation formula is:

[0132] ;

[0133] ;

[0134] SSIM is the structural similarity index, which quantifies the local structural similarity from three aspects: brightness, contrast, and structural similarity degree, imitating the human visual system. The larger the SSIM, the less the image distortion. The specific calculation formula is:

[0135] ;

[0136] where and respectively represent and the averages of and respectively represent and The variance of denotes and the covariance of and are constants to maintain stability and avoid the denominator being zero.

[0137] LPIPS is the Learned Perceptual Image Patch Similarity, which uses the L2 distance of image deep features to measure the similarity of image pairs, and can better reflect the human perception than PSNR and SSIM. The smaller this metric is, the less image distortion there is.

[0138] Abs err is the average absolute error, which is calculated by averaging the absolute depth errors of all pixels and is in meters. Acc is the threshold percentage. Acc(0.01) represents the percentage of pixels with an absolute depth error lower than 0.01 meters, and Acc(0.05) represents the percentage of pixels with an absolute depth error lower than 0.05 meters.

[0139] In the experiment, the number of plane scans is set to 128, the number of randomly sampled rays in a single batch is 1024, and the multi-task loss weights and are set to 2 and 1. The model is trained on a server with a single NVIDIA Titan RTX GPU using the Adam optimizer, and the initial learning rate is Combined with the cosine annealing strategy to dynamically adjust the learning rate, the best performance is achieved after 15 epochs.

[0140] The present invention constructs a photometric feature volume and a geometric feature volume, gradually and orderly fuses multi-modal information, samples rays and fuses the ray context features, decodes the volume density and radiance to render and generate free-view RGB-D images, realizes high-quality 3D reconstruction and 2D rendering, solves the shape-radiance ambiguity problem under as few as three unstructured multi-views, and combines the advantages of explicit and implicit scene reconstruction. The present invention introduces a sparse geometric supervision signal and combines photometric supervision, thereby enhancing the robustness and accuracy of joint appearance and geometric reconstruction. In addition, the present invention can achieve high-quality generalization of the 3D reconstruction task in the environment, and when as few as three unstructured multi-views are captured, the depth estimation performance exceeds that of the advanced 3D reconstruction neural network MVSNet with full-depth map supervision. In the new view generation task, compared with the existing advanced generalizable neural rendering work MVSNeRF, the present invention significantly improves the rendering quality of complex regions such as boundaries and specular reflections, and achieves more accurate new view generation performance. The present invention introduces a multi-modal information fusion module for geometric information, photometric information and semantic information, which gradually fuses the context features of rays, thereby enhancing the generalization ability of the multi-modal neural encoding volume under limited viewpoints.

[0141] Furthermore, for the appearance reconstruction metrics of the LLFF and DTU test datasets, the method proposed in this invention was compared with other generalizable baseline methods, including PixelNeRF, IBRNet, and MVSNeRF. The results are shown in Tables 1 and 2. PixelNeRF discretely fuses multi-view features through average pooling operations, but there is an overfitting problem in cross-scene training, making it difficult to generalize to the LLFF environment. IBRNet significantly improves the RGB appearance reconstruction performance in the generalization environment by fusing features along the rays. However, the scope of its feature fusion is limited to 2D images and 3D rays, so the perceptual similarity of the generated 3D space is relatively low. MVSNeRF uses 3D CNN to fuse the full-environment features on the cost volume, but the lack of region- and object-level perception introduces noise into the rendering process.

[0142] Table 1 Quantitative results of novel view synthesis on the LLFF dataset

[0143]

[0144] Table 2 Quantitative results of novel view synthesis on the LLFF dataset

[0145]

[0146] The results of the quantitative comparison experiments show that the proposed method significantly outperforms the baseline methods on both datasets. The average relative improvements in PSNR, SSIM, and LPIPS reach 1.6%, 1.6%, and 4.6% respectively, verifying that the strategy of combining multi-feature volume fusion and ray context fusion can more reasonably fuse the three-dimensional environmental context information and generate finer novel view images. The performance improvement is the largest in LPIPS, indicating that the photometric feature volume enhanced by semantic priors extracts effective semantic perception information through bidirectional encoding, providing richer semantic knowledge for the environmental generalization of neural radiance fields. In particular, this method complements the unfair comparison with the original NeRF, which is a non-generalizable offline tuning method that uses all the test scene images for training and takes about 10.2 hours of optimization with approximately 200 thousand iterations. It is worth noting that without tuning using the test scene images, the method of this invention outperforms the NeRF method in all metrics, verifying the feasibility of the high-quality generalizable NeRF method.

[0147] The geometric reconstruction metrics of the reference view and the new view are divided into two categories. For the DTU test dataset, the method proposed in the present invention is compared with other generalizable baseline methods, including the neural radiance field methods PixelNeRF, IBRNet, and MVSNeRF, and the classical multi-view depth neural network method MVSNet. The results are shown in Table 3. Like other neural radiance field methods, the method of the present invention only takes RGB images as input data and can freely select new views for geometric reconstruction. MVSNet is trained using real depth maps but can only reconstruct the depth of the reference view.

[0148] Table 3 Quantitative depth reconstruction results on the DTU dataset

[0149]

[0150] The results of the quantitative comparison experiment show that the method of the present invention significantly outperforms all baseline methods for both the input view and the new view. With only three input images, the reconstructed Abs err is 1 cm in the reference view. Compared with the existing generalizable neural radiance field models, it is reduced by 56.5%, and the Acc(0.01) is increased by 13.5%. In addition, the reconstruction performance improvement for the new view is even greater, with Abs err reduced by 62.9% and Acc(0.01) relatively increased by 18.4%.

[0151] It should be noted that different from MVSNet which uses real annotated dense depth maps, the neural radiance field method proposed in the present invention only uses the sparse key points reconstructed by SFM and can still use geometric supervision at both the voxel and ray levels to achieve high-quality depth reconstruction. In generalizable neural networks, this is the first time that a neural radiance field type model has achieved overall performance superior to that of the multi-view stereo network, with the reconstructed Abs err in the reference view reduced by 44.4%.

[0152] Figure 8 Examples of RGB appearance reconstruction results on the LLFF dataset are shown, and the proposed method is qualitatively compared with the state-of-the-art baseline method MVSNeRF. The results show that flickering and artifacts can be observed in the generated views of MVSNeRF, while the method of the present invention can better adapt to environments with different distributions, demonstrating the good generalization of the proposed technology. In low-texture regions with reflection, such as TV screens and shiny painted desktops, the highlight effects generated by the method of the present invention change significantly with the movement of the view and are more realistic in terms of both brightness and shape, demonstrating that the proposed feature volume fusion and ray fusion strategies can effectively help the radiance field model understand the anisotropy of reflective surfaces.

[0153] Figure 9Examples of RGB appearance reconstruction results on the DTU dataset are shown, and the proposed method is qualitatively compared with the state-of-the-art baseline method MVSNeRF. The method of the present invention has strong advantages for free-view generation in low-texture environments. Inside the low-texture regions, both methods can generate reasonable new-view results. However, at the edges of the low-texture regions, the generated results of MVSNeRF have obvious artifacts, while the generated results of the method of the present invention have clear and sharp boundaries. This proves that the neural radiance field model that fuses semantic priors and sparse key-point cloud information has a certain overall object perception ability, and can effectively divide the boundary regions in 3D space even with a small number of input views, improving the quality of new-view rendering.

[0154] Figure 10 Examples of depth reconstruction results on the DTU dataset are shown, and the proposed method is qualitatively compared with the state-of-the-art generalizable neural radiance field method MVSNeRF. The visualization results show that MVSNeRF fails to solve the shape radiance ambiguity problem of neural radiance fields, and its reconstruction results contain a lot of background noise, the edges of the objects are blurred, and the reconstruction accuracy of the overall surface tends to be smooth due to the influence of the feature volume resolution. The method of the present invention extracts effective features from low-texture surfaces and emphasizes the edge information at the object boundaries by anchoring robust key points, thus solving the shape radiance ambiguity problem of neural radiance fields.

[0155] In summary, the present invention can solve the shape radiance ambiguity problem, achieve high-quality 3D reconstruction and 2D rendering, and improve the surface reconstruction accuracy of generalizable neural radiance fields.

Claims

1. A generalizable neural radiation field reconstruction method based on multimodal information fusion, characterized in that: The specific steps include: Step 1: construct photometric feature bodies and geometric feature bodies based on unstructured multi-views, and construct a multimodal neural encoding body by progressively and complementary fusion of the photometric feature bodies and the geometric feature bodies; The process of constructing the photometric feature body is to firstly use the bidirectional fusion backbone network f according to the unstructured multi-view. T Extract image features, use ConvNeXt to extract multi-scale semantic information that provides regional and overall surface features of the target from downsampling by 4, 8, 16, and 32 times, extract shallow local appearance features downsampled by 4 times, and then perform bidirectional feature fusion to encode unstructured multi-views into semantically enhanced photometric feature volumes F i T : F i T =f T (I i ); Where: I i Representing unstructured multiple views; The geometric feature body F is constructed Λ The process is as follows: First, obtain the camera parameters Φ = [K, R, t], where K is the camera intrinsic parameter matrix, R is the rotation matrix of the camera relative to the world coordinate system, and t is the translation vector of the camera relative to the world coordinate system. Then use the homography transformation matrix Ξ i (z) Transform the 2D features of the i-th auxiliary view to the reference view to obtain the photometric feature volume at depth z Among them, K i , R i and t i is the camera parameter of the i-th auxiliary view, K1, R1 and t1 are the camera parameters of the reference view, and n1 is the normal direction of the reference view; Among them, u, v, and 1 are the standardized device coordinates under the reference view angle; To explicitly encode the difference between the spatial point photometric features of unstructured multi-views, the photometric variance feature volume F is calculated. V , the formula is as follows: Among them, Var is the variance function of the photometric characteristics of each spatial point under the reference viewing angle; Then use 3D CNN to analyze the photometric variance feature volume F V Encode and decode, and convert it into an explicit geometric feature volume F through sigmoid operation Λ ; The photometric feature F i T and geometric feature body F Λ Constructing a multimodal neural encoder through progressive complementary fusion L The process is as follows: Calculate the photometric peak feature F M : Among them, MVSMaxPooling calculates the maximum value of the photometric feature of each spatial point under the reference viewing angle; Using a 3D convolutional layer f VC Fused original RGB pixel volume F C and the photometric variance feature F V , using a 3D convolutional layer f MΛ Fusion photometric peak feature F M and geometric feature body F Λ , and finally in the trainable proportional coefficient α MΛ The final multimodal neural encoding body F is generated under the control of L , the formula is as follows: F VC =f VC ([F V ;F C ]); F MΛ =f MΛ ([F M ;F Λ ]); F L =F VC +α MΛ F MΛ ; Among them, F VC is the original RGB pixel volume F C and the photometric variance feature F V The fusion feature body, F MΛ is the peak luminosity feature F M and geometric feature body F Λ The fusion feature body of Step 2: Convert the original RGB pixel volume of the multimodal neural encoding volume and the unstructured multi-view into volume density and radiance; Step 3: based on the constructed multimodal neural encoder, sample light, fuse the context features of the sampled light based on the transformer network, and obtain the light context features; Step 4: Use light context features to decode volume density and radiance, generate free-view RGB-D images based on the decoded volume density and radiance rendering, and then combine photometric supervision with sparse geometric supervision to guide the dense reconstruction of low-texture scenes.

2. The generalizable neural radiation field reconstruction method based on multimodal information fusion according to claim 1 is characterized in that: The photometric peak feature F M contains voxels, and the geometric feature volume F Λ They are fused separately to form a gating mechanism to filter out the noise introduced by non-surface elements in the feature space.

3. The generalizable neural radiation field reconstruction method based on multimodal information fusion according to claim 1, characterized in that: In step 2, the multimodal neural encoding body F L and the raw RGB pixel volume F of the unstructured multi-view C The process of converting to volume density σ and radiance r is to construct the neural radiation field f A , given any 3D position x and viewing direction unit vector v, learn a multimodal neural encoder F L The 3D environment geometry and appearance information encoded in the neural radiance field f A From the multimodal neural encoder F L and the raw RGB pixel volume F of the unstructured multi-view C The volume density σ and radiance r are converted into the corresponding volume density σ and radiance r by continuous interpolation, which can be expressed as follows: σ,r=f A (x,v,F L ,F C ); Where x is the normalized device coordinate under the reference view, v is the unit vector of the reference view coordinate system, and the multimodal neural encoding volume F L Perform trilinear interpolation based on coordinate x.

4. The generalizable neural radiation field reconstruction method based on multimodal information fusion according to claim 3 is characterized by: In step 3, the process of fusion of light context features is to first transform the camera light r S (d) is expressed as: r S (d)=o+dv; Where o is the origin of the ray, d is the distance along the direction of the ray, and v is the unit vector along the direction of the ray; Then, the camera ray is layered and sampled, ranging from the farthest boundary d of the ray origin o f and the nearest boundary d n , d f to d n Divided into K D intervals, randomly select a sample point from each interval, the i-th sampling point d i The formula is: Then, based on the transformer residual network f Trans The multimodal neural encoder F in the light context information L and the original RGB pixel volume F C Fusion, get the light context feature volume F A : F A =f Trans ([F L ;F C ])。 5. The generalizable neural radiation field reconstruction method based on multimodal information fusion according to claim 4 is characterized in that: In step 4, the process of decoding volume density and radiance using light context features is based on a multi-layer perceptron f1 MLP , and Decode the volume density σ and radiance r point by point; the decoding process of the i-th sampling point is expressed as: F i B =f1 MLP (F i A ,x i ); The process of rendering and generating a free-view RGB-D image is to render the RGB image C and the depth image D for the decoded volume density σ and radiance r by a differentiable ray-walking algorithm; wherein the RGB value C of the kth pixel is k and depth value D k The calculation formula is as follows: Among them, τ i Represents volume transmittance, K V Indicates the total number of sampling points on a single ray, d i Represents the distance from the i-th sampling point to the origin, Vd i Indicates the distance from the i-th sampling point to the next point.