High-precision view-angle-dependent appearance reconstruction method based on neural radiation field

By adopting the VD-NeRF method based on neural radiation field in three-dimensional reconstruction, the problem of high-frequency details capturing of objects with complex textures and perspective-dependent characteristics in virtual environments is solved, and high-precision three-dimensional reconstruction and realism are achieved.

CN120107467APending Publication Date: 2025-06-06CHONGQING UNIV OF TECH
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510162236.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

Existing three-dimensional reconstruction algorithms are difficult to accurately capture high-frequency details on the surface of objects with complex texture and perspective-dependent characteristics, especially in virtual environments, resulting in reduced rendering and weaker realism.

Method used

The high-precision viewing angle-dependent appearance reconstruction method based on neural radiation fields is adopted, and the accuracy of three-dimensional reconstruction and the modeling ability of viewing angle-dependent features are improved by building a VD-NeRF network architecture, including frequency domain-enhanced adaptive channel weighting module, viewing angle-dependent lighting feature learning module, attention-based volume rendering module and tone mapping module.

Benefits of technology

It significantly improves the restore ability of image details and the modeling performance of perspective-dependent features, improves the accuracy and rendering quality of three-dimensional reconstruction, and makes the surface of objects in the virtual environment more realistic and approximate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107467A_ABST
    Figure CN120107467A_ABST
Patent Text Reader

Abstract

The invention discloses a high-precision visual angle dependent appearance reconstruction method based on a neural radiation field, and relates to the technical field of artificial intelligence computer vision and three-dimensional reconstruction. The method at least comprises the following steps: S1, carrying out colmap processing on a shot multi-view image of an object to be reconstructed to obtain a corresponding camera pose; s2, building a network architecture used for calculating a non-explicit visual angle dependent color, i.e., VD-NeRF; and S3, based on the established VD-NeRF, measurement verification is carried out on the VD-NeRF by using the test set, model training is completed, and appearance reconstruction is carried out based on the trained model. Compared with other methods, in particular in the aspects of processing image blurring and incapable of accurately capturing the view angle dependent object surface, the VD-NeRF effectively improves the image detail reduction capability and the modeling performance of view angle dependent characteristics by combining the attention mechanism and the NeRF.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence computer vision and three-dimensional reconstruction technology, and specifically to a high-precision view-dependent appearance reconstruction method based on neural radiation fields. Background Art

[0002] The revolutionary virtual reality and augmented reality (VR / AR) technologies provide a true immersive experience and smooth interaction, which is expected to completely change the way we entertain, educate, work and live. The concept of virtual reality aims to create personalized content through a full 3D environment, and combine other technologies to simulate the visual, auditory, tactile and other sensory experiences of the real world, thereby achieving an immersive experience close to reality. In this environment, a high-fidelity 3D model needs to be rebuilt for each object.

[0003] Although most 3D reconstruction algorithms have made significant progress in object and scene reconstruction, they have difficulty accurately capturing fine surface details of objects with complex textures. In other words, they still face significant challenges when performing high-precision visual reconstruction and surface analysis, especially for the reconstruction of objects with perspective-dependent characteristics in virtual environments. The perspective-dependent characteristics make the surface of an object present natural light, shadow, reflection, and material texture at different viewing angles, which is closer to the human eye's perception of the real world. Therefore, improving the reconstruction accuracy of objects with complex textures and accurately presenting the perspective-dependent characteristics will significantly enhance users' realistic perception and immersive experience of objects in virtual environments.

[0004] As an emerging technology, Neural Radiance Field (NeRF) has become one of the focuses in 3D reconstruction research with its accurate visual rendering capabilities. NeRF represents 3D scenes through neural volumes, and can generate realistic novel views through the interaction of camera rays and scenes according to arbitrary observation angles. Its high-fidelity rendering effect and flexibility open up broad application prospects for the production and experience of VR / AR content in the future. However, when NeRF renders surfaces with high-frequency details, the clarity is significantly reduced, and blur often occurs at the edges. In addition, when NeRF processes objects with view-dependent characteristics such as translucent or reflective surfaces, it cannot accurately capture and reproduce the real surface appearance. In recent years, some researchers have tried to explain complex view-dependent appearance and improve rendering quality. For example, Ref-NeRF predicts the normal vector of each point on the ray and excludes points in the opposite direction of the camera through regularization, but the normal prediction of points located inside the surface is still uncertain. ABLE-NeRF adopts attention-based volume rendering and uses light probes to store the illumination of static scenes. However, blurring may still occur in areas with dense high-frequency information.

[0005] (1) The limitations of traditional methods in capturing high-frequency details, directly learning spatial information from images or relying on neural networks to process position inputs, often lead to significant degradation in rendering performance when representing high-frequency changes in color and geometry. This degradation is mainly due to the fact that adjacent sampling points in NeRF usually share similar input features, resulting in over-smoothing of output features.

[0006] (2) The current mainstream methods mainly rely on optimizing rendering quality in the spatial domain. Some researchers use Fourier transform and wavelet transform to seek solutions in the frequency domain, but this is not flexible enough. These methods cannot select the most informative frequency components for restoration, resulting in the reconstructed three-dimensional surface not being realistic enough.

[0007] (3) In virtual environments, in order to present more realistic visual effects, it is necessary to accurately capture the surface of objects with perspective-dependent properties (such as gloss and reflection). General methods still have certain limitations in capturing the real interaction between object surfaces and light.

[0008] Therefore, a new solution to the above problems needs to be proposed. Summary of the invention

[0009] The purpose of the present invention is to provide a high-precision view-dependent appearance reconstruction method based on neural radiation fields, and to achieve high-precision three-dimensional reconstruction by utilizing the three-dimensional representation of neural radiation fields and deep learning technology to train the network, so as to solve the technical problems raised in the background technology.

[0010] To achieve the above object, the present invention provides the following technical solution: a high-precision view-dependent appearance reconstruction method based on neural radiation field, comprising at least the following steps:

[0011] S1: The multi-view images of the object to be reconstructed are processed by colmap to obtain the corresponding camera pose;

[0012] S2: Build a network architecture for calculating non-explicit view-dependent color, namely VD-NeRF, wherein the VD-NeRF at least includes a frequency domain enhanced adaptive channel weighting module, a view-dependent illumination feature learning module, an attention-based volume rendering module and a tone mapping module, wherein the frequency domain enhanced adaptive channel weighting module is a FECAM module, the view-dependent illumination feature learning module includes an EAkA mechanism, and the attention-based volume rendering module is a VD-Transformer module;

[0013] S3: Based on the constructed VD-NeRF, the test set is used to measure and verify the VD-NeRF. The model training is completed, and the appearance is reconstructed based on the trained model.

[0014] Preferably, S1 at least includes the following steps: inputting the multi-view images sampled from the object to be reconstructed into colmap to perform feature extraction, feature matching and sparse reconstruction in sequence, and then exporting the obtained JSON file for subsequent network input.

[0015] Preferably, the construction and application of the VD-NeRF at least includes the following steps:

[0016] By sampling N volumes on the ray, each volume is mapped to a high-dimensional space through position encoding to capture high-frequency changes in the scene;

[0017] Then, light markers are inserted into the point sequence and fed into the Transformer model to compute non-explicit view-dependent colors.

[0018] In order to further improve the network's ability to render surface details, a FECAM module is designed to enhance the feature representation of volume embedding in the frequency domain.

[0019] Next, we design the EAKA mechanism and the VD-Transformer module, where the EAKA mechanism is used to optimize scene lighting and the VD-Transformer module is used for volume rendering.

[0020] The EAKA mechanism is used to optimize the light probes used to store scene lighting information, and the view information is integrated into the light sequence and passed to the VD-Transformer module;

[0021] Through the cross-attention and self-attention mechanisms, the VD-Transformer module can effectively process the embedded volume information, memorize the lighting characteristics of the scene and enhance the perspective dependency, thereby optimizing the final rendering effect and achieving better expressiveness;

[0022] In order to more realistically display highlight details and balance the brightness range while avoiding loss of dark details or overexposure of highlight areas, a tone mapping module is designed. The tone mapping module significantly improves the visual quality and realism of the image.

[0023] Set the loss function to gradually improve the rendering accuracy of the model.

[0024] Preferably, the designing of the FECAM module comprises at least the following steps:

[0025] In order to more effectively capture the detailed features of the 3D sampling points, the FECAM module uses discrete cosine transform (DCT) to more comprehensively characterize the features of the sampling points from different frequency dimensions, avoiding relying solely on low-frequency information extracted by global average pooling (GAP);

[0026] Since the weights of the discrete cosine transform are fixed, they can be pre-calculated and cached before training begins, thereby reducing the computational overhead during training;

[0027] In addition, the output of discrete cosine transform is in real number form, which avoids the need for additional inverse transformation or the introduction of parameters, thereby improving the computational efficiency and simplicity of the model;

[0028] In the framework of NeRF, the FECAM module refines the volume embedding of sampling points. Specifically:

[0029] The volume embedding feature sequence of the input sampling points is divided into n subgroups {t0, t1, ..., tn-1} according to the channel, where each subgroup Ti∈R1xL(i∈{0,1,...,n-1},n=Nt), and these subgroups jointly describe the feature components of the sampling points on different channels;

[0030] For each subgroup, discrete cosine transform is performed, which can be obtained:

[0031]

[0032] For each subgroup Ti (i∈{0,1,…,Nt-1}), it is mapped to the frequency domain by discrete cosine transform;

[0033] The index of the frequency component is j∈{0,1,…,Ls-1}, corresponding to a specific one-dimensional frequency basis function

[0034] It means extracting the lth column data from the feature matrix Ti of the i-th channel, that is, extracting a set of eigenvalues;

[0035] Where l represents the frequency index, Ls is the length of the signal, and this proportional factor is used to adjust the frequency scale, that is, to map the frequency index l to a standardized frequency range;

[0036] Through this process, the volume embedding channel features are converted into frequency components Freq i ∈RL, where L represents the dimension of the frequency component;

[0037] The entire frequency channel vector can be constructed by stacking the individual frequency channels;

[0038] Freq=DCT(T)=stack([Freq 0 ,Freq 1 ,...,Freq n-1 ])

[0039] After obtaining the frequency vector Freq∈RCxL, the FECAM module inputs these frequency features into the channel attention mechanism;

[0040] The goal of the channel attention mechanism is to dynamically adjust the importance of channels based on the frequency components of each channel;

[0041] The entire attention mechanism is expressed by the following formula:

[0042] F c -att=max(W 2 σ(W 1 DCT(T)))

[0043] Where σ is the sigmoid activation function, W 1 and W 2 are the weight matrices of two fully connected layers with bottleneck structures, obtained through learning, and max() is the ReLU activation function;

[0044] Through the interaction between the features of each channel and different frequency components, NeRF can not only extract key features from the frequency domain, but also further improve the ability to capture details through adaptive frequency domain attention weights, thereby enhancing the performance of the model in the rendering and reconstruction process.

[0045] Preferably, the EAKA mechanism is an efficient additive attention mechanism based on KAN, which generates highly consistent and realistic perspective-dependent rendering results by capturing the global illumination characteristics of the scene and combining the perspective direction characteristics;

[0046] Traditional attention mechanisms need to calculate the relationship between all features, which usually leads to a significant increase in computational and storage requirements. Our method adopts an efficient additive attention structure, which significantly reduces the computational complexity by limiting the calculation to only the relevant feature area, and more accurately captures the highlights and reflection characteristics under multiple views;

[0047] In addition, the EAKA mechanism uses the KAN layer to implement the interaction between queries and keys, and uses learnable activation functions and nonlinear weights to improve the model's expressiveness and computational efficiency while maintaining the accuracy and details of the rendering effect;

[0048] The learnable illumination embedding is obtained by using the matrix W q and W k is transformed into a query matrix Q and a key matrix K, where Q, K∈R n xd,W q and n represents the number of light probes, and d is the dimension of the embedding vector;

[0049] Then, the query matrix Q and the learnable parameter vector Wα ∈R d Multiply them together to calculate the attention weight of the query;

[0050] Finally, we get the global attention query vector α∈R n :

[0051]

[0052] It represents the importance of each token in the global context. It is a normalization factor used for scaling to prevent the inner product value from being too large in high-dimensional space;

[0053] Then the query matrix Q is weighted and summed by the calculated attention weights to generate the global query vector

[0054]

[0055] The global query vector q and the key matrix K are combined by element-wise multiplication to generate a global context matrix, which can effectively capture the characteristics of each light probe signal and flexibly learn the relationship between different light sources;

[0056] Compared with the traditional multi-head self-attention (MHSA) method, this calculation method can significantly reduce the computational overhead and has linear complexity in the token length, which is more suitable for processing complex light probe data. We use the KAN layer to process the interaction between queries and keys, thereby effectively extracting the potential information in the light probe features. Compared with the traditional MLP, the weight of each connection in KAN is no longer a simple numerical value, but is parameterized as a learnable spline function, which enables KAN to process high-dimensional data and complex lighting interactions more efficiently, thereby more fully exploring the lighting characteristics in the scene. The final output for:

[0057]

[0058] in, is a normalized query matrix. KAN is used to perform nonlinear transformations on the global context and query information, thereby learning richer feature representations and optimizing relationships in the feature space.

[0059] In this way, the embedded information of light probes can be combined with the global context to produce more accurate rendering results, especially when dealing with complex lighting and reflective surfaces.

[0060] Preferably, the VD-Transformer module introduces an attention mechanism to fuse the volume embedding features of the sampling points, the viewing angle information and the optimized scene illumination information, thereby achieving efficient feature interaction and integration and significantly improving the effect of scene rendering;

[0061] Specifically, the VD-Transformer module uses the learned scene illumination information as the query and the feature sequence of volume embedding as the key and value, and effectively integrates the scene features through the cross-attention mechanism to achieve deep fusion of illumination, viewpoint and volume information;

[0062] Subsequently, the fused features are further used to model the dependencies within the voxel sequence through a self-attention mechanism to enhance the representation capability of the illumination features.

[0063] In the final cross-attention operation, the viewpoint information is used as the query and combined with the self-attention processed features as the key and value to generate a globally consistent final representation, thereby improving the expressiveness of the NeRF model in the process of high-quality image rendering.

[0064] Preferably, the tone mapping module uses a mapping function to convert the physically linear color value Clinear output by VD-NeRF into a CsRGB color value that is more consistent with human eye perception;

[0065] In the model, Clinear is represented by both direct illumination and view-dependent illumination of the object surface;

[0066] By limiting the range of the converted CsRGB ([0,1]), we ensure that the final rendering result has balanced dynamic expression under different lighting conditions, significantly improving the visual quality and realism of the image.

[0067] Preferably, the loss function is:

[0068]

[0069] By minimizing the true color With predicted color The L2 loss between them is used to optimize the network, thereby gradually improving the rendering accuracy of the model.

[0070] Preferably, the measurement verification indicators used in S3 are SSIM and PSNR.

[0071] Compared with the prior art, the present invention has the following beneficial effects:

[0072] Compared with other methods, VD-NeRF effectively improves the image detail restoration capability and the modeling performance of perspective-dependent features by combining the attention mechanism with NeRF, especially in dealing with image blur and the inability to accurately capture the perspective-dependent object surface.

[0073] Specifically, position encoding is first used to map low-dimensional input to high-dimensional space, so that the network can better fit high-frequency changes. To further enhance the rendering quality, the present invention introduces an adaptive channel weighting strategy based on frequency domain enhancement, which can dynamically adjust the channel weights of different frequency components when rendering a new view, thereby selecting the most informative frequency components and ensuring that the rendering results can present local details more accurately.

[0074] In addition, perspective-dependent effects such as gloss and reflection of the surface of objects in the virtual world are crucial to improving rendering quality and enhancing immersion. To address this challenge, this paper designs an efficient additive attention mechanism EAKA based on KAN, which can more accurately learn and capture the lighting features in the scene through efficient query-key interaction.

[0075] These contributions enable VD-NeRF to achieve excellent visual effects when rendering scenes with view-dependent effects in the virtual world, further laying the foundation for the development and application of NeRF in virtual reality and augmented reality (VR / AR). BRIEF DESCRIPTION OF THE DRAWINGS

[0076] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for describing the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying any creative work.

[0077] Figure 1 A flow chart of the overall three-dimensional reconstruction method provided by the present invention;

[0078] Figure 2 A schematic diagram of the VD-NeRF model provided by the present invention;

[0079] Figure 3 A local visualization comparison diagram of the data set A provided by the present invention;

[0080] Figure 4 A local visualization comparison diagram of the data set B provided by the present invention;

[0081] Figure 5 This is a visualization comparison diagram of representative objects on datasets A and B provided by the present invention. DETAILED DESCRIPTION

[0082] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.

[0083] In view of the limitations of most reconstruction algorithms in capturing high-frequency details and the inability to reconstruct the surface of objects with perspective-dependent characteristics, this patent proposes a new rendering framework VD-NeRF that combines NeRF and attention mechanism. The model can accurately render three-dimensional detailed surfaces and scene appearances that change with perspective. VD-NeRF maps low-dimensional space coordinates to high-dimensional space through position encoding, so that the model can better fit high-frequency changes. Unlike the current mainstream methods that mainly rely on optimizing rendering quality in the spatial domain, the present invention explores solutions to achieve improvements in the frequency domain. By converting features to the frequency domain, the model can capture richer frequency information. In order to achieve flexible adaptation to different scene requirements, the present invention introduces a frequency-enhanced channel attention mechanism (FECAM module), which enhances high-frequency components by processing input features in the frequency domain. In addition, the FECAM module integrates the channel attention mechanism, which can dynamically assign weights to each channel and accurately capture the importance of each channel on different frequency components, so that the model can focus on the feature channels with the richest information. By combining this frequency enhancement and adaptive channel weighting strategy, the model is made more flexible and can capture key three-dimensional structural information in different scenes and perspectives, thereby improving the accuracy and detail of three-dimensional reconstruction. In order to present a more realistic visual effect, the present invention converts the lighting information in the scene into a trainable feature representation. This enables the radiation field to flexibly adapt to different lighting conditions and perspective changes. At the same time, the present invention proposes an efficient additive attention mechanism EAKA based on KAN, which enables the model to learn global context information to capture the complex interactions of light in the scene. EAKA dynamically generates a global context vector that reflects the characteristics of lighting and materials, so that the model can consistently express the details of light reflection, even when observed from different perspectives.

[0084] The high-precision view-dependent appearance reconstruction method based on neural radiation field includes at least the following steps:

[0085] S1: The multi-view images of the object to be reconstructed are processed by colmap to obtain the corresponding camera pose, and the multi-view images sampled by the object to be reconstructed are input into colmap to perform feature extraction, feature matching and sparse reconstruction in sequence, and then the obtained JSON file is exported for subsequent network input;

[0086] S2: Build a network architecture for computing non-explicit view-dependent color (see Figure 2), namely VD-NeRF, VD-NeRF at least includes a frequency domain enhanced adaptive channel weighting module, a view-dependent illumination feature learning module, an attention-based volume rendering module and a tone mapping module, the frequency domain enhanced adaptive channel weighting module is the FECAM module, the view-dependent illumination feature learning module includes the EAkA mechanism, and the attention-based volume rendering module is the VD-Transformer module;

[0087] The construction and application of VD-NeRF includes at least the following steps:

[0088] By sampling N volumes on the ray, each volume is mapped to a high-dimensional space through position encoding to capture high-frequency changes in the scene;

[0089] Then, light markers are inserted into the point sequence and fed into the Transformer model to compute non-explicit view-dependent colors.

[0090] In order to further improve the network's ability to render surface details, a FECAM module is designed to enhance the feature representation of volume embedding in the frequency domain.

[0091] Next, we design the EAKA mechanism and the VD-Transformer module, where the EAKA mechanism is used to optimize scene lighting and the VD-Transformer module is used for volume rendering.

[0092] The EAKA mechanism is used to optimize the light probes used to store scene lighting information, and the view information is integrated into the light sequence and passed to the VD-Transformer module;

[0093] Through the cross-attention and self-attention mechanisms, the VD-Transformer module can effectively process the embedded volume information, memorize the lighting characteristics of the scene and enhance the perspective dependency, thereby optimizing the final rendering effect and achieving better expressiveness;

[0094] In order to more realistically display highlight details and balance the brightness range while avoiding loss of dark details or overexposure of highlight areas, a tone mapping module is designed. The tone mapping module significantly improves the visual quality and realism of the image.

[0095] Set the loss function to gradually improve the rendering accuracy of the model.

[0096] Designing a FECAM module includes at least the following steps:

[0097] In order to more effectively capture the detailed features of the 3D sampling points, the FECAM module uses discrete cosine transform (DCT) to more comprehensively characterize the features of the sampling points from different frequency dimensions, avoiding relying solely on low-frequency information extracted by global average pooling (GAP);

[0098] Since the weights of the discrete cosine transform are fixed, they can be pre-calculated and cached before training begins, thereby reducing the computational overhead during training;

[0099] In addition, the output of discrete cosine transform is in real number form, which avoids the need for additional inverse transformation or the introduction of parameters, thereby improving the computational efficiency and simplicity of the model;

[0100] In the framework of NeRF, the FECAM module refines the volume embedding of sampling points. Specifically:

[0101] The volume embedding feature sequence of the input sampling points is divided into n subgroups {t0, t1, ..., tn-1} according to the channel, where each subgroup Ti∈R1xL(i∈{0,1,...,n-1},n=Nt), and these subgroups jointly describe the feature components of the sampling points on different channels;

[0102] For each subgroup, discrete cosine transform is performed, which can be obtained:

[0103]

[0104] For each subgroup Ti (i∈{0,1,…,Nt-1}), it is mapped to the frequency domain by discrete cosine transform;

[0105] The index of the frequency component is j∈{0,1,…,Ls-1}, corresponding to a specific one-dimensional frequency basis function

[0106] It means extracting the lth column data from the feature matrix Ti of the i-th channel, that is, extracting a set of eigenvalues;

[0107] Where l represents the frequency index, Ls is the length of the signal, and this proportional factor is used to adjust the frequency scale, that is, to map the frequency index l to a standardized frequency range;

[0108] Through this process, the volume embedding channel features are converted into frequency components Freq i ∈RL, where L represents the dimension of the frequency component;

[0109] The entire frequency channel vector can be constructed by stacking the individual frequency channels;

[0110] Freq=DCT(T)=stack([Freq 0 ,Freq 1 ,...,Freq n-1 ])

[0111] After obtaining the frequency vector Freq∈RCxL, the FECAM module inputs these frequency features into the channel attention mechanism;

[0112] The goal of the channel attention mechanism is to dynamically adjust the importance of channels based on the frequency components of each channel;

[0113] The entire attention mechanism is expressed by the following formula:

[0114] F c -att=max(W 2 σ(W 1 DCT(T)))

[0115] Where σ is the sigmoid activation function, W 1 and W 2 are the weight matrices of two fully connected layers with bottleneck structures, obtained through learning, and max() is the ReLU activation function;

[0116] Through the interaction between the features of each channel and different frequency components, NeRF can not only extract key features from the frequency domain, but also further improve the ability to capture details through adaptive frequency domain attention weights, thereby enhancing the performance of the model in the rendering and reconstruction process.

[0117] The EAKA mechanism is an efficient additive attention mechanism based on KAN. It captures the global illumination characteristics of the scene and combines the view direction characteristics to generate highly consistent and realistic view-dependent rendering results.

[0118] Traditional attention mechanisms need to calculate the relationship between all features, which usually leads to a significant increase in computational and storage requirements. Our method adopts an efficient additive attention structure, which significantly reduces the computational complexity by limiting the calculation to only the relevant feature area, and more accurately captures the highlights and reflection characteristics under multiple views;

[0119] In addition, the EAKA mechanism uses the KAN layer to implement the interaction between queries and keys, and uses learnable activation functions and nonlinear weights to improve the model's expressiveness and computational efficiency while maintaining the accuracy and details of the rendering effect;

[0120] The learnable illumination embedding is obtained by using the matrix W q and W k is transformed into a query matrix Q and a key matrix K, where Q, K∈R n xd,Wq and n represents the number of light probes, and d is the dimension of the embedding vector;

[0121] Then, the query matrix Q and the learnable parameter vector W α ∈R d Multiply them together to calculate the attention weight of the query;

[0122] Finally, we get the global attention query vector α∈R n :

[0123]

[0124] It represents the importance of each token in the global context. It is a normalization factor used for scaling to prevent the inner product value from being too large in high-dimensional space;

[0125] Then the query matrix Q is weighted and summed by the calculated attention weights to generate the global query vector

[0126]

[0127] The global query vector q and the key matrix K are combined by element-wise multiplication to generate a global context matrix, which can effectively capture the characteristics of each light probe signal and flexibly learn the relationship between different light sources;

[0128] Compared with the traditional multi-head self-attention (MHSA) method, this calculation method can significantly reduce the computational overhead and has linear complexity in the token length, which is more suitable for processing complex light probe data. We use the KAN layer to process the interaction between queries and keys, thereby effectively extracting the potential information in the light probe features. Compared with the traditional MLP, the weight of each connection in KAN is no longer a simple numerical value, but is parameterized as a learnable spline function, which enables KAN to process high-dimensional data and complex lighting interactions more efficiently, thereby more fully exploring the lighting characteristics in the scene. The final output for:

[0129]

[0130] in, is a normalized query matrix. KAN is used to perform nonlinear transformations on the global context and query information, thereby learning richer feature representations and optimizing relationships in the feature space.

[0131] In this way, the embedded information of light probes can be combined with the global context to produce more accurate rendering results, especially when dealing with complex lighting and reflective surfaces.

[0132] The VD-Transformer module introduces an attention mechanism to fuse the volume embedding features of the sampling points, the view information, and the optimized scene illumination information, thereby achieving efficient feature interaction and integration, and significantly improving the scene rendering effect;

[0133] Specifically, the VD-Transformer module uses the learned scene illumination information as the query and the feature sequence of volume embedding as the key and value, and effectively integrates the scene features through the cross-attention mechanism to achieve deep fusion of illumination, viewpoint and volume information;

[0134] Subsequently, the fused features are further used to model the dependencies within the voxel sequence through a self-attention mechanism to enhance the representation capability of the illumination features.

[0135] In the final cross-attention operation, the viewpoint information is used as the query and combined with the self-attention processed features as the key and value to generate a globally consistent final representation, thereby improving the expressiveness of the NeRF model in the process of high-quality image rendering.

[0136] The tone mapping module uses a mapping function to convert the physically linear color value Clinear output by VD-NeRF into a CsRGB color value that is more consistent with human eye perception;

[0137] In the model, Clinear is represented by both direct illumination and view-dependent illumination of the object surface;

[0138] By limiting the range of the converted CsRGB ([0,1]), we ensure that the final rendering result has balanced dynamic expression under different lighting conditions, significantly improving the visual quality and realism of the image.

[0139] The loss function is:

[0140]

[0141] By minimizing the true color With predicted color The L2 loss between them is used to optimize the network, thereby gradually improving the rendering accuracy of the model.

[0142] S3: Based on the constructed VD-NeRF, the test set is used to measure and verify the VD-NeRF. The model training is completed, and the appearance is reconstructed based on the trained model.

[0143] See also Figure 1 Based on the above content, a specific implementation plan is proposed, taking the reconstruction of multi-view images of a single object as an example:

[0144] S1: The multi-view images of the object to be reconstructed are processed by colmap to obtain the corresponding camera pose.

[0145] S2: Designing a network architecture for computing non-explicit view-dependent color.

[0146] S3: Designing a network architecture FECAM for feature representation enhancement in volumetric embeddings.

[0147] S4: Design of a network architecture EAKA for optimizing scene lighting.

[0148] S5: Designing a network architecture for volume rendering VD-Transformer

[0149] S6: Use the test set to measure and verify VD-NeRF, and the model training is completed.

[0150] The specific method of processing the sampled multi-view images in S1 is: input the multi-view images sampled by the object to be reconstructed into colmap to perform feature extraction, feature matching and sparse reconstruction in sequence, and then export the obtained JSON file for subsequent network input.

[0151] The network ideas and specific contents designed in S2 for calculating non-explicit view-dependent colors: Transformer processes the points sampled along the light through the self-attention mechanism. The specific sampling is divided into two stages, with 96 samples per ray in each stage. This method allows the model to selectively focus on the front points related to the current point when rendering, thereby more accurately capturing the effect of view dependence.

[0152] The idea and specific content of S3's feature representation network designed to enhance volume embedding are as follows: By using discrete cosine transform, we transfer the features to the frequency domain for analysis. In addition, using the channel attention mechanism, the model can dynamically adjust the channel weights of different frequency components according to the characteristics of the specific scene, further improving the reconstruction of surface details.

[0153] The idea and specific content of S4's network designed to optimize scene lighting are as follows: Traditional attention mechanisms need to calculate the relationship between all features, which usually leads to a significant increase in computing and storage requirements. Our method adopts an efficient additive attention structure, which significantly reduces the computational complexity by limiting the calculation to only the relevant feature area, and more accurately captures the highlights and reflection characteristics under multiple perspectives. In addition, EAKA uses the KAN layer to realize the interaction between queries and keys, and uses learnable activation functions and nonlinear weights to improve the model's expressiveness and computational efficiency while maintaining the accuracy and details of the rendering effect.

[0154] The idea and specific content of the network designed by S5 for volume rendering are as follows: Volume rendering integrates the color and density information along the light by simulating the process of light passing through the volume field, thereby generating an image observed from a new perspective. The rendering result is a weighted aggregation of the color and density of all sampling points along the light path. This process can be modeled by Transformer to learn how to effectively combine the color of each sampling point with other features in the volume field.

[0155] Among them, the calculation principle of the measurement and evaluation indicators involved in S6 is as follows: Performance is measured by two indicators, SSIM and PSNR. SSIM (Structural Similarity Index) measures image quality by evaluating the brightness, contrast and structure preservation of the image, focusing on the structural consistency of the image. PSNR (Peak Signal-to-Noise Ratio) evaluates the reconstruction accuracy and quality of the image by calculating the ratio of the maximum signal power to the noise power during the image reconstruction process. In NeRF, the two indicators SSIM and PSNR are sufficient to fully evaluate the image quality and can effectively measure the structural fidelity and reconstruction accuracy of the image.

[0156]

[0157] MAX is the maximum possible pixel value for the image (eg, for an 8-bit image, MAX=255).

[0158] MSE is the mean square error between the reconstructed image and the reference image.

[0159] The larger the PSNR value, the smaller the difference between the reconstructed image and the reference image, and the better the image quality.

[0160]

[0161] μx and μy are the average brightness of the two image blocks respectively.

[0162] σx 2 and σy 2 are the contrast (variance) of the two images respectively.

[0163] σxy is the covariance of two images, measuring their structural similarity.

[0164] C1 and C2 are constants used to avoid the denominator being zero.

[0165] Dataset A contains eight different objects, each with 100 images for training and validation, and 200 images for testing, all with a resolution of 800x800. The drums, materials, and boats have obvious view-dependent characteristics, such as reflections and highlights. We compared VD-NeRF with several advanced neural rendering methods on the NeRF Blender dataset, and the results are shown in Table 1. The rendering effect of our model at multiple viewpoints surpasses traditional NeRF methods, the latest advances in the field of view-dependent rendering, and some other methods, such as SDF-based methods and 3D Gaussian models. Figure 3 As shown in the figure, compared with ABLE-NeRF, VD-NeRF has a higher consistency with the Ground Truth in reconstruction results, showing significant visual advantages.

[0166] In Table 2, we show in detail the performance of our model on each object in the A dataset. Even compared with the most advanced models in the past two years, such as ABLE-NeRF and GaussianShader, our method shows significant advantages in PSNR and SSIM. This shows that although these advanced models have improved rendering quality, our method still significantly outperforms them, greatly improving the accuracy of detail reconstruction, and enhancing the performance of view-dependent rendering, obtaining higher quality rendering effects under multiple viewpoints.

[0167] To further evaluate our capabilities in view-dependent surface rendering, we use the B dataset. This dataset contains six different material properties, 100 training images and 200 test images per scene, and the resolution of all images is 800x800. Compared with the A dataset, the objects contained in the B dataset have obvious glossy properties and exhibit complex light reflection and refraction behaviors. The surface characteristics of these objects make the highlights and reflections more significant, which increases the difficulty of rendering. This additional complexity requires the model to simultaneously handle more environmental variables and surface details when generating new viewpoint images. Despite this, our method shows superior performance when compared with other state-of-the-art methods on this dataset, as shown in Table 3.

[0168] We compared with the state-of-the-art ABLE-NeRF on the more challenging B dataset. Although ABLE-NeRF uses the attention mechanism to aggregate multiple points on the light to calculate the final color and adopts the light probe technology to store the scene lighting information. However, during the rendering process, this method still faces the blurring problem and has certain inaccuracies in capturing the view dependency. In comparison, our method shows more realistic rendering results in terms of visual effects, such as Figure 4 .

[0169] In Table 4, we show the PSNR and SSIM metrics of our method for each object on the B dataset.

[0170] Table 1 shows the comparison of VD-NeRF with other advanced methods on Blender Dataset, with the first and second ranked results highlighted in bold and underlined respectively. The rest of the tables follow the same format.

[0171] Table 1

[0172]

[0173] Table 2 PSNR / SSIM of VD-NeRF for each scene on Blender Dataset

[0174]

[0175] Table 3 Comparison of VD-NeRF with other advanced methods on dataset B

[0176]

[0177] Table 4 PSNR / SSIM of VD-NeRF for each scene on the B dataset

[0178]

[0179] It will be apparent to those skilled in the art that the invention is not limited to the details of the exemplary embodiments described above and that the invention can be implemented in other specific forms without departing from the spirit or essential features of the invention. Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description, and it is intended that all variations falling within the meaning and scope of the equivalent elements of the claims be included in the invention. Any reference numeral in a claim should not be considered as limiting the claim to which it relates.

Claims

1. A high-precision view-dependent appearance reconstruction method based on neural radiation fields, characterized by: At least the following steps are included: S1: The multi-view images of the object to be reconstructed are processed by colmap to obtain the corresponding camera pose; S2: Build a network architecture for calculating non-explicit view-dependent color, namely VD-NeRF, wherein the VD-NeRF at least includes a frequency domain enhanced adaptive channel weighting module, a view-dependent illumination feature learning module, an attention-based volume rendering module and a tone mapping module, wherein the frequency domain enhanced adaptive channel weighting module is a FECAM module, the view-dependent illumination feature learning module includes an EAkA mechanism, and the attention-based volume rendering module is a VD-Transformer module; S3: Based on the constructed VD-NeRF, the test set is used to measure and verify the VD-NeRF. The model training is completed, and the appearance is reconstructed based on the trained model.

2. The high-precision view-dependent appearance reconstruction method based on neural radiation field according to claim 1, characterized in that: The S1 at least includes the following steps: inputting the multi-view images sampled from the object to be reconstructed into colmap to perform feature extraction, feature matching and sparse reconstruction in sequence, and then exporting the obtained JSON file for subsequent network input.

3. The high-precision view-dependent appearance reconstruction method based on neural radiation field according to claim 1, characterized in that: The construction and application of the VD-NeRF at least includes the following steps: By sampling N volumes on the ray, each volume is mapped to a high-dimensional space through position encoding to capture high-frequency changes in the scene; Then, light markers are inserted into the point sequence and fed into the Transformer model to compute non-explicit view-dependent colors. In order to further improve the network's ability to render surface details, a FECAM module is designed to enhance the feature representation of volume embedding in the frequency domain. Next, we design the EAKA mechanism and the VD-Transformer module, where the EAKA mechanism is used to optimize scene lighting and the VD-Transformer module is used for volume rendering. The EAKA mechanism is used to optimize the light probes used to store scene lighting information, and the view information is integrated into the light sequence and passed to the VD-Transformer module; Through the cross-attention and self-attention mechanisms, the VD-Transformer module can effectively process the embedded volume information, memorize the lighting characteristics of the scene and enhance the perspective dependency, thereby optimizing the final rendering effect and achieving better expressiveness; In order to more realistically display highlight details and balance the brightness range while avoiding loss of dark details or overexposure of highlight areas, a tone mapping module is designed. The tone mapping module significantly improves the visual quality and realism of the image. Set the loss function to gradually improve the rendering accuracy of the model.

4. The high-precision view-dependent appearance reconstruction method based on neural radiation field according to claim 3, characterized in that: The design of the FECAM module comprises at least the following steps: In order to more effectively capture the detailed features of the three-dimensional sampling points, the FECAM module uses discrete cosine transform to more comprehensively characterize the features of the sampling points from different frequency dimensions, avoiding relying solely on low-frequency information extracted by global average pooling; Since the weights of the discrete cosine transform are fixed, they can be pre-calculated and cached before training begins, thereby reducing the computational overhead during training; In addition, the output of discrete cosine transform is in real number form, which avoids the need for additional inverse transformation or the introduction of parameters, thereby improving the computational efficiency and simplicity of the model; In the framework of NeRF, the FECAM module refines the volume embedding of sampling points. Specifically: The volume embedding feature sequence of the input sampling points is divided into n subgroups {t0, t1, ..., tn-1} according to the channel, where each subgroup Ti∈R1xL(i∈{0,1,...,n-1},n=Nt), and these subgroups jointly describe the feature components of the sampling points on different channels; For each subgroup, discrete cosine transform is performed, which can be obtained: For each subgroup Ti (i∈{0,1,…,Nt-1}), it is mapped to the frequency domain by discrete cosine transform; The index of the frequency component is j∈{0,1,…,Ls-1}, corresponding to a specific one-dimensional frequency basis function It means extracting the lth column data from the feature matrix Ti of the i-th channel, that is, extracting a set of eigenvalues; Where l represents the frequency index, Ls is the length of the signal, and this proportional factor is used to adjust the frequency scale, that is, to map the frequency index l to a standardized frequency range; Through this process, the volume embedding channel features are converted into frequency components Freq i ∈RL, where L represents the dimension of the frequency component; The entire frequency channel vector can be constructed by stacking the individual frequency channels; Freq=DCT(T)=stack([Freq 0 ,Freq 1 ,...,Freq n-1 ]) After obtaining the frequency vector Freq∈RCxL, the FECAM module inputs these frequency features into the channel attention mechanism; The goal of the channel attention mechanism is to dynamically adjust the importance of channels based on the frequency components of each channel; The entire attention mechanism is expressed by the following formula: F c -att=max(W2σ(W1DCT(T))) Among them, σ is the sigmoid activation function, W1 and W2 are two fully connected layer weight matrices with bottleneck structures, obtained through learning, and max() is the ReLU activation function; Through the interaction between the features of each channel and different frequency components, NeRF can not only extract key features from the frequency domain, but also further improve the ability to capture details through adaptive frequency domain attention weights, thereby enhancing the performance of the model in the rendering and reconstruction process.

5. The high-precision view-dependent appearance reconstruction method based on neural radiation field according to claim 4, characterized in that: The EAKA mechanism is an efficient additive attention mechanism based on KAN, which generates highly consistent and realistic perspective-dependent rendering results by capturing the global illumination characteristics of the scene and combining the perspective direction characteristics; The EAKA mechanism uses an efficient additive attention structure, which significantly reduces the computational complexity by limiting the calculation to the relevant feature area, and more accurately captures the highlight and reflection characteristics under multiple perspectives. In addition, the EAKA mechanism uses the KAN layer to implement the interaction between queries and keys, and uses learnable activation functions and nonlinear weights to improve the model's expressiveness and computational efficiency while maintaining the accuracy and details of the rendering effect; The learnable illumination embedding is obtained by using the matrix W q and W k is transformed into a query matrix Q and a key matrix K, where Q, K∈R n xd,W q and n represents the number of light probes, and d is the dimension of the embedding vector; Then, the query matrix Q and the learnable parameter vector W α ∈R d Multiply them together to calculate the attention weight of the query; Finally, we get the global attention query vector α∈R n : It represents the importance of each token in the global context. It is a normalization factor used for scaling to prevent the inner product value from being too large in high-dimensional space; Then the query matrix Q is weighted and summed by the calculated attention weights to generate the global query vector The global query vector q and the key matrix K are combined by element-wise multiplication to generate a global context matrix, which can effectively capture the characteristics of each light probe signal and flexibly learn the relationship between different light sources; Final Output for: in, is a normalized query matrix. KAN is used to perform nonlinear transformations on the global context and query information, thereby learning richer feature representations and optimizing relationships in the feature space. In this way, the embedded information of light probes can be combined with the global context to produce more accurate rendering results, especially when dealing with complex lighting and reflective surfaces.

6. The high-precision view-dependent appearance reconstruction method based on neural radiation field according to claim 5, characterized in that: The VD-Transformer module introduces an attention mechanism to fuse the volume embedding features of the sampling points, the viewing angle information, and the optimized scene illumination information, thereby achieving efficient feature interaction and integration and significantly improving the scene rendering effect; Specifically, the VD-Transformer module uses the learned scene illumination information as the query and the feature sequence of volume embedding as the key and value, and effectively integrates the scene features through the cross-attention mechanism to achieve deep fusion of illumination, viewpoint and volume information; Subsequently, the fused features are further used to model the dependencies within the voxel sequence through a self-attention mechanism to enhance the representation capability of the illumination features. In the final cross-attention operation, the viewpoint information is used as the query and combined with the self-attention processed features as the key and value to generate a globally consistent final representation, thereby improving the expressiveness of the NeRF model in the process of high-quality image rendering.

7. The high-precision view-dependent appearance reconstruction method based on neural radiation field according to claim 6, characterized in that: The tone mapping module uses a mapping function to convert the physical linear color value Clinear output by VD-NeRF into a CsRGB color value that is more consistent with human eye perception; In the model, Clinear is represented by both direct illumination and view-dependent illumination of the object surface; By limiting the range of the converted CsRGB ([0,1]), we ensure that the final rendering result has balanced dynamic expression under different lighting conditions, significantly improving the visual quality and realism of the image.

8. The high-precision view-dependent appearance reconstruction method based on neural radiation field according to claim 7, characterized in that: The loss function is: By minimizing the true color With predicted color The L2 loss between them is used to optimize the network, thereby gradually improving the rendering accuracy of the model.

9. The high-precision view-dependent appearance reconstruction method based on neural radiation field according to claim 1, characterized in that: The measurement verification indicators used in S3 are SSIM and PSNR.

Citation Information

Cited By

  • Improved Shiny-NeRF three-dimensional reconstruction method based on NeRF

    CN120807799A

  • A Shiny-NeRF 3D Reconstruction Method Based on NeRF

    CN120807799B

  • End-side reasoning acceleration method for large model of intelligent industrial robot with body

    CN122390099A

  • A large model end side reasoning acceleration method of embodied intelligent industrial robots

    CN122390099B