NeRF sparse view reconstruction method based on adaptive multi-modal feature fusion
The method addresses NeRF's sparse view reconstruction challenges by integrating depth and semantic features with adaptive fusion and progressive training, enhancing reconstruction quality and adaptability in real-world scenarios.
Patent Information
- Application Number
- CN202510488553.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-15
AI Technical Summary
The reconstruction quality of NeRF technology is degraded under sparse view conditions. The existing methods strictly require the distribution density and coverage of the input view, which is difficult to deal with cross-modal data fusion, and has low computing efficiency, which affects user experience and application effects.
Adaptive multimodal feature fusion method is adopted to extract depth and semantic features, combine a mixed attention mechanism and confidence-guided weight distribution network to perform feature weight fusion, and the consistency loss function harmonizes the conflict between modes, and input it to the enhanced neural radiation field rendering module for hierarchical feature encoding and progressive training.
While maintaining computational efficiency, it significantly improves the expression ability and scene adaptability of fusion features, can handle modal inconsistency and sensor noise problems in real-world scenarios, and generates high-quality three-dimensional scene reconstruction results.
Smart Images

Figure CN120318429A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of three-dimensional scene reconstruction, and particularly to a NeRF sparse view reconstruction method based on adaptive multimodal feature fusion. Background Art
[0002] Under the wave of global digital transformation, three-dimensional scene reconstruction technology is rapidly becoming a key technical support for multiple cutting-edge application fields such as computer graphics, computer vision, the metaverse, and smart cities. From traditional reconstruction methods relying on manual rules to neural implicit representations based on deep learning, three-dimensional reconstruction technology has undergone decades of development and transformation. Early three-dimensional reconstruction mainly relied on classical computer vision techniques such as Structure from Motion (SfM) and Multi-View Stereo (MVS), reconstructing the geometric structure of the scene through feature matching and disparity calculation, and then generating explicit three-dimensional representations such as point clouds, meshes, or voxels. Although these methods are theoretically mature, they often face huge challenges when dealing with complex materials and fine geometric details.
[0003] The emergence of Neural Radiance Fields (NeRF) technology has brought a revolutionary breakthrough to the field of three-dimensional scene reconstruction. As a new type of implicit neural representation method, NeRF represents the geometric structure and appearance features of a three-dimensional scene through an implicit neural network, and can learn and reconstruct high-quality three-dimensional scenes from a series of two-dimensional images. Compared with traditional three-dimensional reconstruction methods based on explicit representations, NeRF directly learns a continuous radiance field function from the input images, and has significant advantages in terms of fidelity, rendering quality, and detail performance.
[0004] The core idea of NeRF is to represent a three-dimensional scene as a continuous five-dimensional function that maps the position coordinates and viewing directions in three-dimensional space to corresponding colors and volume densities. Specifically, NeRF uses a Multilayer Perceptron (MLP) to approximate this complex mapping function, and optimizes the network parameters to make the images rendered from different viewpoints as close as possible to the actual captured images. During the rendering process, NeRF samples multiple points along each ray, calculates the color and density of each point, and then integrates this information through the volume rendering equation to obtain the final pixel color. This method can not only accurately reconstruct the geometric structure of the scene, but also capture complex lighting effects, material properties, and fine surface details.
[0005] At the practical application level, NeRF technology has demonstrated broad application prospects and great commercial value. In the fields of virtual reality and augmented reality, NeRF can provide users with ultra-high-quality immersive experiences, seamlessly integrating virtual scenes with the real world. In cultural heritage protection and museum digitization, NeRF can accurately record and digitize every detail of precious cultural relics and historical buildings, providing digital guarantees for the inheritance of human civilization. In the architecture and interior design industries, NeRF can create highly realistic virtual showrooms to help designers and clients communicate and make decisions better. In the field of e-commerce, NeRF technology can generate interactive 3D displays of products, significantly enhancing the user experience and purchase conversion rate. In film and television production and game development, NeRF can greatly simplify the 3D content creation process, reduce production costs, and shorten the development cycle. In medical image analysis, NeRF can reconstruct accurate 3D organ models from 2D slices such as CT or MRI, assisting doctors in diagnosis and surgical planning, and improving medical efficiency and safety.
[0006] Despite the broad application prospects of NeRF technology, its actual implementation still faces many challenges, one of which is the high standards and requirements for input data. Standard NeRF models usually require densely sampled multi-view images as input. In an academic research environment, these images are typically collected through carefully designed multi-camera arrays or turntable systems to ensure uniform view distribution and complete coverage. However, in a practical application environment, due to various physical constraints, equipment limitations, or economic cost considerations, we often can only obtain a limited number of input images with restricted views. For example, during the digitization of cultural relics, to protect the safety of the cultural relics, the shooting scene may be strictly restricted; in a complex industrial environment, equipment occlusion and space limitations make it difficult to comprehensively shoot certain areas; in the reconstruction of large outdoor scenes, limited by shooting equipment and time costs, it is difficult to obtain complete view coverage; in consumer-grade application scenarios, ordinary users may only take a small number of photos with a mobile phone, making it difficult to ensure a professional view distribution. These practical limitations make sparse view reconstruction a key challenge that NeRF technology must overcome to achieve wide application.
[0007] Although NeRF series methods perform excellently under dense view conditions, under sparse view conditions, the performance of the standard NeRF model will decline sharply, mainly manifested as: serious geometric distortions in unobserved areas; blurred texture and color performance; unnatural jumps when the view changes; inaccurate capture of surface reflection characteristics; incomplete reconstruction of detailed structures, etc. These problems seriously affect the user experience and application effects, restricting the popularization and commercialization of NeRF technology.
[0008] First, existing algorithms have strict requirements for the distribution density and coverage range of input views. Usually, dozens or hundreds of evenly distributed perspective inputs are required. When the number of input images is lower than the critical threshold, the reconstruction quality will show a cliff-like decline. This phenomenon stems from the lack of an explicit constraint mechanism for scene geometry in the standard NeRF framework, resulting in the problem of radiance field ambiguous solutions under sparse view conditions - that is, different spatial configurations may produce the same projection results.
[0009] Second, existing methods have obvious deficiencies in dealing with cross-modal data fusion. Most improvement schemes only focus on optimizing the feature extraction network of RGB images, but ignore the intrinsic value of auxiliary information such as depth sensors, LiDAR point clouds, and semantic segmentation maps. For example, in industrial inspection scenarios, sparse RGB images are difficult to accurately reconstruct the microscopic geometry of metal surfaces, and the failure to effectively fuse high-precision ToF (Time of Flight) depth data will directly lead to inaccurate reflection models. In addition, most existing feature fusion strategies adopt static weighting methods and cannot dynamically adjust the contribution weights of multi-modal data according to scene characteristics, resulting in feature conflicts easily under complex lighting or occlusion conditions.
[0010] Finally, when existing optimization schemes improve the performance of sparse view reconstruction, they often do so at the cost of computational efficiency. For example, the global attention mechanism based on transformers can enhance feature consistency, but it leads to an exponential increase in memory occupancy; while the method using probabilistic radiance field modeling improves the rendering quality of uncertain regions, but greatly increases the inference time. This contradiction between efficiency and accuracy severely restricts the application of this technology in real-time interactive systems.
[0011] Generally speaking, as a new paradigm for three-dimensional scene representation and reconstruction, NeRF technology shows great application potential and technical value. However, high-quality reconstruction under sparse view conditions remains a key problem to be solved urgently. Summary of the Invention
[0012] In order to overcome the deficiencies of the prior art, the object of the present invention is to provide a NeRF sparse view reconstruction method based on adaptive multi-modal feature fusion, which realizes the intelligent fusion of multi-modal features, overcomes the limitations of traditional fusion methods, and provides a solid foundation for high-quality neural radiance field reconstruction under sparse views. Compared with traditional fusion methods, the present invention significantly improves the expression ability and scene adaptability of the fused features while maintaining computational efficiency, especially showing obvious advantages in dealing with complex problems such as modal inconsistency and sensor noise in real-world scenarios.
[0013] To achieve the above object, the present invention provides the following solutions:
[0014] An NeRF sparse view reconstruction method based on adaptive multimodal feature fusion, comprising:
[0015] Extract depth features and semantic features from the target sparse input views, and perform feature enhancement on the depth features and the semantic features to obtain first multimodal features;
[0016] Based on a hybrid attention mechanism, combine channel attention and spatial attention to perform dual selective enhancement on the first multimodal features to obtain second multimodal features;
[0017] Based on a confidence-guided weight assignment network, generate normalized weights according to the modal confidence map, and use the normalized weights to perform feature weighted fusion on each of the second multimodal features to obtain preliminary fusion features;
[0018] Harmonize the feature conflicts between different modalities in the preliminary fusion features through a consistency loss function to obtain final fusion features;
[0019] Input the final fusion features into an enhanced neural radiance field rendering module for hierarchical feature encoding and progressive training to generate a three-dimensional scene reconstruction result.
[0020] Preferably, extracting depth features and semantic features from the target sparse input views, and performing feature enhancement on the depth features and the semantic features to obtain first multimodal features, includes:
[0021] Based on the target sparse input views, use a DPT model based on the Transformer architecture to extract a multi-scale geometric feature pyramid and output a feature map with the original resolution;
[0022] Perform edge-preserving smoothing on the feature map with the original resolution, and map the depth values to the [0, 1] interval through normalization to obtain a normalized feature map;
[0023] Based on the normalized feature map, generate the depth features through Monte Carlo random inactivation technology;
[0024] Based on the SegFormer model, generate a semantic segmentation map and intermediate layer features according to the depth features to obtain the semantic features;
[0025] Introduce cross-view semantic consistency constraints in the semantic features, and ensure the consistency of the semantic labels of the same object under different perspectives through feature projection and matching;
[0026] Combine Canny edge detection and morphological operations to extract the object boundary information of the semantic features and generate a boundary weight map;
[0027] Implement multi-level feature fusion of the semantic features and the depth features through an encoder-decoder architecture; the encoder includes 5 layers of convolution and max pooling; the decoder uses transposed convolution for upsampling.
[0028] Establish dense connections between the encoder and the decoder to form a feature pyramid network structure, and output the first multi-modal feature.
[0029] Apply the combined normalization technique to reduce the distribution difference between the depth features and the semantic features in the first multi-modal feature.
[0030] Preferably, the semantic features include: texture and edge information at 1 / 4 resolution and object-level semantic information at 1 / 16 resolution.
[0031] Preferably, the resolutions of the first multi-modal feature include: 1 / 4, 1 / 8, 1 / 16, and 1 / 32.
[0032] Preferably, the channel attention calculates the channel importance weights through global pooling and a multi-layer perceptron; the expression of the channel attention M_c(F) is: M_c(F) = σ(MLP(AvgPool(F)) + MLP(MaxPool(F))); where F ∈ R C×H×W is the input feature, AvgPool and MaxPool are the global average pooling and max pooling operations respectively, MLP is a two-layer perceptron, and σ is the Sigmoid activation function; the channel attention map contains the importance weights of each channel.
[0033] Preferably, the spatial attention generates a spatial attention map through spatial pooling and a 7×7 convolution; the spatial attention M s (F) has the expression: M s (F) = σ(Conv 7×7 ([AvgPool(F) spatial ; MaxPool(F) spatial )); where and are the average pooling and max pooling results along the channel dimension respectively, Conv 7×7 is a 2D convolution with a kernel size of 7×7, generating the spatial attention map
[0034] Preferably, the double selective enhancement is achieved through cascaded attention operations.
[0035] Preferably, the expression of the consistency loss function L cons is:
[0036]
[0037] Among them, Cons ij (x, y) is the consistency measure between the feature F of modality i i and the feature F of modality j j , ∈ is a constant to prevent division by zero, λ ij is a weight parameter for balancing the consistency between different modality pairs, and H and W are the height and width of the feature map respectively.
[0038] Preferably, the final fused feature is input into an enhanced neural radiance field rendering module for hierarchical feature encoding and progressive training to generate a three-dimensional scene reconstruction result, including:[[]]
[0039] Decompose the final fused feature into four levels: detailed texture, local structure, object parts, and global layout to obtain multiple level features;
[0040] Dynamically adjust the contribution of each level feature through a spatial adaptive gating mechanism;
[0041] Adopt a geometry-appearance decoupled optimization mode, inject multi-modal information in stages based on each level feature, and achieve coarse-to-fine multi-granularity scene modeling through progressive expansion of the network architecture to obtain the three-dimensional scene reconstruction result.
[0042] According to the specific embodiments provided by the present invention, the following technical effects are disclosed by the present invention:
[0043] The present invention provides a NeRF sparse view reconstruction method based on adaptive multi-modal feature fusion, including: extracting depth features and semantic features from the target sparse input view, and performing feature enhancement on the depth features and semantic features to obtain the first multi-modal features; based on a hybrid attention mechanism, performing dual selective enhancement on the first multi-modal features by combining channel attention and spatial attention to obtain the second multi-modal features; based on a confidence-guided weight assignment network, generating normalized weights according to the modality confidence map, and using the normalized weights to perform feature weighted fusion on each second multi-modal feature to obtain a preliminary fused feature; reconciling the feature conflicts between different modalities in the preliminary fused feature through a consistency loss function to obtain a final fused feature; inputting the final fused feature into an enhanced neural radiance field rendering module for hierarchical feature encoding and progressive training to generate a three-dimensional scene reconstruction result. The present invention realizes the intelligent fusion of multi-modal features, overcomes the limitations of traditional fusion methods, and provides a solid foundation for high-quality neural radiance field reconstruction under sparse views. Compared with traditional fusion methods, the present invention significantly improves the expression ability and scene adaptability of the fused features while maintaining computational efficiency, especially showing obvious advantages in dealing with complex problems such as modality inconsistency and sensor noise in real-world scenes. Brief Description of the Drawings
[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0045] Figure 1 It is a flowchart of the method provided by the embodiment of the present invention;
[0046] Figure 2 It is an overall architecture diagram provided by the embodiment of the present invention;
[0047] Figure 3 It is a flowchart of the multi-modal feature extraction module provided by the embodiment of the present invention;
[0048] Figure 4 It is a flowchart of the adaptive feature fusion module provided by the embodiment of the present invention. Detailed Description of the Embodiments
[0049] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0050] The purpose of the present invention is to provide a NeRF sparse view reconstruction method based on adaptive multi-modal feature fusion, which realizes the intelligent fusion of multi-modal features, overcomes the limitations of traditional fusion methods, and provides a solid foundation for high-quality neural radiance field reconstruction under sparse views. Compared with traditional fusion methods, the present invention significantly improves the expression ability and scene adaptability of the fused features while maintaining computational efficiency, especially showing obvious advantages in dealing with complex problems such as modal inconsistency and sensor noise in real-world scenes.
[0051] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the drawings and specific embodiments.
[0052] Figure 1 It is a flowchart of the method provided by the embodiment of the present invention. As Figure 1 shown, the present invention provides a NeRF sparse view reconstruction method based on adaptive multi-modal feature fusion, including:
[0053] Step 100: Extract depth features and semantic features from the target sparse input view, and perform feature enhancement on the depth features and semantic features to obtain the first multimodal features;
[0054] Step 200: Based on the hybrid attention mechanism, combine channel attention and spatial attention to perform double selective enhancement on the first multimodal features to obtain the second multimodal features;
[0055] Step 300: Based on the confidence-guided weight assignment network, generate normalized weights according to the modal confidence map, and use the normalized weights to perform feature weighted fusion on each second multimodal feature to obtain the preliminary fusion features;
[0056] Step 400: Harmonize the feature conflicts between different modalities in the preliminary fusion features through a consistency loss function to obtain the final fusion features;
[0057] Step 500: Input the final fusion features into the enhanced neural radiance field rendering module for hierarchical feature encoding and progressive training to generate a 3D scene reconstruction result.
[0058] The NeRF sparse view reconstruction system based on adaptive multimodal feature fusion proposed by the present invention adopts a modular design concept and is mainly composed of three core modules: a multimodal feature extraction module, an adaptive feature fusion module, and an enhanced neural radiance field rendering module. This modular architecture not only improves the flexibility and scalability of the system but also facilitates customized optimization for specific application scenarios. The overall architecture is as Figure 2 shown.
[0059] Specifically, as Figure 3 shown, the multimodal feature extraction module of the present invention is the front-end processing unit of the system, responsible for extracting rich and complementary prior information from the sparse input view. This module adopts a parallel processing architecture and captures the geometric structure, semantic content, and texture details of the scene from different dimensions simultaneously, laying a solid foundation for subsequent feature fusion and scene reconstruction. The specific implementation includes three modules: a depth feature extraction module, a semantic feature extraction module, and a feature enhancement network.
[0060] Specifically, the depth feature extraction in this embodiment uses a pre-trained monocular depth estimation network to extract depth information and its related features from the RGB image. The specific implementation is as follows:
[0061] First, the DPT (Dense Prediction Transformer) based on the transformer architecture is used as the basic depth estimation network. However, its depth output is not directly used. By modifying the network, a feature output interface is added in the middle layer, enabling it to output multiple intermediate layer feature maps. Then, 4 feature maps with different resolutions (1 / 4, 1 / 8, 1 / 16, 1 / 32 of the original resolution) are extracted from the encoder part of the network, corresponding to the local geometric details in the shallow layer and the global structural information in the deep layer respectively. These features are retained through skip connections to form a multi-scale geometric feature pyramid. Finally, the bilateral filter is applied to the original depth map for edge-preserving smoothing, reducing noise while retaining object boundaries. Then, the depth values are mapped to the [0, 1] interval through normalization to improve training stability.
[0062] In addition, a depth uncertainty estimation branch is additionally introduced in this stage. Through the Monte Carlo dropout technique, 20 forward propagations are performed during the inference stage to calculate the variance of the depth prediction for each pixel point, generating an uncertainty map. This uncertainty map will be used for weight adjustment in the adaptive fusion process, enabling the system to reduce its dependence on unreliable depth predictions.
[0063] Furthermore, the semantic feature extraction in this embodiment integrates a high-performance semantic segmentation model to extract the semantic information of the scene, providing high-level cognitive guidance for scene understanding and reconstruction:
[0064] First, the SegFormer model is used as the semantic segmentation network. The model is pre-trained on the ADE20K and COCO-Stuff datasets and has strong generalization ability. For the extraction results, not only the final semantic segmentation map (class probability map) is extracted, but also the feature maps in the intermediate layer of the segmentation network are retained, including the low-level features (containing texture and edge information) at 1 / 4 resolution and the high-level features (containing object-level semantic information) at 1 / 16 resolution.
[0065] Then, cross-view semantic consistency constraints are introduced. By feature projection and matching, the semantic labels of the same object in different views are ensured to be consistent. Specifically, using the known camera parameters, the semantic features are projected from one view to another, the matching degree is calculated, and it is used for weight calculation in the subsequent fusion stage. Finally, the object boundary information is extracted from the semantic segmentation results by combining canny edge extraction and morphological operations to generate a boundary weight map, which is used to guide the precise modeling of object boundaries during rendering, effectively reducing the texture aliasing phenomenon between different objects.
[0066] Furthermore, in this embodiment, a dedicated feature enhancement network is designed to optimize and enhance the initially extracted multi-modal features: The network architecture adopts a U-Net variant structure, including 5 layers of encoders and 5 layers of decoders. Each layer of the encoder consists of two 3×3 convolutional layers, BatchNorm, and the LeakyReLU activation function, and downsampling is achieved through max pooling with a stride of 2; the decoder uses transposed convolution for upsampling.
[0067] Then, this network is used for multi-scale feature fusion: Dense connections are established between each level of the encoder and decoder to form a Feature Pyramid Network (FPN) structure, realizing complementary fusion of features with different resolutions. Finally, feature maps with 4 resolutions (1 / 4, 1 / 8, 1 / 16, 1 / 32 of the original resolution) are output, and the number of channels for each feature map is 64.
[0068] Finally, feature normalization technology is applied to ensure that the numerical ranges of different modal features are similar. The combination of InstanceNorm and LayerNorm is adopted to effectively reduce the distribution differences within batches and between channels. At this stage, depth features and semantic features are fused to generate a set of enhanced multi-modal feature representations, providing richer prior information for subsequent neural radiance field reconstruction. Specifically, depth features provide geometric structure information of the scene, while semantic features provide object-level recognition and classification information. The combination of the two forms a comprehensive understanding of the scene.
[0069] Through the collaborative work of the above three sub-modules, the multi-modal feature extraction module of the present invention can mine rich complementary information from sparse input views, providing comprehensive prior knowledge for subsequent adaptive feature fusion and high-quality neural radiance field reconstruction, and effectively overcoming the key problems such as geometric uncertainty and texture blur faced by traditional NeRF methods under sparse view conditions.
[0070] As Figure 4 shown, the core innovation of the present invention lies in designing an efficient adaptive feature fusion module, which can intelligently fuse multi-modal information and dynamically adjust the contribution weights of each modality according to the characteristics of different scene regions. Compared with traditional fixed-weight or simple feature concatenation methods, this module significantly improves the fusion efficiency and adaptability, providing key support for high-quality neural scene reconstruction under sparse views.
[0071] The present invention designs a dual hybrid attention network to achieve fine-grained dynamic fusion at the feature level by combining spatial attention and channel attention:
[0072] First is the channel attention module: responsible for learning which types of features are more important, including color, depth, and semantics, and its mathematical expression is:
[0073] M_c(F) = σ(MLP(AvgPool(F)) + MLP(MaxPool(F)))
[0074] where F ∈ R C×H×W is the input feature, AvgPool and MaxPool are global average pooling and max pooling operations respectively, MLP is a two-layer perceptron, and σ is the Sigmoid activation function. The channel attention map contains the importance weights for each channel.
[0075] Then comes the spatial attention module, which focuses on where the features in the space should be enhanced, and is calculated as follows:
[0076] M s (F) = σ(Conv 7×7 ([AvgPool(F) spatial ; MaxPool(F) spatial ))
[0077] where and are the average pooling and max pooling results along the channel dimension respectively, and Conv 7×7 is a 2D convolution with a kernel size of 7×7, generating the spatial attention map
[0078] Finally, there is the attention fusion strategy: cascading the channel attention and spatial attention and applying them to the original features:
[0079]
[0080] where represents element-wise multiplication. This cascading design first emphasizes important channels and then highlights key spatial positions, achieving dual selective attention.
[0081] In addition, to enhance the feature expression ability, the present invention also extends the single attention mechanism to a multi-head design, using 8 parallel attention heads. Each head independently processes a subset of the input features, and then combines the results:
[0082] F multi = Concat(Head1(F), Head2(F), …, Head8(F))W O
[0083] where W O is the projection matrix used to map the cascaded features back to the original dimensional space.
[0084] Specifically, this embodiment also realizes confidence-guided adaptive weight assignment. This part uses a Confidence-Guided Weight Assignment Network (CGWAN) to calculate the dynamic weights of each modality feature:
[0085] First, the modality confidence estimates the confidence map for each modality feature F i (RGB, depth, and semantic features), representing the reliability of the modality at each spatial position: C
[0086] = σ(Conv i (Conv 1×1 (F 3×3 ))) i
[0087] Then, based on the confidence map, normalized weights are generated to ensure that the sum of all modality weights is 1:
[0088]
[0089] where N is the number of modalities, is the weight map of the i-th modality.
[0090] Finally, the weights are used for feature weighted fusion, and the calculated weights are used to perform weighted summation on each modality feature:
[0091]
[0092] Specifically, this embodiment uses a modality consistency constraint mechanism to handle possible inconsistencies between modalities. First, the consistency measure between modality i feature F i and modality j feature F j is defined as:
[0093]
[0094] where ∈ is a small constant to prevent division by zero. The consistency measure Cons ij ∈ [0, 1], and the higher the value, the more consistent the two modalities are.
[0095] Then, during the training process, a modality consistency loss is introduced to promote consistent representations between modalities. The modality consistency loss is expressed as:
[0096]
[0097] where λ ij is a weight parameter that balances the consistency between different modality pairs. This loss function encourages the network to learn to generate mutually consistent feature representations and improve the fusion quality.
[0098] The adaptive feature fusion module of the present invention realizes the intelligent fusion of multimodal features through the organic combination of the above four innovative mechanisms, overcomes the limitations of traditional fusion methods, and provides a solid foundation for high-quality neural radiation field reconstruction under sparse views. Compared with traditional fusion methods, this module significantly improves the expressiveness and scene adaptability of fusion features while maintaining computational efficiency, especially when dealing with complex problems such as modal inconsistency and sensor noise in real-world scenes.
[0099] Furthermore, the rendering module of the present invention is also optimized for sparse scenes based on the standard NeRF framework. The main purpose is to render gradually from coarse to fine, from structure to texture details in stages, including hierarchical feature encoding and progressive training:
[0100] (1) Hierarchical feature encoding enhances geometric reasoning capabilities through multi-scale scene representation. The scene features are decomposed into four levels: detail texture, local structure, object parts, and global layout. Multi-view projection is used to extract cross-view aligned feature information from the input image. Each 3D spatial point is projected to each input view, and feature fusion is performed by combining attention weights based on view direction and geometric confidence. The feature representations of different levels are then processed independently through parallel network branches. A spatially adaptive gating mechanism is used between levels to dynamically adjust feature contributions, so that high-frequency details and macro structures are optimized in a coordinated manner, effectively alleviating the problem of multi-scale feature conflicts and significantly improving edge clarity and geometric consistency in weak texture areas.
[0101] (2) The progressive training strategy improves training stability through staged feature learning and target optimization. In the early stage of training, only basic color features are used to construct the geometric prototype of the scene. As the iteration progresses, auxiliary information such as multi-view depth clues and semantic context are gradually injected. The strength of multimodal feature fusion is adjusted through a controllable gating mechanism to avoid noise interference in early training. The loss function design adopts a geometric-appearance decoupling optimization mode. In the early stage, geometric constraints such as surface smoothness and depth consistency are dominant. As the training progresses, it gradually transitions to an appearance optimization goal with texture detail reconstruction as the core. The network architecture adopts a progressive expansion mode, starting from a shallow basic network. As the feature complexity increases, the network depth is gradually increased and residual connections are introduced. Training stability is ensured through parameter inheritance and progressive unfreezing, realizing multi-granular scene modeling from coarse to fine.
[0102] The key innovative points of this embodiment are as follows:
[0103] (1) Adaptive Multimodal Feature Fusion Architecture: Breaking the dependence of traditional NeRF on single RGB information, constructing a complete technical route of "multimodal feature extraction → adaptive feature fusion → enhanced neural radiance field rendering". By introducing prior knowledge such as depth and semantics, geometric uncertainty and texture ambiguity problems under sparse view conditions are solved.
[0104] (2) Dynamic Weight Allocation Mechanism Based on Hybrid Attention: Innovatively combining spatial attention and channel attention to achieve fine-grained weight allocation at the feature level. Different from traditional fixed weight or simple feature splicing methods, this mechanism can automatically learn the confidence of different regions and different modalities, and dynamically adjust the contribution ratio of each modality accordingly, effectively dealing with the problems of information inconsistency between modalities and noise interference within modalities.
[0105] (3) Modal Consistency Constraint Mechanism: Introducing a dedicated modal consistency loss function to explicitly model and constrain the possible information conflicts between different modalities. This mechanism can detect and reconcile the inconsistencies between different modality data such as depth, semantics, and RGB, significantly improving the robustness of the system in the face of practical problems such as calibration errors and sensor noise, and ensuring the visual coherence of the final reconstruction results.
[0106] (4) Enhanced Neural Radiance Field Rendering Module: Innovatively designing a hierarchical feature encoding and progressive training strategy to break through the limitations of the single representation method of traditional NeRF. Hierarchical feature encoding decomposes the scene into four levels: detailed texture, local structure, object parts, and global layout. Through a spatial adaptive gating mechanism, the feature contributions of each level are dynamically adjusted to coordinate and solve multi-scale feature conflicts; the progressive training strategy realizes stage-by-stage optimization from geometry to appearance, adopting a network architecture progressive expansion and parameter inheritance mechanism to ensure training stability and reconstruction quality.
[0107] The present invention mainly integrates three modality information of RGB, depth, and semantics at present, but this architecture can be flexibly extended to more modalities. For example, normal estimation can be introduced as a supplement to geometric prior to provide surface orientation information; integrating the reflectance and illumination information estimated by the light decoupling network to improve the material and shadow reconstruction quality; or combining optical flow information to enhance the establishment of correspondence between multi-views. For each additional modality, only the corresponding feature extraction network needs to be designed and attention weights are allocated in the fusion module.
[0108] The current design adopts spatial and channel hybrid attention, and other attention variants can be explored. For example, introducing a temporal attention mechanism to process the video input sequence; adopting a Transformer architecture to establish feature dependencies globally; or designing a Graph Attention Network (GAN) to model the relationships between different modality nodes. These alternatives may provide more effective feature fusion capabilities in specific scenarios, but the computational complexity and performance improvement need to be balanced.
[0109] The neural radiance field rendering module in this embodiment can be replaced by other volume rendering methods. For example, adopting a 3D Gaussian-based representation method can significantly improve the rendering speed; using a voxel grid + hash encoding strategy can greatly reduce the memory footprint; or introducing an implicit surface reconstruction method (such as neural SDF) can provide a more accurate geometric representation. These alternatives enable the system to achieve different balances among quality, speed, and memory footprint according to application requirements.
[0110] In the current design of this embodiment, the feature extraction network uses a pre-trained model, and an end-to-end joint optimization strategy can be explored. By designing a specific gradient backpropagation mechanism, allowing the rendering loss gradient to flow to the feature extraction network, the joint fine-tuning of the entire pipeline is achieved. This solution may improve the overall consistency and performance of the system, but requires more computing resources and a carefully designed training strategy to avoid overfitting.
[0111] Customized optimization is carried out for specific application fields. For example, in the field of medical imaging, medical image features such as CT / MRI can be integrated to improve the accuracy of organ reconstruction; in the digital protection of cultural relics, the prior knowledge of material characteristics and historical information can be combined to enhance the restoration of surface details; in the industrial inspection scenario, thermal imaging or ultrasonic data can be fused to improve the perception ability of defects. These domain-specific customization solutions require the targeted design of feature extraction networks and fusion strategies.
[0112] For large-scale scene reconstruction, a distributed parallel processing architecture is designed. The scene is divided into multiple sub-regions, which are processed in parallel by different computing nodes; a feature sharing and boundary consistency constraint mechanism is designed to ensure smooth transitions between sub-regions; a progressive merging strategy is adopted to gradually integrate the local reconstruction results. This architectural variation can significantly expand the system's ability to process ultra-large scenes, but the data transmission bottleneck and consistency maintenance issues need to be addressed.
[0113] The beneficial effects of the present invention are as follows:
[0114] (1) High-quality reconstruction under sparse view conditions: The system can generate high-quality three-dimensional scene reconstruction results with only a small number of input views, effectively solving the problem of the traditional NeRF method's dependence on dense views.
[0115] (2) Strong generalization ability: Thanks to the integration of multi-modal prior knowledge and the adaptive feature extraction mechanism, this system demonstrates excellent generalization ability and can effectively handle various complex scenarios, including indoor and outdoor environments, people, objects, and scenarios with complex lighting and materials.
[0116] (3) Robustness to rough inputs: Based on the modal consistency constraint and the adaptive fusion mechanism, the system has strong robustness to noise, calibration errors, and modal inconsistencies in the input data, greatly improving its practicality in actual application scenarios.
[0117] (4) Scalability and flexibility: The modular design enables the system to have good scalability and can easily integrate new modal information or feature extraction networks to adapt to the ever-evolving technological requirements and application scenarios.
[0118] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other.
[0119] Specific examples are used in this article to elaborate on the principles and implementation methods of the present invention. The descriptions of the above embodiments are only used to help understand the method and its core idea of the present invention. At the same time, for those of ordinary skill in the art, based on the idea of the present invention, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A NeRF sparse view reconstruction method based on adaptive multi-modal feature fusion, characterized in that, Comprising: Extract depth features and semantic features from the target sparse input view, and perform feature enhancement on the depth features and the semantic features to obtain first multi-modal features; Based on a hybrid attention mechanism, perform dual selective enhancement on the first multi-modal features by combining channel attention and spatial attention to obtain second multi-modal features; Based on a confidence-guided weight assignment network, generate normalized weights according to the modal confidence map, and use the normalized weights to perform feature weighted fusion on each of the second multi-modal features to obtain preliminary fusion features; Harmonize the feature conflicts between different modalities in the preliminary fusion features through a consistency loss function to obtain final fusion features; Input the final fusion features into an enhanced neural radiance field rendering module for hierarchical feature encoding and progressive training to generate a 3D scene reconstruction result.
2. The NeRF sparse view reconstruction method based on adaptive multi-modal feature fusion according to claim 1, wherein Extract depth features and semantic features from the target sparse input view, and perform feature enhancement on the depth features and the semantic features to obtain first multi-modal features, including: Based on the target sparse input view, use a DPT model based on the Transformer architecture to extract a multi-scale geometric feature pyramid and output a feature map with the original resolution; Perform edge-preserving smoothing on the feature map with the original resolution, and map the depth values to the interval [0, 1] through normalization to obtain a normalized feature map; Based on the normalized feature map, generate the depth features through Monte Carlo random inactivation technology; Based on the SegFormer model, generate a semantic segmentation map and intermediate layer features according to the depth features to obtain the semantic features; Introduce cross-view semantic consistency constraints in the semantic features, and ensure the consistency of semantic labels of the same object under different perspectives through feature projection and matching; Combine Canny edge detection and morphological operations to extract the object boundary information of the semantic features and generate a boundary weight map; Implement multi-level feature fusion of the semantic features and the depth features through an encoder-decoder architecture; the encoder includes 5 layers of convolution and max pooling; the decoder uses transposed convolution upsampling; Establish dense connections between the encoder and the decoder to form a feature pyramid network structure and output first multi-modal features; Apply a combined normalization technique to reduce the distribution difference between the depth features and the semantic features in the first multi-modal features.
3. The NeRF sparse view reconstruction method based on adaptive multi-modal feature fusion according to claim 1, characterized in that The semantic features include: texture and edge information with a resolution of 1 / 4 and object-level semantic information with a resolution of 1 / 16.
4. The NeRF sparse view reconstruction method based on adaptive multi-modal feature fusion according to claim 1, wherein The resolutions of the first multi-modal features include: 1 / 4, 1 / 8, 1 / 16, and 1 / 32.
5. The NeRF sparse view reconstruction method based on adaptive multi-modal feature fusion according to claim 1, wherein The channel attention calculates the channel importance weights through global pooling and a multi-layer perceptron; the expression of the channel attention M_c(F) is: M_c(F) = σ(MLP(AvgPool(F)) + MLP(MaxPool(F))); where, F ∈ R C×H×W is the input feature, AvgPool and MaxPool are the global average pooling and max pooling operations respectively, MLP is a two-layer perceptron, and σ is the Sigmoid activation function; the channel attention map contains the importance weights of each channel.
6. The NeRF sparse view reconstruction method based on adaptive multi-modal feature fusion according to claim 1, wherein The spatial attention generates a spatial attention map through spatial pooling and 7×7 convolution; the spatial attention M s (F) is expressed as: M s (F) = σ(Conv 7×7 ([AvgPool(F) spatial ; MaxPool(F) spatial )); where and are the average pooling and max pooling results along the channel dimension respectively, Conv 7×7 is a 2D convolution with a kernel size of 7×7, generating the spatial attention map 7. The NeRF sparse view reconstruction method based on adaptive multi-modal feature fusion according to claim 1, wherein, The dual selective enhancement is achieved through cascaded attention operations.
8. The NeRF sparse view reconstruction method based on adaptive multi-modal feature fusion according to claim 1, characterized in that, The consistency loss function L cons has the following expression: Among them, Cons ij (x,y) is the feature F of mode i i and mode j feature F j The consistency measure between ∈ is a constant to prevent division by zero, λ ij is a weight parameter that balances the consistency between different modal pairs, h and W are the height and width of the feature map, respectively, and x and y are temporary variables used to traverse the feature map.
9. The NeRF sparse view reconstruction method based on adaptive multi-modal feature fusion according to claim 1, characterized in that Input the final fusion features into an enhanced neural radiance field rendering module for hierarchical feature encoding and progressive training to generate a 3D scene reconstruction result, including: Decompose the final fusion features into four levels: detailed texture, local structure, object components, and global layout to obtain multiple level features; Dynamically adjust the contributions of each of the level features through a spatial adaptive gating mechanism; Adopt a geometric-appearance decoupled optimization mode, inject multi-modal information in stages based on each of the hierarchical features, and achieve coarse-to-fine multi-granularity scene modeling through progressive expansion of the network architecture to obtain the three-dimensional scene reconstruction result.
Citation Information
Cited By
Target behavior recognition method and device and medium
CN120853266A
A target behavior recognition method, device and medium
CN120853266B
Multi-modal data processing method and device, equipment and medium
CN121211297A
Generative sparse potential color field method for three-dimensional native texture generation
CN121685808A