System and method for synthesizing new view of sparse sampling scene based on multi-source image feature fusion
By employing multi-source image feature fusion and hybrid perception constraints, the problem of geometric and appearance detail degradation in new view synthesis in sparse sampling scenarios is solved, achieving high-fidelity new view synthesis results.
Patent Information
- Application Number
- CN202510010393.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2045-01-03
AI Technical Summary
In sparse sampling scenarios, existing new view synthesis techniques suffer from problems such as loss of geometric details, blurred textures, and unnatural view transitions. The accuracy and robustness of depth prior information are insufficient, and the lack of a unified optimization framework for depth-appearance fusion leads to geometric misalignment and texture distortion.
A multi-source image feature fusion method is adopted. High-dimensional feature information is extracted through a multi-source feature generation module, and feature aggregation and geometric cues are aggregated. Combined with feature mapping module, the image is projected to a new perspective of the target. Color appearance and density feature information are obtained through appearance-density estimation network. Finally, a high-fidelity image is generated using a synthetic rendering module, and hybrid perception constraints are introduced for iterative optimization.
A new high-fidelity view synthesis was achieved in sparse sampling scenarios, improving geometric consistency and appearance fidelity, solving the degradation problem of geometric and appearance details in sparse sampling scenarios, and significantly improving the quality of the generated images.
Smart Images

Figure CN119942282B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the technical field of novel view synthesis, and mainly relates to a novel view synthesis system and method for sparse sampling scenes based on multi-source image feature fusion. Background Technology
[0002] Novel view synthesis is an important research topic in computer vision and graphics, aiming to synthesize novel views from any observation perspective within a scene, using multiple viewpoints as input. This task has wide applications in virtual reality (VR) and augmented reality (AR), film and television production, autonomous driving, and robot vision. While neural rendering techniques, such as neural radiation fields, have achieved high-fidelity novel view synthesis in recent years, they require a large number of input views and stringent shooting standards, severely limiting their widespread adoption in related fields. In particular, in sparsely sampled scenes—where only a small number of sparsely distributed observation views exist—the synthesized novel views often suffer from geometric detail loss, texture blurring, and unnatural view transitions due to model degradation. Therefore, achieving high-fidelity novel view synthesis from sparsely sampled scenes remains a major challenge.
[0003] To address this challenge, existing research can be summarized into three main approaches: (1) Pre-trained models. These models enhance their generation capabilities using a large amount of scene data, even enabling the synthesis of new views with a single view input. However, these methods are highly dependent on large-scale scene data, have high implementation costs, and are prone to artifacts and geometric errors in unfamiliar scene types. (2) Introducing additional constraints. In the process of 3D scene representation or regularization, additional constraints such as appearance constraints and sampling range constraints are introduced to alleviate model degradation. Although this method performs well in simple scenes centered on a single object, its fidelity in synthesizing new views in complex indoor scenes is still insufficient due to the lack of complete geometric information integrity and coherence in sparsely sampled scenes. (3) Depth priors. Depth information can intuitively reflect the geometric details of a scene. Using depth priors as supervision or embedding them into the model input can effectively alleviate the training burden of the model and improve model performance. In practical applications, depth priors can be easily obtained through consumer-grade depth sensors or depth estimation models, which facilitates the promotion and application of this method.
[0004] Although depth-prior methods have shown great potential in the synthesis of new views in sparsely sampled scenes, they still have the following drawbacks: (1) Insufficient accuracy and robustness of depth information. Depth maps generated by consumer-grade depth sensors or depth estimation models often contain noise, missing data, and quantization errors, especially in highly reflective, transparent, or less textured areas. These defects may lead to inaccurate depth-prior information, which in turn affects the representation of scene geometry. (2) Insufficient depth-appearance fusion. In the process of fusing depth priors and appearance features, existing methods usually lack a unified optimization framework, resulting in insufficient coordination between geometric and appearance information. This separate processing approach is prone to causing geometric misalignment or texture distortion problems in the synthesis results. Summary of the Invention
[0005] This invention addresses the problems existing in the prior art by proposing a new scene view synthesis system and method based on multi-source image feature fusion using sparse sampling. The system extracts and generates high-dimensional feature information through a multi-source feature generation module. A feature aggregation module then aggregates this high-dimensional feature information with geometric cues to obtain aggregated feature information. A feature mapping module projects and transforms the aggregated feature information from the observation viewpoint to the target new viewpoint, yielding aggregated feature information under the target new viewpoint. This aggregated feature information under the target new viewpoint is then fed into an appearance-density estimation network to obtain color appearance and density feature information under the target new viewpoint. The density feature information is then fed into an occupancy estimation network to obtain density information under the target new viewpoint. Finally, the density and color appearance information are fed into a synthesis rendering module to obtain the reconstructed visible light and depth images under the target new viewpoint. This invention utilizes the visible light and depth images obtained from actual observations of the target new viewpoint as prior information and establishes hybrid perception constraints as prior supervision. The parameters of the multi-source feature generation module, appearance-density estimation network, and occupancy estimation network are iteratively optimized to ultimately achieve high-fidelity new scene view synthesis and support the generation of visible light images from any viewpoint within the scene. This invention enables the synthesis of high-fidelity new view images from sparsely sampled scenes.
[0006] To achieve the above objectives, the technical solution adopted by the present invention is: a new view synthesis system for sparse sampling scenes based on multi-source image feature fusion, which includes at least a multi-source feature generation module, a feature aggregation module, a feature mapping module, an appearance-density estimation network, an occupancy estimation network, and a synthesis rendering module;
[0007] The multi-source feature generation module adopts a dual-branch architecture, including a visible light-dominated branch, a depth-dominated branch, and a feature fusion network, which is used to generate high-dimensional feature information that integrates scene appearance texture and geometric priors.
[0008] The feature aggregation module uses tensor splicing to aggregate the obtained high-dimensional feature information with geometric clues to obtain aggregated feature information.
[0009] The feature mapping module adopts a homography transformation mapping method, uses camera parameters to calculate the homography transformation matrix between the current observation view and the new target view, projects and transforms the aggregated feature information of the observation view to the new target view, and obtains the aggregated feature information under the new target view.
[0010] The appearance-density estimation network is a multilayer perceptron containing multiple fully connected layers, including at least an appearance feature estimation module, an appearance calculation module, and a density feature estimation module, to obtain color appearance information and density feature information of the target from a new perspective.
[0011] The occupancy estimation network adopts a symmetric encoder-decoder architecture, consisting of a 3D convolutional layer with skip connections and a transposed convolutional layer; density feature information and original occupancy information are input to the occupancy estimation network, and density information of the target under a new perspective is output.
[0012] The synthetic rendering module takes the color appearance information and density information of the target under the new perspective as input, and outputs the visible light image and depth image synthesized under the new perspective of the target.
[0013] As an improvement of the present invention, in the visible light dominant branch of the multi-source feature generation module, the acquired visible light image and depth image are used as input, the visible light image is used to correct the depth image, and the output is the corrected depth image.
[0014] In the depth-dominant branch, the original input depth image and the corrected depth image are used as inputs, and the output is a multi-dimensional feature image.
[0015] In the fusion network, the visible light image after position encoding and the multi-dimensional feature image are used as inputs, and the output is the finally extracted high-dimensional feature information.
[0016] As an improvement of the present invention, the visible light dominant branch and the depth dominant branch of the multi-source feature generation module are both composed of an encoder-decoder network with symmetrical skip connections, containing multiple cascaded nonlinear activation residual blocks; wherein, the position encoding method of the visible light image is Fourier encoding; the fusion network is composed of a batch normalization layer and a ReLU activation function.
[0017] As another improvement of the present invention, the homography transformation matrix in the feature mapping module is specifically as follows:
[0018]
[0019] In the formula, H i→tLet D be the homography transformation matrix from observation viewpoint i to target viewpoint t. t Let I be the sampling depth, I be the identity matrix, and K, R, and T be the camera parameters. The principal axis unit vector of the target view.
[0020] To achieve the above objectives, the present invention also adopts the following technical solution: a method for synthesizing new views of sparsely sampled scenes based on multi-source image feature fusion, comprising the following steps:
[0021] S1. Data preprocessing: Collect visible light images and depth images from different observation angles in the same scene, preprocess the collected image data to ensure that each visible light image from the observation angle has a corresponding depth image and camera parameter data, forming a scene dataset.
[0022] S2, High-dimensional feature extraction: Initialize the multi-source feature generation module, input the visible light image and depth image from the observation perspective in the scene dataset into the multi-source feature generation module, extract and generate high-dimensional feature information;
[0023] S3. Aggregated Feature Acquisition: Using the corrected depth image as geometric cues, the high-dimensional feature information extracted in step S2 is aggregated with the geometric cues through the feature aggregation module to obtain aggregated feature information; through the feature mapping module, the aggregated feature information is projected and transformed to the new target perspective to obtain aggregated feature information under the new target perspective.
[0024] S4. Obtaining color appearance information and density feature information: Input the aggregated feature information of the target under the new perspective obtained in step S3 into the appearance-density estimation network to obtain the color appearance information and density feature information of the target under the new perspective.
[0025] S5. Density information and visible light image acquisition: Input the density feature information obtained in step S4 into the occupancy estimation network to obtain the density information under the new view of the target. Then, input the density information and the color appearance information obtained in step S4 into the synthesis rendering module to obtain the reconstructed visible light image under the new view of the target.
[0026] S6. Iterative optimization: Using the visible light image of the target under the new perspective obtained in step S5 as prior information, a hybrid perception constraint is established as prior supervision. The parameters of the multi-source feature generation module, appearance-density estimation network and occupancy estimation network are iteratively optimized to achieve the synthesis of the new scene view.
[0027] As an improvement of the present invention, step S3 projects and transforms the aggregated feature information of the observation viewpoint to the new target viewpoint to obtain the aggregated feature information under the new target viewpoint, specifically as follows:
[0028] F f,i→t (p) = concat[Ff,i (H i→t (D t )p T (X,Y,Z)]
[0029] In the formula, F f,i→t This represents the aggregated feature information transformed from the observation viewpoint i to the target viewpoint t, where p is the position coordinate of the feature image, and F... f,i H represents the feature information of observation viewpoint i. i→t Let D be the homography transformation matrix from observation viewpoint i to target viewpoint t. t X represents the sampling depth, and X, Y, and Z represent the constructed geometric cues.
[0030] As another improvement of the present invention, step S4 specifically includes the following steps:
[0031] S41: Input the aggregated feature information from multiple observation perspectives to the new target perspective into the appearance-density estimation network to obtain visible light features and weight vectors:
[0032]
[0033] In the formula, w i These are the visible light features and weight vectors, respectively. G0 is the appearance feature estimation module in the appearance-density estimation network, K is the number of observation views, and F... f,i→t p represents the aggregated feature information transformed from the observation viewpoint i to the target viewpoint t, where p is the position coordinate of the feature image.
[0034] S42: Obtain color appearance information based on the light features and weight vector, specifically:
[0035]
[0036] In the formula, C represents the synthesized color appearance information, and G0 represents the appearance calculation module. w i These are the visible light features and weight vectors, respectively; K is the number of observation viewpoints; and p is the position coordinate of the feature image.
[0037] S43: Obtain density feature information; the specific calculation process is as follows:
[0038]
[0039] In the formula, f d For the estimated density feature information, G d For appearance calculation module, w i These are the visible light features and weight vectors, respectively, where K is the number of observation viewpoints and p is the position coordinate of the feature image.
[0040] As another improvement of the present invention, the calculation method for obtaining density information through the occupancy estimation network in step S5 is as follows:
[0041] σ(p)=G occ (ψ(p),f d (p))
[0042] In the formula, σ represents the estimated density information, and G... occ To estimate network occupancy, f d For the estimated density feature information, p is the position coordinate of the feature image, and ψ is the original occupancy information;
[0043] The specific calculation process for obtaining the reconstructed visible light image of the target from a new perspective in the compositing rendering module is as follows:
[0044]
[0045] In the formula, and These represent the visible light image and depth image reconstructed from the new perspective of the target, respectively. render is the composite rendering module using the volumetric rendering method, C is the composite color appearance information, σ is the estimated density information, and p is the image position coordinates.
[0046] As a further improvement of the present invention, the hybrid sensing constraints in step S6 include visible light error constraints, local depth constraints, and global depth constraints, wherein...
[0047] The visible light error constraint is:
[0048]
[0049] In the formula, L c For visible light error constraints, w and h are the image sizes in the scene dataset, and I gt (p) is a visible light image actually observed from the new perspective of the target. The visible light image reconstructed from the new perspective of the target is shown, where p is the image position coordinate.
[0050] The local depth constraint is:
[0051]
[0052] In the formula, For local depth constraints, the patch is an N×N pixel block. For the reconstructed depth image, D is the actual observed depth image, and ε is a redundancy constant considering minute differences;
[0053] The global depth constraint is:
[0054]
[0055] In the formula, For global depth constraints, Gauss(D) represents the discrete distribution of the reconstructed depth image. i R represents the Gaussian mixture model of the actual observed depth image, where R is the number of depth value discretization intervals.
[0056] Compared with the prior art, the present invention has the following beneficial effects:
[0057] (1) This invention innovatively designs a multi-source feature generation module with a dual-branch architecture, including a visible light-dominated branch and a depth-dominated branch, and generates high-dimensional feature information that integrates scene appearance texture and geometric priors through a feature fusion network. The visible light-dominated branch improves the accuracy of geometric information by correcting the depth image, while the depth-dominated branch enhances the geometric representation capability by combining multi-dimensional feature images. At the same time, Fourier position encoding is introduced to alleviate the difficulty of the network in learning high-frequency features, thereby enhancing the ability to represent complex geometric details.
[0058] (2) This invention proposes a feature aggregation and homography transformation mapping method based on geometric cues. It utilizes depth information as geometric cues to aggregate high-dimensional features through tensor concatenation, and calculates the homography transformation matrix using camera parameters to map the feature information of the observation viewpoint to the target new viewpoint. This method fully utilizes geometric cues, significantly improving the geometric consistency and appearance fidelity of new viewpoint synthesis in sparsely sampled scenes.
[0059] (3) This invention proposes a hybrid perception constraint mechanism that includes visible light error constraints, local depth constraints, and global depth constraints. Prior supervision is used to jointly optimize the multi-source feature generation module, appearance-density estimation network, and occupancy estimation network. By using KL divergence constraints on local depth consistency and global depth distribution, the degradation of geometric and appearance details in sparse sampling scenarios is effectively solved, improving the geometric fidelity and detail quality of the synthesized results. Attached Figure Description
[0060] Figure 1 This is a schematic diagram of the structure of the sparse sampling scene new view synthesis system based on multi-source image feature fusion of the present invention;
[0061] Figure 2 These are example images of a single-object-centered scene and a complex indoor scene in Example 2 of this invention;
[0062] Figure 3 This is a schematic diagram of the structure of the multi-source feature generation module in the system of the present invention;
[0063] Figure 4 This is a comparison chart showing the results of a single-scene view synthesis model synthesizing a new view of an instance scene in the test examples of this invention;
[0064] Figure 5 This is a comparison chart of the results of the generalized view synthesis model synthesizing new views for instance scenes in the test examples of this invention;
[0065] Figure 6 This is a flowchart of the steps of the new view synthesis method for sparse sampling scenes based on multi-source image feature fusion according to the present invention. Detailed Implementation
[0066] The present invention will be further illustrated below with reference to the accompanying drawings and specific embodiments. It should be understood that the following specific embodiments are for illustrative purposes only and are not intended to limit the scope of the invention.
[0067] Example 1
[0068] A novel view synthesis system for sparsely sampled scenes based on multi-source image feature fusion, such as Figure 1 As shown, it includes at least a multi-source feature generation module, a feature aggregation module, a feature mapping module, an appearance-density estimation network, an occupancy estimation network, and a synthetic rendering module;
[0069] The multi-source feature generation module adopts a dual-branch architecture, mainly consisting of a visible light-dominated branch, a depth-dominated branch, and a feature fusion network. This generates high-dimensional feature information that integrates scene appearance texture and geometric priors. The visible light-dominated branch takes visible light and depth images from the scene dataset as input, using the visible light image to correct the depth image, outputting the corrected depth image. The depth-dominated branch takes the original depth image and the corrected depth image as input, outputting a multi-dimensional feature image. The fusion network takes the position-encoded visible light image and the multi-dimensional feature image as input, outputting the final extracted high-dimensional feature information. Both the visible light-dominated and depth-dominated branches consist of encoder-decoder networks with symmetric skip connections, containing multiple cascaded nonlinear activation residual blocks. The position encoding method for the visible light image is Fourier encoding, which maps the coordinates of each pixel to a higher-dimensional frequency space, alleviating the learning difficulty of the network for high-frequency feature changes. The fusion network consists of batch normalization layers and a ReLU activation function, providing an effective representation of the joint characteristics of different data sources in the scene.
[0070] The feature aggregation module uses tensor stitching to stitch pixel-aligned depth information as auxiliary information for geometric cues into high-dimensional feature information to achieve feature aggregation. The feature mapping module uses homography transformation mapping to calculate the homography transformation matrix between the current observation view and the new target view using camera parameters. It then projects and transforms the aggregated feature information of the observation view to the new target view to obtain the aggregated feature information under the new target view.
[0071] The appearance-density estimation network is a multilayer perceptron containing multiple fully connected layers, including at least an appearance feature estimation module, an appearance calculation module, and a density feature estimation module. It takes aggregated feature information as input to obtain color appearance information and density feature information of the target from a new perspective.
[0072] The occupancy estimation network adopts a symmetric encoder-decoder architecture, consisting of a 3D convolutional layer with skip connections and a transposed convolutional layer. Density feature information and original occupancy information are input to the occupancy estimation network, and density information of the target under a new perspective is output. The original occupancy information is calculated from the depth image and camera parameters, and the original point cloud of the calculated scene is used as the original occupancy information.
[0073] The compositing rendering module uses volumetric rendering, taking the color appearance information and density information of the target from a new perspective as input, and outputting the visible light image and depth image synthesized from the target's new perspective.
[0074] This invention constructs and initializes a multi-source feature generation module. Visible light and depth images from the observation perspectives in the scene dataset are fed into this module to extract and generate high-dimensional feature information. Using the depth image as geometric cues, the extracted high-dimensional feature information is aggregated with the geometric cues through a feature aggregation module to obtain aggregated feature information. A feature mapping module then projects and transforms the aggregated feature information from the observation perspectives to a new target perspective to obtain aggregated feature information under the new target perspective. This aggregated feature information under the new target perspective is fed into an appearance-density estimation network to obtain color appearance and density feature information under the new target perspective. The density features are then fed into an occupancy estimation network to obtain density information under the new target perspective. Finally, the density and color appearance information are fed into a synthesis rendering module to obtain reconstructed visible light and depth images under the new target perspective. This allows for the subsequent use of the visible light and depth images actually observed from the new target perspective as prior information to establish hybrid perception constraints. After iterative optimization of the parameters of the multi-source feature generation module, the appearance-density estimation network, and the occupancy estimation network, high-fidelity scene new view synthesis is finally achieved, supporting the generation of visible light images from any perspective in the scene. The system of this invention can synthesize high-fidelity new view images from sparsely sampled scenes.
[0075] Example 2
[0076] This example captures visible light images of the same scene from nine different viewing angles. The image size is 800×800 pixels. The example scenes include a single-object-centered scene and a complex indoor scene. Image examples are shown below. Figure 2 As shown.
[0077] A novel view synthesis method for sparsely sampled scenes based on multi-source image feature fusion, using the system described in Example 1, is as follows: Figure 6 As shown, firstly, visible light and depth images from different observation perspectives within the same scene are acquired and preprocessed to form a scene dataset. Then, a multi-source feature generation module is constructed to extract high-dimensional feature information. Next, using the depth image as geometric cues, the feature information is projected onto the target new perspective through a feature aggregation module and a feature mapping module. Subsequently, the aggregated features are input into an appearance-density estimation network to obtain color appearance information and density features, which are then fed into an occupancy estimation network to generate density information. Finally, prior supervision is established using actual observation images to optimize network parameters and achieve high-fidelity new view synthesis. The specific method includes the following steps:
[0078] Step S1: Collect visible light images and depth images from 9 different observation perspectives in the same scene, and ensure that each visible light image from each observation perspective has a corresponding depth image and camera parameter data, and jointly construct a scene dataset suitable for view synthesis under the new perspective of the target; select images from a limited number of perspectives (the number of perspectives should be less than 10) as input perspective data for training the new view synthesis model of the scene, and use the images from the remaining perspectives as test perspective data for performance evaluation of the view synthesis results.
[0079] In this embodiment, six perspectives are selected as input perspectives, and the scene data corresponding to the input perspectives are used as input perspective data for training the new view synthesis model for that scene. The remaining three perspectives are used as test perspectives for performance evaluation of the view synthesis results.
[0080] Step S2: Construct and initialize the multi-source feature generation module. Input the visible light image and depth image of the input viewpoint data into the module to extract and generate high-dimensional feature information.
[0081] like Figure 3 As shown, the multi-source feature generation module adopts a dual-branch architecture, mainly including a visible light-dominated branch, a depth-dominated branch, and a feature fusion network, which are used to generate high-dimensional feature information that integrates scene appearance texture and geometric priors.
[0082] The visible light image and depth image of the input viewpoint data are used as inputs to the multi-source feature generation module. The depth image is corrected by the visible light dominant branch, and then the multi-dimensional feature image is obtained by the depth dominant branch. The calculation process is as follows:
[0083] F o =B dept h (D r D i ) = B dept h (B color (I i D i ),D i )
[0084] In the formula, I i and D i These are the visible light images from the input viewpoint data, D dept h D is the dominant branch of visible light. r For the corrected depth image, B color As a depth-dominant branch, F o This is the extracted multidimensional feature image.
[0085] Position encoding is performed on visible light images, which involves Fourier encoding the coordinates of each pixel and mapping them to a higher-dimensional frequency space. This helps alleviate the learning difficulty of the network for high-frequency feature changes. The calculation process is as follows:
[0086] γ L (p;L)=(sin(2 0 πp),cos(2 0 πp),…,sin(2 L-1 πp),cos(2 L-1 πp))
[0087] In the formula, p represents the position coordinates of the multidimensional feature image, L represents the Fourier series, and in this embodiment, L = 6, γ L This is a position encoding function.
[0088] Using the location-encoded visible light image and the multi-dimensional feature image as input to the fusion network, the final extracted high-dimensional feature information is output. The calculation process is as follows:
[0089] F f (p) = Fusion <F o ,γ L (p;L)>
[0090] In the formula, p represents the position coordinates of the multidimensional feature image, and F... f To extract high-dimensional feature information, F o For a multidimensional feature image, L is the Fourier series, and γ L For positional encoding functions, Fusion represents a fusion network consisting of a batch normalization layer and a ReLU activation function.
[0091] Step S3: Using the corrected depth image as geometric cues, the extracted high-dimensional feature information is aggregated with the geometric cues through the feature aggregation module to obtain aggregated feature information. Then, the aggregated feature information from the observation perspective is projected and transformed to the new target perspective using the feature mapping module to obtain aggregated feature information from the new target perspective.
[0092] The feature aggregation module uses tensor concatenation to stitch pixel-aligned depth information as auxiliary information for geometric cues to high-dimensional feature information to achieve feature aggregation. The construction process of the geometric cues is as follows:
[0093]
[0094] Z = D r
[0095] In the formula, u and v represent the pixel coordinates of the multidimensional feature image, respectively, and c x c y f x f y D represents the camera parameters corresponding to the observation viewpoint in the scene dataset. r The image shows the corrected depth image, with X, Y, and Z representing the constructed geometric cues.
[0096] The homography transformation matrix between the current observation viewpoint and the new target viewpoint is calculated using camera parameters. The calculation process is as follows:
[0097]
[0098] In the formula, H i→t Let D be the homography transformation matrix from observation viewpoint i to target viewpoint t. t Let I be the sampling depth, I be the identity matrix, and K, R, and T be the camera parameters. The principal axis unit vector of the target view.
[0099] The aggregated feature information from the observation perspective is projected and transformed to the new target perspective to obtain the aggregated feature information from the new target perspective. The calculation process is as follows:
[0100] F f,i→t (p) = concat[F f,i (H i→t (D t )p T (X,Y,Z)]
[0101] In the formula, F f,i→t This represents the aggregated feature information transformed from the observation viewpoint i to the target viewpoint t, where p is the position coordinate of the feature image, and F... f,i H represents the feature information of observation viewpoint i. i→tLet D be the homography transformation matrix from observation viewpoint i to target viewpoint t. t X represents the sampling depth, and X, Y, and Z represent the constructed geometric cues.
[0102] Step S4: Input the aggregated feature information of the target from the new perspective into the appearance-density estimation network to obtain the color appearance information and density feature information of the target from the new perspective.
[0103] The appearance-density estimation network includes an appearance feature estimation module, an appearance calculation module, and a density feature estimation module, as detailed below:
[0104] S41: The aggregated feature information from multiple observation perspectives to the new target perspective is fed into the appearance-density estimation network to obtain visible light features and weight vectors. The calculation process is as follows:
[0105]
[0106] In the formula, w i These are the visible light features and weight vectors, respectively. G0 is the appearance feature estimation module in the appearance-density estimation network, K is the number of observation views, and F... f,i→t p represents the aggregated feature information transformed from the observation viewpoint i to the target viewpoint t, where p is the position coordinate of the feature image.
[0107] S42: Color appearance information is obtained by using light features and weight vectors. The calculation process is as follows:
[0108]
[0109] In the formula, C represents the synthesized color appearance information, and G0 represents the appearance calculation module. w i These are the visible light features and weight vectors, respectively, where K is the number of observation viewpoints and p is the position coordinate of the feature image.
[0110] S43: Density feature information is obtained through the density feature estimation module in the appearance-density estimation network. The calculation process is as follows:
[0111]
[0112] In the formula, f d For the estimated density feature information, G d For appearance calculation module, w i These are the visible light features and weight vectors, respectively, where K is the number of observation viewpoints and p is the position coordinate of the feature image.
[0113] Step S5: The density features are fed into the occupancy estimation network to obtain density information of the target from the new perspective. The density information and color / appearance information are then fed into the synthesis rendering module to obtain the reconstructed visible light image and depth image of the target from the new perspective. The specific process is as follows:
[0114] The density features and original occupancy information are fed into the occupancy estimation network to obtain density information from a new perspective of the target. The calculation process is as follows:
[0115] σ(p)=G occ (ψ(p),f d (p))
[0116] In the formula, σ represents the estimated density information, and G... occ To estimate network occupancy, f d The density feature information is estimated, where p is the position coordinate of the feature image and ψ is the original occupancy information, which is obtained by estimating the original point cloud from the depth image and camera parameters in the scene dataset.
[0117] Density and color appearance information are fed into the compositing rendering module to obtain a reconstructed visible light image of the target from a new perspective. The calculation process is as follows:
[0118]
[0119] In the formula, and These represent the visible light image and depth image reconstructed from the new perspective of the target, respectively. render is the composite rendering module using the volumetric rendering method, C is the composite color appearance information, σ is the estimated density information, and p is the image position coordinates.
[0120] Step S6: Using the visible light image and depth image obtained from the actual observation of the target from the new perspective as prior information, a hybrid sensing constraint is established as prior supervision based on this information. The parameters of the multi-source feature generation module, appearance-density estimation network and occupancy estimation network are iteratively optimized to finally achieve high-fidelity scene new view synthesis and support the generation of visible light images from any perspective in the scene.
[0121] The prior-supervised hybrid sensing constraint consists of three parts: visible light error constraint, local depth constraint, and global depth constraint. The calculation process is as follows:
[0122]
[0123] In the formula, L represents the hybrid sensing constraint, L c For visible light error constraints, For local depth constraints, For global depth constraints, λ is a hyperparameter, and in this embodiment, the hyperparameter is set to λ.p =0.2, λ f =0.35.
[0124] The visible light error constraint is the mean square error between the real observed visible light image and the synthetically rendered visible light image, used to minimize the difference between the real observed image and the synthetic image. The calculation process is as follows:
[0125]
[0126] In the formula, L c For visible light error constraints, w and h are the image sizes in the scene dataset, and I gt (p) is a visible light image actually observed from the new perspective of the target. p represents the reconstructed visible light image of the target from a new perspective, where p is the image position coordinate.
[0127] The local depth constraint is calculated from the maximum squared error of the depth values in an N×N pixel block, and the calculation process is as follows:
[0128]
[0129] In the formula, For local depth constraints, the patch is an N×N pixel block. For the reconstructed depth image, D represents the actual observed depth image, and ε is a redundancy constant considering minor differences. In this embodiment, N = 16 and ε = 10. -6 .
[0130] The global depth constraint is calculated from the depth distribution similarity loss of the entire pixel plane. It is obtained by discretizing the depth values into R intervals and calculating the KL divergence between the real depth image and the synthetic depth image. The calculation process is as follows:
[0131]
[0132] In the formula, For global depth constraints, Gauss(D) represents the discrete distribution of the reconstructed depth image. i ) represents the Gaussian mixture model representation of the actual observed depth image, and R is the number of depth value discretization intervals. In this embodiment, R = 24 is set.
[0133] Hybrid sensing constraints are used as prior supervision to iteratively optimize the parameters of the multi-source feature generation module, appearance-density estimation network, and occupancy estimation network. In this embodiment, the Adam optimizer is trained for 80,000 epochs with a batch size of 2048 and a spatial resolution of 64. 3The learning rate is set to 1e-4. After training, high-fidelity scene new view synthesis can be achieved, and visible light images from any viewpoint in the synthesized scene can be supported.
[0134] Test case
[0135] To verify the effectiveness and advantages of the method of this invention, multiple different scenarios were selected for testing. These scenarios should include both single-object-centered scenarios and complex indoor scenarios, such as... Figure 2 As shown in the diagram. Five fixed-position cameras, numbered 1 to 5, are set up in each scene, capturing images of 800×800 pixels. Images from cameras 1 to 4 are used as the input dataset, and images from camera 5 are used as the test dataset. Different models are trained to full convergence using the same input dataset. The trained models are then used to generate the synthesized view from camera 5, and the results are visualized and compared with real images captured by camera 5. This test case selects the four best-performing existing single-view synthesis models for comparison to test the single-view synthesis performance of the models. The results are shown in the diagram. Figure 4 As shown in the figure. Additionally, the three best-performing existing generalized view synthesis models were selected for comparison to test the generalized view synthesis performance of the model. The results are as follows. Figure 5 As shown.
[0136] Figure 4 This paper presents a comparison of the results of synthesizing new views of instance scenes using single-scene view synthesis models. The invention selected four of the best-performing existing single-view synthesis models for comparison: DSNeRF, 3DGS, DINER, and SparseNeRF. The results of synthesizing new views using different models are listed in two types of scenes: single-object-centered scenes and complex indoor scenes. The comparison shows that the new view synthesis results of this invention have higher fidelity, especially in better restoring the surface features of objects with rich texture details, while avoiding blurring and artifacts.
[0137] Figure 5 This is a comparison chart showing the results of generalized view synthesis models synthesizing new views for instance scenes. This invention selected the three best-performing existing generalized view synthesis models for comparison: MVSNERF, IBRNet, and MVSGS. The chart lists the new view synthesis results of different models in two types of scenes: single-object-centered scenes and complex indoor scenes. The comparison shows that in generalized scenes, the new view synthesis results of this invention can generate more detailed scene details, and the synthesized view is closest to the actual observed view, effectively restoring the surface of objects with specular reflections.
[0138] In summary, this invention can synthesize high-fidelity new view images from sparsely sampled scenes and supports the generation of visible light images from any viewpoint in the scene, making it more accurate and efficient.
[0139] It should be noted that the above content merely illustrates the technical concept of the present invention and should not be construed as limiting the scope of protection of the present invention. For those skilled in the art, various improvements and modifications can be made without departing from the principle of the present invention, and all such improvements and modifications fall within the scope of protection of the claims of the present invention.
Claims
1. A novel view synthesis system for sparsely sampled scenes based on multi-source image feature fusion, characterized in that: It includes at least a multi-source feature generation module, a feature aggregation module, a feature mapping module, an appearance-density estimation network, an occupancy estimation network, and a synthetic rendering module; The multi-source feature generation module adopts a dual-branch architecture, including a visible light-dominated branch, a depth-dominated branch, and a feature fusion network, which is used to generate high-dimensional feature information that integrates scene appearance texture and geometric priors. The feature aggregation module uses tensor splicing to aggregate the obtained high-dimensional feature information with geometric clues to obtain aggregated feature information. The feature mapping module adopts a homography transformation mapping method, uses camera parameters to calculate the homography transformation matrix between the current observation view and the new target view, projects and transforms the aggregated feature information of the observation view to the new target view, and obtains the aggregated feature information under the new target view. The appearance-density estimation network is a multilayer perceptron containing multiple fully connected layers, including at least an appearance feature estimation module, an appearance calculation module, and a density feature estimation module, to obtain color appearance information and density feature information of the target from a new perspective. The occupancy estimation network adopts a symmetric encoder-decoder architecture, consisting of a 3D convolutional layer with skip connections and a transposed convolutional layer; density feature information and original occupancy information are input to the occupancy estimation network, and density information of the target under a new perspective is output. The synthetic rendering module takes the color appearance information and density information of the target under the new perspective as input, and outputs the visible light image and depth image synthesized under the new perspective of the target.
2. The sparse sampling scene novel view synthesis system based on multi-source image feature fusion as described in claim 1, characterized in that: In the visible light dominant branch of the multi-source feature generation module, the acquired visible light image and depth image are used as input, the visible light image is used to correct the depth image, and the output is the corrected depth image. In the depth-dominant branch, the original input depth image and the corrected depth image are used as inputs, and the output is a multi-dimensional feature image. In the fusion network, the visible light image after position encoding and the multi-dimensional feature image are used as inputs, and the output is the finally extracted high-dimensional feature information.
3. The sparse sampling scene new view synthesis system based on multi-source image feature fusion as described in claim 2, characterized in that: The visible light dominant branch and depth dominant branch of the multi-source feature generation module are both composed of encoder-decoder networks with symmetrical skip connections, containing multiple cascaded nonlinear activation residual blocks; the position encoding method of the visible light image is Fourier encoding; the fusion network consists of batch normalization layers and ReLU activation functions.
4. The sparse sampling scene novel view synthesis system based on multi-source image feature fusion as described in claim 1, characterized in that: The homography transformation matrix in the feature mapping module is specifically as follows: ; In the formula, From the perspective of observation To the target perspective The homography transformation matrix, For sampling depth, It is the identity matrix. , , For camera parameters, The principal axis unit vector of the target view.
5. A method for synthesizing new scenes from sparse sampling based on multi-source image feature fusion using the system described in claim 1, characterized in that, Includes the following steps: S1. Data preprocessing: Collect visible light images and depth images from different observation angles in the same scene, preprocess the collected image data to ensure that each visible light image from the observation angle has a corresponding depth image and camera parameter data, forming a scene dataset. S2, High-dimensional feature extraction: Initialize the multi-source feature generation module, input the visible light image and depth image from the observation perspective in the scene dataset into the multi-source feature generation module, extract and generate high-dimensional feature information; S3. Feature aggregation: Using the corrected depth image as geometric cues, the high-dimensional feature information extracted in step S2 is aggregated with the geometric cues through the feature aggregation module to obtain aggregated feature information. The feature mapping module projects and transforms the aggregated feature information onto the new perspective of the target, thus obtaining the aggregated feature information under the new perspective of the target. S4. Obtaining color appearance information and density feature information: Input the aggregated feature information of the target under the new perspective obtained in step S3 into the appearance-density estimation network to obtain the color appearance information and density feature information of the target under the new perspective. S5. Density information and visible light image acquisition: Input the density feature information obtained in step S4 into the occupancy estimation network to obtain the density information under the new view of the target. Then, input the density information and the color appearance information obtained in step S4 into the synthesis rendering module to obtain the reconstructed visible light image under the new view of the target. S6. Iterative optimization: Using the visible light image of the target under the new perspective obtained in step S5 as prior information, a hybrid perception constraint is established as prior supervision. The parameters of the multi-source feature generation module, appearance-density estimation network and occupancy estimation network are iteratively optimized to achieve the synthesis of the new scene view.
6. The method for synthesizing new views of sparsely sampled scenes based on multi-source image feature fusion as described in claim 5, characterized in that: Step S3 projects and transforms the aggregated feature information of the observation viewpoint to the new target viewpoint to obtain the aggregated feature information under the new target viewpoint, specifically as follows: ; In the formula, From the perspective of observation Change to target perspective Aggregated feature information, The location coordinates of the feature image. From the perspective of observation Feature information, From the perspective of observation To the target perspective The homography transformation matrix, For sampling depth, , , The geometric cues for construction.
7. The method for synthesizing new views of sparsely sampled scenes based on multi-source image feature fusion as described in claim 5, characterized in that: Step S4 specifically includes the following steps: S41: Input the aggregated feature information from multiple observation perspectives to the new target perspective into the appearance-density estimation network to obtain visible light features and weight vectors: ; In the formula, , These are the visible light features and weight vectors, respectively. This is the appearance feature estimation module in the appearance-density estimation network. For the number of observation angles, From the perspective of observation Change to target perspective Aggregated feature information, The coordinates of the feature image; S42: Obtain color appearance information based on visible light features and weight vectors, specifically: ; In the formula, For the synthesized color appearance information, For appearance calculation module, , These are the visible light features and weight vectors, respectively. For the number of observation angles, The coordinates of the feature image; S43: Obtain density feature information; the specific calculation process is as follows: ; In the formula, For the estimated density feature information, For appearance calculation module, , These are the visible light features and weight vectors, respectively. For the number of observation angles, These are the position coordinates of the feature image.
8. The method for synthesizing new views of sparsely sampled scenes based on multi-source image feature fusion as described in claim 7, characterized in that: The calculation method for obtaining density information through the occupancy estimation network in step S5 is as follows: ; In the formula, For the estimated density information, To estimate network occupancy, For the estimated density feature information, The location coordinates of the feature image. This is the original occupancy information; The specific calculation process for obtaining the reconstructed visible light image of the target from a new perspective in the compositing rendering module is as follows: ; In the formula, and These represent the visible light image and depth image reconstructed from the new perspective of the target, respectively. This is a compositing rendering module that uses a volumetric rendering method. For the synthesized color appearance information, For the estimated density information, These are the image position coordinates.
9. The method for synthesizing new views of sparsely sampled scenes based on multi-source image feature fusion as described in claim 5, characterized in that: The hybrid sensing constraints in step S6 include visible light error constraints, local depth constraints, and global depth constraints, wherein... The visible light error constraint is: ; In the formula, For visible light error constraints, , The image size in the scene dataset. Visible light images actually observed from a new perspective of the target. The visible light image reconstructed from a new perspective of the target. These are the image position coordinates; The local depth constraint is: ; In the formula, For local depth constraints, for pixel blocks, For the reconstructed depth image, This is a depth image based on actual observation. Redundancy constants to account for minute differences; The global depth constraint is: ; In the formula, For global depth constraints, This represents the discrete distribution of the reconstructed depth image. The Gaussian mixture model represents the depth image as observed in reality. The number of intervals for discretizing the depth value.