Sparse sampling scene new view synthesis system and method based on multi-source image feature fusion
By adopting a system of multi-source image feature fusion in sparse sampling scenarios and combining iterative optimization methods of mixed perception constraints, the high fidelity problem of new view synthesis in sparse sampling scenarios is solved, efficient geometric and appearance information fusion is achieved, and the quality of the synthesis results is improved.
Patent Information
- Application Number
- CN202510010393.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-03
AI Technical Summary
In sparse sampling scenarios, existing new view synthesis technology is difficult to achieve high-fidelity results, and often due to model degradation, geometric details, blurred textures, and unnatural perspective transitions.
Using a system based on multi-source image feature fusion, high-dimensional feature information is generated and fused through multi-source feature generation module, feature aggregation module, feature mapping module, appearance-density estimation network, occupation estimation network and synthetic rendering module, and iteratively optimized through hybrid perception constraints to achieve high-fidelity new view synthesis.
It realizes high-fidelity new view synthesis from sparsely sampled scenes, supports the generation of visible light images at any view angle in the scene, and improves the geometric consistency and appearance fidelity of the synthesis results.
Smart Images

Figure CN119942282A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of new view synthesis, and mainly relates to a new view synthesis system and method for a sparse sampling scene based on multi-source image feature fusion. Background Art
[0002] New view synthesis is an important research topic in the field of computer vision and graphics. It aims to synthesize new views from any observation point of view in a scene using views from multiple observation points of the scene as input. This task is widely used in fields such as virtual reality (VR) and augmented reality (AR), film and television production, autonomous driving, and robotic vision. In recent years, although neural rendering technology represented by neural radiance fields can achieve high-fidelity new view synthesis, it requires a large number of input views and strict shooting standards, which seriously limits its promotion and use in practice in related fields. In particular, in sparsely sampled scenarios, that is, when there are only a small number of sparsely distributed observation views, the new view synthesis results usually suffer from problems such as loss of geometric details, blurred textures, and unnatural perspective transitions due to model degradation. Therefore, achieving high-fidelity new view synthesis from sparsely sampled scenes is the main challenge currently faced.
[0003] To address this challenge, existing related research can be summarized into the following three categories: (1) Pre-training models. The model's generative capabilities are enhanced through a large amount of scene data, and new views can even be synthesized under a single view input. However, such methods are highly dependent on large-scale scene data, have high implementation costs, and are prone to artifacts and geometric errors in unseen scene types. (2) Introducing additional constraints. In the process of 3D scene representation or regularization, additional constraints such as appearance constraints and sampling range constraints are introduced to alleviate the model degradation problem. Although this method performs well in simple scenes centered on a single object, due to the lack of complete geometric information integrity and coherence in sparsely sampled scenes, the fidelity of synthesizing new views in complex indoor scenes is still insufficient. (3) Depth prior. Depth information can intuitively reflect the geometric details of the scene. Using depth prior as supervision or embedding it into the model input can effectively alleviate the model's training burden and improve model performance. In practical applications, depth prior can be easily obtained through consumer-grade depth sensors or depth estimation models, which facilitates the promotion and application of this method.
[0004] Although depth prior methods have shown great potential in synthesizing new views of sparsely sampled scenes, they still have the following defects: (1) Insufficient accuracy and robustness of depth information. Depth maps generated by consumer-grade depth sensors or depth estimation models usually have noise, missing and quantization errors, especially in highly reflective, transparent or less textured areas. These defects may lead to inaccurate depth prior information, which in turn affects the expression of scene geometry. (2) Insufficient depth-appearance fusion. In the process of fusing depth prior and appearance features, existing methods usually lack a unified optimization framework, resulting in insufficient coordination between geometric information and appearance information. This separate processing method is prone to cause geometric misalignment or texture distortion problems in the synthesis results. Summary of the invention
[0005] The present invention aims at the problems existing in the prior art and proposes a sparse sampling scene new view synthesis system and method based on multi-source image feature fusion. The multi-source feature generation module extracts and generates high-dimensional feature information, and the feature aggregation module aggregates the extracted high-dimensional feature information with geometric clues to obtain aggregated feature information. The feature mapping module is then used to project the aggregated feature information of the observation perspective to the new perspective of the target to obtain the aggregated feature information under the new perspective of the target. The aggregated feature information under the new perspective of the target is sent to the appearance-density estimation network to obtain the color appearance information and density feature information under the new perspective of the target. The density feature is sent to the occupancy estimation network to obtain the density information under the new perspective of the target, and the density information and the color appearance information are sent to the synthesis rendering module together to obtain the visible light image and depth image reconstructed under the new perspective of the target. The method of the present invention uses the visible light image and the depth image actually observed from the new perspective of the target as prior information, establishes mixed perception constraints as prior supervision, iteratively optimizes the parameters of the multi-source feature generation module, the appearance-density estimation network and the occupancy estimation network, and finally realizes high-fidelity scene new view synthesis and supports the generation of visible light images of any perspective in the scene. The present invention can realize the synthesis of high-fidelity new view images from sparsely sampled scenes.
[0006] In order to achieve the above-mentioned object, the technical solution adopted by the present invention is: a sparse sampling scene new view synthesis system based on multi-source image feature fusion, comprising at least a multi-source feature generation module, a feature aggregation module, a feature mapping module, an appearance-density estimation network, an occupancy estimation network and a synthesis rendering module;
[0007] The multi-source feature generation module: adopts a dual-branch architecture, including a visible light dominant branch, a depth dominant branch and a feature fusion network, for generating high-dimensional feature information that integrates scene appearance texture and geometric priors;
[0008] The feature aggregation module: uses a tensor splicing method to aggregate the obtained high-dimensional feature information and geometric clues to obtain aggregated feature information;
[0009] The feature mapping module: adopts a homography transformation mapping method, uses camera parameters to calculate the homography transformation matrix between the current observation angle and the target new angle, and projects the aggregated feature information of the observation angle to the target new angle to obtain the aggregated feature information under the target new angle;
[0010] The appearance-density estimation network: a multi-layer perceptron comprising multiple fully connected layers, including at least an appearance feature estimation module, an appearance calculation module and a density feature estimation module, to obtain color appearance information and density feature information of the target under a new viewing angle;
[0011] The occupancy estimation network: adopts a symmetrical encoder-decoder architecture, which is composed of a three-dimensional convolutional layer and a transposed convolutional layer with skip connections; inputs density feature information and original occupancy information to the occupancy estimation network, and outputs density information under a new perspective of the target;
[0012] The synthesis rendering module takes the color appearance information and density information of the target under the new viewing angle as input, and outputs the visible light image and depth image synthesized under the new viewing angle of the target.
[0013] As an improvement of the present invention, in the visible light dominant branch of the multi-source feature generation module, the collected visible light image and depth image are used as input, the depth image is corrected using the visible light image, and the corrected depth image is output;
[0014] In the depth-dominant branch, the original input depth image and the corrected depth image are used as input, and the output is a multi-dimensional feature image;
[0015] In the fusion network, the position-encoded visible light image and the multi-dimensional feature image are used as input, and the output is the final extracted high-dimensional feature information.
[0016] As an improvement of the present invention, the visible light dominant branch and the depth dominant branch of the multi-source feature generation module are both composed of a codec network with symmetric jump connections, including multiple cascaded nonlinear activation residual blocks; wherein the position encoding method of the visible light image is Fourier coding; and the fusion network is composed of a batch normalization layer and a ReLU activation function.
[0017] As another improvement of the present invention, the homography transformation matrix in the feature mapping module is specifically:
[0018]
[0019] In the formula, H i→tis the homography transformation matrix from observation view i to target view t, D t is the sampling depth, I is the unit matrix, K, R, T are the camera parameters, is the principal axis unit vector of the target view.
[0020] In order to achieve the above object, the present invention also adopts a technical solution: a new view synthesis method of a sparse sampling scene based on multi-source image feature fusion, comprising the following steps:
[0021] S1. Data preprocessing: Collect visible light images and depth images of different observation angles under the same scene, and preprocess the collected image data to ensure that the visible light image under each observation angle has a corresponding depth image and camera parameter data to form a scene data set;
[0022] S2, high-dimensional feature extraction: Initialize the multi-source feature generation module, input the visible light image and depth image of the observation angle in the scene data set into the multi-source feature generation module, extract and generate high-dimensional feature information;
[0023] S3, aggregate feature acquisition: using the corrected depth image as a geometric clue, the high-dimensional feature information extracted in step S2 is aggregated with the geometric clue through the feature aggregation module to obtain aggregate feature information; through the feature mapping module, the aggregate feature information is projected and transformed to the new target perspective to obtain the aggregate feature information under the new target perspective;
[0024] S4, color appearance information and density feature information acquisition: input the aggregate feature information of the target under the new viewing angle acquired in step S3 into the appearance-density estimation network to acquire the color appearance information and density feature information of the target under the new viewing angle;
[0025] S5, density information and visible light image acquisition: input the density feature information obtained in step S4 into the occupancy estimation network to obtain the density information of the target under the new viewing angle, and then input the density information and the color appearance information obtained in step S4 into the synthesis rendering module to obtain the reconstructed visible light image under the new viewing angle of the target;
[0026] S6, iterative optimization: Using the visible light image of the target under the new perspective obtained in step S5 as prior information, establish hybrid perception constraints as prior supervision, iteratively optimize the parameters of the multi-source feature generation module, appearance-density estimation network and occupancy estimation network to achieve new view synthesis of the scene.
[0027] As an improvement of the present invention, the step S3 projects the aggregated feature information of the observation perspective to the target new perspective to obtain the aggregated feature information under the target new perspective, specifically:
[0028] F f,i→t (p) = concat[Ff,i (H i→t (D t ) T ),(X,Y,Z)]
[0029] In the formula, F f,i→t is the aggregated feature information transformed from the observation perspective i to the target perspective t, p is the position coordinate of the feature image, and F f,i is the characteristic information of observation angle i, H i→t is the homography transformation matrix from observation view i to target view t, D t is the sampling depth, and X, Y, and Z are the geometric clues of the construction.
[0030] As another improvement of the present invention, step S4 specifically includes the following steps:
[0031] S41: The aggregated feature information from multiple observation perspectives to the target new perspective is input into the appearance-density estimation network to obtain the visible light features and weight vector:
[0032]
[0033] In the formula, w i are the visible light features and weight vectors respectively, G0 is the appearance feature estimation module in the appearance-density estimation network, K is the number of observation angles, and F f,i→t is the aggregated feature information transformed from the observation perspective i to the target perspective t, and p is the position coordinate of the feature image;
[0034] S42: Obtain color appearance information based on the light feature and the weight vector, specifically:
[0035]
[0036] Where C is the synthesized color appearance information, G0 is the appearance calculation module, w i are the visible light feature and weight vector respectively, K is the number of observation angles, and p is the position coordinate of the feature image;
[0037] S43: Obtain density feature information. The specific calculation process is as follows:
[0038]
[0039] In the formula, f d is the estimated density feature information, G d is the appearance calculation module, w i are the visible light features and weight vectors respectively, K is the number of observation angles, and p is the position coordinate of the feature image.
[0040] As another improvement of the present invention, in step S5, the calculation method for obtaining density information through the occupancy estimation network is:
[0041] σ(p)=G occ (ψ(p),f d (p))
[0042] In the formula, σ is the estimated density information, G occ is the occupancy estimation network, f d is the estimated density feature information, p is the position coordinate of the feature image, and ψ is the original occupancy information;
[0043] The specific calculation process of obtaining the reconstructed visible light image under the new viewing angle of the target in the synthetic rendering module is as follows:
[0044]
[0045] In the formula, and They respectively represent the visible light image and depth image reconstructed from the new perspective of the target, render is the synthetic rendering module using the volume rendering method, C is the synthesized color appearance information, σ is the estimated density information, and p is the image position coordinate.
[0046] As a further improvement of the present invention, the mixed perception constraint in step S6 includes a visible light error constraint, a local depth constraint and a global depth constraint, wherein:
[0047] The visible light error constraint is:
[0048]
[0049] Where, L c is the visible light error constraint, w and h are the image sizes in the scene dataset, and I gt (p) is the visible light image actually observed at the new viewing angle of the target. is the visible light image reconstructed from the new viewing angle of the target, and p is the image position coordinate;
[0050] The local depth constraint is:
[0051]
[0052] In the formula, is a local depth constraint, patch is an N×N pixel block, is the reconstructed depth image, D is the actual observed depth image, and ε is the redundant constant considering small differences;
[0053] The global depth constraint is:
[0054]
[0055] In the formula, is the global depth constraint, Represents the discrete distribution of the reconstructed depth image, gauss(D i ) represents the Gaussian mixture model representation of the actual observed depth image, and R is the number of discretized intervals of the depth value.
[0056] Compared with the prior art, the present invention has the following beneficial effects:
[0057] (1) The present invention innovatively designs a multi-source feature generation module with a dual-branch architecture, including a visible light dominant branch and a depth dominant branch, and generates high-dimensional feature information that integrates scene appearance texture and geometric priors through a feature fusion network. The visible light dominant branch improves the accuracy of geometric information by correcting the depth image, and the depth dominant branch combines the multi-dimensional feature image to enhance the geometric expression ability. At the same time, the introduction of Fourier position coding alleviates the difficulty of the network in learning high-frequency features, thereby enhancing the representation ability of complex geometric details.
[0058] (2) This paper proposes a feature aggregation and homography transformation mapping method based on geometric cues. It uses depth information as a geometric cue to achieve high-dimensional feature aggregation through tensor splicing, and uses camera parameters to calculate the homography transformation matrix to map the feature information of the observation perspective to the new perspective of the target. This method makes full use of geometric cues and significantly improves the geometric consistency and appearance fidelity of new perspective synthesis in sparsely sampled scenes.
[0059] (3) The present invention proposes a hybrid perceptual constraint mechanism that includes visible light error constraints, local depth constraints, and global depth constraints, and uses prior supervision to jointly optimize the multi-source feature generation module, appearance-density estimation network, and occupancy estimation network. Through the KL divergence constraints of local depth consistency and global depth distribution, the degradation problem of geometric and appearance details in sparsely sampled scenes is effectively solved, and the geometric fidelity and detail quality of the synthesis results are improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] Figure 1 It is a structural schematic diagram of a new view synthesis system for a sparse sampling scene based on multi-source image feature fusion according to the present invention;
[0061] Figure 2 exemplified images of a single object central scene and a complex indoor scene in the example scene of Embodiment 2 of the present invention;
[0062] Figure 3 It is a structural schematic diagram of a multi-source feature generation module in the system of the present invention;
[0063] Figure 4 is a comparison diagram of the results of synthesizing new views of an instance scene by a single scene view synthesis model in a test example of the present invention;
[0064] Figure 5 is a comparison diagram of the results of synthesizing new views of an instance scene by the generalized view synthesis model in the test example of the present invention;
[0065] Figure 6 It is a flowchart of the steps of the new view synthesis method of sparse sampling scene based on multi-source image feature fusion of the present invention. DETAILED DESCRIPTION
[0066] The present invention will be further explained below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the present invention and are not used to limit the scope of the present invention.
[0067] Example 1
[0068] A new view synthesis system for sparsely sampled scenes based on multi-source image feature fusion, such as Figure 1 As shown, it at least includes a multi-source feature generation module, a feature aggregation module, a feature mapping module, an appearance-density estimation network, an occupancy estimation network and a synthetic rendering module;
[0069] The multi-source feature generation module adopts a dual-branch architecture, which mainly includes a visible light dominant branch, a depth dominant branch and a feature fusion network, which is used to generate high-dimensional feature information that integrates the scene appearance texture and geometric priors; the visible light dominant branch takes the visible light image and depth image in the scene data set as input, uses the visible light image to correct the depth image, and outputs the corrected depth image; the depth dominant branch takes the original input depth image and the corrected depth image as input, and outputs a multi-dimensional feature image; the fusion network takes the position-encoded visible light image and the multi-dimensional feature image as input, and outputs the final extracted high-dimensional feature information; both the visible light dominant branch and the depth dominant branch are composed of an encoder-decoder network with symmetric jump connections, which contains multiple cascaded nonlinear activation residual blocks; the position encoding method of the visible light image is Fourier encoding, that is, the coordinate value of each pixel is Fourier encoded and mapped to a higher-dimensional frequency space, which is used to alleviate the network's learning difficulty for high-frequency feature changes; the fusion network consists of a batch normalization layer and a ReLU activation function, which provides an effective representation for the joint characteristics of different data sources in the scene.
[0070] The feature aggregation module adopts tensor splicing to splice pixel-aligned depth information as auxiliary information of geometric clues to high-dimensional feature information to achieve feature aggregation; the feature mapping module adopts homography transformation mapping to calculate the homography transformation matrix between the current observation perspective and the new target perspective using camera parameters, and projects the aggregated feature information of the observation perspective to the new target perspective to obtain the aggregated feature information under the new target perspective.
[0071] The appearance-density estimation network includes a multilayer perceptron with multiple fully connected layers, including at least an appearance feature estimation module, an appearance calculation module and a density feature estimation module, which takes aggregated feature information as input to obtain color appearance information and density feature information of the target under a new perspective.
[0072] The occupancy estimation network adopts a symmetrical encoder-decoder architecture, which is composed of a three-dimensional convolutional layer and a transposed convolutional layer containing jump connections; density feature information and original occupancy information are input into the occupancy estimation network, and density information of the target under a new perspective is output; the original occupancy information is calculated by a depth image and camera parameters, and the original point cloud of the scene is calculated as the original occupancy information.
[0073] The synthetic rendering module adopts volume rendering, takes the color appearance information and density information of the target under the new perspective as input, and outputs the visible light image and depth image synthesized under the new perspective of the target.
[0074] The system of the present invention constructs and initializes a multi-source feature generation module, sends a visible light image and a depth image of an observation perspective in a scene data set to the module, extracts and generates high-dimensional feature information; uses the depth image as a geometric clue, aggregates the extracted high-dimensional feature information with the geometric clue through a feature aggregation module to obtain aggregated feature information, and uses a feature mapping module to project and transform the aggregated feature information of the observation perspective to a new target perspective to obtain aggregated feature information under the new target perspective; sends the aggregated feature information under the new target perspective to an appearance-density estimation network to obtain color appearance information and density feature information under the new target perspective; sends the density feature to an occupancy estimation network to obtain density information under the new target perspective, and sends the density information and color appearance information together to a synthesis rendering module to obtain a visible light image and a depth image reconstructed under the new target perspective, so as to facilitate the subsequent use of the visible light image and the depth image actually observed from the new target perspective as prior information, establish hybrid perception constraints, iteratively optimize the parameters of the multi-source feature generation module, the appearance-density estimation network and the occupancy estimation network, and finally realize high-fidelity scene new view synthesis, and support the generation of visible light images of any perspective in the scene. The system of the present invention can realize synthesis of high-fidelity new view images from sparsely sampled scenes.
[0075] Example 2
[0076] In this example, visible light images of the same scene at 9 different observation angles are collected. The image size is 800×800 pixels. The example scene includes a single object center scene and a complex indoor scene. The image examples are as follows: Figure 2 shown.
[0077] A new view synthesis method for a sparse sampling scene based on multi-source image feature fusion uses the system described in Example 1, such as Figure 6 As shown, first, visible light images and depth images of different observation angles of the same scene are collected and preprocessed to form a scene dataset. Then, a multi-source feature generation module is constructed to extract high-dimensional feature information. Next, the depth image is used as a geometric clue, and the feature information is projected to the target new perspective through the feature aggregation module and the feature mapping module. Subsequently, the aggregated features are input into the appearance-density estimation network to obtain color appearance information and density features, and then sent to the occupancy estimation network to generate density information. Finally, the actual observed images are used to establish prior supervision, optimize network parameters, and achieve high-fidelity new view synthesis. The specific method includes the following steps:
[0078] Step S1: collect visible light images and depth images of 9 different observation perspectives of the same scene, and ensure that each visible light image under each observation perspective has a corresponding depth image and camera parameter data, and jointly construct a scene dataset suitable for view synthesis under a new perspective of the target; select images with limited perspectives (the number of perspectives should be less than 10) as input perspective data for training the new view synthesis model of the scene, and use the images of the remaining perspectives as test perspective data for performance evaluation of the view synthesis results.
[0079] In this embodiment, 6 of the viewpoints are selected as input viewpoints, and the scene data corresponding to the input viewpoints are used as input viewpoint data for training a new view synthesis model for the scene. The remaining 3 viewpoints are used as test viewpoints for performance evaluation of the view synthesis results.
[0080] Step S2: construct and initialize a multi-source feature generation module, input the visible light image and depth image of the input viewing angle data into the module, and extract and generate high-dimensional feature information.
[0081] like Figure 3 As shown in the figure, the multi-source feature generation module adopts a dual-branch architecture, which mainly includes a visible light dominant branch, a depth dominant branch and a feature fusion network, which is used to generate high-dimensional feature information that integrates the scene appearance texture and geometric priors.
[0082] The visible light image and depth image of the input view data are used as the input of the multi-source feature generation module. The depth image is corrected by the visible light dominant branch, and then the multi-dimensional feature image is obtained by the depth dominant branch. The calculation process is as follows:
[0083] F o =B dept h (D r ,D i )=B dept h (B colo r (I i ,D i ),D i )
[0084] In the formula, I i and D i are the visible light images in the input viewing angle data, B dept h The visible light dominant branch, D r is the corrected depth image, B color is the depth-dominant branch, F o is the extracted multi-dimensional feature image.
[0085] Position encoding is performed on the visible light image, that is, the coordinate value of each pixel is Fourier encoded and mapped to a higher-dimensional frequency space to ease the difficulty of the network learning high-frequency feature changes. The calculation process is as follows:
[0086] γ L (p; L) = (sin(2 0 πp),cos(2 0 πp),…,sin(2 L-1 πp),cos(2 L-1 πp))
[0087] Where p is the position coordinate of the multidimensional feature image, L is the Fourier series, and in this embodiment, L=6, γ L is the position encoding function.
[0088] The position-encoded visible light image and multi-dimensional feature image are used as the input of the fusion network, and the final extracted high-dimensional feature information is output. The calculation process is as follows:
[0089] F f (p) = Fusion <F o ,γ L (p; L)>
[0090] Where p is the position coordinate of the multidimensional feature image, F f F is the extracted high-dimensional feature information. o is a multidimensional feature image, L is the Fourier series, γ L is the position encoding function, and Fusion represents a fusion network consisting of a batch normalization layer and a ReLU activation function.
[0091] Step S3: Using the corrected depth image as a geometric clue, the extracted high-dimensional feature information is aggregated with the geometric clues through the feature aggregation module to obtain aggregated feature information, and using the feature mapping module, the aggregated feature information of the observation perspective is projected and transformed to the new perspective of the target to obtain the aggregated feature information under the new perspective of the target.
[0092] The feature aggregation module uses tensor splicing to splice pixel-aligned depth information as auxiliary information of geometric clues to high-dimensional feature information to achieve feature aggregation. The construction process of geometric clues is as follows:
[0093]
[0094] In the formula, u and v represent the pixel coordinates of the multidimensional feature image, c x 、c y 、f x 、f y is the camera parameter corresponding to the observation angle in the scene dataset, D r is the corrected depth image, and X, Y, and Z are the constructed geometric clues.
[0095] The camera parameters are used to calculate the homography transformation matrix between the current observation angle and the target's new angle of view. The calculation process is as follows:
[0096]
[0097] In the formula, H i→t is the homography transformation matrix from observation view i to target view t, D t is the sampling depth, I is the unit matrix, K, R, T are the camera parameters, is the principal axis unit vector of the target view.
[0098] The aggregate feature information of the observation perspective is projected to the target new perspective to obtain the aggregate feature information under the target new perspective. The calculation process is as follows:
[0099] F f,i→t (p) = concat[F f,i (H i→t (D t ) T ),(X,Y,Z)]
[0100] In the formula, F f,i→t is the aggregated feature information transformed from the observation perspective i to the target perspective t, p is the position coordinate of the feature image, and F f,i is the characteristic information of observation angle i, H i→t is the homography transformation matrix from observation view i to target view t, D t is the sampling depth, and X, Y, and Z are the geometric clues of the construction.
[0101] Step S4: Send the aggregated feature information of the target under the new viewing angle to the appearance-density estimation network to obtain the color appearance information and density feature information of the target under the new viewing angle.
[0102] The appearance-density estimation network includes an appearance feature estimation module, an appearance calculation module, and a density feature estimation module. The specific details are as follows:
[0103] S41: The aggregated feature information from multiple observation perspectives to the target new perspective is transmitted to the appearance-density estimation network to obtain the visible light features and weight vector. The calculation process is as follows:
[0104]
[0105] In the formula, w i are the visible light features and weight vectors respectively, G0 is the appearance feature estimation module in the appearance-density estimation network, K is the number of observation angles, and F f,i→t is the aggregated feature information transformed from the observation perspective i to the target perspective t, and p is the position coordinate of the feature image.
[0106] S42: Using the light feature and the weight vector to obtain the color appearance information, the calculation process is as follows:
[0107]
[0108] Where C is the synthesized color appearance information, G0 is the appearance calculation module, w i are the visible light features and weight vectors respectively, K is the number of observation angles, and p is the position coordinate of the feature image.
[0109] S43: Obtain density feature information through the density feature estimation module in the appearance-density estimation network. The calculation process is as follows:
[0110]
[0111] In the formula, f d is the estimated density feature information, G d is the appearance calculation module, w i are the visible light features and weight vectors respectively, K is the number of observation angles, and p is the position coordinate of the feature image.
[0112] Step S5: Send the density feature to the occupancy estimation network to obtain the density information of the target under the new viewing angle, and send the density information and color appearance information to the synthesis rendering module to obtain the reconstructed visible light image and depth image under the new viewing angle of the target. The specific process is as follows:
[0113] The density features and original occupancy information are fed into the occupancy estimation network to obtain the density information of the target under the new perspective. The calculation process is as follows:
[0114] σ(p)=G occ (ψ(p),f d (p))
[0115] In the formula, σ is the estimated density information, G occ is the occupancy estimation network, f d is the estimated density feature information, p is the position coordinate of the feature image, ψ is the original occupancy information, and the original point cloud is estimated by the depth image and camera parameters in the scene dataset.
[0116] The density information and color appearance information are sent to the synthetic rendering module to obtain the reconstructed visible light image under the new viewing angle of the target. The calculation process is as follows:
[0117]
[0118] In the formula, and They respectively represent the visible light image and depth image reconstructed from the new perspective of the target, render is the synthetic rendering module using the volume rendering method, C is the synthesized color appearance information, σ is the estimated density information, and p is the image position coordinate.
[0119] Step S6: Using the visible light image and depth image actually observed from the new perspective of the target as prior information, a hybrid perception constraint is established as a priori supervision based on this, and the parameters of the multi-source feature generation module, the appearance-density estimation network, and the occupancy estimation network are iteratively optimized, ultimately achieving high-fidelity synthesis of new views of the scene and supporting the generation of visible light images of any perspective in the scene.
[0120] The hybrid perception constraint of prior supervision consists of three parts: visible light error constraint, local depth constraint, and global depth constraint. The calculation process is as follows:
[0121]
[0122] Where L is the hybrid perception constraint, L c is the visible light error constraint, is the local depth constraint, is the global depth constraint, λ is a hyperparameter, and in this embodiment, the hyperparameter is set to λ p =0.2,λ f =0.35.
[0123] The visible light error constraint is the mean square error between the real observed visible light image and the synthetic rendered visible light image, which is used to minimize the difference between the real observed image and the view synthesis image. The calculation process is as follows:
[0124]
[0125] Where, L c is the visible light error constraint, w and h are the image sizes in the scene dataset, and I gt (p) is the visible light image actually observed at the new viewing angle of the target. is the visible light image reconstructed from the new perspective of the target, and p is the image position coordinate.
[0126] The local depth constraint is calculated by the maximum square error of the depth value in the N×N pixel block. The calculation process is as follows:
[0127]
[0128] In the formula, is a local depth constraint, patch is an N×N pixel block, is the reconstructed depth image, D is the actual observed depth image, and ε is a redundant constant considering small differences. In this embodiment, N=16 and ε=10 -6 .
[0129] The global depth constraint is calculated by the depth distribution similarity loss of the entire pixel plane. It is obtained by discretizing the depth value into R intervals and calculating the KL divergence of the real depth image and the synthetic depth image. The calculation process is as follows:
[0130]
[0131] In the formula, is the global depth constraint, Represents the discrete distribution of the reconstructed depth image, gauss(D i ) represents the Gaussian mixture model representation of the actually observed depth image, R is the number of discretization intervals of the depth value, and in this embodiment, R=24 is set.
[0132] Using the hybrid perceptual constraints as prior supervision, the parameters of the multi-source feature generation module, the appearance-density estimation network, and the occupancy estimation network are iteratively optimized. In this example, the Adam optimizer is used for training 80,000 rounds, the batch size is 2048, and the spatial resolution is set to 64. 3 , the learning rate is set to 1e-4. After training, high-fidelity synthesis of new scene views can be achieved, and visible light images of any perspective in the synthesis scene can be supported.
[0133] Test Case
[0134] In order to verify the effectiveness and advantages of the method of the present invention, multiple different scenes are selected for testing. The scenes should include two types: single object center scene and indoor complex scene, such as Figure 2 As shown. Five fixed-position shooting cameras are set in each scene, numbered 1 to 5, and the captured images are 800×800 pixels in size. The images taken by cameras 1 to 4 are used as the input data set, and the images taken by camera 5 are used as the test data set. Different models use the same input data set for model training until full convergence. The trained model is used to produce synthetic views of camera 5, and a visual comparison is made with the images actually taken by camera 5. This test case selects the four models with the best performance among the existing single-view view synthesis models for comparison to test the single-view synthesis performance of the model. The results are shown in the figure. Figure 4 At the same time, we also selected the three best performing models among the existing generalized view synthesis models for comparison to test the generalized view synthesis performance of the model. The results are shown in Figure 5 shown.
[0135] Figure 4 The comparison chart of the results of synthesizing new views of example scenes by the single scene view synthesis model is shown. The present invention selects the four models with the best performance among the existing single view view synthesis models as comparison, namely DSNeRF, 3DGS, DINER and SparseNeRF, and lists the new view synthesis results of different models in two types of scenes: single object center scene and indoor complex scene. From the comparison in the figure, it can be seen that the new view synthesis result of the present invention has higher fidelity, especially on the surface of objects with rich texture details, it can better restore the surface features of the object, while avoiding the generation of blur and artifacts.
[0136] Figure 5 This is a comparison chart of the results of new view synthesis of the generalized view synthesis model for the instance scene. The present invention selects the three models with the best performance in the existing generalized view synthesis model for comparison, namely MVSNeRF, IBRNet and MVSGS, and lists the new view synthesis results of different models in two types of scenes: single object center scene and indoor complex scene. From the comparison in the figure, it can be seen that in the generalized scene, the new view synthesis result of the present invention can generate more delicate scene details, the synthesized view is also closest to the actual observation view, and can effectively restore the surface of objects with specular reflection.
[0137] In summary, the present invention can realize the synthesis of high-fidelity new view images from sparsely sampled scenes, and supports the generation of visible light images of any viewing angle in the scene, which is more accurate and efficient.
[0138] It should be noted that the above content only illustrates the technical idea of the present invention and cannot be used to limit the protection scope of the present invention. For ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications all fall within the protection scope of the claims of the present invention.
Claims
1. A new view synthesis system for sparsely sampled scenes based on multi-source image feature fusion, characterized by: At least includes a multi-source feature generation module, a feature aggregation module, a feature mapping module, an appearance-density estimation network, an occupancy estimation network and a synthetic rendering module; The multi-source feature generation module: adopts a dual-branch architecture, including a visible light dominant branch, a depth dominant branch and a feature fusion network, for generating high-dimensional feature information that integrates scene appearance texture and geometric priors; The feature aggregation module: uses a tensor splicing method to aggregate the obtained high-dimensional feature information and geometric clues to obtain aggregated feature information; The feature mapping module: adopts a homography transformation mapping method, uses camera parameters to calculate the homography transformation matrix between the current observation angle and the target new angle, and projects the aggregated feature information of the observation angle to the target new angle to obtain the aggregated feature information under the target new angle; The appearance-density estimation network: a multi-layer perceptron comprising multiple fully connected layers, including at least an appearance feature estimation module, an appearance calculation module and a density feature estimation module, to obtain color appearance information and density feature information of the target under a new viewing angle; The occupancy estimation network: adopts a symmetrical encoder-decoder architecture, which is composed of a three-dimensional convolutional layer and a transposed convolutional layer with skip connections; inputs density feature information and original occupancy information to the occupancy estimation network, and outputs density information under a new perspective of the target; The synthesis rendering module takes the color appearance information and density information of the target under the new viewing angle as input, and outputs the visible light image and depth image synthesized under the new viewing angle of the target.
2. The sparse sampling scene new view synthesis system based on multi-source image feature fusion as claimed in claim 1, characterized in that: In the visible light dominant branch of the multi-source feature generation module, the collected visible light image and depth image are used as input, the depth image is corrected using the visible light image, and the corrected depth image is output; In the depth-dominant branch, the original input depth image and the corrected depth image are used as input, and the output is a multi-dimensional feature image; In the fusion network, the position-encoded visible light image and the multi-dimensional feature image are used as input, and the output is the final extracted high-dimensional feature information.
3. The sparse sampling scene new view synthesis system based on multi-source image feature fusion as claimed in claim 2, characterized in that: The visible light dominant branch and the depth dominant branch of the multi-source feature generation module are both composed of a codec network with symmetric jump connections, including multiple cascaded nonlinear activation residual blocks; wherein the position encoding method of the visible light image is Fourier coding; and the fusion network is composed of a batch normalization layer and a ReLU activation function.
4. The sparse sampling scene new view synthesis system based on multi-source image feature fusion as claimed in claim 1, characterized in that: The homography transformation matrix in the feature mapping module is specifically: In the formula, H i→t is the homography transformation matrix from observation view i to target view t, D t is the sampling depth, I is the unit matrix, K, R, T are the camera parameters, is the principal axis unit vector of the target view.
5. A method for synthesizing new views of sparsely sampled scenes based on multi-source image feature fusion using the system as claimed in claim 1, characterized in that: The steps include: S1. Data preprocessing: Collect visible light images and depth images of different observation angles under the same scene, and preprocess the collected image data to ensure that the visible light image under each observation angle has a corresponding depth image and camera parameter data to form a scene data set; S2, high-dimensional feature extraction: Initialize the multi-source feature generation module, input the visible light image and depth image of the observation angle in the scene data set into the multi-source feature generation module, extract and generate high-dimensional feature information; S3, aggregate feature acquisition: using the corrected depth image as a geometric clue, the high-dimensional feature information extracted in step S2 is aggregated with the geometric clue through a feature aggregation module to obtain aggregate feature information; Through the feature mapping module, the aggregated feature information is projected and transformed to the new target perspective to obtain the aggregated feature information under the new target perspective; S4, color appearance information and density feature information acquisition: input the aggregate feature information of the target under the new viewing angle acquired in step S3 into the appearance-density estimation network to acquire the color appearance information and density feature information of the target under the new viewing angle; S5, density information and visible light image acquisition: input the density feature information obtained in step S4 into the occupancy estimation network to obtain the density information of the target under the new viewing angle, and then input the density information and the color appearance information obtained in step S4 into the synthesis rendering module to obtain the reconstructed visible light image under the new viewing angle of the target; S6, iterative optimization: Using the visible light image of the target under the new perspective obtained in step S5 as prior information, establish hybrid perception constraints as prior supervision, iteratively optimize the parameters of the multi-source feature generation module, appearance-density estimation network and occupancy estimation network to achieve new view synthesis of the scene.
6. The method for synthesizing new views of sparsely sampled scenes based on multi-source image feature fusion according to claim 5, characterized in that: The step S3 projects the aggregated feature information of the observation angle of view to the target new angle of view to obtain the aggregated feature information under the target new angle of view, specifically: F f,i→t (p)=concat[F f,i (H i→t (D t )p T ),(X,Y,Z)] In the formula, F f,i→t is the aggregated feature information transformed from the observation perspective i to the target perspective t, p is the position coordinate of the feature image, and F f,i is the characteristic information of observation angle i, H i→t is the homography transformation matrix from observation view i to target view t, D t is the sampling depth, and X, Y, and Z are the geometric clues of the construction.
7. The method for synthesizing new views of sparsely sampled scenes based on multi-source image feature fusion according to claim 5, characterized in that: The step S4 specifically includes the following steps: S41: The aggregated feature information from multiple observation perspectives to the target new perspective is input into the appearance-density estimation network to obtain the visible light features and weight vector: In the formula, w i are the visible light features and weight vectors respectively, G0 is the appearance feature estimation module in the appearance-density estimation network, K is the number of observation angles, and F f,i→t is the aggregated feature information transformed from the observation perspective i to the target perspective t, and p is the position coordinate of the feature image; S42: Obtain color appearance information based on the light feature and the weight vector, specifically: Where C is the synthesized color appearance information, G0 is the appearance calculation module, w i are the visible light feature and weight vector respectively, K is the number of observation angles, and p is the position coordinate of the feature image; S43: Obtain density feature information. The specific calculation process is as follows: In the formula, f d is the estimated density feature information, G d is the appearance calculation module, w i are the visible light features and weight vectors respectively, K is the number of observation angles, and p is the position coordinate of the feature image.
8. The method for synthesizing new views of sparsely sampled scenes based on multi-source image feature fusion according to claim 7, characterized in that: In step S5, the calculation method for obtaining density information through the occupancy estimation network is: σ(p)0G occ (ψ(p),f d (p)) In the formula, σ is the estimated density information, G occ is the occupancy estimation network, f d is the estimated density feature information, p is the position coordinate of the feature image, and ψ is the original occupancy information; The specific calculation process of obtaining the reconstructed visible light image under the new viewing angle of the target in the synthetic rendering module is as follows: In the formula, and They respectively represent the visible light image and depth image reconstructed from the new perspective of the target, render is the synthetic rendering module using the volume rendering method, C is the synthesized color appearance information, σ is the estimated density information, and p is the image position coordinate.
9. The method for synthesizing new views of sparsely sampled scenes based on multi-source image feature fusion according to claim 5, characterized in that: The mixed perception constraints in step S6 include visible light error constraints, local depth constraints and global depth constraints, wherein: The visible light error constraint is: Where, L c is the visible light error constraint, w and h are the image sizes in the scene dataset, and I gt (p) is the visible light image actually observed at the new viewing angle of the target. is the visible light image reconstructed from the new viewing angle of the target, and p is the image position coordinate; The local depth constraint is: In the formula, is a local depth constraint, patch is an N×N pixel block, is the reconstructed depth image, D is the actual observed depth image, and ε is the redundant constant considering small differences; The global depth constraint is: In the formula, is the global depth constraint, Represents the discrete distribution of the reconstructed depth image, gauss(D i ) represents the Gaussian mixture model representation of the actual observed depth image, and R is the number of discretized intervals of the depth value.
Citation Information
Patent Citations
Sparse sampling-based method and system for generating images from shot images to any viewpoint images
CN114820945A
New view angle synthesis method, device, equipment, medium and computer program product
CN118446909A
Prior-incorporated deep learning framework for sparse image reconstruction by using geometry and physics priors from imaging system
US20230368438A1